Audio signal processor and related methods and computer programs for generating two-channel audio signals using smart distribution of computations to physically separate devices
The system addresses computational and latency issues in binaural synthesis by integrating directivity information and distributing processing tasks, achieving precise localization and externalization of virtual sound sources in multimedia applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BRANDENBURG LABS GMBH
- Filing Date
- 2023-10-24
- Publication Date
- 2026-04-20
Smart Images

Figure 2026512676000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, method, or computer program for audio playback, such as binaural playback via headphones or speakers. In particular, the present invention relates to processing a digital audio signal together with acoustic data describing an acoustic environment. [Background technology]
[0002] Using the latest binaural audio rendering technology, users can simulate and hear virtual sound sources that can be precisely positioned in space. The simulated sound appears to originate from outside the head, a phenomenon known as "externalization." With a proper system, binaurally rendered sound sources can be perceived as having a stable position in space and appear to have acoustic properties similar to those of actual sound sources. This makes them virtually indistinguishable from real sound sources.
[0003] There are several binaural synthesis methods and algorithms that can be used to achieve externalization. They all aim to approximate the filtering effects that sound undergoes along a simulated path to the listener's ear. The combined filter of a system consisting of the sound source, the acoustic effects of a virtual or real environment and its geometric shape, the listener's head and body, and potentially other influences on the sound caused by the environment, is called a binaural chamber impulse response (BRIR).
[0004] The two main components of BRIR are the head-related transfer function (HRTF) and the intra-intra-iron impulse response (RIR). The HRTF encodes the measured or approximated filtering effect of the human head, torso, and outer ear. Therefore, it depends on the listener's head and geometric shape, as well as the relative position and rotation of the head and sound source.
[0005] RIR encodes the filtering effects of a room introduced by the room's geometric shape, namely reflection, diffraction, and sound shadowing. It depends on the room's geometric shape as well as the position and rotation of the listener and sound source within the room. (In this specification, "room" refers to any type of environment, not limited to buildings.)
[0006] Simulating these effects is often done through complex simulations or lighter approximations, which require complex room geometric models to simulate a convincing room impulse response. Depending on the binaural synthesis algorithm used, current state-of-the-art algorithms often have to make trade-offs between computational complexity, the limitations of smaller target system sizes, or the effectiveness of the simulation, often resulting in poorly localized or perfectly in-head localized sound sources.
[0007] Furthermore, these devices require room geometry data of the current room, including reflective surfaces and their absorption and scattering coefficients. This data is particularly difficult to obtain in augmented reality (AR) settings where device use is not limited to a single room. Acquiring it is usually impossible even for trained users, and measuring it automatically is a challenging task.
[0008] Depending on the binaural synthesis algorithms and techniques used, these processes can be highly computationally intensive and time-consuming. However, processing power is often limited by the target device. For example, binaural rendering may be deployed in "true wireless earbuds" or similar smart headphones or wearables that offer very limited processing power to provide adequate battery life.
[0009] These devices are often wirelessly coupled to other devices, such as smartphones, via Bluetooth or similar wireless protocols. However, these connections introduce additional delays by requiring coding, conversion, and wireless transmission. This delay typically far exceeds the maximum motion-to-sound latency required to achieve externalization. Here, motion-to-sound latency describes the time frame required for a binaural audio system to perceive acoustic changes caused by the user's head movement. The exact audibility threshold for motion-to-sound latency varies and depends on the listener, the signal used, and the acoustic characteristics of the environment. A latency of up to 50ms has been determined as a valid threshold, which is inaudible to most users under most circumstances.
[0010] To generate compelling virtual sound sources, binaural signals and binaural filters are typically updated at high speed. Depending on the binaural synthesis method used, this results in computational complexity that is often too large for mobile and wearable devices. Instead, such devices are often cabled to another computing device to handle these calculations.
[0011] The publication "Proof of Concept of a Binaural Renderer with Increased Plausibility" by U. Sloma, et al., DAGA 2023 Hamburg, pages 208-211, describes a proof-of-concept demonstration comparing a real-world speaker setup in a given room with headphone-based rendering. In particular, it includes room acoustic processing, which is handled at runtime. Specifically, the binaural room impulse response (BRIR) is calculated in real time based on a single omnidirectional room impulse response (RIR). A very basic room geometric model, as well as the positions of the sound source and microphones, must be incorporated. From this, the direction of arrival (DOA) of direct sound and early reflections is estimated by a simplified image source model. The RIR is processed in segments and appropriately convolved with a common HRTF filter. Late reverberation is simulated by noise shaping. This algorithm enables 6DoF rotation and translation. Furthermore, a spatial resolution method is described. This method uses one measurement microphone and six electret condenser microphones. The sound field is assumed to consist of a sequence of individual acoustic events, which can be described by the captured RIR and captured DOA. In post-processing, the HRIR is calculated relative to the measurement location using 3DoF rotation and a general-purpose HRTF filter.
[0012] The publication "Creation of Auditory Augmented Reality Using a Position-Dynamic Binaural Synthesis System—Technical Components, Psychoacoustic Needs, and Perceptual Evaluation," S. Werner, et al., Applied Sciences, 2021, 11, 1150, discloses a position-dynamic binaural synthesis system used to synthesize ear signals for a moving listener. The goal is to fuse auditory perception of virtual audio objects with the actual listening environment. For each possible position of the listener in a room, a set of binaural room impulse responses (BRIRs) that match the expected auditory environment is required to avoid room divergence effects. The required spatial resolution of the BRIR positions can be estimated by the spatial auditory threshold. In particular, a specific position-dynamic binaural synthesis system relies on a real-time processing block and convolutional engine including preprocessing of room geometry, spatial resolution of playback, listening position representation, acquired tracking data and processing, as well as a filter creation block including listening position and BRIR synthesis. The result of BRIR synthesis is a binaural filter used by the convolution engine in the real-time processing block for position-dynamic binaural playback. We discuss synthesis techniques for fitting constant reverberation, acoustic shaping, initial time delay gap (ITDG), sound source directivity, and real-time processing.
[0013] The publication "Binauralization of Omnidirectional Room Impulse Responses—Algorithm and Technical Evaluation," C. Porschmann, et al., Proceedings of the 20th International Conference on Digital Audio Effects (DAFx-17), Edinburgh, UK, September 5-9, 2017, pp. 345-352, discloses the binauralization of an omnidirectional room impulse response algorithm for synthesizing a BRIR dataset for dynamic auditoryization based on a single measured omnidirectional room impulse response (RIR). Direct sound, early reflections, and diffuse reverberation are extracted from the omnidirectional RIR and spatialized separately. Spatial information is added according to assumptions about the room geometry and typical characteristics of diffuse reverberation. The initial part of the RIR is described by a parametric model; therefore, changes in listener position are possible. The late reverberation portion is synthesized using binaural noise that fits the energy decay curve of the measured RIR. The direct sound frame begins at the start of the sound and ends 10 ms later. The next time section is allocated to the transition to early reflections and diffuse reverberation. Sections with strong early reflections are determined. Following this procedure, small window sections of the omnidirectional RIR describing the early reflections are extracted. The incident direction of the synthesized reflections is based on a spatial reflection pattern adapted from a shoebox room with asymmetrically positioned sources and receivers. A fixed lookup table including the incident direction is used. This creates a parametric model of the direct sound and early reflections. The amplitude, incident direction, delay, and envelope of each reflection are stored. A binaural representation of the early geometric reflection portion is obtained by convolving each window section of the RIR with each HRIR(φ) of the direction. Interpolation in the spherical region is performed to synthesize the provisional directions between given HRIRs. The initial portion of a single measured omnidirectional RIR contains the direct sound and strong early reflections. For this portion, the incident direction is modeled to reach the listener from an arbitrarily selected direction.The later portion of the RIR is thought to be diffuse and is synthesized by convolving binaural noise into a small section of omnidirectional RIR. This approximates the characteristics of diffuse reverberation. The synthesized BRIR can be adapted to the listener's shift and thus allow for the audibility of a freely selected position within the virtual room.
[0014] Existing BRIR synthesis algorithms have been found to have several drawbacks, including computationally high processing costs, unnatural sound perception by listeners, difficulties in efficiently adapting the system to specific source characteristics, location, or orientation, or the location or orientation of a particular listener, or even the inability to operate the system in real time. Furthermore, a further drawback is the generation of artifacts that reduce the externalization of sound impressions, contributing to an unnatural and unpleasant sensation for the listener. [Prior art documents] [Non-patent literature]
[0015] [Non-Patent Document 1] U. Sloma, et al., DAGA 2023 Hamburg, pages 208-211, publication "Proof of Concept of a Binaural Renderer with Increased Plausibility" [Non-Patent Document 2] Publication “Creation of Auditory Augmented Reality Using a Position-Dynamic Binaural Synthesis System-Technical Components,Psychoacoustic Needs,and Perceptual Evaluation”, S.Werner,et al.,Applied Sciences,2021,11,1150 [Non-Patent Document 3] Publication "Binauralization of Omnidirectional Room Impulse Responses - Algorithm and Technical Evaluation", C. Porschmann, et al., Proceedings of the 20th International Conference on Digital Audio Effects (DAFx-17), Edinburgh, UK, September 5-9, 2017, pp. 345-352
Summary of the Invention
Problems to be Solved by the Invention
[0016] <0Next, specific improvements to the algorithm will be described with respect to seven aspects of the present invention. It should be emphasized that implementations of a single aspect in existing systems already bring about significant improvements beyond the scope of the art. However, it is possible to combine subsets of the seven aspects, or even combine all seven aspects with each other, in order to achieve an improved audio signal processor for generating two-channel audio signals. Therefore, it should be emphasized that the seven aspects described below can be used separately from each other or combined in any way, namely, for example, the third aspect and the fifth aspect can be combined, or the third through seventh aspects can be combined, or the first through fourth aspects and the seventh aspect can be combined, and so on.
[0020] According to a first aspect of the present invention, specific source characteristics, particularly the directivity information of the sound source, are integrated into the two-channel synthesis for the purpose of synthesizing two-channel audio data from single-channel audio data. This integration of sound source directivity information can be performed in particular during the processing of the direct tone (DS) portion of the single-channel audio data, which describes the acoustic environment. However, it is also possible to integrate directivity information that can naturally reproduce sound sources with non-omnidirectional directivity characteristics and integrate it into the processing of the early reflection (ER) portion of the single-channel audio data, or to efficiently integrate directivity information into both direct tone processing and early reflection processing.
[0021] According to a second aspect of the present invention, specific processing of the early reflection (ER) portion of single-channel acoustic data is enhanced. In particular, the early reflection portion is segmented into a plurality of segments, each segment containing a specific reflection. In particular, a plurality of image source positions representing the source of the reflected sound are determined, and these image source positions are associated with the segments using a matching operation of the present invention that depends on the sound arrival time calculated for each image source at the listener position in the initial measurement. Next, matching is performed to associate the sound arrival time of each image source with a specific segment, i.e., a specific reflection within the segment. Thereby, an automated high-quality association of image source positions for different early reflections is obtained. By further integrating the directivity information not only for the direct sound but also for individual image sources, the specific orientation of the image sources can also be considered in order to achieve a more natural sound reproduction.
[0022] According to a third aspect of the present invention, the processing of the early reflection portion of single-channel acoustic data describing the acoustic environment is enhanced by calculating two-channel acoustic data for the early reflection portion, taking into account not only the specular reflection portion describing the clear early reflection but also the diffusion portion describing the diffusion effect within the early reflection portion. It has been found that the "second part" of the room impulse response shows significant early reflections, but does not consist only of these. Instead, even this early reflection portion has a significant diffusion portion with an increasing influence in the process from the start to the end of the early reflection portion, i.e., near the start of the late reverberation portion of the room impulse response. Therefore, by calculating two-channel audio data describing the acoustic environment using the diffusion contribution even in the early reflection portion, for example, by relying also on the diffusion portion, two-channel audio data generated by a sound generator using two-channel acoustic data not only for the specular portion but also for the early reflection portion can be supplied to a speaker, and an artificial acoustic scene can be more naturally auditoryized.
[0023] According to a fourth aspect of the present invention, this relates to an improved calculation of the late reverberation (LR) portion of single-channel acoustic data, such as BRIR or BRTF (binaural room transfer function), and depends on the specific generation of a two-channel late reverberation portion by combining amplitude data derived from single-channel acoustic data with, preferably, a binaural two-channel noise sequence. Thus, the generation of two channels from one channel is achieved by using the same amplitude but different phase values.
[0024] In particular, a preferably binaural noise sequence consisting of two channels is transformed into the spectral domain using a short-time Fourier transform or any other time-domain / frequency-domain transformation algorithm. This yields two spectrograms. Furthermore, the late reverberation portion or a combination of the early reverberation portion and the late reverberation portion of single-channel acoustic data is also transformed into a spectral representation, preferably using the same transformation algorithm. Next, two-channel acoustic data of the environment is derived by relying on the same amplitude, which can also be low-pass filtered for actual generation using two phase spectra, for example, and these two obtained spectrograms are then transformed into the time domain to obtain the diffuse portion of the processed late reverberation portion, preferably the diffuse portion of the processed early reflection portion, as described above with respect to the fourth embodiment. Thus, the particular procedure for calculating the diffuse signal can be applied to the late reverberation portion only, or to the calculation of the diffuse portion of the early reverberation portion only, or to the calculation of both the early reflection portion and the late reverberation portion, as in the preferred embodiment of the present invention. In particular, for the calculation of combined early reflections and late reverberation portions, the calculation of the binaural diffusion portion is performed without any knowledge of any separation between the early reflection portion and the late reverberation portion; therefore, separation into these portions is not necessary at all, and consequently, no such separation between the early reflection portion and the late reverberation portion is necessary in this aspect of the present invention. This method results in a particular saving of computational resources. Furthermore, even higher audio quality is obtained, particularly for the calculation of the late reverberation portion, such that any changes depending on the listener position or source position or orientation do not need to be considered in order to further improve the efficiency of the algorithm. Changes depending on the listener position or source position or orientation also do not need to be considered in the calculation of the diffusion portion of the early reflection portion of the room impulse response.
[0025] According to a fifth aspect of the present invention, the problem is addressed by a method for efficiently and flexibly acquiring high-quality single-channel acoustic data, such as a single-channel chamber impulse response, which is of sufficient quality to obtain high-quality auditory representation. For this purpose, the input interface is configured to acquire a raw representation related to the single-channel acoustic data, and the input interface is further configured to derive the single-channel acoustic data using the raw representation and additional data stored in or accessible by the audio signal processor. Thus, initial measurements relying on natural sounds that can be generated by the user, such as the user clapping their hands or stomping their feet on the floor, or even speech signals, can be used instead of the commonly used sinusoidal sweep signal, which is a very unnatural signal and, of course, cannot be generated by the listener at all.
[0026] Furthermore, initial measurements can be performed using low-quality microphones, such as those found in laptops or mobile phones, and then, based on this raw representation related to single-channel acoustic data, synthesis, or generally the generation of high-quality single-channel acoustic data, can be performed using a database matching process that relies on tests and reference fingerprints, or synthesis can be performed using one or more neural networks that also rely on acquired raw representations such as initial measurements, or geometric data on the acoustic environment, as well as possibly the intended source position and the intended or initial listener position.
[0027] Another procedure in this embodiment involves simply recording sound, such as a musical piece, played by a speaker or multiple speakers in a specific acoustic environment, and retrieving a database, which is usually remote, through some kind of audio fingerprinting process against the original version of the sound played by the speakers. By using the clear or ideal sound played by the speakers and the sound with the effects of room acoustics, a room impulse response or room transfer function, or generally two-channel acoustic data, can be calculated.
[0028] This procedure effectively addresses the problem of having a single-channel intra-intravenous impulse response that is good enough to perform useful calculations of the head-related impulse response based on a specific listener and source location.
[0029] According to a sixth aspect of the present invention, processing tasks can be distributed across multiple different devices having different power sources. This allows most of the tasks to be performed on wearable devices such as headphones, earphones, or in-ear elements, while the second device is a device with a large battery, such as a mobile phone, smartwatch, tablet, or notebook computer or fixed computer.
[0030] In particular, it has been found that the most computationally expensive part is the calculation of late reverberation, and to some extent, the calculation of early reflections. However, it has been found that the update rate of these procedures can be lower compared to the update rate of the direct sound calculation. On the other hand, the direct sound calculation is computationally inexpensive because this part is a short time segment and therefore requires only short filters that can be processed very efficiently.
[0031] Therefore, the processing task of calculating the direct sound portion can be easily performed by low-power devices such as wearable devices, and more demanding tasks can be performed by a separate second device. The resulting transmission delay is not a problem because a lower update ratio is sufficient for computationally more aggressive calculations, namely the calculation of the early reflection portion, and especially the calculation of the late reverberation portion, which, depending on the specific acoustic environment, has a considerable length of time in which the room impulse response is considered. In particular, in reverberation chambers such as churches, the late reverberation portion can extend over several seconds of diffuse reverberation.
[0032] According to the seventh aspect, it was found that after calculating the individual parts, particular attention must be paid to separating the room impulse response from the combination. In particular, in order to have a high-quality system that allows different parts (DS, ER, LR) to be calculated by individual processes and combines the results without suffering from audio quality problems arising from the separation into individual parts and the combination of the individually calculated results, specific extensions of the corresponding parts at separation times, such as between the direct sound and the early reflection part or between the early reflection part and the late reverberation part, must be performed to obtain the overlap range at the corresponding separation time. Furthermore, in order to avoid any artifacts and to enable seamless processing that needs to be done in a very short time, at least one extension part is windowed using a window function that takes sample extension, i.e., overlap, into account. A particular window that has been shown to be very useful for the purposes of RIR processing is a Tukey window with lobes having a width of 2n, where n is a specific number of samples used for extension of the part.
[0033] Alternatively, in the overlap portion, the Tukey window is selected such that an overlap of n=16 samples is preferably observed. The overlap may be in the range of 8 to 32 samples. The remaining samples maintain 100 percent amplitude, i.e., a window coefficient of 1. Thus, there is a Tukey window with a small number of samples (e.g., 16) as a lobe that results in seamless transition characteristics between the two corresponding parts of DS and ER, and / or ER and LR.
[0034] A further issue related to this embodiment is the integration of an initial time delay gap (ITDG), which can preferably be performed within this overlap range between the direct sound portion and the early reflection portion. Thus, since the overlap range is typically larger than the maximum ITDG movement range, forward and backward movement of the ITDG is not an issue. Therefore, even if the overlap is no longer ideal, when movement relative to the ITDG is performed, this nevertheless proves to be sufficiently accurate.
[0035] This invention describes an unprecedented system for the auditory representation of binaural audio. It synthesizes precisely positionable virtual sound sources that appear disjointed in the physical environment surrounding the user, using the acoustic and geometric properties of the environment. The binaurally rendered sound sources can be perceived as having a stable position in space and appear to originate from outside the head, a phenomenon referred to as "externalization." In this system, the virtual sound sources can be perceived as indistinguishable from real sound sources. This is achieved by combining a sound source filter (directional transfer function - DTF), the acoustic influence of the environment (room impulse response - RIR), and the listener's head and body (head-related transfer function - HRTF) to obtain a binaural room impulse response (BRIR). Processing of the binaural signal in response to the user's motion and acoustic environment enables the externalization of sound and interaction with the system. Applications of the described system and method include multimedia applications, including digital audio playback, virtual reality, and augmented reality.
[0036] In its most basic embodiment, the system consists of a single device containing all necessary sensors, components, and sound transducers. The system includes the necessary components in a headphone or earplug form factor, and all processing can be performed directly on the device. In other embodiments, the system operates on a distributed device. The disclosed system consists of three main functional components that work together to generate binaural signals in real time. The first component provides an omnidirectional RIR, such as an RIR recorded from an omnidirectional speaker, or using an omnidirectional microphone, or preferably using both omnidirectional elements, which has desired acoustic characteristics and includes relevant acoustic cues of the environment. This includes, in particular, the frequency-dependent energy distribution of reverb over time. In one embodiment, the RIR provider uses a speaker and an omnidirectional microphone to hold qualitative in-situ measurements of the RIR. Supplementally, the system can estimate (psycho)acoustic parameters of low-quality RIR measurements or ambient noise, synthesize an RIR from these parameters, or select a suitable high-quality RIR from a database. The system also incorporates, for example, machine learning techniques to support parameter estimation. If necessary, multiple RIRs can be blended to improve transitions between different acoustic environments, such as in combined rooms.
[0037] The second component is a binaural synthesizer, which acquires the RIR, adds binaural cues, and converts the RIR into a BRIR. The binaural synthesizer further receives room geometric information as input. In one embodiment, the room geometric information consists of the geometric shape of a shoebox that approximates the user's actual environment by fitting it to a rectangular room consisting of six surfaces. This provides estimates of acoustic reflective surfaces in the environment, particularly the floor, ceiling, and walls closest to the listener. While the simplification by the room geometric shape of the shoebox already yields good results, improvements can be made from a more accurate geometric model of the room. The RIR itself is processed in segmented segments motivated by fundamental research in the field of psychoacoustics. Direct sound represents the first sound wave that reaches the listener directly. Here, the influence is given by the HRTF and DTF, as well as the distance laws of sound propagation. These cues can be applied straight ahead. For reflections in the room represented by the RIR, there is a transition from specular to diffuse. A given RIR is combined with phase information from a binaural noise sequence to produce a diffuse layer of the BRIR. The early reflection segment is divided into blocks, each assigned an estimate of the ratio between specular reflection energy and diffuse energy. The blocks are then convolved with HRTF and optionally DTF to obtain a directional segment layered with snippets from the diffuse segment, depending on the index. After combining the three segments, BRIR is completed.
[0038] The binaural synthesizer is connected to a position sensor, which can determine the user's head rotation relative to a reference system, and further, its position. This pause information (where "pause" represents the listener's position and orientation, or the source's position and orientation) is provided in real time by a position tracking system. A virtual source pause is provided by a preset and optionally changes over time as a moving source. The binaural synthesizer is connected to a system that generates a measured or synthesized HRTF corresponding to the direction of arrival. Similarly, part of the system unravels the directivity transfer function (DTF) of the sound source depending on the relative position. As with the HRTF, the DTF can be derived from a measurement or synthesis process. The synthesized BRIR is sent to an Auralizer, where it is convolved with the audio signal in real time. For this, state-of-the-art block-based real-time convolution methods can be used.
[0039] The resulting binaural audio signal is then played back through headphones, although crosstalk-canceling speakers may also be used. To maintain the plausibility of the externalized sound source, the BRIR needs to be periodically resynthesized with current position data. In some embodiments, the three segments can be computed at different speeds while maintaining an immersive experience. The described system represents novelty in the field of binaural synthesis. It enables the experience of realistic virtual spatial sound.
[0040] This document describes a system for reasonable binaural playback of digital audio. It enables the auditory representation of virtual sound sources (sound sources that do not exist in the user's actual listening environment) by finding and incorporating room impulse responses (RIRs) similar to those belonging to actual rooms, without directly measuring them.
[0041] RIR is an impulse response that describes the combined filtering effect of the sound source, sink, and the precise influence of the room (environment) on the acoustic signal for a particular configuration of these elements. Therefore, the measured RIR depends on the position and spectral characteristics of both the source and sink, among other influences. Similarly, binaural room impulse response (BRIR) describes the filtering effect of the source, the local influence of the environment, and the influence of human anatomical structures (e.g., the outer ear, head shape, and torso).
[0042] The following describes a solution for deriving BRIR for any configuration of related components, and even for new positions of virtual sound sources, and for using them in binaural synthesis and rendering in real-time scenarios. This allows the rendered sound source to be stably externalized and perceived outside the head.
[0043] This solution divides the problem into three parts, which are components of the processing chain.
[0044] 1. Derive a new RIR from available audio recordings, taking room acoustics into account.
[0045] 2. Use the RIR to extrapolate the BRIR for the specific configuration of the listener, source, and the room being heard.
[0046] 3. Playback of binaural audio to the device user.
[0047] All necessary processing steps can be performed on a single device, and all required subsystems can be combined. However, in its most basic form, it consists of two systems connected by a network.
[0048] Furthermore, it should be noted that the three parts described above can be applied independently of each other, and the other two corresponding parts are not implemented as described but are implemented through alternative solutions. Alternatively, for the most favorable result, the three parts can be implemented together. Alternatively, only two of the three parts can be combined, but the remaining parts are not implemented as described but are implemented through alternative solutions.
[0049] The first system comprises at least one microphone (or an array of microphones), at least one processor and playback device capable of delivering binaural audio, such as headphones or speakers, and a device capable of measuring the position and movement of a user (head) in the environment, such as an IMU or an optical tracking system. The second system comprises at least one processor and non-temporary memory.
[0050] This invention describes a system for the auditory representation of binaural audio. It uses information about the user's physical environment in the form of an impulse response or reverberation audio signal. It acquires its acoustic properties and synthesizes a well-externalized, precisely locatable virtual sound source that appears to originate from the physical environment around the user.
[0051] The audio rendering system of the present invention allows users to simulate and hear virtual sound sources that can be precisely located in space. The simulated sound appears to originate from outside the head, a phenomenon referred to as "externalization." With a suitable system, binaurally rendered sound sources can be perceived as having a stable location in space and appear to have acoustic properties similar to those of actual sound sources. This makes them virtually indistinguishable from actual sound sources.
[0052] This effect is achieved by precisely controlling the sound that reaches the user's eardrum. Typically, two speakers are used, each nearly reproducing (or audibly reproducing, "making audible") the sound that reaches the listener's ears. The reproduced audio signal can then be played directly into the ears using headphones. Alternatively, the crosstalk-canceled speakers may be used further away from the user's ears.
[0053] The embodiment uses psychoacoustic knowledge to reduce the computational complexity of the system and to enable distributed computation of binaural synthesis on devices connected by transmission channels that introduce a greater delay into the signal processing than otherwise permissible.
[0054] BRIR combines multiple filter effects. They can be split at any point, resulting in any number of subfilters. These can then be reassembled by summing the individual parts with respect to the individual delays, or by convolving the filters with the complete or partial signal and summing the resulting signal with respect to the individual delays. The same basic segmentation and summing process is also valid if some or all of the audibly perceived signal is not processed by convolving BRIR with the signal, but instead simulated directly, i.e., by using a delay network-based method.
[0055] By using psychoacoustic knowledge about how different parts of a binaural filter are perceived differently, rendering systems can be designed to compute less critical parts of the filter less frequently and distribute those computations across devices.
[0056] In one form, the system consists of a single device capable of synthesizing and auditoryizing binaural signals in real time. This includes at least two speakers, each capable of playing sound for one ear, i.e., all types of common headphones or crosstalk-canceling speakers.
[0057] The system further includes one or more position sensors capable of determining the user's head rotation relative to a reference system. (This is commonly referred to as 3-degrees-of-freedom or 3DoF tracking.) In a different embodiment, the system instead includes one or more position sensors capable of determining their position relative to the reference system, in addition to the user's head rotation relative to the reference system. (This is commonly referred to as 6-degrees-of-freedom or 6DoF tracking.) The system can process binaural filters or directly simulate the auditoryized signal by using one or more suitable binaural synthesis algorithms. It is not dependent on one particular method of auditoryization. Different embodiments of the system may use different binaural synthesis algorithms.
[0058] In this embodiment, the binaural synthesis algorithm used must be able to compute filters for the direct sound path and room reverb separately. Auditory representation of the direct sound path is typically achieved by block convolution of the audio signal with filters that approximate the filtering effects of the user's head, ears, and torso on the sound source at a given position and distance (HRTF).
[0059] The processing of these filters must encode not only changes in sound intensity and other cues, but also correct changes in interauricular time difference (ITD) and interauricular level difference (ILD). Human listeners are relatively sensitive to even small changes in these values, and therefore, these changes must be computed with good spatial and temporal resolution. However, these filters are relatively short, and deriving them usually involves only a few processing steps.
[0060] Room reverb simulates the filtering effect on sound caused by the environmental geometry of sound that does not travel directly from the sound source to the user's ears. This includes reflection, refraction, absorption, and resonance effects. Such reverb filtering is expected to be much longer than short direct tone filtering. Numerous processes, algorithms, and systems can handle proper binaural reverb, including image source algorithms, ray tracing, parametric reverberators, and many delay network-based techniques.
[0061] In this embodiment, the system utilizes the fact that human listeners are more sensitive to changes in the direct tone filter and less sensitive to changes in the reverberation filter. The signal processor is programmed to compute the direct tone filter much faster than the reverberation filter. This allows the system to minimize the clearly audible jump in the audible sound when the filter is replaced and the sense of externalization increases, while avoiding a complete filter update. Updating these filters or signal portions encoding the direct tone path at a rate of approximately 188 Hz has proven to be a reasonable default for such a system, although in different embodiments of the system, lower refresh rates (such as 94 Hz or 50 Hz, or even lower than 15 Hz) may be feasible. The reverb filter is computed at a much lower rate, typically up to 1 / 10 of the direct tone processing rate, depending on the acoustic characteristics of the environment and the user.
[0062] Either a signal processor or another processor is configured as an aggregator. In some embodiments where the binaural synthesis method used returns a continuous stream of blocky binaural audio signals, this aggregator simply sums the blocks supplied by the direct and reverberation processing paths and functions as a signal aggregator. This requires that the blocks being summed correspond to the same point in time or include control data that identifies the time frames they correspond to. Alternatively, the aggregator can be configured to sum two subfilters with respect to their time delays, as determined by the algorithm. Thus, it reconstructs a full BRIR filter from the results of the individual processors and functions as a filter aggregator. The filters can then be used to convolve the audio signal blocks using a modern real-time (blocky) convolution method. The aggregator always keeps a full BRIR filter in memory. Thus, the BRIR can be partially updated at the individual rates of the individual processors processing the subfilters. The resulting signal blocks contain the combined binaural signals for the direct and reverberation paths. They are then passed to a speaker signal generator and played on the system's speakers. The speaker can be any speaker, such as a speaker within a wearable device, a crosstalk-canceling speaker, or a speaker with some kind of sound separation element in between. This makes it possible to create binaural audio with the same level of externalization and perceptual quality as individual algorithms, while significantly reducing processing requirements.
[0063] A further part or aspect of the solution takes a previously derived RIR as input and synthesizes a BRIR from it. It uses additional metadata for the synthesis process, such as available location data for both the room, listener, and sound source. The system tracks the user's position relative to the audible source and the actual room using a tracking system consisting of one or more sensors, such as an IMU or optical tracking device. It receives metadata about the virtual sound source location and a set of (individual or general) HRTFs. More metadata, such as real or virtual room geometry, sound source directivity, sound source boundaries, etc., may be optionally supplied. For processing, the system can divide the received RIR into arbitrary time segments, which can be processed in parallel with different algorithms and at different intervals. In one embodiment, the RIR is divided into three parts, including direct sound, early reflections, and late reverberation. The direct sound segment is truncated to include a portion of the RIR that contains sound transmitted directly from the source to the receiver but does not include the first reflection that reaches the receiver. The late reverberation segment can begin at some point, after which a single strong reflection is no longer perceptible. The segments are windowed appropriately, for example, using overlapping Tukey windows, so that they can be reconstructed later. The relative positions of the listener and source determine the direction of incidence of the direct sound, which is used to select a fitting HRTF from the set and convolve it with the direct sound segment of the RIR, either directly or interpolated, channel by channel.
[0064] For the entire length of the two reverberation sections, the pseudo-diffuse RIR is calculated by modeling the frequency-dependent energy envelope of the RIR on binaural white noise (a signal with uniformly distributed energy across all frequency bands but containing the phase information of the fully diffused field of the BRIR), while preserving the phase information of the high-density reflection pattern. This can be done by separating the frequency bands using a complete reconstruction filter bank, determining the low-pass envelope for each band, and multiplying the noise signal by it. Alternatively, the RIR and binaural noise can be converted to the time domain, for example, by applying the amplitude of the RIR to the noise while preserving the phase, and then converting it back to the time domain using STFT. The resulting windowed pseudo-diffuse section is then used by the system as the late reverberation of the BRIR.
[0065] The early reflection segments of the RIR are further windowed into sub-windows that may or may not correspond to the locations of one or more early reflections. Similar to direct sound, each detected sub-segment is assumed to have an incidence direction if it corresponds to an early reflection. This arrival direction is derived from a room model of appropriate complexity using an algorithm such as an image source algorithm, or is statistically selected. The HRTF is selected or interpolated based on its direction and convolved with the sub-segments. To overcome the sparseness of this method, the system blends the pseudo-diffuse portion with the fully directional ("spectral") portion to simulate reflections arriving at similar times, and / or diffusion from the nonlinear portion of the RIR.
[0066] To achieve this, a function is used to determine the coefficient of diffusivity in each window, and the diffusive and specular reflection portions of each subsegment are linearly interpolated. A suitable function is formed from the energy ratio of the low pass-mean energy in the small window around the signal to the ratio of the low pass-mean energy in the large window around the signal, and thus the ratio of local energy to short-term mean energy can be approximated as a predictor of the masking effect.
[0067] The resulting subsegments are then windowed and reassembled. Depending on the signal and HRTF used, further post-processing such as diffusion field or headphone equalization may be applied. One embodiment of the system may even use additional knowledge or metadata of the room to pre-process the RIR to adjust for the room's characteristics. For example, the energy decay of late reverberation can be adjusted, or further adjusted by inferring the arrival time of early reflections from a reflection model. The resulting BRIR is then convolved with the audio signal using block convolution to produce real-time audibility.
[0068] This solution further minimizes room acoustic divergence overall and can be tailored to the user using a personalized HRTF. This makes it usable in auditory augmented reality scenarios.
[0069] Next, preferred embodiments will be described with reference to the attached drawings. [Brief explanation of the drawing]
[0070] [Figure 1] This is an overall diagram showing the preferred rationale for the seven embodiments. [Figure 2] A preferred implementation of a two-channel synthesizer is shown, illustrating the procedures of the first to fourth and seventh aspects of the present invention. [Figure 3a] The preferred procedure of the first and / or second embodiment is shown below. [Figure 3b] This table shows what needs to be updated under specific conditions. [Figure 4a] This shows the amplitude representation of the intra-intravenous impulse response / intra-intravenous transfer function in three dimensions. [Figure 4b] This shows the room impulse response when the emission direction is as shown in Figure 4a, i.e., emission in front of the sound source. [Figure 4c]Figure 4b shows the directional transfer function of the directional impulse response. [Figure 5a] A preferred implementation of the first embodiment is shown. [Figure 5b] Further parts of the preferred procedure according to the first embodiment are shown. [Figure 5c] Further processing according to the first embodiment is shown. [Figure 5d] Further steps according to the first aspect are shown. [Figure 6a] This shows a three-dimensional sphere for determining / selecting head-related transfer functions or head-related impulse responses. [Figure 6b] The left and right HRIRs are shown when the user is in the front / left position shown in Figure 6a. [Figure 6c] Figure 6b shows the left and right HRTFs corresponding to the HRIR. [Figure 7] A preferred implementation of a second aspect of the present invention is shown. [Figure 8a] This demonstrates the generation of an image-based sound source up to the first reflection. [Figure 8b] The procedure for a second aspect of the present invention is shown below. [Figure 9] A preferred implementation of the second embodiment is shown. [Figure 10a] An embodiment of a third aspect of the present invention is shown. [Figure 10b] Further embodiments of the third aspect are shown. [Figure 11a] Another preferred implementation of the third embodiment is shown. [Figure 11b] A preferred embodiment of the combination of specular reflection and diffusion portions according to a third aspect is shown. [Figure 12a] This shows the initial time delay gap. [Figure 12b] The application of the initial time delay gap in the third or seventh embodiment is shown. [Figure 12c] Further reference is made to the initial time delay gap (ITDG) in the implementations according to the third and seventh aspects. [Figure 13a] A preferred implementation of the fourth embodiment is shown. [Figure 13b] An embodiment of a fourth aspect of the present invention is shown. [Figure 13c] A preferred implementation of the fourth embodiment is shown. [Figure 13d] Further steps according to a fourth aspect of the present invention are shown. [Figure 14a] A fifth aspect or other different implementation forms that are particularly relevant to the other aspects are shown. [Figure 14b] A fifth aspect or other different implementation forms that are particularly relevant to the other aspects are shown. [Figure 14c] A fifth aspect or other different implementation forms that are particularly relevant to the other aspects are shown. [Figure 14d] A fifth aspect or other different implementation forms that are particularly relevant to the other aspects are shown. [Figure 14e] A fifth aspect or other different implementation forms that are particularly relevant to the other aspects are shown. [Figure 15] A sixth aspect of the present invention shows an implementation configuration of hardware required for the first device on the one hand and for the second device on the other. [Figure 16] A preferred implementation of a fifth aspect of the present invention is shown. [Figure 17a] An actual embodiment of the fifth aspect is shown. [Figure 17b] Other embodiments of the fifth aspect are shown. [Figure 17c] Further embodiments of the fifth aspect are shown. [Figure 18] Further embodiments of the fifth or seventh aspect are shown. [Figure 19] This is a schematic diagram of one embodiment of the sixth aspect. [Figure 20] Further embodiments of a sixth aspect of the present invention are shown. [Figure 21a] Different embodiments of the audio sound generator are shown. [Figure 21b] Different embodiments of the audio sound generator are shown. [Figure 22a] Further embodiments of a sixth aspect of the present invention are shown. [Figure 22b] Further embodiments of a sixth aspect of the present invention are shown. [Figure 22c] Further embodiments of a sixth aspect of the present invention are shown. [Figure 22d] Further embodiments of a sixth aspect of the present invention are shown. [Figure 22e] Further embodiments of a sixth aspect of the present invention are shown. [Figure 22f] Further embodiments of a sixth aspect of the present invention are shown. [Figure 23a] This invention illustrates an implementation in which a sound generator uses a complete two-channel acoustic dataset to generate a two-channel audio signal. [Figure 23b] An alternative embodiment is shown in which the same audio signal is convolved with individual 2-channel data segments, and the individual reproduced binaural audio signals are combined with each other. [Figure 24a] An embodiment according to a seventh aspect is shown. [Figure 24b] Further processing according to a seventh aspect is shown. [Figure 25] This shows another implementation of the seventh aspect, which integrates ITDG coordination. [Figure 26] A preferred implementation of ITDG adjustment according to the seventh or third aspect of the present invention is shown. [Modes for carrying out the invention]
[0071] Figure 1 shows an input interface 100 that can receive multiple inputs, as described later, and provides single-channel acoustic data describing an acoustic environment. The single-channel acoustic data can be a room impulse response or room transfer function, or any other description of an acoustic environment such as a room, an open room, or a semi-open room. The acoustic environment may also be the environment outside the room, depending on the context. Typically, the acoustic environment includes reflective objects such as the walls of the room or furniture, or absorbing objects such as people in the room or curtains in the room, or any other "acoustic objects."
[0072] The audio signal processor further comprises a two-channel synthesizer for synthesizing two-channel acoustic data from single-channel acoustic data using listener position or orientation, as shown in Figure 1. The result of the two-channel synthesizer 200 is two-channel acoustic data such as a binaural room impulse response or binaural room transfer function or, optionally, any other arbitrary two-channel impulse response or transfer function. Other descriptions from impulse responses or transfer functions can also be applied as acoustic data, such as specific parameterizations.
[0073] Two-channel acoustic data is input to a sound generator to produce a two-channel audio signal from an audio signal, which is typically a monaural signal as shown in Figure 1, and two-channel acoustic data received from the two-channel combiner 200 in Figure 1. The input interface may also be referred to herein as an RIR provider. In this specification, a two-channel combiner is also referred to as a binaural combiner, and a sound generator is also referred to as an auralizer. Nevertheless, both descriptions mean the same thing, namely, an RIR provider is generally an input interface, a binaural combiner is a general two-channel combiner, and a sound generator is a general auralizer.
[0074] The 2-channel combiner 200 is configured to separate single-channel acoustic data into at least two parts consisting of a direct sound portion, an early reflection portion, and a late reverberation portion, and the 2-channel combiner 200 is configured to process at least two parts individually to generate 2-channel acoustic data for each part.
[0075] This is shown in Figure 2. In block 210, the single-channel acoustic data is separated into at least two parts. Block 220 shows direct sound processing. Block 230 shows early reflection processing, and block 240 shows late reverberation processing. As shown in Figure 250, all three 2-channel acoustic data from each part are combined by aggregation or combination of 2-channel acoustic data. Furthermore, it should be noted that block 250 covers two alternatives that can generally be implemented. The first alternative is that the individual parts of the BRIR are aggregated into a full BRIR, and then the full BRIR is applied to the audio signal by convolution, as shown in Figure 3a. The convolution is performed by the sound generator 300 in Figure 1, which also receives the audio signal.
[0076] An alternative embodiment shown in block 250 of Figure 2 is also shown in Figure 23b. Here, the aggregation of individual parts of BRIR does not occur. Instead, the binaural audio data is calculated by convolving each part separately into an audio signal so that three streams of binaural audio data are obtained, and then combining the three individual streams, binaural audio 1, binaural audio 2, and binaural audio 3. This allows for processing and aggregation of the audio signal in the sound generator 300, as shown in Figure 1.
[0077] Furthermore, it should be noted that, according to different embodiments of the present invention, it is not always necessary to process all three parts. Instead, in the first embodiment, it is sufficient to separate the single-channel acoustic data into only two parts: the direct sound portion and, for example, the rest of the RIR. For the purposes of the second embodiment of the present invention, where the image source is associated with individual segments, separation into three parts is required because the early reflection portion is located between the direct sound portion and the late reverberation portion. Separation into three parts is also useful for the purposes of the third embodiment, which refers to a specific merging of specular reflection and diffuse portions. However, for the purposes of the fourth embodiment, which relates to a specific calculation based on binaural noise, separation in only two parts is required, with the first part containing the direct sound and early reflection, and the second part containing the late reverberation. For the purposes of the fifth embodiment, no partitions are required at all, and any auditory processing can be performed that requires single-channel acoustic data describing the acoustic environment, because the fifth embodiment relates to providing a room impulse response rather than how it is further processed. However, the fifth aspect can, of course, be combined with all the other aspects, and therefore, in certain embodiments, the fifth aspect can also be used with a separation of two or three parts, as shown in block 210. According to the sixth aspect, separation into at least two parts is required, because it is preferable that direct sound processing is performed on a wearable device and processing of the remaining part is performed on a second device. If three devices are used, separation into three parts is required. The seventh aspect relating to the separation of RIR and the combination of separated parts is sufficient with separation into two parts, and it is preferable that the seventh aspect is also applicable when separation into three parts is performed. The same applies to the introduction of an initial time delay gap, which can be similarly done according to the present invention when there is no separation between the early reflection portion and the late reverberation portion.
[0078] In the preferred embodiment shown in Figure 2, direct sound processing depends on source directivity, initial source or sink data from initial measurements, current listener data, and / or current source data. In this regard, note in the following text that current listener data refers to listener position, listener orientation, or both, also referred to as listener “pause.” The same applies to source data. Source data may be source position, source rotation, or both. Specifically, source rotation can be advantageously considered for non-omnidirectional sources using source directivity information according to the first or second aspect of the present invention.
[0079] The early reflection processing in block 230 depends on initial data such as the listener's position and / or orientation, geometric data regarding the acoustic environment, and typically the association between the image source and the early reflection. Furthermore, source directivity can also be considered in the early reflection processing in block 230.
[0080] The late reverberation processing 240 relies on two-channel noise data, shown as two arrows in Figure 2, to illustrate the transformation of the late reverberation portion, which is a single-channel portion, to the two output channels shown at the bottom of block 240 in Figure 2.
[0081] A preferred embodiment provides a binaural synthesis system that uses RIR as highly simplified room geometric data as input, instead of complex geometric data. The goal is to synthesize virtual sound sources that appear to originate from any location around the user. These are stable and fixed, and respond to the listener's movement, just as real sound sources do. Applications of the described system and method include multimedia applications, including digital audio playback, virtual reality, and augmented reality.
[0082] The processing of binaural signals in response to the user's motion and acoustic environment enables the externalization of sound and interaction with the system. The described device includes another system that transmits audio content and all necessary metadata described below to the described system.
[0083] In its most basic embodiment, the system consists of a single device containing all the necessary sensors, components, and sound transducers. Such a device has two speakers, one for each ear, used to reproduce a binaural signal to the user. The system includes the necessary components in a headphone or earplug form factor, and all processing can be performed directly on the device.
[0084] The disclosed system consists of three main functional components that work together to generate binaural signals in real time. Different embodiments of the system may include different implementations of these components, but their purpose remains the same.
[0085] The first component is the RIR provider. The purpose of this component is to provide an omnidirectional RIR that has desired acoustic characteristics and includes associated acoustic cues for a real or virtual environment or a modified version thereof. This includes, in particular, the frequency-dependent energy distribution of reverb over time. The exact characteristics and cues encoded by the RIR depend heavily on the system's embodiment. The operating mechanism of the component is also dependent. In one embodiment, the RIR provider is connected to a single microphone and speaker. It includes non-temporary memory to hold one or more RIRs.
[0086] RIR can be recorded by any state-of-the-art measurement method capable of providing a good measurement of room acoustic effects. This can be done, for example, by playing an exponential sinusoidal sweep or minimum-length sequence on a speaker and recording the reverberation audio using an omnidirectional microphone within the critical distance of the sound source. The RIR can then be calculated from the reverberation recording and input signal using deconvolution. The recorded RIR is stored in memory.
[0087] The second component is a binaural synthesizer, which receives RIR from an RIR provider as input. At this point, the recorded room impulse response contains important monaural cues for room acoustic effects. Binaural information necessary for spatial hearing and externalized perception by the user must be added to the RIR and converted into BRIR. The binaural synthesizer further receives room geometric information as input. In one embodiment, the room geometric information consists of the geometric shape of a shoebox that approximates the user's actual environment by fitting a rectangular room with six surfaces to the real environment. The width, depth, and height of this shoebox room are provided by the user of the system, but the surfaces must match the main acoustic reflective surfaces in the real environment, particularly the floor, ceiling, and walls closest to the listener. The second component is further connected to one or more position sensors that can determine the user's head rotation relative to a reference system. (This is commonly referred to as 3-degrees-of-freedom or 3DoF tracking.)
[0088] In another embodiment, this instead comprises one or more position sensors that can determine the user's position relative to a reference system, in addition to the user's head rotation relative to the reference system. (This is commonly referred to as 6-degree-of-freedom or 6DoF tracking.) The user's position and rotation in the reference coordinate system are provided in real time by the position tracking system. The position and rotation of a virtual sound source can be provided by a pre-configured setup. In some embodiments, the sound source's position may be periodically changed by an external system representing a moving sound source. In some embodiments, an offset can be periodically added to the user's position and rotation, for example, to simulate the movement of a user representation in a virtual world.
[0089] The binaural synthesizer is further connected to a system capable of approximating an HRTF corresponding to a given relative position between the sound source and the user. In one embodiment, such a system may include a dataset of measured or synthesized HRTFs and the relative position vectors (referred to as direction of arrival or DOA) between the source and sink that they correspond to. Given a DOA as input, it then selects a single HRTF whose corresponding DOA best matches the input DOA, i.e., by maximizing the scalar product of two unit vectors. In other embodiments, the relative position can be taken into account to synthesize a suitable HRTF.
[0090] Similarly, a binaural synthesizer is connected to a system that can approximate the directional transfer function (DTF) of the sound source, which is a relative position-dependent filtering effect of the sound source. Such a system can function by selecting the best match from a database, like an HRTF system, or by synthesizing DTFs.
[0091] Next, the subject matter of the present invention according to the first aspect is shown in Figure 3a. Figure 3a shows a coordinate system with an origin 420 and a listener 100 at the listener's position, which is assumed to be at the origin of the listener's head, with the listener's head being a sphere. Furthermore, a source 410 is shown that is directed away from the listener in the main emission direction 430. For the source, non-omnidirectional transmission characteristics are assumed, for example, as shown in Figure 4a, which shows the amplitude of three-dimensional directional information. As shown in 431, the amplitude behind the sound source, i.e., in Figure 4a, the exemplary speaker, is smaller than the amplitude 440 in front of the speaker.
[0092] Furthermore, in the example of Figure 4a, by placing the listener in front of the speaker, the directional impulse response shown in Figure 4b is obtained. The directional transfer function, i.e., the directional impulse response when converted to the spectral domain, is shown in Figure 4c. Therefore, it can be seen that the sound sources in Figures 4a, 4b, and 4c have strong non-omnidirectional directivity, and furthermore, even when facing the front, they have specific room impulses that exhibit a significant nonlinear frequency response when converted to the spectral domain. Note that the phase is not shown in Figure 4c, but the directional impulse response is a complex directional transfer function.
[0093] According to a second aspect of the present invention, a two-channel combiner, as shown in Figure 1, is configured to determine the directivity information of a sound source for a specific listener position and the source position and / or orientation of the sound source. Furthermore, the two-channel combiner is configured to use the directivity information in the calculation of two-channel acoustic data for the direct tone portion, as shown by the corresponding input to block 220 in Figure 2.
[0094] In particular, the 2-channel synthesizer 200 is configured to determine two head-related data channels from source position or orientation and listener position or orientation, in addition to directional information, and to use the two head-related data channels and directional information to calculate the 2-channel acoustic data for the direct sound portion. Specifically, referring to Figure 3a, the DOA vector 421 is shown as the difference between the listener vector 422 and the source vector 423.
[0095] Furthermore, Figure 3a further shows the emission direction vector 424 oriented in the opposite direction as the arrival direction vector 421. Typically, the directivity information of a sound source is given with respect to the primary emission direction 430. Therefore, the rotation of the sound source 410 must be considered in order to select the correct directivity information from a dataset of multiple directivity information about a sphere around the sound source 410, which is usually related to the primary emission direction or to an arbitrary reference point that is usually different from the origin of the world coordinate system to which the source and listener location vectors are given. Thus, in the exemplary figure, the rotation of the primary emission direction 430 with respect to the DOA vector 421 or DoE vector 424 is approximately 90°, so the 2-channel combiner 200 then determines the directivity information given for an orientation angle of 90° with respect to the primary emission direction 430 in Figure 3a.
[0096] Typically, the directional information is given as DIR for each orientation angle and elevation angle, and in a preferred embodiment of the present invention, there are approximately 540 DIR datasets of a sphere measured at corresponding 10-degree differences or 10-degree increments in both the orientation angle and elevation angle.
[0097] Alternatively, this information can also be provided via a directional transfer function (having amplitude and phase or real and imaginary parts) for a given orientation / elevation angle.
[0098] Alternatively, since the two-channel synthesizer provides a full dataset from which directional information can be selected that identifies the correct orientation of the source relative to the listener, the directional information can also be synthesized or actually computed using specific parameters of a particular class of sources. Furthermore, the directional information can be given at a lower resolution than that shown exemplarily with a 10-degree bidirectional resolution. Interpolation of the selected room impulse response can also be performed depending on the current situation being heard. In a further embodiment, if room impulse responses for a particular sound source are not available, these sound sources can be synthesized or measured and stored in a specific memory accessible by the two-channel synthesizer.
[0099] Figure 3b shows a table indicating what needs to be updated under specific conditions of listener movement and source movement. Naturally, if both the listener and source are stationary, no changes may be made to the previous situation. If the source is stationary and only the listener rotates, the room impulse response or directional information remains unchanged; the listener's rotation only affects the process by which a new head-related impulse response or head-related transfer function must be selected that takes the ear's position relative to the source in account for the rotation. Another interesting point is that if only the source rotates and the listener remains stationary, a new room impulse response needs to be calculated, but the head-related impulse response remains the same. In all other examples shown in Figure 3b, both the room impulse response and the head-related impulse response change for the specific situations shown in the table titled "What to Update?".
[0100] Figure 5a shows a preferred implementation of a particular embodiment. Typically, the direct sound portion of the room impulse response is used only for energy calculation in block 221 and not thereafter. Instead, the first portion of the room impulse response provided by the input interface is replaced with the corresponding directional impulse response selected as described with respect to Figure 3a. In the preferred implementation, the energy is related to the frontal DOE. Measurements for determining the three-dimensional directivity information are performed with the microphone on the sound axis.
[0101] Typically, the room impulse response dataset described with respect to Figure 4b does not have the same energy as the initial portion of the room impulse response provided by the input interface 100. Therefore, depending on the source position or orientation and the orientation of the listener position, raw directional information is determined from a database or the like using a specific angle, as shown in 222 of Figure 5a. Alternatively, the room impulse response can also be synthesized depending on the source position and orientation and a specific angle derived from the listener position.
[0102] In block 223, the energy of the raw directional information is calculated. In block 224, the scaling factor is calculated by dividing the direct sound portion energy by the total directional energy. If the distance between source 410 and listener 400 is the same, the new directional information is scaled by the scaling factors from blocks 224 and 226.
[0103] Alternatively, if the distance between the listener 400 and the sound source 410 changes due to the movement of either the source 410 or the listener 400, a further scaling factor is calculated in block 225, or the scaling factor in block 224 is adapted. In particular, the volume of source 410 must be reduced when the source moves away from the listener relative to the initial measurement situation, i.e., when the original room impulse response provided by the input interface is measured. A further scaling factor is then reduced. However, if the movement of the source or the listener results in a smaller distance relative to the initial distance between the source and the listener, the scaling factor must be increased. For the purpose of increasing or decreasing the scaling factor, the distance law of sound is applied. The scaling factor for distance correction is done, for example, using the distance law of sound or a similar procedure. The maximum amplification factor is limited to avoid the direct sound becoming too loud as the listener position approaches the sound source position.
[0104] In block 227, the direction of arrival of the direct sound is determined, as shown in Figure 3a. Next, based on the DOA, the correct HRIR or HRTF is selected, as shown in block 228. In block 229, single-channel directional information, scaled by a scaling factor that may be modified due to the changed distance, is convolved with the HRIR. In particular, since the HRIR has two channels and the directional information has one channel, the single channel is convolved with the left channel of the HRIR to obtain the first channel of the result in block 229, and the DIR is convolved with the right channel of the HRIR to obtain the right channel of the result in block 229, which is the two-channel acoustic data of the direct sound portion.
[0105] Therefore, the 2-channel combiner 200 is configured to determine the rotation of the source source if the source location vector of the source and the listener location vector of the listener are emission directions, and to derive directional information from a database of directional information sets, which are typically associated with a specific angle related to the main emission direction or a particular sound source emission direction.
[0106] In contrast, the direction of arrival of the listener's position or orientation is calculated from the sound source's source location vector, the listener's listener location vector, and the listener's rotation.
[0107] Figure 5c shows a more preferred implementation of how the head-related impulse response is convolved with the directional impulse response shown in block 229. For this purpose, block 261 shows that the directional impulse response and the two-channel HRIR are padded with zeros and then transformed into a spectral domain to obtain three spectra, the first spectrum being the directional transfer function, the second spectrum being the left HRTF, and the third spectrum being the right HRTF.
[0108] Next, as shown in block 263, the DFT spectrum and HRTF L The spectra are multiplied, resulting in the DFT spectrum and HRTF spectrum. R The two spectra are multiplied. The output of block 263 is two spectra transformed into the time domain. Next, the phase delay introduced by convolution, i.e., transform, multiply, and inverse transform is removed in block 265, both channels are truncated to their original lengths before padding in block 261, and finally, in block 267, windowing is performed, for example, by Tukey windowing.
[0109] To acquire the direct tone portion of the 2-channel audio signal, the following procedure can be performed. The RIR is first pre-processed to achieve consistent alignment between different inputs to the system. This allows for later blending between different input RIRs. Alignment is performed by detecting the direct tone using an appropriate modern technology algorithm. Finding the direct tone can rely on maximum peak detection, for example, if it is guaranteed that the direct tone of the input always coincides with the highest peak. In more complex scenarios, a more robust, modern technology direct tone detection can be chosen.
[0110] Next, the first sample of the impulse response is cut or expanded by a zero-value sample so that the detected direct tone sample index matches a predetermined sample index offset from the start of the impulse response. Binaural synthesis assumes that the RIR can be divided into three separate filters: direct tone (DS), early reflection (ER), and late reverberation (LR). The input RIR is then further preprocessed by separating it into these three separate subfilters. The transition between DS and ER can be selected to maximize the distance between the detected DS peak and the first reflection. The transition between ER and LR is selected to match the perceived mixing time of a given acoustic environment. This can be calculated or estimated by algorithms of the latest technology. In some embodiments, the transition between ER and LR can be earlier than the perceived mixing time to reduce computational complexity.
[0111] Next, the intervals between the three segments are expanded by n samples at each transition time, resulting in an overlap of 2n samples. A suitable window function is then selected to allow for a nearly complete reconstruction of the filter segments. For example, the Tukey window function can be chosen, with a lobe width of 2n samples.
[0112] Binaural synthesis assumes that the acoustic effects of a room can be largely separated into so-called specular components, where strong geometric reflections behave like ray and diffuse components. The specular component can be derived from models such as the Image Source Model (ISM) or from raycasting-based simulation methods. The diffuse component is used under the assumption that a portion of the signal can be approximated by a diffuse sound field with high reflection density and a nearly uniform distribution. This allows for the preservation of the time and phase relationships of the diffuse field while modeling the energy distribution of the reverberation portion of the RIR.
[0113] The DS segment includes a combined filtering effect between the sound source and the user's outer ear and body. All of these filters depend on the relative position of the virtual sound source and the sync (listener's ear). Both the position and rotation (of the source and sync) are provided as input to the binaural synthesizer.
[0114] Given relative positions, appropriate HRTF and DTF filters are selected from each subsystem. The filters are padded to double their lengths, then convolved with each other by multiplying them, and inversely transformed into the time domain using the Inverse Fast Fourier Transform (IFFT). Depending on the filters used, the introduced phase delay can be eliminated by temporally shifting the filters before truncating and windowing them, so that the binaural direct-sound filters have the same length as the original filters and the direct-sound center index corresponds to the same sample index as the direct-sound of the original RIR.
[0115] Figure 5d shows a preferred embodiment for determining the emission direction in block 268. It is assumed that the database is organized by specific emission directions. In block 269, matching with the test direction of emission in block 268 is performed, and in block 270, directional information for the best matching DoE is selected.
[0116] In block 271, an alternative is presented. Instead of finding the best matching DoE and selecting directional information from the database based on this DoE, two or more directional information entries with the closest DoE entries are selected, and interpolation is performed as shown in block 271.
[0117] Regarding the further alternatives shown in block 274, the directional information can also be synthesized using a model or neural network and a model based on the test DoE determined by block 268. Regarding the DoE, referring to Figure 3a, it is shown that the DoE is oriented in the opposite direction to the DOA, but the DoE is not related to the origin 420 of the coordinate system in Figure 3a, but rather to the primary emission direction 430. Thus, the DoE reflects a situation where the rotation of the source is applied to the vector DOA, reversing its direction. Of course, other alternatives with other relationships to other coordinate systems can be implemented.
[0118] As outlined in block 260, to obtain a padded function, it is preferable to perform padding to double the length by the directional impulse response and the first and second HRIRs and combine the padding function by convolving in the time domain or by using frequency domain multiplication. Furthermore, the phase adjustment shown in block 265 ensures that the correct time delay is maintained from the zero sample index to a specific index where the first direct tone portion is typically located, thereby ensuring that the construction after the full BRIR is always dependent on the defined circumstances.
[0119] Figure 6a shows a sphere for the purpose of illustrating the concept of HRTF or HRIR. In particular, the diagram in Figure 6a shows the user positioned in front of / to the left of the source. The corresponding left and right HRIR functions are shown in Figure 6b, and it is clear that the left HRIR is significantly stronger than the right HRIR, and that the left HRIR contribution occurs before the right HRIR contribution. This is evident from the fact that sound from the sound source and position shown in Figure 6a reaches the left ear before the right ear, and the amplitude of the sound reaching the right ear is attenuated by the head.
[0120] The corresponding frequency-domain response is shown in Figure 6c, which indicates that at frequencies below 1 kHz, the main effects are amplitude difference and above 1 kHz, and in particular at higher frequencies, the right HRIR exhibits a significant notch filter effect.
[0121] Next, a second aspect of the present invention will be described with reference to Figure 7 and subsequent figures. According to the second aspect, as shown in block 231, the two-channel combiner, in particular the early reflection processing block 230, is configured to segment the early reflection portion into multiple segments. For example, although only four segments 294 are shown in Figure 10b, it is possible to segment up to 50 segments or more. Of course, fewer segments can also be used. In one embodiment, there is a block with an overlap of 256 samples and 128 samples. The number of segments is determined by the length of the early reflection portion (direct sound up to the mixing time) of approximately 7700 samples. Dividing this number by an advance angle value of 128 per segment gives approximately 60 segments. However, this number can vary depending on the length of the early reflection portion, the advance value, and other potential parameters used.
[0122] Furthermore, as shown in block 232, preferably a geometric model of the room, such as a shoebox model, is used to determine multiple image source locations. The image source locations represent the source locations of reflected sound. Furthermore, the association of image source locations with segments is performed using a matching operation. In the matching operation, as shown in block 233, the sound arrival time from each image source to the listener position is calculated. Preferably, the initial listener position is used in this calculation, and as a result, the initial listener position, i.e., the listener position when the RIR was provided by the input interface, is input. Next, the image source locations are associated with the corresponding segments that best match the arrival time of a particular image source, as shown in block 234.
[0123] Therefore, the arrival time of sound from each image source position to the initial listener position is compared to a time index within a specific segment. Typically, since segments have a certain width, for a segment, the time index at the center of the segment is compared to the arrival time. If the arrival time of an image source position is equal to a time index associated with the segment, such as the time index at the center of the segment, this image source position is associated with this segment for further calculations, such as calculating the direction of arrival for this segment. Typically, image source positions are calculated for room models up to a certain order. Several primary image source positions, which are first reflections, are shown in Figure 8a. In particular, Figure 8a shows listener 400 at the initial listener position and source 410 at the initial source position. The construction of four image sources for (part of) primary reflections results in image source positions 1, 2, 3, and 4 for image sources 431 to 434. Note that floor reflections and room reflections, which also belong to primary reflections, are not shown in the two-dimensional Figure 8a. Secondary reflection can also refer to the physical effect where a reflection is constructed, travels to the listener's head, is reflected off a second wall, and then reaches the listener again.
[0124] Therefore, depending on the complexity of the geometric model, a certain number of image sources are determined in relation to their positions and associated with the corresponding segments. For example, if 50 segments are used to segment the early reflection portion, it is sufficient to determine image source positions up to an order that yields 50 sources. However, this can be very complex, and to conserve computational resources, a preferred way to do this is to compute image source positions only up to a certain order that yields fewer than 50 image source positions. The remaining image source positions can be randomly selected, as shown in block 235. Thus, if a particular segment is found to yield unmatched image sources associated with this segment, a random position is associated with this segment, or a random direction of arrival, and therefore a randomly selected HRIR, is used in further computation for processing this segment.
[0125] The results of this procedure are shown in Table 236 at the bottom of Figure 7, where the first three segments are associated with source positions 2, 1, and 4, respectively. Additionally, when a segment is counted from the direct sound / early reflection boundary to the early reflection / reverberation boundary, there are typically one or more segments at the end of the segment, which do not have discrete image source positions but are associated with random source positions or receive random HRIRs during processing of this segment.
[0126] As outlined, the two-channel segment is configured to determine multiple image source positions using geometric data relating to the initial source position and initial sink position of the initial measurement, as well as the acoustic environment. In particular, the image source configuration shown in Figure 8a is preferred.
[0127] In a preferred embodiment of the present invention, the two-channel combiner 200 is configured to detect prominent reflections, and from these detected prominent reflections, overlap segments are constructed as shown in block 280. The procedures in blocks 281 to 283 are performed for the purpose of detecting prominent reflections. In block 281, the average energy per sample of a small window sliding over the early reflection portion is calculated. In block 282, the average energy per sample of a larger window is calculated again as it slides over the early reflection portion. In block 283, both average energies are compared sample by sample to determine whether the average energy per sample in the small window is greater than the average energy per sample in the large window by, for example, a third specific threshold. This results in segmentation of the early reflection portion.
[0128] In block 284, the determination of the direction of arrival information for each segment is performed. Preferably, directional information as described above with respect to the first embodiment, particularly Figure 3a, can also be performed. This allows for consideration of a specific orientation of the image source relative to the listener, in particular, that the image source IS1 shown in 431 is oriented away from the listener 400.
[0129] In this embodiment, the two-channel combiner is configured to determine the directivity information of the image source for the listener position and the image source position or orientation, and to use the directivity information in the calculation 220 of the two-channel acoustic data for the early reflection portion. Preferably, the directivity information for each image source is derived from the same set of directivity information determined for the direct sound portion, or the orientation of the image source is determined by the image source model, and the directivity information is determined and used for a predetermined subset of segments of the early reflection portion, which in a particular embodiment includes fewer than 10 segments, preferably only two segments. The remaining segments can be calculated without the directivity information of the image source. Other steps for the calculation of directivity can be performed as outlined in Figure 5a, where, for segments for which directivity information is considered, the actual RIR segments are replaced by directivity information weighted by an energy scaling factor as determined by block 224, but using the energy of the corresponding reflection segment. For simplicity, it is preferable not to apply distance correction as in block 225, but nevertheless, this can be done when the listener approaches an image source that is far from the corresponding image source responsible for the reflection under consideration.
[0130] In block 285, each determined segment is padded to a specific length, particularly a length present in the HRIR database, and in block 286, each segment is convolved with the corresponding DOA-associated HRIR of the segment, as indicated by two connecting lines between 284 and 286. This procedure yields two-channel acoustic data of the specular reflection portion within the segment. When processing only the specular reflection portion of the early reflection portion of the room impulse response, the results of block 286 can be used for further processing. However, when the second and third embodiments are combined, the diffuse portion is also processed for the early reflection portion. This will be discussed later with reference to Figure 10a.
[0131] The ER segment is assumed to consist of a specular component and a diffuse component. Generally, the first part of the ER is expected to be primarily specular, as it contains strong primary reflections. The later parts of the ER segment are expected to contain more matching reflections and diffuse them. ER synthesis first further segments the ER portion of the RIR into smaller segments. In some embodiments, this is done by detecting perceptually prominent reflections and selecting a window of at least head-associated impulse response (HRIR) sample counts around them.
[0132] Such windows containing reflections can be detected by heuristics, such as comparing the average energy per sample within the window to a sample count n, and comparing the average energy per sample within a larger window to a sample count m around the first window. For a window of size m = 2n, a common heuristic is that reflections are considered significant, with the average energy per sample being 6 dB higher than the average energy of the surrounding window. These windows are then assumed to contain significant reflections.
[0133] In some embodiments, this technique can be generalized by assuming a continuous, regular grid of reflection windows, each having the same sample count and overlap. This effectively quantizes the assumed reflection time of the incident light onto the grid. Each detected reflection window is assumed to consist of a partially specular and a diffuse portion, while the remainder is assumed to be completely diffuse. Each reflection window is assigned a diffusion coefficient, which approximates how much the reflection diffuses. The exact diffusion coefficient can be determined using a heuristic or formula. Different embodiments of the system may use different methods to determine the coefficient. A possible heuristic is a small window E in dB units. s Total energy and large window E in dB units l Given the total energy, the diffusion coefficient α is given by α = (Es / E l Based on the early heuristic used to find prominent reflections, it can be calculated as (+6dB) / 12dB.
[0134] A diffusion coefficient α of 1 or greater means that the reflection is perfectly specular, and α of 0 or less means that the reflection is perfectly diffused. Therefore, the value of α is limited to the range [0,1]. Similar to the DS portion, the omnidirectional specular reflection portion of the reflection window needs to be convolved with the HRTF. For this purpose, the DS processing step and the HRTF from the same HRTF provider can be used.
[0135] The required DOA used to obtain the HRTF is calculated using the image source method. Based on the provided geometric room information, the image source positions are calculated. The best candidate is then selected by comparing the sound incidence time at the sink for each image source and comparing it to the incidence time at the reflection window. The best-matching image source is selected, and the normalized vector between it and the sink is assumed to be the DOA for specular reflection.
[0136] In other embodiments, the DOA may instead be determined by so-called spatial resolution methods, such as by other means including statistical distribution heuristics or by analyzing arrival times on the microphone array. The binaural specular reflection portion is then calculated by convolving the windowed reflection segment with an HRTF, for example by appropriately padding the segment and multiplying it by the frequency domain HRTF, before inversely transforming the result and removing the introduced phase delay if necessary. This gives the specular reflection window w s This is generated.
[0137] The binaural diffuse portion is obtained by selecting the same subwindow, but this time it originates from the composite diffuse filter. This diffuse reflection snippet is then multiplied by a Hann window of size n, and wd gives
[0138] Next, the diffuse reflection window w d and the specular reflection window w s are linearly combined when the diffusion coefficient α is given, and the resulting window w bin is combined according to the following equation.
[0139] w bin = α * w s + (1 - α) * w d (Equation 1) Therefore, with respect to Figure 8b, a preferred implementation is that, for initialization, the necessary signals are loaded. The first signal is the room impulse response provided by the input interface or single-channel acoustic data. The second signal is binaural noise to be used later according to the third or fourth embodiment. Furthermore, the HRTF dataset and, if necessary, the average HRTF amplitude response are loaded. Furthermore, compensation filters from microphones and headphones can be applied, as shown with respect to the seventh embodiment, and the directional transfer function of speakers can be applied, as described with respect to the first and second embodiments. Furthermore, headphone compensation can be applied to the HRTF before selecting a specific HRTF, and as a result, an already compensated HRTF can be selected in response to DOA. Next, the position and rotation of the recording constellation are stored as the initial listener position or orientation or initial sink position or orientation and initial source position or orientation. Then, the image source model calculation as shown in Figure 8a is performed, and an order is selected based on the assumed mixing time of the early reflection portion and the late reverberation portion so that a sufficient image source is obtained, for example, 160 ms is selected for the room impulse response. However, to simplify this procedure, fewer image source locations can be used, and segments for which no relevant image source locations were received during the matching process are generally associated with random locations or random data. If no matching image source locations for a particular reflection are found during the matching process, it is also preferable to perform this procedure to randomly associate a specific HRTF with the segment.
[0140] Next, a third aspect of the present invention will be described with reference to Figure 10a. In particular, the two-channel combiner is configured to calculate the specular reflection portion for segment n, or generally for the early reflection portion, as discussed with respect to the second aspect and shown in block 237. Furthermore, the two-channel acoustic data of the early reflection portion is also calculated using the diffusion portion, as shown in block 238, which describes the diffusion effect within the early reflection portion. Both blocks 237 and 238 receive a single-channel early reflection portion for segment n, as provided by block 210 in Figure 2. Both blocks 237 and 238 output two channels of binaural data, which are combined correspondingly in block 239 so that the first and second channels of segment n are obtained, and this two-channel data of the early reflection portion not only represents the specular reflection effect of separate early reflections as in the prior art procedure, but also describes the diffusion portion, which greatly contributes to a natural and pleasant sound impression for the listener.
[0141] In particular, the two-channel synthesizer 200 is configured to calculate the diffusion portion using a combination of the early reflection portion of single-channel acoustic data and a two-channel noise sequence as input to block 238. Preferably, this two-channel noise sequence is a binaural noise sequence measured when a specific noise signal is emitted by a speaker at a specific position relative to an artificial head and the full HRTF is detected by two microphones present on the artificial head. Such binaural noise can be measured or synthesized, and alternatively, if this is impractical for any reason, even two different noise sequences can be used to binauralize the late reverberation portion of a room impulse response.
[0142] In a preferred implementation as shown in Figure 11b, a weighted sum of the specular and diffuse portions is performed, with the weighting coefficients determined, for example, as shown in blocks 290a and 290b, and the actual sum of the weighted contributions is performed in block 290c. For this purpose, also see Figure 10b, which shows an exemplary room impulse response that can be measured or synthesized in block 291. However, it should be noted that room impulse response 291 is not a true room impulse response because, for explanatory purposes, early reflections are enhanced with respect to direct sound. In particular, the impulse response includes the specular portion 291 and the diffuse portion 293. In 292, this figure shows specular reflection and is therefore not a true snippet from 291. The same is true for block 293, which shows a kind of diffuse portion, but it is not directly derived from room impulse response 291 because the scale of early reflections and direct sound is modified. Block 294 shows seven overlapping segments in a situation where the early reflection portion is separated into only seven segments. However, more segments can be used as well, and in a typical implementation, 50 segments or even 60 or more segments can be used.
[0143] The same segment is applied to the diffuse portion following the windowing operation applied to the diffuse portion, allowing the diffuse portion to be effectively combined with the windowed specular portion.
[0144] Therefore, in order to calculate the weighted sum 290, the windowed diffuse portion is used, the windowed specular portion is processed using the image source model 295 as described above, and additionally, the HRFT provider 296 provides the correct HRTF for each segment, followed by a padding operation 297, a subsequent convolution operation 298 with selected HRTFs from block 296, and finally, delayed compensation 299 is applied to obtain the correct mixture of the specular and diffuse portions for each early reflection segment.
[0145] In the preferred embodiment shown in Figure 11b, a further correction 290d is performed to address the situation where, near the boundary between the direct sound and the early reflection, the specular reflection should dominate the diffuse portion, i.e., have a stronger influence than the diffuse portion. On the other hand, at the other end of the early reflection portion, i.e., at the boundary between the early reflection portion and the late reverberation portion, the diffuse portion must be dominant over the specular reflection portion.
[0146] Typically, the Directionality-to-Diffuse Ratio-Based Scale (DTD) in blocks 290a and 290b is known to already result in situations where either the specular or diffuse portion must be dominant. However, to avoid any unnatural situations, Correction 290d is applied in specific ways, such as by providing a maximum or minimum amount for each or multiple segments, or by applying a specific curve to the scale determined as shown in blocks 290a and 290b.
[0147] Depending on the implementation, the directivity-to-diffuse ratio can be used as a threshold or as a smooth transition from 0 to 1. Perfect specular or prominent reflection is given when the first window has twice the energy of the second window. A perfectly diffuse segment is obtained when the average energy of the first window is 0.5 times the energy of the second window, and all values between 0 and 1 are also possible. These values are preferably used as weighting coefficients or when determining the weighting coefficients of a weighted combination of specular and diffuse portions.
[0148] Furthermore, see Figure 11a, which shows the procedure for calculating a particularly preferred direction versus diffusion ratio (DTD). The early reflection portion is cut into overlapping portions. Then, the pre-gain coefficient is determined to employ the directional transfer energy, with the aim of applying the directional transfer function not only to the direct sound portion but also to the early reflection portion as described above with respect to the first and second embodiments. The early reflection is then sliced into blocks or segments having a block size of 256 samples and a hop size of 128 samples, after which a zero-padding operation is performed on 512 samples, thereby after which a Fourier transform is applied to each block. Next, the direction versus diffusion ratio is determined by determining the amount of energy in each block, comparing it to a moving average, and determining the relationship between the geometric reflection and diffusion portions in decibels. For this purpose, the energy shown in Figure 11a is converted to smoothed energy by a moving average operation.
[0149] A preferred procedure can also be performed as follows: For example, the diffuse component can be derived by taking the reverberation portion of the RIR and a binaural noise sequence of the same length as the RIR reverb as input. (Here, the binaural noise sequence refers to a type of white noise that exhibits the same phase characteristics and interaural correlation as the recording of the diffuse sound field.) The combination of the diffuse portion and the specular portion is preferably performed according to equation (1) above.
[0150] Figure 13a illustrates the subject matter of the present invention according to a fourth embodiment. This embodiment refers to an improved calculation of the diffusion portion of the early reflection portion or the late reverberation portion, or only the late reverberation portion, or both, using the amplitude spectrum of the early reflection portion and / or the late reverberation portion and the phase spectrum of the two-channel (binaural) noise. The two-channel combiner 200 in Figure 1 is configured to use the amplitude spectrum of single-channel acoustic data without the early reflection portion or direct tone portion and the phase spectrum of the first channel noise to obtain the first channel of two-channel acoustic data, and to calculate the two-channel diffusion portion of the single-channel acoustic data without the early reflection portion or direct tone portion using the amplitude spectrum of the single-channel acoustic data without the early reflection portion or direct tone portion and the phase spectrum of the second channel noise.
[0151] In particular, the first channel nose phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence. This is shown by block 530, which shows the calculation of the amplitude spectrum, and block 532 in Figure 13a, which shows the calculation of the phase spectrum of the two-channel (binaural) noise.
[0152] The transformation of the single-channel data derived by block 520, or preferably following the smoothing of the amplitude spectrum of block 531, is transformed into a second channel result by adding the phase of the first channel phase spectrum to the smoothed amplitude of the spectrum of block 531 to obtain a first channel result, and by adding the second channel phase spectrum of block 531 to the preferably smoothed amplitude spectrum of block 532 of the coupler 533 to obtain a second channel of a two-channel diffuse portion for the late reverberation portion, or for the early reflection portion and the late reverberation portion, or in other words, for single-channel acoustic data without a direct tone portion that is assumed to be non-diffusive and therefore does not make a receiving and diffusion contribution. It should be noted that the "addition" of amplitude and phase, in a concrete mathematical sense, is a multiplication as shown in block 444 of Figure 13b, i.e., the multiplication of the amplitude spectrogram and the phase spectrogram |RTF|*e^(angle(binauralNoise1)) and |RTF|*e^(angle(HbinauralNoise2)), where H represents the transformation to the spectral domain.
[0153] As shown in Figure 13b, a monaural RIR is provided in block 440, and the absolute value of an STFT spectrogram consisting of a sequence of vectors is obtained as shown in block 442. A binaural noise sequence 441 is also provided, which undergoes corresponding spectrogram processing by time-frequency conversion, and the phase angle of each channel of the binaural noise is obtained as shown in block 443, and then, preferably following a smoothing operation in block 531, the phase angle is combined with the corresponding amplitude as shown in block 444.
[0154] The smoothing operation in block 531 has the advantage that this smoothing along the frequency direction of the amplitude spectrum naturally avoids any peaks that may arise due to the inverse Fourier transform when phase manipulation is performed in the spectral domain, as in the present invention, in each spectrum of the sequence of vectors covering, for example, the early reflection and late reverberation portions. On the other hand, the procedure of calculating the spectrogram and simply "adding" the phase of the binaural sequence to the (smoothed) spectrogram is a computationally easy procedure that does not require a considerable amount of computational resources. Furthermore, this late reverberation processing has been found to have a particularly pleasant sound for the listener because, due to its quality, it is possible to use the same late reverberation 2-channel acoustic data of the acoustic environment regardless of whether the source position or orientation or the listener position or orientation changes. This situation leads to the important result that the update ratio for calculating the late reverberation portion can be made significantly smaller (usually only one or two orders), as shown in relation to the sixth aspect of the present invention, which further reduces the computational resources required and makes it possible to distribute the processing task to different elements.
[0155] Figure 13c shows a further implementation of the procedure in Figures 13a and 13b. In block 445, the overlap-block transform is applied to the late-reverberation chamber impulse response, or to both the early reflection portion and the late-reverberation portion. This preferably yields a first spectrogram in which low-pass filtering is performed on each amplitude spectrum over frequency, as shown in block 531.
[0156] In block 447, it is preferable to further perform low-pass filtering over time, i.e., across two or more adjacent blocks, and for the same frequency bin, but within adjacent blocks, i.e., using temporally adjacent frequency bins related to the same frequency. A similar transformation 446 is performed on the time-domain binaural noise sequence to obtain a second and a third spectrogram, and in block 449, the phases of the second and third spectrograms are added to the low-pass filtered spectra in the spectral and time domains.
[0157] Next, the result of block 449 is converted to Cartesian form as shown in block 450, and then converted back to the time domain in block 451. In block 452, overlap and summation procedures are performed, and finally, in block 453, truncation and windowing, as well as overlap with the early reflection portion, is performed only for the diffuse signal for the late reverberation, and then 2-channel acoustic data for the reverberation portion of the RIR and the late reverberation portion at the output of block 453, which is the same length as the RIR reverb, is taken as input.
[0158] (Here, a binaural noise sequence refers to a type of white noise that exhibits the same phase characteristics and interaural correlation as a recording of a diffuse sound field.) Next, the two filters can be transformed into a time-frequency representation using a Short-Time Fourier Transform (STFT). The STFT parameters can be selected to allow for a nearly complete reconstruction, for example, by using a half-overlapping Hann window. The complex-valued block spectrum is then transformed into polar form, separating each frequency bin into amplitude and phase components. The complex representation of the diffuse component is then constructed for each channel of noise by pairwise combinations of each bin of the complex amplitude of the transformed RIR block and the complex phase of the transformed noise sequence block. Amplitude and phase recombination is performed for each bin of the equivalent-length block.
[0159] In some embodiments of the described system, the binaural synthesizer is further configured to perform low-pass filtering between the amplitudes of the frequency bins during this processing step. A typical low-pass filter may be a moving average filter corresponding to equivalent bins of one-third octaves, but other configurations are also possible. This reduces artifacts introduced by the combination of the two amplitude and phase parts of two different transfer functions.
[0160] Furthermore, a low-pass filter can be applied between temporally adjacent blocks of the recombination transfer function so that bins corresponding to the same frequency are low-pass filtered between blocks. A typical configuration for this is a moving average filter over three values (or time blocks), assuming a block size of 512 samples at 48000 kHz. The exact parameters also depend on the individual embodiment.
[0161] Next, the new filter is transformed into a Cartesian form and then into the time domain using an inverse STFT with the same parameters used for the forward transformation. The resulting filter is a binaural diffuse reverb filter with the combined lengths of the ER and LR segments and two channels. To be clear, the beginning of this diffuse segment is used as one of the two layers of the ER segment, and the latter diffuse segment is the input to the LR segment.
[0162] Next, Figures 14a to 14e are shown, which can be used in each of the first to fourth embodiments of the present invention, and in particular in the fifth embodiment of the present invention, enabling the efficient provision of the required room impulse response. For this purpose, the input interface is referred to as RIR provider 100, which receives, for example, a microphone signal of initial measurement as input.
[0163] This room impulse response is transferred to a binaural synthesizer, which calculates a binaural room impulse response based on geometric data about the room, the required HRTF, and positional data about the sound source and user, as required for the calculation of the image source. The binaural room impulse response can then be audibly generated by an AuraLizer 300 or a sound generator using the audio signal to obtain two output speaker signals that can be rendered by headphones, earphones, in-ear devices, or discrete speakers.
[0164] In Figure 14d, the microphone signal is measured, and the RIR provider 100 calculates parameters, or generally, calculates a fingerprint from the microphone signal, accesses a specific room impulse response database 110, the database responds with a matching room impulse response, which is then transferred by block 100 to a two-channel combiner or binaural combiner 200.
[0165] In Figure 14c, the RIR provider 100 generates a set of parameters from the microphone signal and transfers these parameters to a database or to an RIR synthesizer located at a position to synthesize RIRs based on the parameters. Thus, block 120 can have two functions compared to block 110. Furthermore, an RIR modifier 130 is provided that modifies the RIR in some way for the purpose of implementing a specific desired sound effect or room effect.
[0166] Figure 14d shows the procedure when using only the RIR synthesizer 140 and without using a database for the purpose of auditory representation of room acoustic effects.
[0167] Figure 14e shows a more preferred method for providing a specific room impulse response. This procedure relies on acoustic measurements, which may be microphone signals or from another source. In block 101, dimensionality reduction is performed to obtain a simplified representation, which may be a fingerprint that can be derived from the acoustic measurements by, for example, a set of parameters, or generally by other procedures different from parameterizing this signal for, for example, psychoacoustic parameters.
[0168] Furthermore, an RIR database 110 is provided, which also includes a dimensionality reduction block 111 for regenerating simplified representations. These simplified representations, along with other simplified representations for other RIRs stored in the RIR database 110, are then input to block 112, which minimizes distance and finds the best-matching RIR. This best-matching RIR is identified from block 112, and this information is sent to block 113, which loads the RIR from the RIR database 110, and then binauralization is performed. The block binauralization in Figure 14e combines the functions of blocks 200 and 300 in Figure 1 or other figures.
[0169] Furthermore, Figure 15 shows specific hardware according to a preferred implementation of a sixth aspect of the present invention. In particular, the first device 901 comprises one or more microphones 911, one or more processors 912, a memory 930, and a position tracking system 914 for tracking the listener's position, which also refers to the listener's orientation, i.e., collectively the listener's posture. Furthermore, the first device 901 may include a speaker 915.
[0170] In this embodiment, the second device 902 comprises a processor 921 and memory 922 and is connected to the first device 901 via a network.
[0171] Part 1 of the solution infers the RIR (for a specific virtual source listener configuration) from the available audio data recorded in the listener listening environment. The actual configuration in which this audio was recorded is possible, but it does not need to match the configuration of the virtual configuration.
[0172] System 1 is configured to record local sound field measurements, including the room acoustic effects of the user's actual environment, either sequentially or once. In some embodiments, this is done by measuring the RIR between a microphone(s) and any actual sound source in the room, for example, using an exponential sinusoidal sweep. This is particularly feasible if the system is calibrated only once for the listening environment. The RIR thus recorded may or may not be feasible for audible representation of a virtual sound source due to belonging to a different sound source or due to limitations in the system's acquisition section, such as the limited bandwidth of a transducer, which may not match the bandwidth of the virtual sound source being audible [A].
[0173] The recorded audio data is transmitted to a second system via a network [B]. Memory maintains a database of pre-recorded high-quality omnidirectional RIRs of different rooms with varying acoustic characteristics. The database may (but does not necessarily) include one or more measurements from actual listening rooms.
[0174] The purpose of this system is to process the transmitted audio data and select a RIR that best matches an unknown RIR in the real room at the current listener position, and to send it back to the first system for further processing and binaural synthesis [C].
[0175] This second system is configured to reduce the high-dimensional, temporal, or temporal-frequency representation of audio data to a lower-dimensional representation. In some embodiments, this is achieved by transforming the data with a well-trained neural network. Such a network can be trained, for example, on the task of classifying RIRs into single room classes other than those existing within a database. The coefficients of the network's layers and the latent space they form are then selected as the lower-dimensional representation of the data to be computed for both pre-recorded and ad-hoc measured RIRs [D]. These coefficients are then used to find the best-matching RIR based on minimizing a suitable distance metric in this dimensionality-reduced space. The acquired RIRs are then returned to the first system via the network. This process of acquiring RIRs is repeated at specific intervals to reflect large changes in room acoustics, for example, when a listener moves to an area with substantially different reflections or different acoustic environments. When a new RIR that minimizes the distance metric to a new data point is found, it is selected. The system maintains a short history of the RIRs used and allows for gradual mixing between changing RIRs.
[0176] Further embodiments are shown below. 1.[A] Instead of recording impulse responses, the system is configured to record and process user-defined self-generated sounds such as clapping or speaking, so that the RIR can be inferred without the need for sinusoidal sweeps.
[0177] 2.[A] Instead of recording the impulse response, the system is configured to record and process a well-defined sound, such as music, so that the RIR can be estimated without requiring a sinusoidal sweep.
[0178] 3.[A] Instead of recording impulse responses, the system is configured to record and process general sound fields independently of a specific class of sound or sound, so that the RIR can be automatically estimated without requiring any user input for sinusoidal sweep or calibration.
[0179] 4.[D] Instead of using the latent space of a neural network for dimensionality reduction, we use appropriate digital signal processing supported by psychoacoustic models to create a low-dimensional space suitable for RIR matching.
[0180] 5.[C] Instead of selecting from a pre-recorded set of RIRs, the neural network is trained to synthesize new RIRs directly. These RIRs may be combinations of existing RIRs or not.
[0181] 6.[C] Instead of selecting from a pre-recorded set of RIRs, a second neural network is trained to synthesize new RIRs from dimensionality-reduced representations. These RIRs may be combinations of existing RIRs or not.
[0182] 7.[C] Instead of selecting from a pre-recorded set of RIRs, train a neural network to synthesize new RIRs from the outputs of the embodiments described in 4.
[0183] 8. [B] System 1 is expanded using non-temporary memory to store the RIR database. All processing is performed on System 1.
[0184] In some embodiments, the RIR provider does not need to be manually configured with RIR. Instead, the system may use speakers and microphones that may not meet the qualitative requirements of broadband RIR measurement; i.e., transducers may have nonlinear or interfered frequency responses across the audible range, or they may be integrated into the same chassis. Instead of directly measuring RIR, the system is configured to measure low-quality RIR, which may or may not be usable for binaural synthesis. The measured RIR is not used directly as input for binaural synthesis. Instead, acoustic or psychoacoustic parameters are derived from the measured RIR. For example, the reverberation time per band (RT60) or energy decay curve (EDC), the directness to reverberation (DRR), or other parameters can be calculated. The exact parameters calculated depend on the embodiment.
[0185] Next, the calculated parameters are used to find a similar RIR from a database of pre-recorded RIRs suitable for binaural synthesis. The pre-recorded RIRs, along with the selected set of acoustic parameters, are stored in non-temporary memory.
[0186] When the RIR provider is set up in a new acoustic environment and a low-quality RIR is measured, these parameters are calculated and compared to a pre-recorded dataset. The best-matching pre-recorded RIR is selected from the database and used as the input RIR for the binaural synthesizer.
[0187] In some embodiments, instead of directly finding the RIR with the best matching parameters, a psychoacoustic weighting function is used that specifies weight coefficients for the influence of each parameter.
[0188] In some other embodiments, the parameters used to find the best-matching RIR are other acoustic or psychoacoustic parameters. Alternatively, the measured RIR is represented by a set of parameters computed by transforming the data with a well-trained neural network. Such a network can be trained, for example, on the task of classifying RIRs into single-room classes other than those existing within a database. The coefficients of the network's layers and the latent space they form are then selected as a low-dimensional representation of the data computed for both pre-recorded and ad-hoc measured RIRs. These coefficients are then used to find the best-matching RIR based on minimizing the distance metric in this dimensionality-reduced parameter space.
[0189] If conventional RIR measurement (i.e., using exponential sinusoidal sweep convolution) is not feasible, some embodiments of the system can utilize a class of sound, i.e., the sound of human applause or human speech, to make assumptions about the RIR or to derive the RIR from reverberation audio directly recorded by one or more microphones. To achieve this, selected parameters are either directly derived from the reverberation audio or an intermediate approximation of the RIR is derived.
[0190] Some embodiments of systems used in two or more acoustic environments need to adapt to changing room acoustics. This is achieved by modifying the RIR transmitted from the RIR provider to the binaural synthesizer. Depending on the embodiment of these RIR updates, these updates can be performed, for example, periodically at a fixed rate, or when significant acoustic changes necessitate an update.
[0191] To achieve a gradual, inaudible transition between two input RIRs, the RIR provider is configured to gradually interpolate between the two filters. A suitable algorithm for achieving this type of interpolation is, for example, linear interpolation in the time domain or frequency domain.
[0192] In some embodiments, when the device is worn and used in multiple environments, such as when listening to music while on the move, it is not possible to pre-configure the system for one or more room acoustic environments. Here, one or more microphones of the device continuously or periodically record sounds from the acoustic environment. The RIR provider is configured to detect one of many classes of sound and to derive an intermediate representation of the RIR. Even when the acoustic characteristics of a room may change rapidly from time to time, such as when entering or leaving a room, the system can be configured to gradually integrate and adjust the detected room acoustic characteristics to improve the stability of the result.
[0193] Instead of retrieving RIRs from a pre-recorded database, state-of-the-art room acoustic simulation techniques can be used to generate omnidirectional RIRs for binaural synthesis. The algorithms used can simulate a good approximation of actual room acoustic effects given a limited time frame and set of input parameters. Since the phase relationship of the BRIRs is modeled by the binaural synthesizer, a good approximation of the actual frequency-dependent distribution of energy over time must be particularly important. Depending on the room acoustic simulation method used, the RIR provider is configured to calculate the input parameters required for the simulation.
[0194] Some embodiments of the described system utilize an extended version of the binaural synthesis method, which can compute specular reflections using room acoustic modeling, such as an image source model, instead of processing individual specular reflections from the recorded RIR. Here, the provided room geometric information is used to determine the arrival time and direction of arrival of individual specular reflections. The acoustic absorption effect of the reflective surface from which the reflection occurs can be included as input to the room geometric data. Alternatively, some embodiments may choose to estimate the absorption coefficient of the walls by analyzing the initial reflection at the arrival of the ISM reflection by deriving filters based on reflection and direct sound windows assumed to be nearly linear over the audible frequency range. This modification allows for the computation of higher density specular reflections and potentially improves localizability at the expense of more computationally intensive work.
[0195] Figure 16 shows a preferred implementation of a fifth embodiment relating to smart determination of the room impulse response from a raw representation associated with single-channel acoustic data. In particular, the input interface 100 of the device shown in Figure 1 is configured to acquire a raw representation associated with acoustic data, as shown in block 150. Furthermore, the input interface 100 is configured to acquire single-channel acoustic data by using the raw representation acquired in block 150 and additional data stored by or accessible by the audio signal processor, the single-channel acoustic data is then transferred to the two-channel combiner 200.
[0196] Exemplary, the input interface is configured to acquire initial measurements of raw single-channel acoustic data as a raw representation in order to derive a test fingerprint of the raw single-channel acoustic data, as shown in block 101 of Figure 17a. Based on this test fingerprint, a pre-stored database 110 having a corresponding set of reference fingerprints is accessed, each reference fingerprint associated with high-resolution single-channel acoustic data, the high-resolution single-channel acoustic data having a higher resolution than the initial measurements. Furthermore, from the pre-stored database 110, high-resolution single-channel acoustic data having the reference fingerprint that best matches the test fingerprint is retrieved, as shown in block 113 of Figure 17a.
[0197] Alternatively, high-resolution single-channel acoustic data can also be synthesized from test fingerprints or raw single-channel acoustic data, typically using additional geometric data or only geometric description data, as shown in block 140 of Figure 17a, which demonstrates the direct synthesis of single-channel acoustic data. For this purpose, block 140 receives a raw representation, such as obtained by block 150, or a test fingerprint, such as calculated by block 101, and as additional data, data for room simulation, data about neural network information, and block 140 implements the neural network or model data as additional data in this case. Thus, when performing the alternative to direct synthesis, database 110 is not required. Furthermore, block 101 is configured to derive a test fingerprint as at least one set of the following parameters RT 60, EDC, DRR, and the reference fingerprint further includes at least one of the following parameters RT 60, EDC, DRR.
[0198] Figure 17b shows further steps for calculating the room impulse response or room transfer function of an acoustic environment. In block 150, sound pieces, such as songs played by speakers in the acoustic environment, are recorded as a raw representation of block 150 in Figure 17a. In block 155, sounds are identified using a kind of audio fingerprint system accessed as shown in blocks 155 and 156 by receiving a test fingerprint and returning a reference fingerprint that identifies or matches the song. In block 157, a remote music database can typically be accessed using the reference fingerprint or song identification information, and in block 158, songs played in the acoustic environment by one or more speakers are retrieved, but without the room acoustic effects imprinted on them, and only the clean version played by the speakers. In block 159, the RIR or RTF of the acoustic environment is calculated using the songs recorded in the environment, using the clean versions of the songs, i.e., without room effects such as those provided by the music database 157.
[0199] Further implementations are shown in Figure 17c, where initial measurements or data are acquired in block 150, and in block 112, a test fingerprint indicating the acoustic environment class is calculated, for example, by a neural network or other procedure. Then, based on the room class, the matching RIR can be retrieved from a pre-stored database as shown in block 152, or it can be synthesized using selected room classes as shown in block 153. Room classes may include closed rooms, open environments, large rooms, small rooms, rooms with significant attenuation, reverberation rooms, etc.
[0200] A further implementation of the present invention is that the user generates a natural sound, as shown in Figure 17a, 160. Such a natural sound is clapping, speaking, or any transient sound that a listener may make. This avoids the generation of unpleasant measurement sounds, such as sine sweeps, in the room. Based on this sound, the (low-resolution) RIR is recorded as a microphone signal and processed by any of the procedures shown in Figures 17a to 17c to obtain a high-resolution room impulse response from this raw representation for further processing.
[0201] Next, preferred embodiments according to a sixth aspect of the present invention will be described. The auditory representation of a direct sound path is typically achieved by block convolution of an audio signal with filters that approximate the filtering effects of the user's head, ears, and torso on a sound source at a given location and distance (HRTF). Processing these filters requires encoding not only changes in sound intensity and other cues, but also correct changes in interaural time difference (ITD) and interaural level difference (ILD). Human listeners are relatively sensitive to even small changes in these values, and therefore, these changes must be computed with good spatial and temporal resolution. However, these filters are relatively short, and deriving them usually involves only a few processing steps. Room reverb simulates the filtering effects on sound caused by the environmental geometry of sound that does not travel to the direct part from the sound source to the user's ears. This includes reflection, refraction, absorption, and resonance effects.
[0202] Such reverb filters are expected to be much longer than short direct tone filters. Numerous processes, algorithms, and systems can handle appropriate binaural reverb, including image source algorithms, ray tracing, parametric reverberators, and numerous delay network-based techniques. The exact form of implementation of the reverberator is irrelevant to the invention as long as it can simulate sufficiently externalized hearing. In this embodiment, the system uses the fact that human listeners are more sensitive to changes in the direct tone filter and less sensitive to changes in the reverberation filter. The signal processor is programmed to compute the direct tone filter much faster than the reverberation filter. This allows the system to minimize the clearly audible jump in the audible sound when the filters are swapped to increase the sense of externalization, while avoiding a complete filter update. Updating these filters or signal portions encoding the direct tone path at a rate of approximately 188 Hz has proven to be a reasonable default for such a system, although lower refresh rates (such as 94 Hz or 50 Hz) may be feasible in different embodiments of the system. The reverb filter is calculated at a much lower rate, typically up to 1 / 10th of the direct sound processing rate, depending on the environment and the user's acoustic characteristics.
[0203] Either a signal processor or another processor is configured as an aggregator. In some embodiments where the binaural synthesis method used returns a continuous stream of blocky binaural audio signals, this aggregator simply sums the blocks supplied by the direct and reverberation processing paths and functions as a signal aggregator. This requires that the blocks being summed correspond to the same point in time or include control data that identifies the time frames they correspond to. Alternatively, the aggregator can be configured to sum two subfilters with respect to their time delays, as determined by the algorithm. Thus, it reconstructs a full BRIR filter from the results of the individual processors and functions as a filter aggregator. The filters can then be used to convolve the audio signal blocks using a state-of-the-art real-time (blocky) convolution method.
[0204] The aggregator always maintains a full BRIR filter in memory. Therefore, the BRIR can be partially updated at the individual rates of the individual processors handling the partial filters. The resulting signal blocks contain combined binaural signals for the direct and reverberation paths. These are then passed to a speaker signal generator and played back on the system speakers. This allows for the audibility of binaural audio with a similar level of externalization and perceptual quality to that of individual algorithms, while significantly reducing processing requirements.
[0205] In different embodiments, the processing of the reverberation tail may be further divided into separate processing paths that compute separate filters for early and late reverberations on one or more processors on the same device. This takes advantage of the fact that strong early reflections often help human listeners locate sounds. These reflections often undergo strong and instantaneous changes, especially as the listener moves through the environment. While humans are less sensitive to these changes than to changes in direct sound, these early reflections carry a lot of energy, and if the filter calculation is slow, it can lead to poor externalization, mislocalization, or audible jumps. On the other hand, the largest portion of a typical reverberation tail includes dense overlapping reflections with relatively low energy. These late reverberations change relatively slowly. In some environments, they may consist mainly of the diffuse portion of the reverberation tail, meaning they are constant throughout the sound field.
[0206] In this embodiment, different reverberators can be used to process the late reverberation portion at even lower rates. Depending on the acoustic environment, the system can be adjusted to keep the late reverb constant, refresh it at low rates such as 1 Hz, or process it on demand when substantial changes in room acoustics are detected. In some embodiments, the algorithm used can provide a mixture of binaural filters and binaural signals. Here, the aggregation stage can be divided into a filter aggregator + convolution and a signal aggregator, first combining the partial filters, convolving the reconstructed filters with the audio signal, and then adding the binaural signal to obtain a complete binaural signal.
[0207] In this embodiment, the system is divided into four or more reverberators, and the BRIR is segmented into four or more segments. This can be used to further process portions of the reverberation tail with varying complexities. For example, a precise geometric algorithm can be used for primary reflections, while subsequent reflections are processed probabilistically, and the late reverberation tail is processed as in the previous embodiment.
[0208] In this embodiment, one direct sound processor and one reverberation processor are implemented separately on two devices, forming a system with the same capabilities as Embodiment 1, suitable for distributed synthesis of binaural signals and auditory representation of binaural audio on wearable devices.
[0209] The first device is a wearable device and includes sensors, transducers, and one or more processors, as in 1. These processors are configured to synthesize a direct-sound binaural filter or binaural signal directly on the device at a sufficiently high refresh rate. Processing of the direct-sound portion is performed directly on the device, avoiding transmission over a wireless channel. The wearable device includes an aggregator and speaker signal generator necessary for filtering and / or signal aggregation, as in 1. It further includes subsystems for wireless transmission and reception of audio and control data. The second device includes a processor configured to compute a reverberation filter or signal. It further includes subsystems for wireless transmission and reception of audio and control data.
[0210] In an embodiment in which the algorithm used in the second device synthesizes a binaural reverberation filter, the sensor data and control data required by the algorithm used are transmitted by the wearable device. The partial filter is processed and sent back wirelessly to the first device. The processed filter is then sent to an aggregator, which reassembles a complete representation of the BRIR, which is held in memory as in 1. The complete BRIR is then convolved and played back on the speaker as in 1.
[0211] In an embodiment where the algorithm used in the second device directly synthesizes the binaural reverberation signal, the audio signal is streamed along with sensor and control data required by the algorithm used. The reverberation processor then synthesizes the binaural signal based on the algorithm used, and the binaural signal is sent directly back to the wearable system on the wireless channel. The data return includes the necessary control data that enables the binaural signal block to determine the corresponding time frame. The audio signal sent to the processor that computes the direct sound path is delayed by a configurable delay of the same length as the transmission delay introduced by at least two wireless transmissions to and from the second device. The aggregator then combines signal blocks corresponding to the same audio signal block specified by the time data specified in the control data stream.
[0212] In a further system, additional reverberators that calculate late reverberation are distributed across separate devices. This second additional device includes a processor configured with a late reverberation algorithm and a subsystem for wireless or wired transmission to connected devices. In some embodiments, the additional reverberator device may be wirelessly connected to a wearable device. In other cases, it may be connected to the first additional device. The total delay of the selected transmission channels can be less than the target refresh interval at which the late reverberation signal or filter is processed. Due to the low delay requirement, the second additional device configured to process late reverberation can be connected over an IP network with a longer delay, such as the internet. In a further system, the additional reverberators are distributed across any number of additional devices.
[0213] Instead of being completely self-contained, some embodiments of the described system connect to another device using a wired or wireless connection. The connected device uses the connection to transmit the necessary audio data and metadata to the system. This allows a device such as a computer or smartphone to be connected to the system and used as a device for making spatial audio content audible.
[0214] Some embodiments of the device can provide positional data to the system using a so-called 3-degrees-of-freedom (3DoF) tracking system that only measures the user's head rotation. Similarly, some embodiments can transmit only 3DoF tracking data and limited translational or acceleration data to the system. In these embodiments, the system can be used to make a virtual audio scene audible to the user using a sound source present in space. As the user rotates their head, the sound source stabilizes in one position. As the user (and device) moves, the virtual sound source appears to move with the user because it is concentrated around the user. If any form of translational or acceleration data is available, it can be used to allow small head translations within a limited radius (i.e., 50 centimeters), which is useful for externalization and localization. Larger movements are not reproduced. These embodiments of the system are particularly useful for making classic spatial audio content, music, movies, and audio dramas audible, where the user is not intended to freely explore or leave the virtual acoustic scene.
[0215] Different embodiments of the device may utilize a 6-degrees-of-freedom (6DoF) tracking system that measures the user's absolute head rotation and position. In these embodiments, the user can freely navigate through and enter and exit a virtual acoustic environment. This is particularly useful for audible AR content, games, navigation content, and human-machine interaction scenarios. The entire system of sensors, RIR providers, binaural synthesizers, and auralizers may be distributed across multiple devices. For example, in particularly small form factors such as earphones, it may be necessary to distribute parts of the system across other devices. In this case, the position tracking sensors and sound transducers remain in the wearable, while the RIR providers, binaural synthesizers, and auralizers are distributed across one or more devices. In embodiments where this applies, motion-to-sound latency requirements must still be maintained.
[0216] In some embodiments of the system, the RIR provider can be configured to provide the RIR with different characteristics that do not partially or completely match the actual acoustic environment parameters. Alternatively, the system can be extended by an RIR modifier component, which receives the RIR input from the RIR provider and modifies it to change specific acoustic parameters of the RIR to match them, so that the modified RIR has these desired qualities and parameters. This can be done to change these room acoustic parameters to a desired level, for example, to make the listening room appear less reverberant for a more pleasant listening experience. Alternatively, this can be done to make the sound of a room sound more like another room, i.e., when listening to a concert, to make the sound of the current virtual acoustic environment sound more like a concert hall for aesthetic purposes. For example, a longer late reverberation tail (LR) can be audibly perceived by selecting an RIR that exhibits similar parameters but shows a longer reverberation time. Alternatively, the original (input) RIR of the LR can be resampled and thereby stretched by a certain amount, resulting in a longer reverberation time while keeping other perceptually relevant parts such as DS and ER intact.
[0217] Figure 19 shows a preferred embodiment of a sixth aspect of the present invention. In particular, the device shown in Figure 1 is separated into a first device and a second device. Specifically, the two-channel combiner 200 consists of two physically separate devices 901 and 902, as also shown in Figure 15. The first device 901 of the two physically separate devices is configured to process the direct sound portion, as shown in block 916 and in block 220 of Figure 2. For this purpose, processing requires listener position or rotation. Furthermore, the second device 902 of the two physically separate devices is configured to process at least one of the early reflection portion and the late reverberation portion. This block is shown in 923 and implements either or both of the functions of blocks 230 and 240 of Figure 2.
[0218] The two devices are connected via a transmission interface 918 of the first device and a transmission interface 925 of the second device. This transmission interface is preferably a wireless interface and operates according to, for example, the Bluetooth standard. Furthermore, one of the consequences of the separation of the two physically separated devices is that the first device 901 has its own power supply 917 and the second device 902 also has its own power supply 924.
[0219] Preferably, as shown in Figure 19, the first device is configured to update the 2-channel acoustic data for the direct tone portion more frequently than the second device updates the 2-channel audio data for at least one of the early reflection portion and the late reverberation portion. In the figure, it is preferable to have direct tone portion updates above 15 Hz, i.e., more than 15 updates per second, preferably more than 20 updates per second, and more preferably more than 50 updates per second. The update rate for the early reverberation portion is preferably in the range of 5 Hz to 15 Hz, and the update rate for the late reverberation portion may be in the range of 0.5 Hz to 5 Hz. Thus, portions requiring lower update rates appear to be handled by the second device 902. On the other hand, it has been found that these portions require very high processing power for long filters that require lower update rates. Therefore, the second device is implemented to be significantly stronger and more powerful than the first device in terms of computation and battery power. The first device may be an earphone device, a headphone device, an in-ear device, or any other wearable device that typically has limited battery power. However, the second device can be a high-power device such as a mobile phone, smartwatch, laptop computer, tablet, or a stationary computer connected to a power source and typically also connected to a large-area network such as the internet. Preferably, as already outlined in relation to Figure 15, the first device not only comprises a processing block for the direct sound portion 916, but also a microphone for recording acoustic measurements for RIR provision, further comprising a function for sound rendering as shown by the sound generator 300 in Figure 1, and further comprising a speaker if the device is, for example, a headphone device. Alternatively, if a Bluetooth signal is provided to the speaker, the speaker may be separated from the device 901 which has a communication interface rather than an actual speaker, for example.
[0220] Figures 21a and 21b show the starting point of an embodiment according to the sixth aspect shown in Figures 22a to 22f. In particular, in Figure 21a, the input interface comprises a microphone array consisting of one or more microphones indicated by 911, a user input tracking system 914, and potentially additional sensors 919. The binaural processor of block 200 comprises a direct sound processor and a reverberator for generating a binaural signal, the signal having a two-channel binaural signal which is then aggregated by a signal aggregator 310 and subsequently processed by a signal generator. The function of the signal aggregator is the same as in Figure 23b.
[0221] Conversely, Figure 21b has a similar implementation, but the binaural filter portion is aggregated as shown in the filter aggregator blocks 250 and 300, and the result of the filter aggregation is processed by the sound generator 300 or "Auralizer" of Figure 21b, which implements the procedure schematically shown in Figure 23a.
[0222] According to the invention as defined in the sixth aspect, reverberation processing is implemented in the second device, and direct tone processing is implemented in the first device 901. In addition, the functions of the input interface 100, the signal aggregator 310, and the signal generator 300 are also implemented in the first device 901 of Figure 23a, which implements the processing alternative of Figure 22a. Figure 22b is similar to Figure 22a, but has an alternative form of signal processing of Figure 23b. Figure 22c shows a further embodiment that differs from the embodiments of Figures 22a and 22b in that a second additional device 903 is provided. In particular, in this embodiment, the second device 903 typically processes the late reverberation processing of block 240 in Figure 2, the first additional device 902 performs the early refraction processing of block 230 in Figure 2, and the direct tone processor of the first device performs the direct tone processing 220 in Figure 2. Here too, the signal aggregator 310 aggregates the convolved audio signals individually, as shown in the alternative in Figure 23b. Figure 22d is similar to Figure 22c, but here it has the function of filter aggregation in line with the processing alternative in Figure 23a.
[0223] Figure 22e shows a further implementation where three or more additional devices are provided. Such additional devices 904 can be implemented to perform initialization tasks, for example, calculating the image source and image source location, so as to minimize battery usage of the wearable device. The additional devices 904 then receive the microphone signal and initial measurement data and perform other initialization procedures, such as image source location processing and correct room impulse response determination, using a database or the like, because these tasks are performed even less frequently than the calculation of the late reverberation portion. Further distribution of processing tasks to even more additional devices is also useful. Figure 22e again has the processing alternative shown in Figure 23b, while Figure 22f has the processing alternative shown in the filter aggregation shown in Figure 23a.
[0224] Figure 18 shows a preferred implementation of the processing according to the sixth embodiment, but this procedure can be applied to any other embodiment. In block 801, early single-channel acoustic data or early raw representation acquired by block 150 of the fifth embodiment is obtained.
[0225] In step 803, a new raw representation is acquired in response to control 802 providing an activation signal to block 803 at regular intervals or in response to a detected event, i.e., when the user moves from one room to another, resulting in an update of the entire room impulse response rather than an update of the user or listener position. In block 804, to determine whether an update is needed, the new raw representation is compared to an earlier raw representation, or the new single-channel acoustic data is compared to the earlier single-channel acoustic data.
[0226] In block 805, if the deviation is determined to be close to the threshold or update condition, new single-channel acoustic data 806 should be determined. Blend-over from early data to new data is used to gradually change from one RIR to the next; alternatively, the new data is used directly if it is not significantly different from the early data. In block 808, the early data in storage is overwritten with the current data, and in the next step by block 801, the current single-channel acoustic data or current raw representation is there.
[0227] Next, Figure 20 is shown to illustrate the processing of two-channel synthesis with different update rates. In block 930, the two-channel acoustic data currently used for early reflections and late reverberations is stored. In block 931, it is assumed that the two-channel acoustic data for the direct tone portion has been updated. In block 932, it is determined whether new data for the early reflection portion or the late reverberation portion is available. If this question is answered, the new data is used for sound generation along with the new direct tone data. However, in block 933, if it is determined that new data for the early reflection portion or the late reverberation portion is not available, the stored data for the early reflection portion or the late reverberation portion is used along with the new direct tone data. Therefore, the fact that the two-channel acoustic data for the portion with a reduced update rate is always stored means that this data can be easily used along with the new updated direct tone portion that requires a higher update rate.
[0228] Next, preferred embodiments of the present invention relating to a seventh aspect concerning improved separation of single-channel acoustic data and improved combination of two-channel acoustic data are shown.
[0229] In Figure 24a, block 600 refers to the preprocessing of the entire room impulse response, such as determining from a database, measurements, or a synthesis process, that the direct tone portion of the room impulse response is at a predetermined sample index. This preprocessed room impulse response is then transferred from the input interface 100 to block 210, which forms a two-channel synthesizer, particularly the separation. In block 601, the separation time between the direct tone portion and the early refracted tone portion is determined, for example, midway between the maximum value of the direct tone portion and the maximum value of the first early reflection. Additionally or alternatively, the separation time between the early deflection portion and the late reverberation portion is determined, for example, at the mixing time, or, for the purpose of saving computational resources, at a specific predetermined time before the mixing time.
[0230] In block 602, at least one of two adjacent parts is extended by a certain number of samples of the corresponding other part. For example, if a directional transfer function or directional impulse response is used in the direct tone part, the direct tone part is removed and does not need to be extended or subsequently windowed by block 603. However, if a directional transfer function is not used, or if the direct tone part of the RIR is used for any reason, the processing in blocks 602 and 603 also applies to the part of the direct tone part at the initial separation time. In block 603, at least the first early reflection part, the last early reflection part, and the first late reverberation part are windowed using a window function that takes extension into account, such as a Tukey window. Thus, in the output of block 603, there is a windowed first early reflection part, a windowed last early reflection part, and a windowed first late reverberation part.
[0231] Figure 24b shows the procedure for combining individually processed data, as shown in item 250 of Figure 2. For this purpose, before combining, each part is processed in an individual way as shown in block 604, the method as shown in relation to items 220, 230 and / or 240 of Figure 2. Next, overlap summation is performed between the 2-channel direct sound portion and the 2-channel early reflection portion in block 250a, overlap summation is performed between the 2-channel early reflection portion and the 2-channel late reverberation portion in block 250b, and finally, post-processing is performed in block 605 to obtain complete 2-channel acoustic data for use with the sound generator 300 in Figure 1.
[0232] In a preferred implementation as shown in Figure 25, the direct sound portion is generated, for example, using a directional impulse response or a directional transfer function + associated head-related impulse response. The result is expanded with n samples using the same procedure as described above for Figure 24a, and in block 603, windowing is performed, for example, using a Tukey window.
[0233] Furthermore, as shown in block 605, each segment within the early reflection portion is processed with overlap, all sequences are overlap-added, and in block 610, it is preferable to adjust the initial time delay gap. Following the adjustment of the initial time delay gap, channel-by-channel overlap-adding is performed as shown in block 606 to finally obtain aggregated 2-channel data of the acoustic environment.
[0234] Figure 26 shows a preferred implementation of the procedure performed in block 610 of Figure 25 for the purpose of adjusting the initial time delay gap. In block 611, the initial source position and initial sink position are used to determine the initial source-sink distance or initial propagation time and are further positioned at the image source position for the first reflection.
[0235] In block 612, the current source position and the current listener position are used to calculate the current distance or corresponding propagation time, along with the image source position for the first reflection. In block 613, the difference in distance or the difference in corresponding propagation time, e.g., delta ITDG, is calculated in block 630, and in block 640, ITDG is adjusted by shifting the early reflection portion (which usually already has a "connected late reverberation portion") more towards the direct sound or further away from the direct sound. For example, if the listener is close to the sound source, ITDG is larger than the initial time delay gap, so the early reflection portion is shifted away from the direct sound portion. Thus, the overlap is no longer perfectly matched, which can be explained by padding with some samples to have a perfect overlap at the beginning of the ER portion.
[0236] However, if the listener is further from the sound source compared to the initial measurement conditions, the ITDG will be smaller and the delta will be negative. In this case, the early reflection portion is shifted closer to the direct sound portion, which is managed by simply truncating several samples prior to the early reflection portion, so that these samples do not overlap and add up with the direct portion in block 606 in Figure 25 after ITDG adjustment in block 610.
[0237] Therefore, in order to maintain a plausible perception of distance, the initial time delay gap (ITDG) must be appropriate to the listener's pause during synthesis. This acoustic feature describes the gap between the direct sound and the first reflection. Thus, the timing relationship between DS and ER must be adapted. In a basic embodiment of a binaural synthesizer, this is achieved by temporally shifting the ER segment, since the system is designed to keep the DS portion in place. Utilizing an image source model, the ITDG can be calculated by obtaining the propagation time of the image source closest to the listener's position and subtracting this from the propagation time of the direct sound. This is done for the source-sync constellation of the initial conditions and the new constellation being synthesized. The difference between the two ITDG values gives how much the ER segment needs to be shifted to represent the new situation. For example, if the listener is closer to the sound source compared to the initial constellation, the ITDG will be larger, and therefore the ER segment will be shifted slightly away from DS. In other embodiments, this mechanism can be derived directly from the image source model by positioning individual reflections in relation to the direct sound.
[0238] Next, embodiments of the present invention relating to the first aspect are summarized, but the reference numerals in parentheses should not be considered to limit the scope of the embodiments.
[0239] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, Single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), The directional information of the sound source is determined regarding the listener's position and the source position or orientation of the sound source (222), An audio signal processor configured to use directional information in the calculation of 2-channel acoustic data for the direct sound portion (220).
[0240] The audio signal processor according to Embodiment 1, wherein the 2.2 channel combiner (200) is configured to determine two head-related data channels from source position and listener position or orientation in addition to directional information (227, 228), and to use the two head-related data channels and directional information for the calculation of 2 channel acoustic data for the direct sound portion (229).
[0241] 3. The audio signal processor according to Embodiment 1 or 2, wherein the 3.2 channel combiner (200) is configured to determine the direction of emittance information and the rotation of the sound source from the source location vector (423) of the sound source and the listener location vector (422) of the listener (222), and to derive directional information from a database of directional information sets, the directional information sets being associated with specific source emittance direction information.
[0242] 4. The apparatus according to Example 2 or 3, wherein the 4.2 channel combiner (200) is configured to derive the direction of arrival (421) of the listener position or orientation using the source location vector (423) of the sound source and the listener location vector (422) of the listener, as well as the listener's rotation.
[0243] 5. The directivity information is a directivity impulse response or a directivity transfer function, or the two head-related data channels are a first head-related impulse response or a first head-related transfer function and a second head-related impulse response or a second head-related transfer function, or the source emission direction information includes an angle or an index of a database, the apparatus according to any one of the preceding embodiments.
[0244] 6. The 2-channel synthesizer (200) determines a directivity impulse response as the directivity information, determines a first head-related impulse response and a second head-related impulse response as the two head-related data channels, combines the directivity impulse response and the first head-related impulse response, and combines the directivity impulse response and the second head-related impulse response using time-domain convolution or frequency-domain multiplication, the apparatus according to any one of the preceding embodiments.
[0245] 7. The 2-channel synthesizer (200) performs a padding operation (261) using the directivity impulse response and the first and second head-related impulse responses to obtain a padded function, converts the padded function to the frequency domain (262), multiplies the frequency-domain directivity information and the frequency-domain head-related data channels (263) to obtain two frequency-domain data channels, converts the two frequency-domain data channels to the time domain (264) to obtain a time-domain data portion of a direct sound portion of 2-channel acoustic data, the apparatus according to Embodiment 6.
[0246] 8. The 2-channel synthesizer (200) adjusts the phase of the 2-channel acoustic data (265) by removing the phase shift introduced by convolution so that the time-domain representation of the 2-channel acoustic data has a length equal to the length of the direct sound portion of the single-channel acoustic data that describes the acoustic environment, and truncates the phase-adjusted 2-channel acoustic data (266), an audio signal processor according to Example 6 or 7.
[0247] 9. The 2-channel synthesizer (200) determines an energy-related scale from the direct sound portion (221), determines an energy-related scale from the raw directivity information determined with respect to the listener position or orientation and the source position (223), and is configured to scale the raw directivity information (226) using a scaling value (224) derived from the energy-related scale to derive the determined directivity information, an apparatus according to any one of the preceding embodiments.
[0248] 10. The 2-channel synthesizer (200) determines distance scaling information from the distance between the source position and the listener position (225), and is configured to take into account the distance in the calculation of the 2-channel acoustic data of the direct sound portion (226), an apparatus according to any one of the preceding embodiments.
[0249] 11. The 2-channel synthesizer (200) is configured to generate amplified 2-channel acoustic data for the direct sound portion when the actual distance is smaller than the distance in the initial situation where the single-channel acoustic data was determined, and generate attenuated 2-channel acoustic data for the direct sound portion when the actual distance is greater than the distance in the initial situation, a signal processor according to Example 10.
[0250] 12. The apparatus according to any one of the prior art, wherein the 2-channel combiner (200) is configured to combine directional information and head-related impulse responses as head-related channel data using padding (261) that increases the length of both filters, multiply both filters in the spectral domain (263), convert the two multiplication results into the time domain (264), and remove introduced phase (265) such that the resulting center index is similar to the center index of the direct sound portion of single-channel acoustic data describing the acoustic environment.
[0251] 13. The audio signal processor according to Example 8 or 12, wherein the 2-channel combiner is configured to apply distance scaling information (226) to the result of phase rejection in the time domain.
[0252] 14. The audio signal processor according to any one of the prior art, wherein the 14.2 channel combiner (200) is configured to update the calculation of the direct tone portion (222) more frequently than the calculation of the early reflection portion (232) or the calculation of the late reverberation portion (240).
[0253] 15. A system comprising, or configured to access, a data set of directional information for multiple angles with respect to a given sound emission direction (430) of a sound source, distributed in a cylinder or sphere around the sound source location. The 2-channel combiner either derives a directional information dataset (269, 270, 271) based on the directional information dataset closest to the sound emittance direction (268) determined for the listener position, sound source position, and orientation, or derives two or more directional information datasets (271) having reference information closest to the determined sound emittance direction, and obtains directional information by interpolating between the two or more directional information datasets, or An audio signal processor according to any one of the prior embodiments, configured to synthesize directional information using a determined sound emittance direction and a sound source directional model (272).
[0254] A method for generating a 16.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, To synthesize, The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), Determining the directional information of the sound source regarding the listener's position and the source position or orientation of the sound source (222), A method comprising using directional information in the calculation (220) of two-channel acoustic data of the direct sound portion.
[0255] 17. A computer program, when executed on a computer or processor, for performing the method described in Example 16.
[0256] Next, embodiments of the present invention relating to a second aspect are summarized, but the reference numerals in parentheses should not be considered to limit the scope of the embodiments.
[0257] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, The system is configured to separate single-channel acoustic data into at least two parts, consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and to process at least two of these parts individually to generate two-channel acoustic data for each portion (220, 230, 240), The two-channel combiner (200) segments the early reflection portion into multiple segments (231), Multiple image source locations representing the source locations of reflected sound are determined (232), A matching operation, the matching operation comprising: calculating the sound arrival time of each image source to the listener position; and associating the image source position with a corresponding segment having a time delay in the corresponding segment that best matches the sound arrival time of the corresponding image source position (234), thereby associating the image source position with a segment, An audio signal processor configured to compute two-channel acoustic data of direct sound using the image source location associated with a segment.
[0258] 2. The audio signal processor according to Embodiment 1, wherein the 2.2 channel combiner (200) is configured to determine multiple image source positions (232) using initial measured initial source positions and initial sink positions for generating single channel acoustic data and geometric data of the acoustic environment.
[0259] 3. The audio signal processor according to Example 1 or 2, wherein the 2-channel combiner (200) is configured to determine the image source position using an image source method that models specular reflection in the acoustic environment.
[0260] 4. The channel synthesizer (200) is configured to use random or predetermined arrival direction data or two-channel head-related data for reflections within a segment that do not have an image source position up to a predetermined order or do not have a sound arrival time within a predetermined matching range for the time delay of reflections within the segment (235), according to any one of the previous embodiments of the audio signal processor.
[0261] 5. The two-channel synthesizer is configured to detect prominent reflections within the early reflection portion and arrange segments around each prominent reflection, the segments having a predetermined length corresponding to the length of the head-related impulse response or dividing the early reflection portion into a regular grid of reflection segments each having a sample count and an overlap to an adjacent segment, according to any one of the previous embodiments of the audio signal processor.
[0262] 6. The two-channel synthesizer is configured to detect prominent reflections by comparing the first average energy per sample within a first window with the second average energy per sample within a second window (283), the sample count of the second window being greater than the sample count of the first window, and a prominent reflection being determined when the first average energy is greater than the second average by a predetermined amount, according to the audio signal processor described in Example 5.
[0263] 7. The predetermined amount is between 3 dB and 9 dB, or the sample count of the first window is preferably at least 0.25 times smaller than the sample count of the second window, according to the audio signal processor described in Example 6.
[0264] 8. An audio signal processor according to any one of the prior embodiments, wherein the 2-channel combiner (200) is configured to determine direction of arrival information for each segment from the listener position and image source position associated with each segment (284), and to combine the early reflection portion within the segment with two head-associated data channels associated with the direction of arrival information to obtain at least a portion of the 2-channel acoustic data of the segment.
[0265] 9. The 2-channel combiner (200) pads the segments to the length of the 2-channel acoustic data in the time domain (285), The padded segment is converted to the frequency domain, and each channel of the frequency domain head-related 2-channel data is multiplied by the frequency domain padded segment to obtain the frequency domain 2-channel acoustic data of the segment. An audio signal processor according to any one of the prior embodiments, configured to convert frequency domain 2-channel data of a segment into a time domain.
[0266] The audio signal processor according to Example 9, wherein the 10.2-channel combiner is configured to remove the phase delay introduced from 2-channel acoustic data in the time domain.
[0267] 11. An audio signal processor according to any one of the prior arts, wherein the 2-channel combiner (200) is configured to generate 2-channel acoustic data for each segment from a combination (239) of a specular reflection portion derived using an image source position associated with the segment and a diffuse portion of the corresponding segment.
[0268] 12. An audio signal processor according to any one of the prior embodiments, wherein the single-channel acoustic data describing the acoustic environment is a room impulse response or a room transfer function, or the two-channel acoustic data is a binaural two-channel head-associated transfer function or a binaural two-channel head-associated transfer function.
[0269] 13. The audio signal processor according to any one of Examples 2 to 11, wherein the 2-channel combiner is configured to maintain the image source position for the listener position at the initial sync position and for listener positions different from the initial sync position, or for the initial source position or for source positions different from the initial source position, or to maintain the association between segments and image source positions for source positions at the initial source position or for source positions different from the initial source position.
[0270] 14. An audio signal processor according to any one of the prior embodiments, wherein the 2-channel synthesizer is configured to determine the directional information of the image source with respect to the listener position and the image source position or orientation, and to use the directional information for the calculation of 2-channel acoustic data of the early reflection portion (220).
[0271] 15. The audio signal processor according to Example 14, wherein the directional information of each image source is derived from the same set of directional information determined for the direct sound portion, or the orientation of the image sound source is determined by the image source model.
[0272] 16. The audio signal processor according to Example 14 or 15, wherein directional information is determined and used for a predetermined subset of segments within the early reflection portion.
[0273] 17. The audio signal processor according to Example 14 or 15, wherein a predetermined subset of segments within the early reflection portion includes fewer than 10 segments, preferably only 2 segments.
[0274] A method for generating an 18.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a 2-channel audio signal from an audio signal and 2-channel acoustic data, and the generation is The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), Segmenting the early reflection portion into multiple segments (231), Determining multiple image source locations that represent the source locations of reflected sound (232), A matching operation, the matching operation comprising: calculating the sound arrival time of each image source to the listener's position; and associating the image source position with a corresponding segment having a time delay within the corresponding segment that best matches the sound arrival time of the corresponding image source position (234), thereby associating the image source position with a segment using a matching operation, A method comprising calculating two-channel acoustic data of direct sound using the image source location associated with a segment.
[0275] 19. A computer program, when executed on a computer or processor, for performing the method described in Example 18.
[0276] Next, embodiments of the present invention relating to a third aspect are summarized, but the reference numerals in parentheses should not be considered to limit the scope of the embodiments.
[0277] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, The system is configured to separate single-channel acoustic data into at least two parts, consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and to process at least two of these parts individually to generate two-channel acoustic data for each portion (220, 230, 240), An audio signal processor configured to compute two-channel acoustic data of an early reflection portion (230) using a specular reflection portion that describes separate early reflections and a diffuse reflection portion that describes the diffuse effects within the early reflection portion.
[0278] 2. The audio signal processor according to Example 1, wherein the 2.2-channel combiner is configured to calculate the diffusion portion using a combination of the early reflection portion of single-channel acoustic data and a 2-channel noise sequence (238).
[0279] 3. The audio signal processor according to Example 1 or 2, wherein the 2-channel combiner (200) is configured to perform a weighted sum (239, 290) of the specular reflection portion (292) and the diffuse portion (293), the weights for the weighted sum are determined by a diffusion coefficient indicating the degree to which the early reflection portion segment of the single-channel acoustic data is diffuse.
[0280] 4.2 The channel synthesizer (200) is configured to determine the diffusivity coefficient from the ratio of a first mean of the energy per sample in a first window having sample count n to a second mean of the energy per sample in a second window having sample count m around the first window. The audio signal processor according to any one of the prior art embodiments, wherein if the ratio plus a first predetermined number is divided by a second predetermined number, the portion is considered to be perfectly mirrored if it is 1 or greater than 1, or if it is 0 or less than 0, the portion is considered to be perfectly diffused, and the second predetermined number is at least 3 dB greater than the first predetermined number, or within the range of 1.5 to 2.5 times the value of the first predetermined number.
[0281] 5.2 The audio signal processor according to any one of the prior arts, wherein the channel combiner is configured to segment the early reflection portion into multiple segments and to calculate the specular reflection portion and the diffuse reflection portion for each segment.
[0282] 6. The weights for weighted summation are further determined by the position of the segments of the early reflection portion relative to the direct sound portion and the late reverberation portion, thereby strengthening the weight of the specular reflection portion relative to the segment closer to the direct sound portion and strengthening the weight of the diffuse portion closer to the late reverberation portion, as described in any one of Examples 3 to 5.
[0283] 7. The audio signal processor according to Example 6, wherein the weights are determined such that the specular reflection portion of a segment temporally closer to the direct sound portion has a greater weight than the specular reflection data of a segment temporally closer to the late reverberation portion, or the diffusion data of a segment temporally closer to the direct sequence portion has a smaller weight than the specular reflection data of a segment temporally closer to the direct sound portion, or the weight of the specular reflection data of a segment is determined using the diffusion scale of the segment, and the weight of the diffusion data of a segment is determined using the diffusion scale of the corresponding specular reflection data of the segment.
[0284] 8. The multi-channel synthesizer (200) is, The specular reflection portion in the two channels is calculated using the direction of arrival data corresponding to the listener's position or orientation and source position of the early reflection portion, and the convolution of the direction of arrival data of the single-channel acoustic data and the head-related data channel associated with the early reflection portion. The diffusion portion is calculated using a combination of 2-channel binaural noise data and the early reflection portion of single-channel acoustic data. An audio signal processor according to any one of the prior embodiments, configured to combine a specular reflection portion and a diffusion portion.
[0285] 9.2-channel combiner (200) The specular reflection portion (292) in multiple segments of the premature reflection portion is calculated to obtain first channel specular reflection data for multiple segments and second channel specular reflection data for multiple segments. To obtain the first channel diffusion segment data and the second channel diffusion segment data, the diffusion portion within the same multiple segments of the early reflection portion is calculated. For each segment, the first channel of the segment's early reflection data is obtained by combining the segment's first channel specular reflection data and the segment's first channel diffuse specular reflection data. The audio signal processor according to Example 8, configured to combine second channel specular reflection data and second channel spread data of a segment in order to obtain a second channel of early reflection data of a segment.
[0286] 10. The audio signal processor according to Example 9, wherein the multichannel combiner (200) is configured to perform a linear combination using a first weight coefficient and a second weight coefficient, the sum of the first and second weight coefficients being substantially 1.
[0287] 11.2 Channel Synthesizer (200) When calculating the first and second channel specular reflection segment data, a window function is used to window the overlapping segments of the early reflection portion of the single-channel acoustic data. Using a similar window function, the overlapping segments of the first channel spread segment data are windowed. Using a similar window function, the overlapping segments of the second channel spread segment data are windowed. Audio signal processing according to Example 9 or 10, configured to obtain two-channel acoustic data of the early reflection portion by performing a weighted sum of corresponding first channel specular reflection segment data and first channel diffuse segment data with second channel specular reflection segment data and second channel diffuse segment data.
[0288] 12. The audio signal processor according to Example 11, wherein the 2-channel combiner is configured to overlap and add the resulting sequence data of segments for each channel in order to obtain 2-channel audio data of the early reflection portion.
[0289] 13. The audio signal processor according to Embodiment 12, wherein the 2-channel combiner is configured to take into account the initial time delay gap depending on the source position and listener position by temporally shifting the result of an overlap addition operation with the segment relative to the direct sound portion in order to acquire the early reflection portion in the timing relationship with the direct sound portion (610).
[0290] A method for generating a 14.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, To synthesize, The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), A method comprising (230) calculating two-channel acoustic data of an early reflection portion using a specular reflection portion that describes a separate early reflection and a diffuse portion that describes the diffuse effect within the early reflection portion.
[0291] 15. A computer program, when executed on a computer or processor, for performing the method described in Example 14.
[0292] Next, embodiments of the present invention relating to a fourth aspect are summarized, but the reference numerals in parentheses should not be considered to limit the scope of the embodiments.
[0293] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, The system is configured to separate single-channel acoustic data into at least two parts, consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and to process at least two of these parts individually to generate two-channel acoustic data for each portion (220, 230, 240), An audio signal processor configured to use the amplitude spectrum of an early reflection portion or single-channel acoustic data without a direct tone portion or late reverberation portion, and a first channel noise phase spectrum for obtaining the first channel of the two-channel acoustic data, and to use the amplitude spectrum of an early reflection portion or single-channel acoustic data without a direct tone portion or late reverberation portion, and a second channel noise phase spectrum to calculate the two-channel diffuse portion of the early reflection portion or single-channel acoustic data without a direct tone portion or late reverberation portion.
[0294] 2. The audio signal processor according to Example 1, wherein the first channel noise phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence.
[0295] 3.2 The audio signal processor according to Example 1 or 2, wherein the channel combiner (200) is configured to calculate a first spectrogram of single-channel acoustic data without the early reflection portion, or the direct tone portion or late reverberation portion of the single-channel acoustic data, a second spectrogram of the first channel noise phase spectrum, and a third spectrogram of the second channel noise phase spectrum (530).
[0296] 4. The audio signal processor according to Embodiment 3, wherein the 2-channel combiner is configured to calculate a first spectrogram as the amplitude spectrum of the first sequence, calculate a second spectrogram as the phase spectrum of the second sequence, calculate a third spectrogram as the phase spectrum of the third sequence, combine the amplitude spectrum of the first sequence and the phase spectrum of the second sequence to obtain the first channel of the 2-channel spread portion, and combine the amplitude spectrum of the first sequence and the phase spectrum of the third sequence to obtain the second channel of the 2-channel spread portion.
[0297] 5. The audio signal processor according to Embodiment 3 or 4, wherein the second channel combiner is configured to use overlapping segments and window functions for each segment in the calculation of the first, second, and third spectrograms.
[0298] 6. The audio signal processor according to Example 4 or 5, wherein the first spectrogram, the second spectrogram, and the third spectrogram are calculated as complex spectra and converted to polar representations.
[0299] 7. The 2-channel combiner is configured to low-pass filter the amplitude spectra of the amplitude spectrum sequence (448), so that the low-pass filtered amplitude spectrum of the first sequence is combined with the phase spectra of the second and third sequences, as described in any one of Examples 4 to 6.
[0300] 8. The low-pass filter is a moving average filter, as described in Example 7 of the audio signal processor.
[0301] 9. The audio signal processor according to Example 8, wherein the moving average filter extends over sizes between 0.1 octaves and 0.75 octaves.
[0302] 10. The audio signal processor according to any one of Examples 3 to 9, wherein the multichannel combiner is configured to perform low-pass filtering on each spectrum in a first sequence of the spectrum of a first spectrogram, or in the spectrogram of the diffusion portion of the first channel or the second channel of the second channel (447), so that the frequency bins of adjacent spectra related to the same frequency are low-pass filtered.
[0303] 11. The audio signal processor according to Example 10, wherein the low-pass filter for low-pass filtering is a moving average filter having a number of inputs between 2 and 6.
[0304] 12. The audio signal processor according to Embodiment 4, wherein the 2-channel combiner (200) is configured to convert the first channel of the 2-channel spread portion and the second channel of the 2-channel spread portion into the time domain (450) in order to obtain an overlap block of the first channel and the second channel.
[0305] 13. The apparatus according to Example 12, wherein the 2-channel synthesizer is configured to perform overlap addition (452) of overlapping time-domain blocks on one side for the first channel and on the other side for the second channel in order to obtain the diffusion portion of the 2-channel representation.
[0306] 14. The apparatus according to any one of the prior art, wherein the 2-channel synthesizer is configured to use only the diffuse portion of the late reverberation as 2-channel acoustic data, or to use a combination of the diffuse portion and the specular reflection portion as 2-channel acoustic data of the early reflection portion.
[0307] A method for generating a 15.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, To synthesize, The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), A method comprising: using the amplitude spectrum of an early reflection portion or single-channel acoustic data without a direct sound portion or late reverberation portion, and a first channel noise phase spectrum for obtaining the first channel of two-channel acoustic data; and using the amplitude spectrum of an early reflection portion or single-channel acoustic data without a direct sound portion or late reverberation portion, and a second channel noise phase spectrum, calculating the two-channel diffuse portion of an early reflection portion or single-channel acoustic data without a direct sound portion or late reverberation portion.
[0308] 16. A computer program, when executed on a computer or processor, for performing the method described in Example 15.
[0309] Next, embodiments of the present invention relating to a fifth aspect are summarized, but the reference numbers in parentheses should not be considered to limit the scope of the embodiments.
[0310] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, An audio signal processor is configured to have an input interface (100) that takes a raw representation (150) of single-channel acoustic data and to derive single-channel acoustic data (151) using the raw representation and additional data stored in or accessible by the audio signal processor.
[0311] 2. The input interface (100) is, As a raw representation, initial measurements of raw single-channel acoustic data were obtained (150), A test fingerprint is derived (101), a reference fingerprint, each reference fingerprint is associated with high-resolution single-channel acoustic data, the high-resolution single-channel acoustic data has a higher resolution than the initial measurement, and access is made to a pre-stored database having an associated set of reference fingerprints. The audio signal processor according to Example 1, configured to retrieve high-resolution single-channel acoustic data having a reference fingerprint that best matches a test fingerprint from a pre-stored database (113), or to synthesize high-resolution acoustic data from a test fingerprint from initial measurements of raw single-channel acoustic data or from geometric parameters (140).
[0312] 3. The input interface (100) is As a raw representation, we obtained initial measurements of raw single-channel acoustic data. Derive a test fingerprint, The audio signal processor described in Example 1, configured to synthesize single-channel acoustic data from test fingerprints or initial measurements of raw single-channel acoustic data (140).
[0313] 4. The audio signal processor according to Example 1, wherein the raw representation is a geometric description of the acoustic environment, and the input interface (100) is configured to perform an acoustic chamber simulation to derive single-channel acoustic data from the geometric description.
[0314] 5. The input interface (100) is configured to determine at least one of the following parameters RT60, EDC, and DRR as a test fingerprint: The reference fingerprint is the audio signal processor according to Example 1, which includes at least one of the following parameters: RT60, EDC, and DRR.
[0315] 6. An audio signal processor according to any one of the prior arts, wherein the input interface (100) is configured to apply a psychoacoustic weighting function to a calculated fingerprint in order to obtain a fingerprint for accessing a pre-stored database (110) or for performing direct synthesis (140).
[0316] 7. An audio signal processor according to any one of Examples 1 to 3, wherein the input interface (100) is configured to derive a fingerprint using a trained neural network, or to perform direct synthesis (140) using a trained neural network from a raw representation associated with single-channel acoustic data.
[0317] 8. The audio signal processor according to Embodiment 1, wherein the input interface (100) is configured to compute a test fingerprint using a trained neural network, the trained neural network is trained to classify single-channel acoustic data into a single room class, and the input interface (100) is configured to synthesize prototype single-channel acoustic data of a fingerprint indicating a matched room class (153) or to retrieve prototype single-channel acoustic data of a matched room class from a pre-stored database (152).
[0318] 9. The input interface (100) is, The test fingerprint is derived such that it has a smaller dimension than the raw single-channel acoustic data. From a pre-stored database, a low-dimensional reference fingerprint is derived using the same procedure as the derivation of the test fingerprint. An audio signal processor according to any one of Examples 1 to 5, configured to select single-channel acoustic data having a reference fingerprint that minimizes the distance to a test fingerprint.
[0319] 10. An audio signal processor according to any one of the prior embodiments, wherein the input interface (100) is configured to use a natural sound that can be generated by a listener for initial measurement.
[0320] 11. The audio signal processor according to Example 10, wherein the natural sound is a transient sound that can be generated by clapping, speaking, or by a listener.
[0321] 12. The input interface (100) is, Record audio played by one or more speakers in an acoustic environment (150), The voice identification process is used to determine the voice (155,156), Access a database that has at least an approximation of the representation of sound played by one or more speakers, without being affected by the acoustic environment (157), The audio signal processor according to Embodiment 1, configured to determine single-channel acoustic data using recorded audio and audio acquired from a database (159).
[0322] 13. The audio signal processor according to Example 8, wherein the input interface (100) includes a second trained neural network for generating single-channel acoustic data from a test fingerprint computed by a first trained neural network.
[0323] 14. An audio signal processor according to any one of the prior embodiments, wherein the input interface (100) comprises a speaker and a microphone embedded in the mobile device, and the input interface (100) is configured to perform initial measurements using the speaker and microphone, or using only the microphone embedded in the mobile device.
[0324] 15. An audio signal processor according to any one of the prior art embodiments, wherein the input interface (100) is configured to receive new single-channel acoustic data at regular intervals or at specific events, and, if the deviation exceeds a deviation threshold, compare the new single-channel acoustic data with the original single-channel acoustic data and replace the original single-channel acoustic data with the new single-channel acoustic data, or compare a new initial measurement with an earlier initial measurement, or compare a new test fingerprint with an earlier test fingerprint, or compare a new raw representation with an earlier raw representation.
[0325] 16. An audio signal processor according to any one of the prior embodiments, wherein the input interface (100) is configured to store a history of early single-channel acoustic data in order to enable blending from early single-channel acoustic data to new single-channel acoustic data.
[0326] 17. The audio signal processor according to Example 16, wherein the blend includes linear interpolation in the time or frequency domain between the preceding single-channel acoustic data and the following single-channel acoustic data.
[0327] A method for generating an 18.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, Synthesis is a method comprising obtaining a raw representation related to single-channel acoustic data (150) and deriving single-channel acoustic data using the raw representation and additional data stored in or accessible by an audio signal processor (151).
[0328] 19. A computer program, when executed on a computer or processor, for performing the method described in Example 18.
[0329] Next, embodiments of the present invention relating to the sixth aspect are summarized, but the reference numbers in parentheses should not be considered to limit the scope of the embodiments.
[0330] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, The system is configured to separate single-channel acoustic data into at least two parts, consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and to process at least two of these parts individually to generate two-channel acoustic data for each portion (220, 230, 240), An audio signal processor comprising two physically separated devices (901, 902), wherein the first of the two physically separated devices (901) is configured to process at least one of the direct sound portion and the early reflection portion (220, 230), and the second of the two physically separated devices (903) is configured to process at least one of the early reflection portion and the late reverberation portion (230, 240), and the first device (901) and the second device (902) are connected via a transmission interface (918, 925) and have separate power supplies (917, 924).
[0331] 2. The audio signal processor according to Embodiment 1, wherein the first device (901) is configured to update the two-channel acoustic data of the direct sound portion or the early reflection portion more frequently than the second device (902) updates the two-channel acoustic data of at least one of the early reflection portion and the late reverberation portion.
[0332] 3. The audio signal processor according to Embodiment 1, wherein the transmission interface (918, 925) is configured to operate according to a wireless transmission protocol.
[0333] 4. An audio signal processor according to any one of the prior embodiments, wherein the first device (901) is a wearable device further comprising an input interface (100) and a sound generator (300), and the second device (102) is a mobile or fixed device separate from the wearable device.
[0334] 5. The audio signal processor according to any one of the prior embodiments, wherein the wearable device (901) is an earphone device, a headphone device, or an in-ear device, and the mobile device or fixed device is a mobile phone, a smartwatch, a tablet, a notebook computer, or a fixed computer.
[0335] 6. An audio signal processor according to any one of the prior embodiments, wherein the first device (901) comprises a user tracking system (914) and is configured to transmit data relating to the user's position or orientation to a second device (902).
[0336] 7. The 2-channel synthesizer (200) is configured to separate single-channel acoustic data into three parts: direct sound, early reflections, and late reverberation (210). An audio signal processor according to any one of the prior art embodiments, wherein two-channel acoustic data for the direct sound portion is generated by a first device (220), two-channel acoustic data for the early reflection portion is generated by a second device (902) (230), or two-channel acoustic data for the late reverberation portion is generated by a third device (903) (240), the third device (903) being separate from the first device (901) and the second device (902).
[0337] 8. The audio signal processor according to Embodiment 6, wherein the second device (902) is a mobile phone having access to the Internet, and the third device (903) is a remote computer connected to the mobile device via the Internet, and the two-channel audio data for the late reverberation portion is not updated more frequently than the two-channel acoustic data for the early reverberation portion.
[0338] 9. An audio signal processor according to any one of the prior arts, wherein the second device (902) is configured to receive user position or orientation from the first device (901), provide two-channel acoustic data of the early reflection portion and / or late reverberation portion, and transmit the two-channel acoustic data of the early reflection portion and / or late reverberation portion to the first device.
[0339] 10. The second device is configured to receive user position or orientation and audio signals from the first device and to provide at least two-channel acoustic data of the early reflection portion. An audio signal processor according to any one of the prior embodiments, wherein the sound generator (300) is distributed between a first device (901) and a second device (902), the first device being configured to generate a two-channel audio signal of the direct sound portion, the second device being configured to generate a two-channel audio signal of at least the early reflection portion, and the second device being configured to transmit the two-channel audio signal of the early reflection portion to the first device.
[0340] 11. An audio signal processor according to any one of the prior embodiments, wherein the first device (901) is configured to delay the two-channel acoustic data of the direct sound portion by a delay value that covers the delay incurred by transmission to and from the second device.
[0341] 12. The first device (901) has a memory for storing second acoustic data of the early reflection portion and / or the late reverberation portion. An audio signal processor according to any one of the prior embodiments, wherein a two-channel combiner (200) or sound generator (300) is configured to use stored two-channel acoustic data in the calculation of complete two-channel acoustic data when updated two-channel data for the direct sound portion is available, but updated two-channel acoustic data for the early reflection portion or late reverberation portion is not available (933) due to different update speeds between the first device (901) and the second device (902).
[0342] 13. The audio signal processor according to any one of the prior embodiments, wherein the sound generator (300) is configured to aggregate 2-channel acoustic data from each part to obtain complete 2-channel acoustic data, and to combine the complete 2-channel acoustic data with a multi-channel audio signal to obtain a 2-channel audio signal, or to combine the 2-channel acoustic data from each part with an input audio signal to obtain partial 2-channel audio signals from each part, and to aggregate the partial 2-channel audio signals to obtain a 2-channel audio signal.
[0343] 14.2 The audio signal processor according to any one of the prior embodiments, wherein the channel analyzer is configured to update the 2-channel acoustic data of each part at different rates, such that the direct sound part is updated more frequently than the rest, or the early reflection part is not updated more frequently than the direct sound part, or the late reverberation part is not updated more frequently than the rest of the 2-channel acoustic data of the acoustic environment.
[0344] 15. The audio signal processor according to any one of the prior arts, wherein the second device comprises a computer or reverberation network for generating or processing two-channel audio data of early reflections and / or late reverberation portions, or the update ratio of the direct tone portion is greater than 15 Hz, the update ratio of the early reflection portion is greater than 5 Hz and less than 15 Hz, or the update ratio of the late reverberation portion is greater than 0.5 Hz and less than 5 Hz.
[0345] A method for generating a 16.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, To synthesize, This includes separating single-channel acoustic data into at least two parts consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and processing at least two parts individually for each part to generate two-channel acoustic data (220, 230, 240), The synthesis method comprises using two physically separated devices (901, 902), the first of the two physically separated devices (901) processing at least one of the direct sound portion and the early reflection portion (220, 230), the second of the two physically separated devices (903) processing at least one of the early reflection portion and the late reverberation portion (230, 240), and the first device (901) and the second device (902) being connected via a transmission interface (918, 925) and having separate power supplies (917, 924).
[0346] 17. A computer program, when executed on a computer or processor, for performing the method described in Example 16.
[0347] Next, embodiments of the present invention relating to the seventh aspect are summarized, but the reference numerals in parentheses should not be considered to limit the scope of the embodiments.
[0348] 1. An audio signal processor for generating 2-channel audio signals, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from single-channel audio data using listener position or rotation, The system includes a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, The 2-channel combiner (200) is, Single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), The separation time within single-channel acoustic data is determined between the direct sound portion and the early reflection portion or between the early reflection portion and the late reverberation portion (601), To achieve overlap in the separation time, extend at least one of the two parts of the separation time by a certain number of samples (602), An audio signal processor configured to window at least one extended portion using a specific window function that takes sample extension into account (603).
[0349] 2. The number of overlapping samples is obtained from each of the other parts, as described in Example 1 of the audio signal processor.
[0350] 3. The audio signal processor according to Example 1 or 2, wherein the window function is a Tukey window having lobes with width 2n, where n is a specific number of samples, and the two parts are each expanded by n samples.
[0351] 4.2 The audio signal processor according to any one of the prior arts, wherein the channel combiner is configured to determine the separation time between the direct sound portion and the early reflection portion, such that the distance of the separation time is substantially midway between the direct sound peak and the first early reflection peak, or to determine the separation time between the early reflection portion and the late reverberation portion as the perceived mixing time of the acoustic environment, or a predetermined time before the perceived mixing time.
[0352] 5.2 The channel combiner is configured to perform overlap addition operations between the first channel of the 2-channel acoustic data for the direct sound portion, the first channel of the 2-channel acoustic data for the early reflection portion, and the first channel of the 2-channel acoustic data for the late reverberation portion, following individual processing of the corresponding portions (220, 230, 240). The apparatus according to any one of the prior arts, wherein the two-channel combiner is configured to perform an overlap addition operation using the second channel of the two-channel acoustic data for the direct sound portion, the second channel of the two-channel acoustic data for the early reflection portion, and the second channel of the two-channel acoustic data for the late reverberation portion, following separate processing of the corresponding portions (220, 230, 240).
[0353] 6.2 The audio signal processor according to any one of the prior arts, wherein the channel combiner (200) preprocesses single-channel acoustic data by detecting the direct sound index in the temporal representation of the single-channel acoustic data (600), and is configured to cut or extend the beginning portion of the temporal representation of the single-channel acoustic data with zero-value samples so that the detected time index matches a predetermined sample index offset from the beginning of the single-channel acoustic data.
[0354] 7. An audio signal processor according to any one of the prior embodiments, wherein the acoustic data describing the acoustic environment is a room impulse response or room transfer function, or the 2-channel acoustic data is a binaural 2-channel head-associated impulse response or a binaural 2-channel head-associated transfer function.
[0355] 8. The audio signal processor according to any one of the prior embodiments, wherein the sound generator (300) is configured to aggregate 2-channel acoustic data from each part to obtain complete 2-channel acoustic data, and to combine the complete 2-channel acoustic data with the input audio signal to obtain a 2-channel audio signal, or to combine the 2-channel acoustic data from each part with the input audio signal to obtain partial 2-channel audio signals from each part, and to aggregate the partial 2-channel audio signals to obtain a 2-channel audio signal.
[0356] 9.2 The audio signal processor according to any one of the prior art, configured to determine the current distance of the listener's position to the source position relative to the initial distance of the initial generation of single-channel acoustic data (611), and to adjust the time period between the two-channel acoustic data of the direct sound portion and the two-channel acoustic data of the early reflection portion such that the time period is extended if the current distance is less than the initial distance, or contracted if the current distance is greater than the initial distance (614).
[0357] 10. The audio signal processor according to Example 9, wherein the multichannel synthesizer is configured to add zero samples to the early reflection portion before overlap addition when the time period is extended, or to remove excess samples when the time period is shortened.
[0358] 11. The audio signal processor according to Embodiment 9 or 10, wherein the multichannel synthesizer (200) is configured to determine a first initial time delay gap for an initial measurement of single-channel acoustic data provided by an input interface (100) (611), determine a second initial time delay gap for the current listener position and the current sound source position (612), calculate the difference between the first initial time delay gap and the second initial time delay gap (630), and adjust the first initial time delay gap by shifting the early reflection portion by the calculated difference (614).
[0359] 12. The audio signal processor according to Embodiment 11, wherein the multichannel synthesizer is configured to determine a second initial time delay gap by the difference between the propagation time of the first reflection from the image source position associated with the first segment of the early reflection portion to the current listener position and the propagation time from the current sound source position to the current listener position (612).
[0360] 13.2 The audio signal processor according to any one of the prior embodiments, wherein the channel analyzer is configured to store early-generated 2-channel acoustic data of early reflection and late-reverberation portions to which a window function is applied, and to retrieve the stored 2-channel acoustic data in order to propose a channel overlap addition operation (606) with other newly updated portions of the 2-channel acoustic data.
[0361] A method for generating a 14.2-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing 2-channel audio data from single-channel audio data using listener position or rotation, This includes generating a two-channel audio signal from an audio signal and two-channel acoustic data, To synthesize, The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of the early reflection portion and the late reverberation portion (210), and at least two parts are processed individually to generate 2-channel acoustic data for each part (220, 230, 240), Determining the separation time within single-channel acoustic data between the direct sound portion and the early reflection portion or between the early reflection portion and the late reverberation portion (601), To achieve overlap in the separation time, extend at least one of the two parts of the separation time by a certain number of samples (602), A method comprising windowing at least one extended portion using a specific window function that takes sample extensions into consideration (603).
[0362] 15. A computer program, when executed on a computer or processor, for performing the method described in Example 14.
[0363] It should be noted that in this specification, all of the alternative forms or embodiments described above, and all of the embodiments defined by the independent claims in the following claims, may be used individually, i.e., without alternative forms or embodiments other than those intended by the alternative forms, purposes or independent claims. However, in other embodiments, two or more alternative forms or embodiments or independent claims may be combined with each other, and in other embodiments, all of the embodiments or alternative forms and all of the independent claims may be combined with each other.
[0364] While some embodiments have been described in the context of the apparatus, it is clear that these embodiments also represent descriptions of the corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, embodiments described in the context of method steps also represent descriptions of the corresponding blocks, items, or functions of the corresponding apparatus.
[0365] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementations may be made using digital storage media such as floppy disks, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which store electronically readable control signals and cooperate (or can cooperate) with a programmable computer system so that each method is performed. Some embodiments of the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is performed. Generally, embodiments of the present invention may be implemented as a computer program product having program code that operates to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier. Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable carrier or a non-temporary storage medium. In other words, one embodiment of the methods of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer. Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods of the present invention is recorded. Accordingly, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods of the present invention. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, for example, over the Internet. Further embodiments include processing means configured or adapted to perform one of the methods of the present invention, such as a computer or a programmable logic device.Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed. In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0366] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the pending claims and not by the specific details presented as part of the description and explanation of the embodiments herein.
Claims
1. An audio signal processor for generating a two-channel audio signal, An input interface (100) for providing single-channel acoustic data describing the acoustic environment, A two-channel synthesizer (200) for synthesizing two-channel audio data from the single-channel audio data using the listener's position or rotation, The system includes a sound generator (300) for generating the two-channel audio signal from the audio signal and the two-channel acoustic data, The aforementioned two-channel combiner (200) The single-channel acoustic data is separated into at least two parts, consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and the at least two parts are processed individually to generate two-channel acoustic data for each part (220, 230, 240). The two-channel combiner comprises two physically separated devices (901, 902), the first of the two physically separated devices (901) configured to process at least one of the direct sound portion and the early reflection portion (220, 230), the second of the two physically separated devices (903) configured to process at least one of the early reflection portion and the late reverberation portion (230, 240), the first device (901) and the second device (902) connected via a transmission interface (918, 925) and having separate power supplies (917, 924), an audio signal processor.
2. The audio signal processor according to claim 1, wherein the first device (901) is configured to update the two-channel acoustic data of the direct sound portion or the early reflection portion more frequently than the second device (902) updates the two-channel acoustic data of at least one of the early reflection portion and the late reverberation portion.
3. The audio signal processor according to claim 1, wherein the transmission interfaces (918, 925) are configured to operate in accordance with a wireless transmission protocol.
4. The audio signal processor according to any one of the prior claims, wherein the first device (901) is a wearable device further comprising the input interface (100) and the sound generator (300), and the second device (102) is a mobile device or fixed device separate from the wearable device.
5. The audio signal processor according to any one of the prior claims, wherein the wearable device (901) is an earphone device, a headphone device, or an in-ear device, and the mobile device or fixed device is a mobile phone, a smartwatch, a tablet, a notebook computer, or a fixed computer.
6. The audio signal processor according to any one of the prior claims, wherein the first device (901) comprises a user tracking system (914) and is configured to transmit data relating to the user's position or orientation to the second device (902).
7. The two-channel combiner (200) is configured to separate the single-channel acoustic data into three parts: partial direct sound, early reflection, and late reverberation (210). The audio signal processor according to any one of the prior claims, wherein the two-channel acoustic data of the direct sound portion is generated by the first device (220), the two-channel acoustic data of the early reflection portion is generated by the second device (902) (230), or the two-channel acoustic data of the late reverberation portion is generated by the third device (903) (240), the third device (903) being separate from the first device (901) and the second device (902).
8. The audio signal processor according to claim 6, wherein the second device (902) is a mobile phone having access to the Internet, and the third device (903) is a remote computer connected to the mobile device via the Internet, and the two-channel audio data of the late reverberation portion is not updated more frequently than the two-channel acoustic data of the early reverberation portion.
9. The audio signal processor according to any one of the prior claims, wherein the second device (902) is configured to receive the user position or orientation from the first device (901), provide the two-channel acoustic data of the early reflection portion and / or the late reverberation portion, and transmit the two-channel acoustic data of the early reflection portion and / or the late reverberation portion to the first device.
10. The second device is configured to receive the user position or orientation and the audio signal from the first device and to provide the two-channel acoustic data for at least the early reflection portion. The audio signal processor according to any one of the prior claims, wherein the sound generator (300) is distributed between a first device (901) and a second device (902), the first device is configured to generate the two-channel audio signal of the direct sound portion, the second device is configured to generate the two-channel audio signal of at least the early reflection portion, and the second device is configured to transmit the two-channel audio signal of the early reflection portion to the first device.
11. The audio signal processor according to any one of the prior claims, wherein the first device (901) is configured to delay the two-channel acoustic data of the direct sound portion by a delay value that covers the delay incurred by transmission to and from the second device.
12. The first device (901) has a memory for storing the second acoustic data of the early reflection portion and / or the late reverberation portion, The audio signal processor according to any one of the prior claims, wherein the two-channel combiner (200) or the sound generator (300) is configured to use the stored two-channel acoustic data in calculating complete two-channel acoustic data when updated two-channel data for the direct sound portion is available, and updated two-channel acoustic data for the early reflection portion or the late reverberation portion is not available (933) due to different update speeds between the first device (901) and the second device (902).
13. The audio signal processor according to any one of the prior claims, wherein the sound generator (300) is configured to aggregate the two-channel sound data from each part to obtain complete two-channel sound data, and to combine the complete two-channel sound data with a multi-channel audio signal to obtain the two-channel audio signal, or to combine the two-channel sound data from each part with the input audio signal to obtain partial two-channel audio signals from each part, and to aggregate the partial two-channel audio signals to obtain the two-channel audio signal.
14. The audio signal processor according to any one of the prior claims, wherein the two-channel analyzer is configured to update the two-channel acoustic data of each portion at different rates, the direct sound portion being updated more frequently than the remaining portion, or the early reflection portion being updated less frequently than the direct sound portion, more frequently than the late reverberation portion, or the late reverberation portion being updated less frequently than the remaining portion of the two-channel acoustic data of the acoustic environment.
15. The audio signal processor according to any one of the prior claims, wherein the second device comprises a computer or reverberation network for generating or processing the two-channel audio data of the early reflections and / or the late reverberation portion, or the update ratio of the direct sound portion is greater than 15 Hz, the update ratio of the early reflection portion is greater than 5 Hz and less than 15 Hz, or the update ratio of the late reverberation portion is greater than 0.5 Hz and less than 5 Hz.
16. A method for generating a two-channel audio signal, To provide single-channel acoustic data that describes the acoustic environment, Synthesizing two-channel audio data from the single-channel audio data using the listener's position or rotation, This includes generating the two-channel audio signal from the audio signal and the two-channel acoustic data, The aforementioned synthesis is The single-channel acoustic data is separated into at least two parts, each consisting of a direct sound portion and at least one of an early reflection portion and a late reverberation portion (210), and the at least two parts are processed individually to generate two-channel acoustic data for each part (220, 230, 240), The synthesis method comprises using two physically separated devices (901, 902), wherein the first device (901) of the two physically separated devices processes at least one of the direct sound portion and the early reflection portion (220, 230), and the second device (903) of the two physically separated devices processes at least one of the early reflection portion and the late reverberation portion (230, 240), and the first device (901) and the second device (902) are connected via a transmission interface (918, 925) and have separate power supplies (917, 924).
17. A computer program for performing the method according to claim 16, when executed on a computer or processor.