Audio signal processor and related method and computer program for generating two-channel audio signal using specific integration of noise sequences
By integrating sound source directional information in binaural audio systems, processing early reflections and late reverbs in segments, and assigning computing tasks between devices, computing complexity and delay problems in the prior art are solved, and efficient and natural virtual sound source generation is achieved.
Patent Information
- Application Number
- CN202380088583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-24
- Filing Date
- 2023-10-24
- Publication Date
- 2025-08-19
AI Technical Summary
Existing binaural audio rendering systems have challenges in computing complexity, latency and processing capabilities, making it difficult to generate high-quality virtual sound sources in real time on devices with limited resources, especially on devices such as augmented reality and true wireless earbuds, resulting in unnatural sound perception and high computational cost.
By integrating the directional information of the sound source from single-channel acoustic data, processing the early reflection part in segments, processing the late reverb part in combination with the dual-channel noise sequence, and assigning computing tasks between different devices, using psychological acoustic knowledge to optimize window functions and delay management, achieving efficient dual-channel acoustic data generation.
It realizes efficient generation of high-quality virtual sound sources on limited resource equipment, ensures sound externalization effect, reduces artifacts, and improves auditory nature and computing efficiency.
Smart Images

Figure CN120513645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus, a method or a computer program for audio reproduction, such as binaural reproduction, via headphones or loudspeakers. In particular, the invention relates to the processing of digital audio signals and acoustic data describing an acoustic environment. Background Art
[0002] State-of-the-art binaural audio rendering systems allow users to simulate and listen to virtual sound sources that are precisely localized in space. The simulated sounds appear to originate outside the head, a process known as "externalization." With an appropriate system, binaurally rendered sound sources can be perceived at stable locations in space and appear to have acoustic properties similar to those of real sound sources. This can make them virtually indistinguishable from real sound sources.
[0003] There are many binaural synthesis methods and algorithms that can be used to achieve externalization. What they all have in common is that they aim to approximate the filtering effects that sound undergoes on its simulated path to the listener's ears. The combined filter of this system, consisting of the sound source, the acoustic influences and geometry of the virtual or real environment, the listener's head and body, and potential other influences on the sound caused by the environment, is called the binaural room impulse response (BRIR).
[0004] The two main components of BRIR are the Head Related Transfer Function (HRTF) and the Room Impulse Response (RIR). HRTF encodes the measured or approximate filtering effects of the human head, torso, and outer ear. Therefore, it depends on the geometry and shape of the listener's head, as well as the relative position and rotation of the head and the sound source.
[0005] RIR encodes the filtering effects of the room—the reflections, diffractions, and shadowing of sounds introduced by the room's geometry. It depends on the room's geometry and the position and rotation of the listener and sound source within it. (A room is any environment, not just a building.)
[0006] Emulating these effects is typically done through complex simulations or more lightweight approximations, which require complex room geometry models to simulate convincing room impulse responses. Depending on the binaural synthesis algorithm used, current state-of-the-art algorithms often have to strike a trade-off between computational complexity, limiting the lower bound size of the target system, or the effectiveness of the simulation, which often results in sound sources that are difficult to localize or completely head-localizable.
[0007] Furthermore, these devices require room geometry data for the current room, including reflective surfaces, their absorption and scattering coefficients. This data is difficult to obtain, especially in augmented reality (AR) settings, where the device's use is not limited to a single room. Acquiring it is often infeasible, and measuring it automatically is a difficult task even for trained users.
[0008] Depending on the binaural synthesis algorithms and techniques employed, these processes can be very computationally intensive and time-consuming. However, the processing power of the target device is often limited. For example, binaural rendering can be deployed on "True Wireless Earbuds" or similar smart headphones or wearable devices, which only offer very limited processing power to provide sufficient battery life.
[0009] These devices typically couple wirelessly to other devices, such as smartphones, via Bluetooth or similar wireless protocols. However, these connections require encoding, conversion, and over-the-air transmission, introducing additional latency. This latency is often far greater than the desired maximum motion-to-sound delay required to achieve externalization. Motion-to-sound delay here describes the time frame required for a binaural audio system to auditorily detect the acoustics caused by the user's head movement. The exact audible threshold of motion-to-sound delay varies and depends on the listener, the signal being used, and the acoustic properties of the environment. A delay of up to 50 milliseconds has been determined to be an effective threshold and, in most cases, is inaudible to most users.
[0010] To produce convincing virtual sound sources, binaural signals and binaural filters are typically updated at this high rate. Depending on the binaural synthesis method employed, this results in computational complexity that is often too high for mobile and wearable devices. Instead, such devices are typically connected by cables to another computing device that handles these calculations.
[0011] U. Sloma et al. describe a proof-of-concept demonstration in the publication "Proof of Concept of a Binaural Renderer with Increased Plausibility," DAGA 2023 Hamburg, pages 208-211, comparing a realistic loudspeaker setup in a given room with a headphone-based rendering. In particular, room acoustic processing is included and performed at runtime. Specifically, a binaural room impulse response (BRIR) is computed in real time based on a single omnidirectional room impulse response (RIR). This requires capturing a very basic model of the room geometry, along with the locations of the sound sources and microphones. From this, the direction of arrival (DOA) of the direct sound and early reflections is estimated using a simplified image source model. The RIR is processed segmentally and appropriately convolved with a generic HRTF filter. Late reverberation is simulated through noise shaping. The algorithm allows for 6DoF rotations and translations. A spatial decomposition method is also discussed. This method uses a measurement microphone and six electret condenser microphones. The sound field is assumed to consist of a series of individual acoustic events, which can be described using the captured RIRs and captured DOAs. In post-processing, the HRIR of the measured position is calculated using 3DoF rotation and a general HRTF filter.
[0012] S. Werner et al., in the publication "Creation of Auditory Augmented Reality Using a Position-Dynamic Binaural Synthesis System – Technical Components, Psychoacoustic Needs, and Perceptual Evaluation," published in Applied Sciences, vol. 11, 2021, p. 1150, disclose a position-dynamic binaural synthesis system for synthesizing ear signals for a moving listener. The goal is to merge the auditory perception of virtual audio objects with the real listening environment. For each possible listener position in the room, a set of binaural room impulse responses (BRIRs) consistent with the intended auditory environment is required to avoid room divergence effects. The required spatial resolution of the BRIR positions can be estimated using spatial auditory perception thresholds. Specifically, the position-dynamic binaural synthesis system relies on preprocessing of the room geometry, the spatial resolution of the reproduction, a listening position representation, a real-time processing block including acquisition and processing of tracking data and a convolution engine, and a filter creation block including listening position and BRIR synthesis. The results of the BRIR synthesis are binaural filters, which are used by the convolution engine in the real-time processing block for position-dynamic binaural playback. Constant reverberation, acoustic shaping, synthesis methods that adapt to the initial time delay gap (ITDG), sound source directionality, and real-time processing are discussed.
[0013] Publication "Binauralization of Omnidirectional Room Impulse Responses–Algorithm and Technical Evaluation" (C. et al., in Proceedings of the 20th International Conference on Digital Audio Effects (DAFx-17), Edinburgh, UK, September 5-9, 2017, pp. 345-352) disclose a duality of omnidirectional room impulse response algorithms that synthesizes a BRIR dataset for dynamic auralization based on a single measured omnidirectional room impulse response (RIR). Direct sound, early reflections, and diffuse reverberation are extracted from the omnidirectional RIR and spatialized separately. Spatial information is added based on assumptions about the room geometry and typical properties of diffuse reverberation. The early part of the RIR is described by a parametric model. Therefore, modifications to the listener position can be taken into account. The late reverberation part is synthesized using binaural noise that is adapted to the energy decay curve of the measured RIR. The direct sound frame starts at the onset of the sound and ends 10 milliseconds later. The following time periods are assigned to early reflections and the transition to diffuse reverberation. Segments with strong early reflections are determined. Following this procedure, small window segments of the omnidirectional RIR are extracted, describing the early reflections. The incidence direction of the synthetic reflections is based on the spatial reflection pattern adapted to a shoebox-type room with asymmetrically positioned sources and receivers. A fixed lookup table containing the incidence direction is used. In this way, a parametric model of the direct sound and early reflections is created. The amplitude, incidence direction, delay and envelope of each reflection are stored. By comparing each window segment of the RIR with the Convolution is performed to obtain a binaural representation of the early geometric reflection portion. To synthesize the transition directions between given HRIRs, interpolation is performed in the spherical domain. The early part of a single measured omnidirectional RIR contains the direct sound and strong early reflections. For this part, the incident direction is modeled as arriving at the listener from an arbitrarily chosen direction. The late part of the RIR is considered diffuse and is synthesized by convolving binaural noise with a small segment of the omnidirectional RIR. In this way, the characteristics of diffuse reverberation are approximated. The synthesized BRIR can adapt to the shift of the listener, allowing the auditoryization of freely chosen positions in the virtual room.
[0014] It has been found that existing BRIR synthesis algorithms have several drawbacks that make the processing computationally expensive, result in an unnatural sound perception for the listener, present problems in efficiently adapting the system to specific source characteristics, positions, or orientations, or to specific listener positions or orientations, and may even prohibit the system from running in real time. Furthermore, an additional drawback may be the creation of artifacts that reduce the externalization of the sound impression, resulting in an unnatural and unpleasant perception for the listener. Summary of the Invention
[0015] The object of the present invention is therefore to provide an improved audio signal processing concept starting from single-channel acoustic data describing the acoustic environment and generating audio sound generation dependent on two-channel acoustic data of a specific setup consisting of the acoustic environment and one or more sources and a listener.
[0016] This object is achieved by an audio signal processor, an audio signal processing method and a computer program according to claim 1 .
[0017] Aspects of the invention start with single-channel acoustic data describing an acoustic environment and produce an audio sound generation that relies on two-channel acoustic data for a specific setup consisting of the acoustic environment, one or more sources, and a listener.
[0018] Subsequently, specific improvements to the algorithm are described for each of the seven aspects of the present invention. It is emphasized that implementing a single aspect in the current existing system has significantly improved the prior art. However, subsets of the seven aspects can also be combined, or even all seven aspects can be combined with each other to achieve an improved audio signal processor for generating a two-channel audio signal. Therefore, it is emphasized that the seven aspects described below can be used separately from each other or can be combined in any manner, that is, for example, the third and fifth aspects can be combined, or the third to seventh aspects can be combined, or the first to fourth aspects can be combined, etc.
[0019] According to a first aspect of the invention, specific source characteristics, and in particular the directionality information of the sound source, are integrated into a two-channel synthesis for synthesizing two-channel acoustic data from single-channel acoustic data. This integration of the sound source directionality information can be performed in particular in the processing of the direct sound (DS) portion of the single-channel acoustic data describing the acoustic environment. However, the integration of the directionality information that allows the natural reproduction of sound sources with non-omnidirectional directionality characteristics can also be integrated in the processing of the early reflection (ER) portion of the single-channel acoustic data, or the directionality information can even be integrated into both the direct sound processing and the early reflection processing in an efficient manner.
[0020] According to a second aspect of the present invention, a specific processing of the early reflection (ER) portion of single-channel acoustic data is enhanced. In particular, the early reflection portion is segmented into a plurality of segments, each of which includes specific reflections. In particular, a plurality of image source positions representing the sources of the reflected sound are determined, and these image source positions are associated with the segments using a matching operation according to the present invention, which relies on the sound arrival time at the listener position calculated for each image source in the initial measurement. Matching is then performed to associate the sound arrival time of each image source with a specific segment, i.e., with a specific reflection in the segment. In this way, an automatic and high-quality association of image source positions with different early reflections is achieved. By additionally integrating directional information not only for the direct sound, but also for the individual image sources, the specific orientation of the image sources can also be taken into account, resulting in a more natural sound reproduction.
[0021] According to a third aspect of the present invention, the processing of the early reflection portion of single-channel acoustic data describing an acoustic environment is enhanced by calculating two-channel acoustic data for the early reflection portion using not only the specular portion describing different early reflections but also the diffuse portion describing the influence of diffuse reflections in the early reflection portion. It has been found that although the "second part" of the room impulse response shows prominent early reflections, it does not consist solely of these. On the contrary, even this early reflection portion has a significant diffuse portion, and the diffuse portion has an even increasing influence from the beginning of the early reflection portion to the end of the early reflection portion (i.e., near the beginning of the late reverberation portion of the room impulse response). Therefore, by calculating two-channel acoustic data describing an acoustic environment using the diffuse contribution even in the early reflection portion, it is also possible to better produce a natural auditory perception of artificial sound scenes, for example, when feeding two-channel audio data generated by a sound generator through headphones or through loudspeakers, the sound generator using two-channel acoustic data for the early reflection portion relies not only on the specular portion but also on the diffuse portion.
[0022] According to a fourth aspect of the present invention, it relates to an improved calculation of the late reverberation (LR) portion of mono-channel acoustic data, such as a BRIR or BRTF (Binaural Room Transfer Function), which relies on a specific generation of a two-channel late reverberation portion by combining amplitude data derived from the mono-channel acoustic data and preferably a binaural two-channel noise sequence. Thus, the generation of two channels from one channel is accomplished by using the same amplitude but different phase values.
[0023] In particular, the preferred binaural noise sequence consisting of two channels is converted to the spectral domain using a short-time Fourier transform or any other time / frequency domain conversion algorithm. This produces two spectrograms. Furthermore, the late reverberation portion of the single-channel acoustic data, or the combination of the early reverberation portion and the late reverberation portion, is preferably also converted to a spectral representation using the same transformation algorithm. The two-channel acoustic data of the environment is then derived by relying on the same amplitude, which can also be low-pass filtered, for example, against the two phase spectra actually generated. The two resulting spectrograms are then converted to the time domain to obtain the processed late reverberation portion and, preferably, also the diffuse portion of the processed early reflection portion, as previously discussed with respect to the fourth aspect. Therefore, the specific procedure for calculating the diffuse reflection signal can be applied only to the late reverberation portion, or only to the calculation of the diffuse portion of the early reverberation portion, or, as in the preferred embodiment of the present invention, to the calculation of both the early reflection portion and the late reverberation portion. In particular, for the calculation of the combined early reflection and late reverberation components, there is no need to separate these components at all, as the calculation of the binaural diffuse reflection component is performed without knowledge of any separation of the early reflection and late reverberation components, making any such separation between the early reflection and late reverberation components unnecessary for this aspect of the invention. This approach can save certain computing resources. Furthermore, high audio quality is achieved, which is sufficient to make it unnecessary to take into account any variations depending on the listener position or source position or orientation, in particular for the calculation of the late reverberation component, further improving the efficiency of the algorithm. For the calculation of the diffuse reflection component in the early reflection component of the room impulse response, there is also no need to take into account any variations depending on the listener position or source position or orientation.
[0024] According to a fifth aspect of the invention, the problem of how to efficiently and flexibly obtain high-quality single-channel acoustic data, such as a single-channel room impulse response of sufficient quality to obtain high-quality auditoryization, is solved. To this end, the input interface is configured to obtain a raw representation related to the single-channel acoustic data, and the input interface is additionally configured to derive the single-channel acoustic data using the raw representation and additional data stored in or accessible to the audio signal processor. Thus, by relying on initial measurements of natural sounds that a user can generate (such as a user clapping his or her hands or stomping his or her feet on the floor), or even using a speech signal instead of the commonly used sine-swept signal, which is a very unnatural signal and, of course, would not be generated by a listener at all.
[0025] Furthermore, the provision of the initial measurements can be performed by a low-quality microphone, for example included in a laptop or mobile phone etc., and then, based on this raw representation associated with the single-channel acoustic data, the synthesis can be done using a database matching process relying on test and reference fingerprints, or generally the generation of high-quality single-channel acoustic data can be done, or the synthesis can be done with a single or multiple neural networks that rely on the obtained raw representations, such as initial measurements or even only geometric data about the acoustic environment and, possibly, expected source positions and expected or initial listener positions.
[0026] Another process in this regard is to simply record a sound, such as a piece of music played by a speaker or speakers in a specific acoustic environment, and look up the original version of the sound played by the speaker in a usually remote database through some kind of audio fingerprinting process. By using the clear or ideal sound played by the speaker and the sound with the effects of the room acoustics, the room impulse response or room transfer function can be calculated, or generally two-channel acoustic data can be calculated.
[0027] This process effectively solves the problem of having a single-channel room impulse response that is good enough to perform a useful calculation of the head-related impulse response based on a specific listener and source position.
[0028] According to a sixth aspect of the invention, processing tasks can be distributed among several different devices with different power sources. This allows the majority of the tasks to be performed on a wearable device (such as headphones, earbuds, in-ear components, etc.), while a second device is a device with a large battery, such as a mobile phone, smartwatch, tablet, or laptop or stationary computer.
[0029] In particular, it has been found that the most computationally expensive part is the calculation of the late reverberation, and to some extent, also the calculation of the earlier reflections. However, it has been found that the update rate of these routines can be much lower than the update rate of the direct sound calculation. On the other hand, the calculation of the direct sound is computationally cheap because it is only a very short part in time and therefore only requires short filters that can be processed very efficiently.
[0030] Thus, the processing task of calculating the direct sound portion can be easily performed by a low-power device (such as a wearable device), while the more demanding tasks are performed by a separate second device. The resulting transmission delay is not a problem, because the lower update rate is sufficient for the more computationally intensive calculations, namely the calculation of the early reflection portion and, in particular, the calculation of the late reverberation portion, which, depending on the specific acoustic environment, can have a considerable time duration when considering the room impulse response. In particular, in reverberant rooms such as churches, the late reverberation portion can last for several seconds of diffuse reverberation.
[0031] According to a seventh aspect, it has been found that particular attention must be paid to the separation of the room impulse response and the combination of the individual components after calculation. In particular, in order to achieve a high-quality system that allows, on the one hand, the calculation of the different components (DS, ER, LR) by individual processes and the combination of the results without the audio quality problems caused by the separation into individual components and the combination of the separately calculated results, a specific expansion of the corresponding components at the separation time (for example, between the direct sound and the early reflection component or between the early reflection component and the late reverberation component) must be performed to obtain an overlapping range at the corresponding separation time. In addition, in order to avoid any artifacts and, moreover, to allow seamless processing, which must be completed in a relatively short time, at least one of the expanded components is windowed using a window function that takes into account the sample expansion (i.e., overlap). A specific window that has proven to be very useful for the purpose of RIR processing is the Tukey window, which has lobes of width 2n, where n is the specific number of samples used in the expansion of the component.
[0032] Alternatively, the Tukey window is selected so that an overlap of preferably n=16 samples occurs in the overlapping portion. The overlap can also range from 8 samples to 32 samples. The remaining samples retain 100% of their amplitude, i.e., for example, the windowing factor is 1. Thus, a Tukey window with a small number (e.g., 16) of samples as lobes produces a seamless transition feature between the two portions of DS and ER and / or ER and LR, respectively.
[0033] Another issue related to this is the integration of the initial time delay gap (ITDG), which can preferably be performed within the overlap between the direct sound portion and the early reflection portion. Therefore, moving the ITDG back and forth is not a problem, as the overlap will generally be larger than the maximum ITDG movement area. Therefore, even if the overlap is no longer ideal, it is still found to be sufficiently accurate when performing the movement relative to the ITDG.
[0034] This invention describes an unprecedented system for the auditoryization of binaural audio. It uses the acoustic and geometric properties of the environment to synthesize precisely positioned virtual sound sources that appear to be seamlessly embedded in the physical environment surrounding the user. The binaurally rendered sound sources can be perceived at stable positions in space and appear to originate from outside the head, a process known as "externalization." Through this system, virtual sound sources can be perceived as indistinguishable from real sound sources. This is achieved by combining the filters of the sound sources (directional transfer function - DTF), the acoustic effects of the environment (room impulse response - RIR), and the listener's head and body (head-related transfer function - HMTF) to obtain a binaural room impulse response (BRIR). Processing of the binaural signals in response to user motion and the acoustic environment enables externalization of sound and interactivity with the system. Applications of the described system and method include digital audio reproduction and multimedia applications including virtual reality and augmented reality.
[0035] In its most basic embodiment, the system consists of a single device that contains all necessary sensors, components, and acoustic transducers. The system may include the necessary components in a headphone or earbud form factor, with all processing performed directly on the device. In other embodiments, the system operates on a distributed device. The disclosed system consists of three main functional components that work together to create binaural signals in real time. The first component provides omnidirectional RIRs, such as those recorded from an omnidirectional loudspeaker or with an omnidirectional microphone, or preferably with two omnidirectional elements, that have the desired acoustic properties and contain relevant acoustic cues from the environment. This includes, in particular, the frequency-dependent energy distribution of the reverberation over time. In one form, the RIR provider performs qualitative in situ measurements of the RIRs using a loudspeaker and an omnidirectional microphone. In addition, the system can estimate (psycho)acoustic parameters of low-quality RIR measurements or ambient noise and synthesize RIRs from these parameters, or select suitable higher-quality RIRs from a database. The system also incorporates machine learning methods, for example, to support parameter estimation. If necessary, multiple RIRs can be blended to improve transition regions between different acoustic environments (such as coupled rooms).
[0036] The second component is a binaural synthesizer, which takes the RIR and adds binaural cues to it, converting the RIR into a BRIR. The binaural synthesizer also receives room geometry information as input. In an embodiment, the room geometry consists of a shoebox geometry, which approximates the user's real environment by fitting a rectangular room consisting of six surfaces. This provides an estimate of the acoustically reflective surfaces in the environment, particularly the floor, ceiling, and walls close to the listener. While the simplified shoebox room geometry already produces good results, improvement can be achieved with a more accurate room geometry model. Inspired by fundamental research in psychoacoustics, the RIR itself is segmented into multiple segments for processing. The direct sound describes the first sound wave that directly reaches the listener. Here, the corresponding HRTF and DTF, as well as the influence of the distance law of sound propagation, are given. These cues can be directly applied. For reflections in the room represented by the RIR, there is a transition from specular reflection to diffuse reflection. The given RIR is combined with phase information from the binaural noise sequence to generate the diffuse reflection layer of the BRIR. The early reflection segment is divided into blocks that are assigned an estimate of the ratio between specular and diffuse energy. These blocks are convolved with the HRTF and optionally the DTF to obtain the directional portion, which is layered with the diffuse portion at the corresponding index. After combining the three segments, BRIR is complete.
[0037] The binaural synthesizer is connected to a position sensor that can determine the rotation of the user's head and its position relative to a reference frame. This pose information ("pose" represents listener position and listener orientation or source position and source orientation) is provided in real time by a position tracking system. The virtual sound source pose is provided by a preset, optionally moving sound source that changes over time. The binaural synthesizer is connected to a system that generates a measured or synthesized HRTF corresponding to the arrival direction. Similarly, part of the system deploys the directional transfer function (DTF) of the sound source according to the relative position. Like the HRTF, the DTF may be obtained from a measurement or synthesis process. The synthesized BRIR is sent to the auralizer, where it is convolved with the audio signal in real time. For this purpose, a state-of-the-art block-by-block real-time convolution method can be used.
[0038] The resulting binaural audio signal is then played back through headphones, but crosstalk-cancelling speakers can also be used. To maintain the illusion of externalized sound sources, the BRIR needs to be periodically resynthesized using the current position data. In some embodiments, the three segments can be calculated at different rates while maintaining an immersive experience. The described system represents a novelty in the field of binaural synthesis. It makes it possible to experience realistic virtual spatial sound.
[0039] A system for authentic binaural reproduction of audio is described. It allows the auralization of virtual sound sources (sources not present in the user's real listening environment) by finding and incorporating room impulse responses (RIRs) similar to those belonging to real rooms, without requiring direct measurement.
[0040] The RIR is an impulse response that describes the exact effect of the source, the combined filtering effects of the receiver, and the room (environment) on the acoustic signal (for a specific configuration of these components). Therefore, the measured RIR depends on the location and spectral characteristics of the source and receiver, among other influences. Similarly, the binaural room impulse response (BRIR) describes the filtering effects of the source, the local influence of the environment, and the influence of human anatomy (such as the outer ears, head shape, and torso).
[0041] The following describes a solution for deriving BRIRs for arbitrary configurations of the involved components, or even new positions of virtual sound sources, and using them for binaural synthesis and rendering in real-time scenarios. This allows rendered sound sources to be stably perceived as externalized outside the head.
[0042] The solution divides the problem into three parts, which are part of the processing chain:
[0043] 1. Derive new RIRs from available audio recordings that take into account room acoustics;
[0044] 2. Use the RIRs to infer the BRIRs for the specific configuration of listeners, sources, and the room to be auralized;
[0045] 3. Play back the binaural audio to the user of the device.
[0046] All necessary processing steps can be performed on a single device, combining all necessary subsystems. However, in its most basic form, it consists of two systems connected by a network.
[0047] It should also be mentioned that the three parts described above can be used independently of each other, with the corresponding other two parts not being implemented as described but via alternative solutions. Alternatively, for the most preferred results, the three parts can be implemented together. Alternatively, only two of the three parts can be combined, with the remaining parts not being implemented as described but via alternative solutions.
[0048] The first system consists of at least one microphone (or microphone array), at least one processor, and a playback device capable of delivering binaural audio (such as headphones or speakers), as well as a device capable of measuring the position and movement of the user (head) in the environment (such as an IMU or optical tracking system). The second system consists of at least one processor and non-volatile memory.
[0049] This paper describes a system for binaural audio auralization. It uses information about the user's physical environment in the form of impulse responses or reverberant audio signals. It extracts its acoustic properties to synthesize well-externalized, precisely positioned virtual sound sources that appear to originate from the user's physical surroundings.
[0050] The audio rendering system of the present invention allows users to simulate and listen to virtual sound sources that are precisely localizable in space. The simulated sounds appear to originate outside the head, a process known as "externalization." With an appropriate system, binaurally rendered sound sources can be perceived at stable locations in space and appear to have acoustic properties similar to those of real sound sources. This can make them virtually indistinguishable from real sound sources.
[0051] This effect is achieved by precisely controlling the sound reaching the user's eardrum. Typically, two speakers are used, each of which closely reproduces (or sonifies, "makes audible") the sound reaching the listener's ear. Headphones can be used to play the reproduced audio signal directly at the ear. Alternatively, speakers can be used further away from the user's ear to cancel crosstalk.
[0052] Embodiments use psychoacoustic knowledge to reduce the computational complexity of the system and allow for distributed computation of binaural synthesis across devices connected by transmission channels that add greater latency to signal processing than would otherwise be acceptable.
[0053] BRIRs combine multiple filtering effects. They can be split at any point in time, resulting in any number of sub-filters. They can then be reassembled by summing the parts relative to their respective delays, or by convolving the filter with the full signal or parts of it and summing the resulting signals relative to their respective delays. The same basic segmentation and summing process is also valid when part or all of the auditory signal is not processed by convolving the BRIR with the signal but is directly simulated (i.e., by using a delay network-based approach).
[0054] By using psychoacoustic domain knowledge about how different parts of the binaural filter are perceived differently, it is possible to design a rendering system that computes the less important parts of the filter at lower frequencies and distributes these computations across devices.
[0055] In one form, the system consists of a single device capable of synthesizing and auralizing binaural signals in real time, including at least two speakers, each capable of reproducing the sound for one ear, such as any type of ordinary headphones or speakers that cancel crosstalk.
[0056] The system also includes one or more position sensors capable of determining the rotation of the user's head relative to the reference system. (This is often referred to as three degrees of freedom, or 3DoF tracking.) In various embodiments, the system includes one or more position sensors capable of determining the rotation of the user's head relative to the reference system, plus their position relative to the reference system. (This is often referred to as six degrees of freedom, or 6DoF tracking.) The system can process binaural filters or directly simulate auditory signals by employing one or more appropriate binaural synthesis algorithms. It is not dependent on one particular auditory method. Different embodiments of the system may use different binaural synthesis algorithms.
[0057] In this embodiment, the binaural synthesis algorithm used must be able to calculate filters for the direct sound path and the room reverberation separately. The auralization of the direct sound path is usually achieved by convolving the audio signal block by block with a filter that approximates the filtering effect (HRTF) of the user's head, ears, and torso relative to the sound source at a given position and distance.
[0058] The processing of these filters requires encoding the correct changes in inter-aural time difference (ITD) and inter-aural level difference (ILD), as well as changes in sound intensity and other cues. Human listeners are relatively sensitive to small changes in these values, which is why it is necessary to calculate these changes with good spatial and temporal resolution. However, these filters are relatively short, and deriving them usually involves only a small number of processing steps.
[0059] Room reverberation simulates the filtering effects of sound caused by the geometry of the environment, affecting the portion of the sound that travels directly from the source to the user's ears. This includes reflection, refraction, absorption, and resonance. This reverberation filter is expected to be much longer than a short direct sound filter. Many processes, algorithms, and systems are capable of producing adequate binaural reverberation, including image source algorithms, ray tracing, parametric reverberators, and many delay network-based approaches.
[0060] In this embodiment, the system takes advantage of the fact that human listeners are more sensitive to changes in the direct sound filter and less so to changes in the reverberation filter. The signal processor is programmed so that it calculates the direct sound filter much faster than it calculates the reverberation filter. This allows the system to minimize clearly audible jumps in the auditory sound when changing filters and increase the sense of externalization while avoiding full filter updates. Updating these filters, or the portion of the signal encoding the direct sound path, at a rate of about 188 Hz has proven to be a reasonable default for such a system, but lower refresh rates (such as 94 Hz or 50 Hz, or even lower and in excess of 15 Hz) may be feasible in different embodiments of the system. The reverberation filter is calculated at a much lower rate, typically at most one-tenth the direct sound processing rate, depending on the acoustics of the environment and the user.
[0061] The signal processor or another processor is configured as an aggregator. In certain embodiments, the binaural synthesis method employed returns a continuous, block-by-block stream of binaural audio signals. The aggregator simply sums the blocks provided by the direct and reverberant processing paths, acting as a signal aggregator. This requires that the blocks to be summed correspond to the same point in time or contain control data identifying the time frame to which they correspond. Alternatively, the aggregator can be configured to sum the two partial filters with respect to an algorithmically determined time delay. Thus, it reconstructs a complete BRIR filter from the individual processor results and acts as a filter aggregator. This filter can then be used to convolve the audio signal blocks using state-of-the-art real-time (block-by-block) convolution methods. The aggregator always stores a complete BRIR filter in its memory. Therefore, the BRIR can be partially updated at the respective rates of the individual processors processing the partial filters. The resulting signal block contains the combined binaural signals from the direct and reverberant paths. These are then passed to the speaker signal generator for playback through the system's speakers. The speakers can be speakers in a wearable device, speakers with crosstalk cancellation, or any other speakers (e.g., speakers with some kind of sound separation element placed in between). This enables auralization of binaural audio with similar levels of externalization and perceptual quality as a single algorithm, while significantly reducing processing requirements.
[0062] Another part or aspect of the solution receives as input the previously derived RIR and synthesizes a BRIR from it. It uses additional metadata, such as available positional data for the room, listener, and sound sources, for the synthesis process. The system uses a tracking system consisting of one or more sensors (such as an IMU or optical tracking device) to track the user's position relative to the source to be auralized and the real room. It receives metadata about the position of the virtual sound sources and a set of (individual or common) HRTFs. Further metadata, such as the real or virtual room geometry, sound source directionality, sound source boundaries, etc., can optionally be provided. For processing, the system can segment the received RIR into arbitrary time segments that can be processed in parallel using different algorithms and at different intervals. In one embodiment, the RIR is segmented into three parts: direct sound, early reflections, and late reverberation. The direct sound segment is truncated in such a way that it contains the portion of the RIR that contains the sound transmitted directly from the source to the receiver, but does not contain the first reflection reaching the receiver. The late reverberation segment may begin at a point after which no single strong reflection is perceived. These segments are appropriately windowed, for example using overlapping Tukey windows, so that they can be reconstructed later. The relative position of the listener and the source determines the direction of incidence of the direct sound, which is used to convolve the HRTF for each channel with the direct sound segments of the RIR, either directly or by interpolation, by selecting a fitted HRTF from a set.
[0063] For the full length of both reverberation parts, a pseudo-diffuse RIR is calculated by modeling the frequency-dependent energy envelope of the RIR onto binaural white noise (a signal with uniformly distributed energy in all frequency bands but the phase information of the perfect diffuse field of the BRIR), while preserving the phase information of the high-density reflection pattern. This can be done by using a perfect reconstruction filter bank to separate the frequency bands, determining a band-wise low-pass envelope, and multiplying the noise signal with it. Alternatively, the RIR and binaural noise can be converted to the time domain, for example by using an STFT, before applying the amplitude of the RIR to the noise while preserving the phase and converting it back to the time domain. The resulting pseudo-diffuse part (windowed accordingly) is used by the system as the late reverberation of the BRIR.
[0064] The early reflection segments of the RIR are further windowed into sub-windows that may or may not correspond to the positions of single or multiple early reflections. Similar to the direct sound, if each detected sub-segment corresponds to an early reflection, it is assumed to have an incidence direction. This arrival direction is either derived from a room model of appropriate complexity using an algorithm such as the image source algorithm, or selected statistically. Based on this direction, an HRTF is selected or interpolated and convolved with the sub-segment. To overcome the sparsity of this approach, the system mixes a pseudo-diffuse portion with the omnidirectional ("specular") portion to simulate the diffuse and / or nonlinear portion of the RIR for reflections arriving at similar times.
[0065] To do this, a function is used to determine the albedo for each window, linearly interpolating between the diffuse and specular parts of each subsegment. An appropriate function can be formed from the energy ratio of the low-pass average energy in a small window around the signal to the low-pass average energy in a large window around the signal, thereby approximating the ratio of local energy to short-term average energy as a predictor of masking effects.
[0066] The resulting sub-segments are then windowed and reassembled. Depending on the signal and HRTF used, further post-processing such as diffuse field or headphone equalization may be applied. Embodiments of the system can pre-process the RIR using additional knowledge of the room or even metadata to adjust for the room's characteristics. For example, the late reverberation energy decay can be adjusted, or the arrival times of early reflections can be further adjusted by inferring them from a reflection model. The resulting BRIR is convolved with the audio signal using block-wise convolution, resulting in real-time auralization.
[0067] The solution further minimizes room acoustic divergence overall and can be customized to the user using a personalized HRTF, allowing it to be used in auditory augmented reality scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Preferred embodiments are then discussed with reference to the accompanying drawings, in which:
[0069] Figure 1 It is a general diagram indicating the preferred basis of seven aspects;
[0070] Figure 2 A preferred implementation of a two-channel synthesizer is shown, which illustrates the processes of the first to fourth and seventh aspects of the present invention;
[0071] Figure 3a The preferred process of the first aspect and / or the second aspect is shown;
[0072] Figure 3b A table is shown for indicating content that must be updated under certain conditions;
[0073] Figure 4a shows the magnitude representation of the room impulse response / room transfer function in three dimensions;
[0074] Figure 4b It shows that when the emission direction is Figure 4a The room impulse response when the sound is directed in the direction indicated by (i.e., radiated in front of the sound source);
[0075] Figure 4c Shown Figure 4b Directional transfer function of the directional impulse response;
[0076] Figure 5a A preferred embodiment of the first aspect is shown;
[0077] Figure 5b shows another part of a preferred process according to the first aspect;
[0078] Figure 5c shows further processing according to the first aspect;
[0079] Figure 5d Another process according to the first aspect is shown;
[0080] Figure 6a A three-dimensional sphere for determining / selecting a head-related transfer function or a head-related impulse response is shown;
[0081] Figure 6b Shows when the user is in Figure 6a Left HRIR and right HRIR in the front / left position as shown;
[0082] Figure 6c Shown Figure 6b The left HRTF and right HRTF of the corresponding HRIR;
[0083] Figure 7 shows a preferred embodiment of the second aspect of the present invention;
[0084] Figure 8a shows the generation of an image sound source before the first order reflection;
[0085] Figure 8b The process of the second aspect of the present invention is shown;
[0086] Figure 9 A preferred embodiment of the second aspect is shown;
[0087] Figure 10a An embodiment of the third aspect of the present invention is shown;
[0088] Figure 10bAnother embodiment of the third aspect is shown;
[0089] Figure 11a Another preferred embodiment of the third aspect is shown;
[0090] Figure 11b shows a preferred embodiment of a combination of a mirror portion and a diffuse reflection portion according to the third aspect;
[0091] Figure 12a The initial time delay gap is shown;
[0092] Figure 12b The application of the initial time delay gap of the third aspect or the seventh aspect is shown;
[0093] Figure 12c Reference is also made to the initial time delay gap (ITDG) in the embodiments according to the third and seventh aspects;
[0094] Figure 13a A preferred embodiment of the fourth aspect is shown;
[0095] Figure 13b An embodiment of the fourth aspect of the present invention is shown;
[0096] Figure 13c A preferred embodiment of the fourth aspect is shown;
[0097] Figure 13d Another process according to the fourth aspect of the present invention is shown;
[0098] Figure 14a -e shows different embodiments specifically related to the fifth aspect or other aspects;
[0099] Figure 15 shows an embodiment of the hardware required for the first device (on the one hand) and the second device (on the other hand) according to the sixth aspect of the present invention;
[0100] Figure 16 shows a preferred embodiment of the fifth aspect of the present invention;
[0101] Figure 17a A real embodiment of the fifth aspect is shown;
[0102] Figure 17b Another embodiment of the fifth aspect is shown;
[0103] Figure 17c Another embodiment of the fifth aspect is shown;
[0104] Figure 18 Another embodiment of the fifth or seventh aspect is shown;
[0105] Figure 19 shows a schematic representation of an embodiment of the sixth aspect;
[0106] Figure 20 Another embodiment of the sixth aspect of the present invention is shown;
[0107] Figure 21a -b shows different embodiments of audio sound generators;
[0108] Figure 22a -f shows another embodiment of the sixth aspect of the present invention;
[0109] Figure 23a shows an embodiment aspect of the invention in which a sound generator uses a complete two-channel acoustic dataset to generate a two-channel audio signal:
[0110] Figure 23b An alternative embodiment is shown in which the same audio signal is convolved with separate two-channel data segments and the separate regenerated binaural audio signals are combined with each other;
[0111] Figure 24a An embodiment according to a seventh aspect is shown;
[0112] Figure 24b shows further processing according to the seventh aspect;
[0113] Figure 25 Another embodiment of the seventh aspect is shown, wherein ITDG adjustment is integrated;
[0114] Figure 26 A preferred embodiment of the ITDG adjustment according to the seventh aspect or the third aspect of the present invention is shown. DETAILED DESCRIPTION
[0115] Figure 1 An input interface 100 is shown, which can receive several inputs, which will be described later, and provide single-channel acoustic data describing the acoustic environment. The single-channel acoustic data can be a room impulse response or a room transfer function, or any other description of the acoustic environment (such as a room, an open room, or a semi-open room). Depending on the situation, the acoustic environment can also be the environment outside the room. Typically, the acoustic environment will include reflective objects, such as room walls, furniture, etc., or absorbent objects, such as people in the room, curtains in the room, or any other "acoustic objects".
[0116] The audio signal processor further includes a two-channel synthesizer for synthesizing two-channel acoustic data from the single-channel acoustic data using the listener position or orientation, such as Figure 1The result of the two-channel synthesizer 200 is two-channel acoustic data, such as a binaural room impulse response or a binaural room transfer function, or any other two-channel impulse response or transfer function as appropriate. Other descriptions of the impulse response or transfer function can also be applied as acoustic data, such as specific parameterizations, etc.
[0117] The two-channel acoustic data is input to the sound generator, which is used to generate the sound from the audio signal (usually Figure 1 The mono signal shown in Figure 1 The dual-channel synthesizer 200 generates a dual-channel audio signal from the dual-channel acoustic data received. In this specification, the input interface may also be referred to as an RIR provider. In addition, in this specification, the dual-channel synthesizer is also referred to as a binaural synthesizer, and the sound generator is also referred to as an auralizer. However, both descriptions mean the same thing, that is, the RIR provider is generally an input interface, the binaural synthesizer is a general-purpose dual-channel synthesizer, and the sound generator is a general-purpose auralizer.
[0118] The two-channel synthesizer 200 is configured to separate single-channel acoustic data into at least two parts consisting of a direct sound part, an early reflection part, and a late reverberation part, and the two-channel synthesizer 100 is configured to process the at least two parts separately to generate two-channel acoustic data of each part.
[0119] This is Figure 2 . In block 210, the single channel acoustic data is separated into at least two parts. Block 220 shows the direct sound processing. Block 230 shows the early reflection processing, and block 240 shows the late reverberation processing. As shown in 250, all three two-channel acoustic data of each part are combined by aggregating or combining the two-channel acoustic data. In addition, it should be noted that block 250 covers two alternatives that can be generally performed. The first alternative is to aggregate the various parts of the BRIR into a complete BRIR, and then apply the complete BRIR to the audio signal by convolution, as shown in FIG. Figure 3a As shown. The convolution is performed by Figure 1 The sound generator 300 is executed, and the sound generator 300 also receives the audio signal.
[0120] Figure 2 The alternative embodiment shown in block 250 is also shown in Figure 23b Here, no aggregation of the individual parts of the BRIR occurs. Instead, each part is convolved with the audio signal separately, resulting in three binaural audio data streams, and the binaural audio data is then calculated by combining the three individual streams binaural audio 1, binaural audio 2, and binaural audio 3. Thus, Figure 1As shown, the processing of the audio signal and the aggregation of the audio signal are performed by the sound generator 300 .
[0121] Furthermore, it should be noted that, according to the various aspects of the present invention, it is not always necessary to process three parts. Instead, for the first aspect, it is sufficient to separate the mono-channel acoustic data into only two parts: a direct sound part, such as an RIR, and a residual part. For the purposes of the second aspect of the present invention, image sources are associated with the respective segments, and separation into three parts is necessary, as the early reflection part is positioned between the direct sound part and the late reverberation part. Separation into three parts is also useful for the purposes of the third aspect, which deals with the specific merging of the specular and diffuse parts. However, for the purposes of the fourth aspect, which involves specific binaural noise-based calculations, separation into two parts is only necessary: the first part includes the direct sound and early reflections, and the second part includes the late reverberation. For the purposes of the fifth aspect, no division is required at all, and any auralization processing requiring a description of the acoustic environment can be performed, as the fifth aspect deals with the provision of a room impulse response, not its further processing. However, the fifth aspect can of course be combined with all other aspects, and thus, in certain embodiments, the fifth aspect may also utilize the separation into two or three parts, as indicated in block 210. According to the sixth aspect, since the direct sound processing is preferably performed on the wearable device and the remaining processing is performed on the second device, it is necessary to separate into at least two parts. When three devices are used, it is necessary to separate into three parts. The seventh aspect relates to the separation of RIR and the combination of the separated parts. Separation into two parts is sufficient, and when separated into three parts, the seventh aspect is also preferably applicable. When there is no separation between the early reflection part and the late reverberation part, the same is true for introducing an initial time delay gap, which can also be achieved according to the present invention.
[0122] exist Figure 2 In the preferred embodiment shown, direct sound processing relies on source directivity, initial source or receiver data from initial measurements, current listener data, and / or current source data. In this context, it should be noted that in the following text, current listener data is referred to as listener position, listener orientation, or both, also referred to as listener "pose." The same applies to source data. Source data can be source position or source rotation, or both. In particular, according to the first or second aspects of the present invention, source rotation can be advantageously taken into account even for non-omnidirectional sources using source directivity information.
[0123] The early reflection processing in block 230 depends on the listener position and / or orientation, as well as the geometric data of the acoustic environment, and usually also on initial data, such as the association of image sound sources with early reflections. In addition, source directivity can also be considered in the early reflection processing in block 230.
[0124] Late reverberation processing 240 depends on Figure 2 The two arrows in the figure show the dual-channel noise data to show the late reverberation part (single channel part) Figure 2 The conversion of the two output channels is shown in the lower portion of block 240 .
[0125] The preferred embodiment provides a binaural synthesis system that uses RIR and very simplified room geometry data as input, rather than complex geometry. It aims to synthesize virtual sound sources that appear to originate from arbitrary locations around the user. Like real sound sources, they are stably anchored and react to the listener's movements. Applications of the described system and method include digital audio reproduction and multimedia applications, including virtual reality and augmented reality.
[0126] The processing of binaural signals, reacting to user movements and the acoustic environment, enables the externalization of sound and interactivity with the system. The described device contains another system that sends the audio content and all required meta-information (described later) to the described system.
[0127] In its most basic embodiment, the system consists of a single device that contains all necessary sensors, components, and acoustic transducers. This device has two speakers, one for each ear, to reproduce binaural signals to the user. The system may include the necessary components in a headphone or earbud form factor, with all processing performed directly on the device.
[0128] The disclosed system consists of three main functional components that work together to create binaural signals in real time. Different embodiments of the system may include different implementations of these components, but their purpose remains the same.
[0129] The first component is the RIR Provider. Its purpose is to provide an omnidirectional RIR with the desired acoustic properties and incorporate relevant acoustic cues from the real or virtual environment, or a modified version thereof. This specifically includes the frequency-dependent energy distribution of the reverberation over time. The exact characteristics and cues encoded by the RIRs depend largely on the system's implementation. This also applies to the component's operating mechanism. In one form, the RIR Provider is connected to a single microphone and speaker. It contains non-volatile memory that stores one or more RIRs.
[0130] The RIR can be recorded using any state-of-the-art measurement method that provides a good measure of the room's acoustics. For example, this can be achieved by playing an exponential sine sweep or minimum-length sequence through a loudspeaker and recording the reverberant audio using an omnidirectional microphone at a critical distance from the sound source. The RIR can be calculated from the reverberant recording and the input signal by deconvolution. The recorded RIR is stored in memory.
[0131] The second component is a binaural synthesizer, which takes the RIR from the RIR provider as input. At this point, the recorded room impulse response contains important monophonic clues about the room acoustics. Binaural information is necessary for the user's spatial hearing and externalized perception and needs to be added to the RIR to convert the RIR into a BRIR. The binaural synthesizer also receives room geometry information as input. In an embodiment, the room geometry information consists of a shoebox geometry that approximates the user's real environment by fitting a rectangular room consisting of six surfaces into the real environment. The width, depth, and height of this shoebox room are provided by the user of the system, and the surfaces should coincide with the main acoustic reflective surfaces in the real environment, specifically the floor, ceiling, and walls close to the listener. The second component is also connected to one or more position sensors that can determine the user's head rotation relative to the reference system. (This is often referred to as three degrees of freedom or 3DoF tracking.)
[0132] In various embodiments, this is included in one or more position sensors that are capable of determining the rotation of the user's head relative to a reference system, plus their position relative to the reference system. (This is often referred to as six degrees of freedom, or 6DoF, tracking.) The position and rotation of the user in the reference coordinate system are provided in real time by the position tracking system. The position and rotation of the virtual sound source can be provided by a preset configuration. In some embodiments, the position of the sound source can be periodically changed by an external system, representing a moving sound source. In some embodiments, an offset to the user's position and rotation can be periodically added, for example to allow a simulated user representation to move in the virtual world.
[0133] The binaural synthesizer is also connected to a system that is capable of approximating HRTFs corresponding to a given relative position between a sound source and a user. In one form, such a system may contain a dataset of measured or synthesized HRTFs and their corresponding relative position vectors between the source and receiver (called the direction of arrival or DOA). Given the DOA as input, it then selects a single HRTF whose corresponding DOA best matches the input DOA, i.e., by maximizing the scalar product of the two unit vectors. In other embodiments, given a relative position, a suitable HRTF can be synthesized.
[0134] Likewise, a binaural synthesizer is connected to a system that can approximate the directional transfer function (DTF) of a sound source, which is a filtering effect that depends on the relative position of the sound source. Such a system can work by selecting the best match from a database or synthesizing the DTF, such as an HRTF system.
[0135] Then, Figure 3a The subject matter of the invention is shown according to a first aspect. Figure 3aA coordinate system is shown, wherein the origin 420 is located at the listener position, which is assumed to be the origin of the listener's head when the listener's head is assumed to be a sphere. In addition, a source 410 is shown, whose main emission direction 430 is away from the listener. It is assumed that the source has a non-omnidirectional transmission characteristic, such as Figure 4a As shown, Figure 4a The magnitude of the directional information in three dimensions is shown. As shown in 431, the sound source (at Figure 4a The amplitude behind the speaker (illustrated in the figure) is smaller than the amplitude 440 in front of the speaker.
[0136] In addition, Figure 4a In the example, the listener is placed in front of the loudspeaker, resulting in Figure 4b The directional impulse response is shown in Figure 2. The directional transfer function, that is, the directional impulse response when converted to the spectral domain, is shown in Figure 2. Figure 4c Therefore, it can be seen that Figure 4a 、 Figure 4b 、 Figure 4c The sound source in the image has a strong non-omnidirectional directivity and, in addition, there is a certain room impulse even in the front, which shows a significant nonlinear frequency response when converted to the spectral domain. It should be noted that Figure 4c The phase is not shown in , but the directional impulse response produces a complex directional transfer function.
[0137] According to the second aspect of the present invention, Figure 1 The illustrated two-channel synthesizer is configured to determine the directionality information of the sound source for a specific listener position and source position and / or orientation of the sound source. In addition, the two-channel synthesizer is configured to use the directionality information when calculating the two-channel acoustic data of the direct sound portion, such as Figure 2 The corresponding inputs to block 220 are shown.
[0138] In particular, the two-channel synthesizer 200 is configured to determine two head-related data channels from the source position or orientation and the listener position or orientation in addition to the directionality information, and use the two head-related data channels and the directionality information in the calculation of the two-channel acoustic data of the direct sound part. In particular, referring to Figure 3a , the DOA vector 421 is shown as the difference between the listener vector 422 and the sound source vector 423.
[0139] also, Figure 3aAlso shown is an emission direction vector 424, which is opposite to the direction of arrival direction vector 421. Typically, the directional information of a sound source is given relative to a main emission direction 430. Therefore, in order to select the correct directional information from among several data sets of directional information for a sphere around the sound source 410 (the directional information being related to the main emission direction or to an arbitrary reference point, typically different from the origin of a world coordinate system, in which the position vectors of the source and the listener are given), the rotation of the sound source 410 must be taken into account. Thus, since in the example the main emission direction 430 is rotated by about 90° relative to the DOA vector 421 or the DoE vector 424, the two-channel synthesizer 200 will determine the rotation relative to Figure 3a The 90° azimuth angle of the main emission direction 430 gives the directionality information.
[0140] Typically, the directional information is given as a DIR for each azimuth and elevation angle, and in a preferred embodiment of the present invention, there are actually approximately 540 DIR data sets for a sphere, which are measured in corresponding ten degree differences or ten degree increments in the azimuth and elevation directions.
[0141] Alternatively, this information may also be provided via a directivity transfer function (with magnitude and phase or real and imaginary parts) for a specific azimuth / elevation angle.
[0142] Alternatively, by providing a complete data set, the two-channel synthesizer can select the identified directional information for the correct orientation of the source relative to the listener. Certain parameters for certain categories of sources can also be used to synthesize or actually calculate the directional information. Furthermore, the directional information can be provided at a lower resolution than the exemplarily shown ten-degree resolution in two directions. Interpolation can also be performed on the selected room impulse response based on the current situation to be auralized. In another embodiment, when the room impulse response of a sound source is not available, these sources can be synthesized or measured and stored in a memory accessible to the two-channel synthesizer.
[0143] Figure 3b A table is shown which indicates what has to be updated under the specific conditions of listener movement and source movement. Of course, when both listener and source are stationary no changes have to be performed with respect to the previous situation. When the source is stationary and the listener only rotates then the room impulse response or the directionality information will not change, the rotation of the listener only affects the process by which a new head-related impulse response or head-related transfer function has to be selected which takes into account the rotational position of the ear relative to the source. Another interesting point is that when only the source rotates and the listener remains stationary then a new room impulse response has to be calculated but the head-related impulse response remains unchanged. Figure 3bIn all other cases shown, the room impulse response and the head-related impulse response are changed for the specific cases indicated in the table titled "What to Update?"
[0144] Figure 5a A preferred implementation of a particular embodiment is shown. Typically, the direct sound portion of the room impulse response is only used for the energy calculation in block 221, but not used later. Instead, the initial portion of the room impulse response provided by the input interface is replaced by the corresponding directional impulse response, as described with respect to Figure 3a In a preferred embodiment, the energy is related to the frontal DOE. The measurements for determining the three-dimensional directional information are performed in such a way that the microphone is located on the acoustic axis.
[0145] Usually, as about Figure 4b As discussed, the data set of the room impulse response will differ from the energy of the first part of the room impulse response provided by the input interface 100. Therefore, depending on the source position or orientation and the listener position orientation, the original directivity information is determined, for example from a database, using a specific angle, such as Figure 5a As shown at 222. Alternatively, a room impulse response may be synthesized based on a specific angle derived from the source positions and orientations and the listener position.
[0146] In block 223, the energy of the original directional information is calculated. In block 224, a scaling factor is calculated by dividing the direct sound portion energy by the total directional energy. In the case where the distance between source 410 and listener 400 is the same, the new directional information is scaled using the scaling factors from blocks 224 and 226.
[0147] Alternatively, when the distance between listener 400 and sound source 410 changes due to movement of source 410 or listener 400, another scaling factor is calculated in block 225, or the scaling factor of block 224 is adapted. Specifically, when the source moves away from the listener relative to the initial measurement situation, i.e., when the original room impulse response provided by the input interface was measured, the loudness of source 410 must be reduced. Consequently, another scaling factor is reduced. However, when movement of the source or listener results in a smaller initial distance between the source and the listener, the scaling factor must be increased. To increase or decrease the scaling factor, the distance law of sound is applied. The scaling factor for distance correction is implemented using, for example, the distance law of sound or a similar process. The maximum amplification factor is limited to prevent the direct sound from becoming too loud as the listener position moves closer to the sound source.
[0148] In block 227, if Figure 3aAs shown, the direction of arrival of the direct sound is determined. Then, based on the DOA, the correct HRIR or HRTF is selected, as shown in block 228. In block 229, the single-channel directional information, scaled by a scaling factor that may be modified due to distance changes, is convolved with the HRIR. Specifically, the HRIR has two channels, and the directional information has a single channel. Therefore, the single channel is convolved with the left channel of the HRIR to obtain the first channel of the result of block 229, and the DIR is convolved with the right channel of the HRIR to obtain the right channel of the result of block 229, which is the two-channel acoustic data of the direct sound portion.
[0149] Therefore, the two-channel synthesizer 200 is configured to determine the emission direction of the source position vector of the source and the listener position vector of the listener as well as the rotation of the source and to derive directional information from, for example, a database of directional information sets, wherein the directional information sets are associated with specific angles that are typically related to a main emission direction or an emission direction of a specific source.
[0150] In contrast, the direction of arrival of the listener's position or orientation is calculated from the source position vector of the sound source and the listener's listener position vector as well as the listener's rotation.
[0151] Figure 5c A further preferred implementation is shown of how to convolve the head-related impulse response shown in block 229 with the directional impulse response. To this end, block 261 indicates that both the directional impulse response and the two-channel HRIR are padded with zeros and then transformed into the spectral domain to obtain three spectra, where the first spectrum is the directional transfer function, the second spectrum is the left HRTF, and the third spectrum is the right HRTF.
[0152] Then, as shown in block 263, the DFT spectrum and HRTF L Multiply the spectrum and combine the DFT spectrum with the HRTF R The output of block 263 is two spectra transformed into the time domain. Phase delays introduced, for example, by convolution (i.e., transformation, multiplication, and inverse transformation) are then removed in block 265, and both channels are truncated to their original lengths before being padded in block 261, and finally, windowed, for example, using a Tukey window, in block 267.
[0153] The following process can be performed to obtain a two-channel audio signal for the direct sound portion. The RIRs are first pre-processed to achieve consistent alignment between the different inputs to the system. This can later be mixed between the different input RIRs. The alignment is achieved by detecting the direct sound using a suitable state-of-the-art algorithm. If it is guaranteed that the input direct sound always coincides with the highest peak, then finding the direct sound can rely on maximum peak detection, for example. In more complex scenarios, a more robust, state-of-the-art direct sound detection can be chosen.
[0154] The first sample of the impulse response is then cut or extended by a zero-valued sample so that the detected direct sound sample index coincides with a predefined sample index offset from the start of the impulse response. The binaural synthesis method assumes that the RIR can be divided into three separate filters, which can be processed separately, namely the direct sound (DS), early reflections (ER) and late reverberation (LR). The input RIR is then further preprocessed by separating it into these three separate partial filters. The transition between DS and ER can be selected in such a way that it maximizes the distance between the detected DS peak and the first reflection. The transition between ER and LR is selected in such a way that it coincides with the perceptual mixing time of a given acoustic environment. It can be calculated or estimated by a state-of-the-art algorithm. In some embodiments, the transition between ER and LR can be earlier than the perceptual mixing time, thereby reducing computational complexity.
[0155] Then, at their corresponding transition times, the three segment intervals are extended by n samples so that they overlap by 2n samples. An appropriate window function is chosen, which allows for a near-perfect reconstruction of the filter segments later. For example, a Tukey window function can be chosen with a lobe width of 2n samples.
[0156] Binaural synthesis methods assume that the room acoustics can be largely separated into so-called specular components, assuming that strong geometric reflections behave like light, and diffuse components. The specular component can be derived from models such as the Image Source Model (ISM) or simulation methods based on ray casting. The diffuse component is used under the assumption that part of the signal can be approximated as a roughly uniformly distributed diffuse sound field with a high reflection density. It can preserve the time and phase relationships of the diffuse field while modeling the energy distribution of the reverberant part of the RIR.
[0157] The DS segment contains the combined filtering effects of the sound source and the user's outer ears and body. Both filters depend on the relative position of the virtual sound source and the sink (the listener's ear). Both the position and rotation (of the source and sink) are provided as input to the binaural synthesizer.
[0158] Given the relative positions, appropriate HRTF and DTF filters are selected from the corresponding subsystems. These filters are padded to double their length and then convolved with each other by multiplying them and converting back to the time domain using an inverse Fast Fourier Transform (IFFT). Depending on the filters used, the introduced phase delay can be eliminated by shifting the filters in time before truncation and windowing so that the length of the binaural direct sound filters is the same as the original filter length and the direct sound center index corresponds to the same sample index as the original RIR direct sound.
[0159] Figure 5d A preferred embodiment is shown in which the transmission direction is determined in block 268. Assuming that the database is organized according to certain transmission directions, a match is performed with the test transmission directions of block 268 in block 269, and in block 270 the directional information of the best matching DoE is selected.
[0160] In block 271 , an alternative is shown. Instead of finding the best matching DoE and selecting directional information from the database based on that DoE, two or more directional information with the closest DoE entries are selected and interpolation is performed, as shown in block 271 .
[0161] As another alternative shown in block 274, the directionality information can also be synthesized using a model or neural network and a model based on the test DoE determined by block 268. Figure 3a , indicating that although the DoE points in the opposite direction of the DOA, the DoE Figure 3a The origin of the coordinate system 420 in is not related to the main emission direction 430. Therefore, the DoE reflects the situation where the rotation of the source is applied to the vector DOA and the direction is reversed. Of course, other alternatives with other relationships to other coordinate systems can be implemented.
[0162] As described in block 260, padding is preferably performed using the directional impulse response and the first and second HRIRs to double the length to obtain a padded function, and the padded function is combined by convolution in the time domain or multiplication in the frequency domain. In addition, the phase adjustment indicated in block 265 ensures that the correct time delay from the zero sample index to the index where the first direct sound portion is usually located is maintained, so that the subsequent construction of the complete BRIR always depends on the defined situation.
[0163] Figure 6a Spheres are shown to illustrate the concept of HRTF or HRIR. In particular, Figure 6a The diagram in FIG shows that the user is located in front of / left of the source. The corresponding left and right HRIR functions are as follows Figure 6bAs shown in Figure 2, it is obvious that the left HRIR is significantly stronger than the right HRIR, and the contribution to the left HRIR occurs before the contribution of the right HRIR. Figure 6a Sound at the position shown reaches the left ear before the right ear, and the amplitude of the sound reaching the right ear is attenuated by the head.
[0164] The corresponding frequency domain response is Figure 6c As shown, it indicates that at frequencies below 1 kHz, the main effect is the amplitude difference, while at frequencies above 1 kHz, especially higher frequencies, the right HRIR shows an obvious notch filtering effect.
[0165] Then, refer to Figure 7 The following figures describe the second aspect of the present invention. According to the second aspect, the two-channel synthesizer, in particular the early reflection processing block 230 is configured to split the early reflection part into a plurality of segments, as shown in block 231. Exemplarily, Figure 10b Only four segments 294 are shown, but the segmentation can be performed up to fifty segments or more. Of course, fewer segments can also be used. In an embodiment, there is a block with 256 samples and 128 samples overlapping. The number of segments is obtained by the length of the early reflection part (direct sound until the mixing time) with approximately 7700 samples. This number is divided by the pre-value of 128 per segment, resulting in approximately 60 segments. However, this number can vary depending on the length of the early reflection part, the pre-value and potential other parameters used.
[0166] Furthermore, as indicated by block 232, a plurality of image source positions are preferably determined using a geometric model of the room (e.g., a shoebox model). The image source positions represent the source positions of the reflected sounds. Furthermore, a matching operation is used to associate the image source positions with the segments. In the matching operation, as indicated by block 233, the arrival time of the sound from each image source at the listener position is calculated. Preferably, the initial listener position is used for this calculation, thereby inputting the initial listener position, i.e., the listener position when the RIR is provided by the input interface. Then, as indicated by block 234, the image source positions are associated with the corresponding segments that best match the arrival time of the particular image source.
[0167] Therefore, the arrival time of the sound from each image source position to the initial listener position is compared to a time index within a segment. Typically, segments have a certain width, and therefore, for each segment, the time index in the middle of the segment is compared to the arrival time. When the arrival time of an image source position is equal to the time index associated with the segment, such as the time index in the middle of the segment, then the image source position is associated with the segment for further calculations, such as calculating the direction of arrival for that segment. Typically, image source positions are calculated for a room model up to a certain order. Figure 8aIn , some first-order image source positions are indicated as first reflections. In particular, Figure 8a A listener 400 is shown at an initial listener position and a source 410 is shown at an initial source position. The construction of four image sources for (partial) first order reflections results in image source positions 1, 2, 3, 4 of image sources 431 to 434. Note that the two-dimensional Figure 8a Floor and room reflections, which are also first-order reflections, are not shown. Second-order reflections can also be constructed and refer to the physical effect of a reflection reaching the listener's head, propagating, reflecting off a second wall, and then reaching the listener again.
[0168] Therefore, depending on the complexity of the geometric model, a certain number of image source positions are determined and associated with corresponding segments. For example, when fifty segments are used to segment the early reflection portion, it is sufficient to determine the image source positions until an order of fifty sources is reached. However, this can be very complex, and to save computational resources, a preferred way of doing this is to only calculate image source positions up to a certain order that produces fewer than 50 image source positions. As shown in block 235, the remaining image source positions can be selected in a random manner. Therefore, if a segment is found to produce an image source that does not match the one associated with that segment, a random position is associated with that segment, or in further computations, a random arrival direction is associated with it, and thus, a randomly selected HRIR is used to process that segment.
[0169] The result of this process is Figure 7 This is shown in Table 236 at the bottom, which shows that the first three segments are associated with source positions 2, 1, and 4, respectively, and when counting segments from the direct sound / early reflection boundary to the early reflection / reverberation boundary, there are usually one or more segments at the end of the segment that do not have a discrete image source position, but are associated with a random source position when processing the segment, or receive a random HRIR.
[0170] As outlined, the dual-channel segmentation is configured to determine multiple image source positions using initially measured initial source positions and initial sink positions and geometric data of the acoustic environment. In particular, Figure 8a As shown, the image source method is preferred.
[0171] In a preferred embodiment of the present invention, the dual-channel synthesizer 200 is configured to detect significant reflections and, from these detected significant reflections, construct overlapping segments, as shown in block 280. For the purpose of detecting significant reflections, the processes in blocks 281-283 are performed. In block 281, the average energy per sample of a small window sliding over the early reflection portion is calculated. In block 282, the average energy per sample of a larger window sliding over the early reflection portion is again calculated. In block 283, the two average energies are compared sample by sample, and it is determined whether the average energy per sample in the small window is greater than the average energy per sample in the large window, for example, a third specified threshold. This process produces a segmentation of the early reflection portion.
[0172] In block 284, the determination of the direction of arrival information for each segment is performed. Preferably, the above steps with respect to the first aspect, in particular Figure 3a This allows taking into account a specific orientation of an image source relative to a listener, in particular, for example, image source IS1 shown at 431 facing away from the listener 400 .
[0173] In this embodiment, the two-channel synthesizer is configured to determine directional information of the image sound source with respect to the listener position and the image source position or orientation of the image sound source, and use the directional information in the calculation 220 of the two-channel acoustic data of the early reflection sound portion. Preferably, the directional information of each image source is derived from the same set of directional information determined for the direct sound portion, or wherein the orientation of the image sound source is determined by an image source model, and in a specific embodiment, the directional information is determined and used for a predetermined subset of the segments in the early reflection portion, the subset comprising less than ten segments, preferably only two segments. The remaining segments can then be calculated without any directional information of the image source. As Figure 5a As shown, other procedures for calculating directivity may be performed, wherein for segments for which directivity information is taken into account, the actual RIR segments are replaced by directivity information weighted by the energy scaling factor determined by block 224, but using the energy of the corresponding reflection segments. For simplicity, distance corrections such as in block 225 are preferably not applied, but may still be done when the listener is close to the image source or far away from the corresponding image source responsible for the considered reflections.
[0174] In block 285, each determined segment is padded to a certain length, and in particular, to the existing length of the HRIR database, and in block 286, each segment is convolved with the HRIR associated with the corresponding DOA of the segment, as indicated by the two connecting lines between 284 and 286. This process will produce two-channel acoustic data for the specular part of the segment. In the case of processing only the specular part of the early reflection part of the room impulse response, the result of block 286 can be used for further processing. However, when the second aspect is combined with the third aspect, the diffuse part of the early reflection part is also processed. This will be discussed later. Figure 10a Discuss this.
[0175] It is assumed that an ER segment consists of a specular component and a diffuse component. In general, the first part of the ER is expected to be mostly specular because it contains strong first-order reflections. The later part of the ER segment is expected to contain more coincident reflections, making it more diffuse. ER synthesis first further divides the ER part of the RIR into smaller segments. In some embodiments, this is achieved by detecting perceptually significant reflections and selecting a window around them with at least a head-related impulse response (HRIR) sample count.
[0176] Such windows containing reflections can be detected by heuristic methods, such as comparing the energy average of each sample in the window to the sample count n, and comparing the energy average of each sample in the larger window to the sample count m surrounding the first window. For windows of size m = 2n, a common heuristic method might be that a reflection is considered significant if its average energy per sample is 6 dB higher than the average energy of the surrounding windows. These windows are then assumed to contain significant reflections.
[0177] In some embodiments, this approach can be generalized by assuming a continuous, regular grid of reflection windows, each with the same sample count and overlap. This effectively quantifies the assumed reflection times incident on the grid. Each detected reflection window is assumed to consist partially of specular and diffuse reflection parts, while the remaining part is assumed to be completely diffuse. Each reflection window is assigned a diffuse reflection coefficient that approximates how diffuse the reflection is. A heuristic method or formula can be used to determine the exact diffuse reflection coefficient. Different embodiments of the system may use different methods to determine the coefficient. One possible heuristic method may be based on the heuristic method previously used to find significant reflections, namely, the total energy E of a given small window s (in dB) and the total energy E of the large window l (unit is dB), the diffuse reflection coefficient α can be calculated as α=(E s / E l +6dB) / 12dB.
[0178] A reflectance coefficient α greater than or equal to 1 means that the reflection is completely specular, while α less than or equal to 0 means that the reflection is completely diffuse. Therefore, the value of α is restricted to the range [0,1]. Similar to the DS part, the fully directional specular part of the reflection window needs to be convolved with an HRTF. For this purpose, it may use an HRTF from the same HRTF provider as the DS processing step.
[0179] The required DOA for acquiring the HRTF is calculated using the image source method. Image source positions are calculated based on the provided geometric room information. The best candidate is then selected by comparing the sound incidence time of each image source at the receiver and comparing it to the incidence time of the reflection window. The best matching image source is selected, and the normalized vector between it and the receiver is assumed to be the DOA for specular reflection.
[0180] In other embodiments, the DOA can be determined by other means, such as statistical distribution heuristics or by analyzing the arrival times on a microphone array, so-called spatial decomposition methods. The binaural specular portion is then calculated by convolving the windowed reflection segments with the HRTF, for example by padding the segments appropriately and multiplying them with the HRTF in the frequency domain, then converting the result back and removing the introduced phase delay if necessary. This generates the specular window w s .
[0181] The binaural diffuse part is obtained by selecting the same subwindow, but this time it is from the synthesized diffuse filter. This diffuse fragment is then multiplied by a Hann window of size n to obtain w d .
[0182] Then, given the diffuse reflection coefficient α, the diffuse reflection window w d and mirrored window w s Linear combination, thus obtaining the window w according to the following formula bin
[0183] w bin =α*w s +(1-α)*w d (Equation 1)
[0184] Therefore, about Figure 8b, a preferred implementation is that for initialization, the required signals are loaded. The first signal is the room impulse response or single channel acoustic data provided by the input interface. The second signal is the binaural noise used later according to the third or fourth aspect. In addition, the HRTF data set is loaded and, if necessary, the average HRTF amplitude response is also loaded. In addition, as shown in relation to the seventh aspect, compensation filters from the microphone and headphones and the directional transfer function of the loudspeaker as discussed in relation to the first and second aspects can be applied. In addition, before selecting certain HRTFs, headphone compensation can be applied to the HRTFs so that in response to the DOA, an already compensated HRTF can be selected. The position and rotation of the recording constellation are then saved as the initial listener position or orientation or the initial receiver position or orientation and the initial source position or orientation. Then, proceed as shown in relation to the seventh aspect. Figure 8a The image source model is calculated as shown, and in order to ensure sufficient image sources, the order is selected based on the assumed mixing time between the early reflection component and the late reverberation component in the room impulse response, for example, 160 milliseconds. However, to simplify the process, a smaller number of image source positions can be used, and segments that do not receive an associated image source position during the matching process are typically associated with a random position or random data. When a matching image source position is not found for a certain reflection during the matching process, the process of randomly associating a certain HRTF for the segment is also a preferred method.
[0185] Then, refer to Figure 10a The third aspect of the present invention is described. In particular, the two-channel synthesizer is configured to calculate the specular portion of segment n, or, in general, the specular portion of the early reflection portion, as discussed with respect to the second aspect, as shown in block 237. In addition, the two-channel acoustic data of the early reflection portion is also calculated using the diffuse reflection portion, as shown in block 238, which describes the diffuse reflection effect in the early reflection portion. Figure 2 As shown in block 210 of FIG. , blocks 237 and 238 receive a single-channel early reflection portion of segment n. The two blocks 237 and 238 output two channels of binaural data, and the two channels are correspondingly combined in block 239 to obtain a first channel and a second channel of segment n, and the two channel data of the early reflection portion not only represent the specular effects of different early reflections in the prior art process, but also take into account the diffuse reflection portion that significantly contributes to the natural and pleasant sound impression of the listener.
[0186] In particular, the two-channel synthesizer 200 is configured to calculate the diffuse reflection portion using a combination of the early reflection portion of the single-channel acoustic data and a two-channel noise sequence input into block 238. Preferably, the two-channel noise sequence is a binaural noise sequence measured when a loudspeaker emits a specific noise signal at a specific position relative to the artificial head, and the complete HRTF is detected by two microphones located in the artificial head. Such binaural noise can be actually measured, or alternatively, it can be synthesized, or if this is not feasible for some reason, even two different noise sequences can be used for binauralization of the late reverberation portion of the room impulse response.
[0187] In such Figure 11b In the preferred embodiment shown, a weighted addition of the specular and diffuse portions is performed, wherein the weighting factors are determined, for example, as indicated in blocks 290a, 290b, and the actual addition of the weighted contributions occurs in block 290c. Figure 10b , Figure 10b An exemplary room impulse response, which may be measured or synthesized, is shown at block 291. Note, however, that the room impulse response 291 is not a true room impulse response, as early reflections have been enhanced relative to the direct sound for purposes of illustration. Specifically, the impulse response includes a specular portion 291 and a diffuse portion 293. In 292, the inset shows specular reflections and is therefore not a true segment from 291. The same is true for block 293, which shows a diffuse portion that is not directly derived from the room impulse response 291 because the scaling of the early reflections and direct sound has been modified. For the case where the early reflection portion is divided into only seven segments, block 294 shows seven overlapping segments. However, more segments may be used, and in a typical implementation, 50 segments, or even 60 or more, may be used.
[0188] After applying the windowing operation to the diffuse part, the same segmentation is applied to the diffuse part so that the diffuse part can be combined nicely with the windowed specular part.
[0189] Therefore, for the calculation of the weighted sum 290, the windowed diffuse part is used and the windowed specular part is processed using the image source model 295 as described previously, wherein, in addition, the HRFT provider 296 provides the correct HRTF for each segment, followed by a filling operation 297, a subsequent convolution operation 298 with the selected HRTF from block 296, and finally a delay compensation 299 is applied, thereby obtaining the correct mixture of the specular and diffuse parts for each early reflection segment.
[0190] exist Figure 11bIn the preferred embodiment shown, a further correction 290d is performed to account for the situation where the specular component should dominate the diffuse component (i.e., should have a stronger influence than the diffuse component) near the boundary between the direct sound and the early reflections. On the other hand, at the other end of the early reflection component, i.e., at the boundary between the early reflection component and the late reverberation component, the diffuse component should dominate the specular component.
[0191] Typically, it has been found that the measurements based on directional to diffuse reflectance (DTD) in blocks 290a, 290b already produce a situation where either the specular portion or the diffuse portion should be dominant. However, to avoid any unnatural situation, a correction 290d is applied in some way, such as providing a maximum or minimum amount for each segment or segments, or by applying some kind of curve to the measurements determined as shown in blocks 290a, 290b.
[0192] Depending on the embodiment, the directional to diffuse reflectance ratio can be used as a threshold or as a smooth transition from 0 to 1. When the energy in the first window is twice the energy in the second window, a completely specular or significant reflection is obtained. When the average energy in the first window is 0.5 times the energy in the second window, and all values between 0 and 1 are possible, a completely diffuse reflection segment is obtained. These values are preferably used as weighting factors, or weighting factors for determining a weighted combination of the specular and diffuse reflection parts.
[0193] In addition, reference Figure 11a , Figure 11a A process for calculating a particularly preferred Directional to Diffuse Reflectance (DTD) is shown. The early reflection portion is cut into overlapping portions. A pre-gain factor is then determined to transfer energy in a directionally manner, with the goal of applying the directional transfer function not only to the direct sound portion, but also to the early reflection portion, as previously discussed with respect to the first and second aspects. The early reflections are then cut into blocks or segments with a block size of 256 samples and a hop count of 128 samples, and a zero padding operation is then performed on 512 samples, whereby a Fourier transform is then applied to each block. The Directional to Diffuse Reflectance is then determined by determining the amount of energy of each block, comparing it to a moving average, and determining the relationship between the geometric reflection and diffuse reflection portions (in decibels). To this end, Figure 11a The energy shown is converted into smooth energy through moving average operation.
[0194] The preferred process can also be performed as follows. For example, by taking as input a binaural noise sequence of the same length as the reverberant portion of the RIR and the reverberation of the RIR, the diffuse component can be derived. (The binaural noise sequence here refers to a white noise that exhibits the same phase characteristics and interaural correlation as the recording of the diffuse sound field.) The combination of the diffuse and specular components is preferably performed according to equation (1) above.
[0195] Figure 13a The subject matter of the invention according to a fourth aspect is shown. This aspect relates to an improved calculation of the diffuse portion of the early reflection portion or the late reverberation portion or only the late reverberation portion or both portions using the magnitude spectrum of the early reflection portion and / or the late reverberation portion and using the phase spectrum of the two-channel (binaural) noise. Figure 1 The dual-channel synthesizer 200 is configured to use the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part and the first channel noise phase spectrum of the first channel for obtaining the dual-channel acoustic data, and use the amplitude spectrum of the early reflection part and the amplitude spectrum of the single-channel acoustic data without the direct sound part and the second channel noise phase spectrum to calculate the dual-channel diffuse reflection part of the early reflection part or the dual-channel diffuse reflection part of the single-channel acoustic data without the direct sound part.
[0196] In particular, the first channel noise phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence. Figure 13a Block 530 in shows the calculation of the magnitude spectrum and block 532 shows the calculation of the phase spectrum of the two-channel (binaural) noise.
[0197] The conversion of the single-channel data obtained by block 520 or preferably after smoothing of the amplitude spectrum in block 531 is transformed into the second channel result by adding the phase of the first channel phase spectrum to the smoothed amplitude of the spectrum in block 531 to obtain the first channel result, and by adding the second channel phase spectrum of block 532 to the preferably smoothed amplitude spectrum of block 531 in combiner 533 to obtain the second channel of the two-channel diffuse portion of the late reverberation part or the early reflection part plus the late reverberation part, or in other words, to obtain the second channel of the two-channel diffuse portion of the single-channel acoustic data without the direct sound part (which is assumed to be non-diffuse and therefore does not receive and diffuse reflection contribution). It should be noted that in a specific mathematical sense, the "addition" of amplitude and phase is a multiplication, as Figure 13b As shown in block 444, the multiplication of the magnitude spectrum and the phase spectrum is: |RTF|*e^(angle(binauralNoise1)) and |RTF||*e^(angle(HbinauralNoise2)), where H represents the transformation to the spectral domain.
[0198] like Figure 13b As shown, a mono RIR is provided in block 440, and the absolute value of the STFT spectrogram consisting of a series of spectra is shown in block 442. A binaural noise sequence 441 is also provided, and its corresponding spectrogram is processed by time-frequency conversion, as shown in block 443, taking the phase angle of each channel of the binaural noise, and then, as shown in block 444, preferably after the smoothing operation in block 531, combining the phase angle with the corresponding amplitude.
[0199] The smoothing operation in block 531 has the advantage that this smoothing along the frequency direction of the magnitude spectrum naturally avoids any peaks that might occur due to the inverse Fourier transform when performing phase manipulation in the spectral domain, as is the case with the present invention, in each spectrum of the spectral sequence covering, for example, the early reflection portion and the late reverberation portion. On the other hand, the process of calculating the spectrogram and simply "adding" the phase of the binaural sequence to the (smoothed) spectrogram is a computationally simple process that does not require a large amount of computing resources. Furthermore, it has been found that this late reverberation processing has a pleasant sound to the listener, which is particularly useful because, due to its quality, the same late reverberation two-channel data of the acoustic environment can be used regardless of changes in source position or orientation or listener position or orientation. This has the significant consequence that the update rate for the calculation of the late reverberation portion can be significantly reduced (typically by one or even two orders of magnitude), which further reduces the required computing resources and also allows the processing tasks to be distributed to different elements, as will be shown with reference to the sixth aspect of the present invention.
[0200] Figure 13c Shown Figure 13a and 13b Another embodiment of the process in . In block 445, an overlapping block transform is applied to the late reverberation room impulse response or to both the early reflection part and the late reverberation part. This produces a first spectrogram, wherein, preferably, low-pass filtering is performed over frequency in each magnitude spectrum as shown in block 531.
[0201] In block 447, a low-pass filtering is preferably additionally performed over time (i.e., over two or more adjacent blocks and with respect to the same frequency bins, but in adjacent blocks, i.e., with temporally adjacent frequency bins associated with the same frequency). A similar transformation 446 is performed on the time-domain binaural noise sequence to obtain a second spectrogram and a third spectrogram, and in block 449 the phases of the second and third spectrograms are added to the frequency-domain and time-domain low-pass filtered spectra.
[0202] Then, as shown in block 450, the result of block 449 is transformed into a Cartesian format and then inversely transformed into the time domain in block 451. In block 452, an overlap and add process is performed, and finally in block 453, truncation and windowing as well as overlap with the earlier reflection part are performed, which is only performed on the diffuse reflection signal of the late reverberation, and then the two-channel acoustic data of the late reverberation part is obtained at the output of block 453.
[0203] A binaural noise sequence of the same length as the reverberation portion of the RIR is used as input. (The binaural noise sequence here refers to a white noise that exhibits the same phase characteristics and interaural correlation as the recording of the diffuse sound field.)
[0204] The two filters can then be transformed into a time-frequency representation using a short-time Fourier transform (STFT). The parameters of the STFT can be chosen in such a way that it allows for near-perfect reconstruction, for example by using a semi-overlapping Hann window. The complex-valued block-wise spectrum is then converted to polar form, separating each frequency bin into its amplitude and phase components. A complex representation of the diffuse component is then constructed for each channel of the noise by pairwise combining the complex amplitude of each bin of the transformed RIR block with the complex phase of the transformed noise sequence block.
[0205] In some embodiments of the described system, the binaural synthesizer is further configured to perform low-pass filtering between the amplitudes of the frequency bins during this processing stage. A typical low-pass filter may be a moving average filter, corresponding to intervals equivalent to 1 / 3 octave, but other configurations are possible. This reduces artifacts introduced by combining the two amplitude and phase components of the two different transfer functions.
[0206] Furthermore, a low-pass filter can be applied between temporally adjacent blocks of the reconstructed transfer function, so that intervals corresponding to the same frequency are low-pass filtered between the blocks. A typical configuration is a moving average filter of 3 values (or time blocks), assuming a block size of 512 samples at 48000 kHz. The exact parameters also depend on the individual embodiment.
[0207] The new filter is then converted back to Cartesian form and transformed back to the time domain using an inverse STFT with the same parameters as used for the forward transform. The resulting filter is a binaural diffuse reverberation filter with the length of the combined ER and LR sections and two channels. Specifically: the beginning of this diffuse section is used as one of the two layers of the ER segmentation, and the later diffuse section is the input to the LR segmentation.
[0208] Then, it is shown Figures 14a to 14e, which can be used in each of the first to fourth aspects of the present invention, and in particular, for the fifth aspect of the present invention, which allows the required room impulse response to be effectively provided. To this end, the input interface is referred to as an RIR provider 100, and the RIR provider receives, for example, an initially measured microphone signal as input.
[0209] The room impulse response is forwarded to a binaural synthesizer, which calculates the binaural room impulse response based on geometric data about the room (e.g., geometric data required for the calculation of the image source), the required HRTF, and position data about the sound source and the user. The binaural room impulse response can then be auditoryized by an auralizer 300 or sound generator using an audio signal to obtain two output speaker signals that can be rendered by headphones, earbuds, in-ear devices, or discrete speakers.
[0210] exist Figure 14d In the embodiment of the present invention, the microphone signals are measured, the RIR provider 100 calculates parameters or, in general, a fingerprint from the microphone signals, a certain room impulse response database 110 is accessed and the database replies with a matching room impulse response which is then forwarded by the block 100 to the two-channel synthesizer or binaural synthesizer 200.
[0211] exist Figure 14c In the embodiment of the present invention, an RIR provider 100 generates a set of parameters from the microphone signal and forwards these parameters to a database or to an RIR synthesizer that can synthesize an RIR based on these parameters. Therefore, compared with block 110, block 120 can have two functions. In addition, an RIR modifier 130 is provided, which modifies the RIR in some way for the purpose of achieving some desired sound effect or room effect.
[0212] Figure 14d A process is shown using only the RIR synthesizer 140 without a database for the purpose of auralization of the room acoustics.
[0213] Figure 14e Another preferred way of providing a specific room impulse response is shown. The process relies on acoustic measurements, which can be microphone signals or can come from other sources. In block 101, dimensionality reduction is performed to obtain a simplified representation, which can be, for example, a set of parameters or, in general, a fingerprint, which can be derived from the acoustic measurements by a process other than parameterizing the signal (e.g., psychoacoustic parameters).
[0214] In addition, an RIR database 110 is provided, which also includes a dimensionality reduction block 111 for generating a simplified representation again, which is then input into a block 112 along with other simplified representations of other RIRs stored in the RIR database. Block 112 minimizes the distance and finds the best matching RIR. The best matching RIR is identified from block 112 and this information is sent to block 113, which loads the RIR from the RIR database 100 and then performs binauralization. Figure 14e The binaural combination of the blocks Figure 1 or the functionality of blocks 200 and 300 in other figures.
[0215] also, Figure 15 Specific hardware of a preferred embodiment according to the sixth aspect of the present invention is shown. Specifically, the first device 901 includes one or more microphones 911, one or more processors 912, memory 930, and a position tracking system 914 for tracking the position of a listener, where the position of the listener is also referred to as the listener's orientation, i.e., collectively referred to as the listener's posture. Furthermore, the first device 901 may include a speaker 915.
[0216] In this embodiment, the second device 902 includes a processor 921 and a memory 922 , and is connected to the first device 901 via a network.
[0217] Part 1 of the solution infers the RIR (for a particular virtual source-listener configuration) from available audio data recorded at the listener's listening environment. The real configuration in which this audio was recorded may, but does not necessarily, match that of the virtual configuration.
[0218] System 1 is configured to record, either continuously or in one go, measurements of the local sound field, including the room acoustics of the user's real environment. In some embodiments, this is achieved by measuring the RIR between the microphone and any real sound source in the room, for example using an exponential sine sweep. This is particularly useful when the system is calibrated only once for the listening environment. The RIR recorded in this manner may or may not be suitable for sonification of virtual sound sources, either because they belong to different sources or because they may not match the virtual sound source to be sonified due to limitations of the capture portion of the system, such as the limited bandwidth of the sensor. [A].
[0219] The recorded audio data is sent to the second system via the network [B]. The memory stores a database of pre-recorded high-quality omnidirectional RIRs for different rooms with varying acoustic properties. The database may (but need not) include one or more measurements of the actual listening room.
[0220] The purpose of this system is to process the transmitted audio data and select the RIR, i.e. the unknown RIR that best fits the real room at the current listener position, which is sent back to the first system for further processing and binaural synthesis [C].
[0221] This second system is configured to reduce the high-dimensional, temporal, or time-frequency representation of the audio data to a lower-dimensional representation. In some embodiments, this is achieved by transforming the data using a fully trained neural network. Such a network can, for example, be trained on the task of classifying RIRs into individual room categories (categories not residing in the database). The coefficients of the network layers and the resulting latent space are then selected as the lower-dimensional representation of the data, computed for the pre-recorded RIRs and the ad hoc measured RIRs [D]. These coefficients are then used to find the best-matching RIR based on minimizing a suitable distance metric in this reduced-dimensional space. The obtained RIRs are then transmitted back to the first system via the network. The process of obtaining RIRs is repeated at regular intervals to reflect major changes in the room acoustics, such as when the listener moves into an area with significantly different reflections or a different acoustic environment. When a new RIR is found that minimizes the distance metric to the new data point, it is selected. The system maintains a short-term history of the used RIRs, allowing it to gradually converge between the changing RIRs.
[0222] Further examples are given below:
[0223] 1. [A] These systems do not record impulse responses, but are instead configured to record and process well-defined, self-generated sounds of the user, such as clapping or speaking, allowing the RIR to be inferred without the need for a sine sweep.
[0224] 2. [A] These systems do not record impulse responses, but are configured to record and process well-defined sounds such as music, allowing the RIR to be inferred without the need for a sine sweep.
[0225] 3. [A] These systems do not record impulse responses, but are instead configured to record and process generic sound fields (not tied to a specific class or sound of sound), allowing the RIR to be automatically inferred without the need for a sine sweep or any user input for calibration.
[0226] 4. [D] Use appropriate digital signal processing supported by psychoacoustic models to create a lower dimensional space suitable for matching RIRs, rather than using the latent space of a neural network for dimensionality reduction.
[0227] 5. [C] Train a neural network to directly synthesize new RIRs, rather than selecting from a set of pre-recorded RIRs. These RIRs may or may not be a combination of existing RIRs.
[0228] 6. [C] Train a second neural network to synthesize new RIRs from the reduced representation, rather than selecting from a set of pre-recorded RIRs. These RIRs may or may not be a combination of existing RIRs.
[0229] 7. [C] Train a neural network to synthesize new RIRs from the output of the embodiment described in 4, rather than selecting from a set of pre-recorded RIRs.
[0230] 8. [B] System 1 extends the non-transient storage for the RIR database. All processing is done on System 1.
[0231] In some embodiments, the RIR provider does not have to be manually configured with RIR. Instead, the system can use a speaker and a microphone, which may not meet the qualitative requirements of broadband RIR measurement, that is, the transducer may have a nonlinear or blocked frequency response in the audible range, or they may be installed in the same chassis. It does not measure RIR directly, but is configured to measure low-quality RIR, which may or may not be used for binaural synthesis. The measured RIR is not directly used as input for binaural synthesis. Instead, acoustic or psychoacoustic parameters are derived from the measured RIR. For example, the band reverberation time (RT60) or the energy decay curve (EDC), the direct to reverberant ratio (DRR) or other parameters can be calculated. The exact parameters calculated depend on the embodiment.
[0232] The calculated parameters are then used to find similar RIRs from a database for the pre-recorded RIRs, which are suitable for binaural synthesis. The pre-recorded RIRs are stored in a non-transitory memory along with the selected set of acoustic parameters.
[0233] When the RIR provider is set up in a new acoustic environment and a low-quality RIR is measured, these parameters are calculated and compared with a pre-recorded dataset. The best matching pre-recorded RIR is selected from the database and used as the input RIR for the binaural synthesizer.
[0234] In some embodiments, rather than directly finding the RIR with the best matching parameters, a psychoacoustic weighting function is employed that specifies a weighting coefficient for the influence of each parameter.
[0235] In some other embodiments, the parameters used to find the best matching RIR are other acoustic or psychoacoustic parameters. Instead, the measured RIR is represented by a set of parameters that are calculated by transforming the data with a well-trained neural network. Such a network can be trained, for example, on the task of classifying RIRs into categories of individual rooms (which categories are not resident in the database). The coefficients of the network layers and the latent space they form are then selected as a lower dimensional representation of the data, which is calculated for the pre-recorded RIRs and the self-measured RIRs. These coefficients are then used to find the best matching RIR based on minimization of a distance metric in the reduced dimensionality parameter space.
[0236] When traditional RIR measurements (i.e., using exponential sine swept deconvolution) are not feasible, some embodiments of the system can make assumptions about the RIR based on the type of sound, i.e., human clapping or human speech, or derive the RIR from reverberant audio recorded directly with one or more microphones. To achieve this, the selected parameters are either derived directly from the reverberant audio or an intermediate approximation of the RIR.
[0237] Some embodiments of the system used in more than one acoustic environment need to adjust to changing room acoustics. This is achieved by changing the RIRs sent from the RIR provider to the binaural synthesizer. Depending on the embodiment of these RIRs, updates can be performed periodically, for example at a fixed rate, or when significant acoustic changes require an update.
[0238] In order to achieve an inaudible gradual change between the two input RIRs, the RIR provider is configured to gradually interpolate between the two filters.A suitable algorithm to achieve such interpolation is, for example, linear interpolation in the time or frequency domain.
[0239] In some embodiments, it is not possible to preconfigure the system for one or more room acoustic environments, such as when the device is worn and used in multiple environments, such as when listening to music while traveling. Here, one or more microphones of the device continuously or periodically record sounds from the acoustic environment. The RIR provider is configured to detect one of many classes of sounds, and it is configured to derive an intermediate representation of the RIR. Even though the room acoustics may sometimes change rapidly, such as when entering or leaving a room, the system can be configured to gradually integrate the detected room acoustics and gradually adjust to improve the stability of the results.
[0240] Instead of obtaining RIRs from a database of pre-recorded RIRs, state-of-the-art room acoustics simulations can be used to generate omnidirectional RIRs for binaural synthesis. Given a finite time frame and a set of input parameters, the employed algorithm is able to simulate a good approximation of real room acoustics. Since the binaural synthesizer models the phase relationship of the BRIRs, the algorithm must specifically reproduce a good approximation of the real, frequency-dependent energy distribution over time. Depending on the room acoustics simulation method employed, the RIR provider is configured to calculate the input parameters required for the simulation.
[0241] Some embodiments of the described system use an extended version of the binaural synthesis method that can use room acoustic modeling (such as an image source model) to calculate specular reflections, rather than processing individual specular reflections from recorded RIRs. Here, the provided room geometry information is used to determine the arrival time and arrival direction of each specular reflection. The acoustic absorption effect of the reflective surface from which the reflection originates can be included as an input to the room geometry data. Alternatively, some embodiments may choose to estimate the absorption coefficient of the wall by analyzing the initial reflection at the arrival time of the reflection of the ISM, by deriving a filter based on a reflection window and a direct acoustic window (assuming it is close to linear in the audible frequency range). This modification allows a higher density of specular reflections to be calculated, thereby potentially improving localizability, but at the cost of more computational requirements.
[0242] Figure 16 A preferred implementation of the fifth aspect is shown, which relates to intelligently determining a room impulse response from a raw representation associated with single-channel acoustic data. In particular, Figure 1 The input interface 100 of the illustrated device is configured to obtain a raw representation associated with the acoustic data, as indicated at 150. Furthermore, the input interface 100 is configured to derive single-channel acoustic data using the raw representation obtained in block 150 and using additional data stored by or accessible to the audio signal processor to obtain single-channel acoustic data, which is then forwarded to the two-channel synthesizer 200.
[0243] Exemplarily, the input interface is configured to obtain an initial measurement of the original single-channel acoustic data as a raw representation to derive a test fingerprint of the original single-channel acoustic data, such as Figure 17a Based on the test fingerprint, a pre-stored database 110 having an associated set of reference fingerprints is accessed, wherein each reference fingerprint is associated with a higher resolution single channel acoustic data, wherein the high resolution single channel acoustic data has a higher resolution than the initial measurement. Figure 17a As shown in block 113 , high-resolution single-channel acoustic data having a reference fingerprint that best matches the test fingerprint is retrieved from the pre-stored database 110 .
[0244] Alternatively, high-resolution single-channel acoustic data can also be synthesized from test fingerprints or from original single-channel acoustic data, usually using additional geometric data or using only geometric description data, such as Figure 17a , which illustrates a direct synthesis of single-channel acoustic data. To this end, block 140 receives the original representation obtained by block 150 or the test fingerprint calculated by block 101, as well as room simulation data, data on neural network information (in this case, block 140 implements the neural network) or model data as additional data. Therefore, when performing the alternative of direct synthesis, database 110 is not required. In addition, block 101 is configured to derive the test fingerprint as a set of at least one of the following parameters: RT 60, EDC, DRR, and wherein the reference fingerprint also includes at least one of the following parameters: RT 50, EDC, DR.
[0245] Figure 17b Another process for calculating the room impulse response or room transfer function of an acoustic environment is shown. In block 150, a sound clip (such as a song played by a speaker in an acoustic environment) is recorded as Figure 17a The original representation of block 150 is shown in block 155. In block 155, an audio fingerprint system is used to identify the sound, which is accessed by receiving a test fingerprint and returning an identification of the music piece or a matching reference fingerprint, as shown in blocks 155 and 156. In block 157, a usually remote music database can be accessed using the reference fingerprint or the identification of the music piece, and in block 158, a song played by one or more speakers in the acoustic environment is retrieved, not a version with the room acoustics imprinted on it, but a clean version played by the speakers. In block 159, the RIR or RTF of the acoustic environment is calculated using the song recorded in the environment and using the clean version of the song (i.e., without any room effects provided by the music database 157).
[0246] Figure 17c Another implementation is shown in FIG, where initial measurements or data are obtained in block 150 and a test fingerprint indicative of the class of the acoustic environment is calculated, for example, by a neural network or other process, in block 112. Then, based on the room class, matching RIRs may be retrieved from a pre-stored database, as shown in block 152, or may be synthesized using the selected room class, as shown in block 153. The room class may be an enclosed room, an open environment, a large room, a small room, a room with significant damping, a reverberant room, etc.
[0247] Another implementation of the present invention is that the user generates natural sounds, such as Figure 17a160. This natural sound is applause, speech or any transient sound that the listener can produce. This avoids generating unpleasant measurement sounds in the room, such as sine sweeps. Based on this sound, a (low resolution) RIR is then recorded as a microphone signal and passed through Figures 17a-17c The raw representation is processed using any of the processes shown to obtain a high-resolution room impulse response from this raw representation for further processing purposes.
[0248] Subsequently, preferred embodiments according to the sixth aspect of the present invention are discussed. The auditory perception of the direct sound path is typically achieved by block-by-block convolution of filters with the audio signal, which approximate the filtering effect (HRTF) of the user's head, ears, and torso relative to the sound source at a given position and distance. The processing of these filters requires encoding the correct changes in inter-aural time difference (ITD) and inter-aural level difference (ILD), as well as changes in sound intensity and other clues. Human listeners are relatively sensitive to small changes in these values, which is why it is necessary to calculate these changes with good spatial and temporal resolution. However, these filters are relatively short, and deriving them usually involves only a few processing steps. Room reverberation simulates the filtering effects of sound, which are caused by the geometry of the environment because the sound does not propagate from the sound source to the user's ears on the direct part. This includes reflection, refraction, absorption, and resonance effects.
[0249] This reverberation filter is expected to be much longer than the short direct sound filter. Many processes, algorithms, and systems are capable of processing adequate binaural reverberation, such as image source algorithms, ray tracing, parametric reverberators, and many delay network-based approaches. The exact implementation of the reverberator is irrelevant to the present invention, as long as it can simulate good externalized sonification. In this embodiment, the system exploits the fact that human listeners are more sensitive to changes in the direct sound filter and less so to changes in the reverberation filter. The signal processor is programmed to calculate the direct sound filter much faster than the reverberation filter. This allows the system to minimize audible jumps in the sonified sound and increase the sense of externalization when changing filters, while avoiding full filter updates. Updating these filters, or encoding the signal portion of the direct sound path, at a rate of approximately 188 Hz has proven to be a reasonable default for this system, but lower update rates (such as 94 Hz or 50 Hz) may be feasible in different embodiments of the system. The reverberation filter is calculated at a much lower rate, typically up to one-tenth the direct sound processing rate, depending on the acoustic characteristics of the environment and the user.
[0250] The signal processor or another processor is configured as an aggregator. In some embodiments, the binaural synthesis method used returns a continuous block-by-block binaural audio signal stream, and the aggregator simply sums the blocks provided by the direct and reverberation processing paths and acts as a signal aggregator. This requires that the blocks to be summed correspond to the same time point or contain control data identifying the time frame to which they correspond. Alternatively, the aggregator can be configured to sum the two partial filters relative to a time delay determined by the algorithm. It thus reconstructs the complete BRIR filter from the individual processor results and acts as a filter aggregator. The filter can then be used to convolve the audio signal blocks using state-of-the-art real-time (block-by-block) convolution methods.
[0251] The aggregator always holds the complete BRIR filter in its memory. Thus, the BRIR can be partially updated at the individual rates of the individual processors processing parts of the filter. The resulting signal block contains the combined binaural signals of the direct and reverberation paths. These are then passed to the speaker signal generator for playback through the system's speakers. This enables binaural audio auralization with a similar level of externalization and perceptual quality as the separate algorithms, while significantly reducing processing requirements.
[0252] In various embodiments, the processing of the reverberation tail can be further separated into separate processing paths, with separate filters calculated for early reflections and late reverberation on one or more processors on the same device. This exploits the fact that strong early reflections often help human listeners localize sounds. These reflections tend to undergo strong and short-lived variations, especially as the listener moves around their environment. While humans are less sensitive to these variations than to variations in the direct sound, these early reflections carry a significant amount of energy, and slow filter calculations can produce undesirable externalization, localization errors, or audible jumps. On the other hand, the largest portion of a typical reverberation tail consists of densely overlapping reflections with relatively low energy. These late reverberations vary relatively slowly. For some environments, they may consist primarily of the diffuse portion of the reverberation tail, meaning they are constant across the entire sound field.
[0253] In this embodiment, a different reverberator can be used to process the late reverberation portion at a lower rate. Depending on the acoustic environment, the system can be tuned to keep the late reverberation constant, refresh at a low rate such as 1 Hz, or process it on demand when a significant change in the room acoustics is detected. In some embodiments, the algorithm used can provide a mixture of binaural filters and binaural signals. Here, the aggregation stage can be divided into a filter aggregator plus a convolver and a signal aggregator, which first combines the partial filters, convolves the reconstructed filters with the audio signal, and then sums the binaural signals to obtain the complete binaural signal.
[0254] In this embodiment, the system is divided into four or more reverberators, splitting the BRIR into four or more segments. This can be used to further process portions of the reverberation tail with varying degrees of complexity. For example, a precise geometric algorithm can be used for first-order reflections, while later-order reflections can be processed randomly, and the late reverberation tail can be processed as in the previous embodiment.
[0255] In this embodiment, the direct sound processor and reverberation processor of 1. are implemented on two devices respectively, forming a system with the same functions as in Example 1, which is suitable for distributed synthesis of binaural signals and auditoryization of binaural audio on wearable devices.
[0256] The first device is a wearable device that includes sensors, transducers, and one or more processors as shown in Figure 1. These processors are configured to synthesize direct sound binaural filters or binaural signals directly on the device at a sufficiently high refresh rate. The processing of the direct sound portion is completed directly on the device, avoiding transmission through a wireless channel. The wearable device includes an aggregator and a speaker signal generator required for aggregation of filters and / or signals, as shown in Figure 1. It also includes a subsystem for wireless transmission and reception of audio and control data. The second device includes a processor configured to calculate reverberation filters or signals. It also includes a subsystem for wireless transmission and reception of audio and control data.
[0257] In an embodiment where an algorithm employed on a second device synthesizes a binaural reverberation filter, the sensor data and control data required by the employed algorithm are sent by the wearable device. Partial filters are processed and wirelessly transmitted back to the first device. The processed filters are then sent to an aggregator, which reassembles a complete representation of the BRIR and stores it in memory, as shown in 1. The complete BRIR is then convolved and played back through the speaker, as shown in 1.
[0258] In an embodiment where an algorithm employed on the second device directly synthesizes a binaural reverberation signal, the audio signal is streamed along with the sensor data and control data required by the employed algorithm. A reverberation processor then synthesizes a binaural signal based on the employed algorithm, which is returned directly to the wearable system over a wireless channel. The data return contains the necessary control data, which allows the time frame to which the binaural signal block corresponds to be determined. The audio signal sent to the processor that calculates the direct acoustic path is delayed by a configurable delay that is at least as long as the transmission delay introduced by the two wireless transmissions to and from the second device. An aggregator then combines signal blocks corresponding to the same audio signal block specified by the time data specified in the control data stream.
[0259] In another system, an additional reverberator that calculates late reverberation is distributed on another device. This second additional device includes a processor configured with a late reverberation algorithm and a subsystem for wireless or wired transmission to the connected device. In some embodiments, the additional reverberator device can be wirelessly connected to the wearable device. In other cases, it can be connected to the first additional device. The total latency of the selected transmission channel can be less than the target refresh interval for processing the late reverberation signal or filter. Due to low latency requirements, the second additional device configured to process the late reverberation can be connected via an IP network with greater latency, such as the Internet. In another system, the additional reverberators are distributed across any number of additional devices.
[0260] Some embodiments of the described system do not operate completely independently, but rather can be connected to another device using a wired or wireless connection. The connected device uses the connection to send the required audio data and metadata to the system. This allows devices such as computers or smartphones to be connected and used with the system, acting as auralizers for spatial audio content.
[0261] Some embodiments of the device may use a so-called three-degree-of-freedom (3DoF) tracking system, which measures only the user's head rotation to provide positional data to the system. Similarly, some embodiments may send only 3DoF tracking data and limited translation or acceleration data to the system. In these embodiments, the system can be used to auralize a virtual audio scene, where sound sources exist in space relative to the user. When the user rotates only their head, the sound sources remain stable in one position. As the user (and the device) moves, the virtual sound sources appear to move with them because they are centered around the user. When some form of translation or acceleration data is available, it can be used to allow small head translations within a limited radius (e.g., 50 centimeters), which facilitates externalization and localization. Larger movements are not reproduced. These embodiments of the system are particularly useful for auralizing classic spatial audio content, music, movies, and sound dramas, where the user does not intend to freely explore the virtual acoustic scene or leave it.
[0262] Different embodiments of the device may use a six degrees of freedom (6DoF) tracking system that measures the user's absolute head rotation and position. In these embodiments, the user can freely navigate the virtual acoustic environment, passing through and leaving it. This is particularly useful for auditoryization of AR content, games, navigation content, and human-computer interaction scenarios. The entire sensor, RIR provider, binaural synthesizer, and auditoryizer may be distributed across multiple devices. For example, a particularly small form factor, such as an earbud, may require that parts of the system be distributed to another device. In this case, the position tracking sensor and acoustic transducer remain on the wearable device, while the RIR provider, binaural synthesizer, and auditoryizer are distributed across one or more devices. In embodiments in this case, the motion-to-sound latency requirements must still be adhered to.
[0263] Some embodiments of the system can configure the RIR provider to provide RIRs with different characteristics that partially or completely mismatch the parameters of the actual acoustic environment. Alternatively, the system is extended by a RIR modifier component that receives the RIR input from the RIR provider and modifies it to change certain acoustic parameters of the matching RIR so that the modified RIR has these desired qualities and parameters. This can modify these room acoustic parameters to a desired level, for example, making the listening room appear less reverberant, resulting in a more pleasant listening experience. Alternatively, this can be done to make the room sound more like a different room, i.e., when listening to a concert, making the listener's current virtual acoustic environment sound more like the concert hall for aesthetic purposes. For example, a longer late reverberation tail (LR) can be auditorily realized by selecting an RIR that exhibits similar parameters but has a longer reverberation time. Alternatively, the LR original (input) RIR can be resampled so that it is stretched by a certain amount, resulting in a longer reverberation time, while keeping other perceptually relevant components such as the DS and ER intact.
[0264] Figure 19 The preferred embodiment of the sixth aspect of the present invention is shown. In particular, Figure 1 The device shown is separated into a first device and a second device. In particular, the dual-channel synthesizer 200 is configured by two physically separated devices 901, 902, such as Figure 15 The first device 901 of the two physically separate devices is configured to process the direct sound portion, as shown in blocks 916 and Figure 2 For this purpose, the processing requires the listener position or rotation. In addition, the second device 902 of the two physically separate devices is configured to process at least one of the early reflection part and the late reverberation part. This block is shown at 923 and implements Figure 2 One or both of the functions of blocks 230 and 240.
[0265] The two devices are connected to each other via a transmission interface 918 of the first device and 925 of the second device. The transmission interface is preferably a wireless interface and operates according to, for example, the Bluetooth standard. In addition, a result of the separation of the two physically separate devices is that the first device 901 has its own power supply 917 and the second device 902 also has its own power supply 924.
[0266] Preferably, if Figure 19 As shown, the first device is configured to update the two-channel acoustic data of the direct sound part more frequently than the two-channel audio data of at least one of the early reflection part and the late reverberation part of the second device. In the figure, it is preferred to have an update of the direct sound part above 15Hz, that is, more than 15 updates per second, preferably more than 20 updates per second, and even more preferably more than 50 updates per second. The update rate of the early reverberation part is preferably in the range between 5Hz and 15Hz, and it is sufficient to update the late reverberation part in the range between 0.5Hz and 5Hz. Therefore, the parts that seem to require a lower update rate are processed in the second device 902. It has been found that it is these parts that require significantly higher processing power because long filters, on the other hand, require a lower update rate. Therefore, the second device is implemented to be significantly stronger and more powerful than the first device in terms of computing and battery power. The first device can be an earbud device, a headphone device, an in-ear device, or any other wearable device that generally has limited battery power. However, the second device may be a high-power device such as a mobile phone, smartwatch, laptop, tablet, or even a stationary computer connected to a power line and typically also connected to a large area network such as the Internet. Figure 15 As outlined, the first device comprises not only a processing block for the direct acoustic portion 916, but also a microphone for recording the acoustic measurements provided by the RIR and, in addition, also comprises the following: Figure 1 The sound rendering function shown in the sound generator 300 and, in addition, a speaker, for example when the device is a headphone device. Alternatively, the speaker can also be separated from the device 901 with a communication interface instead of an actual speaker, when the speaker is provided with, for example, a Bluetooth signal.
[0267] Figure 21a and Figure 21b Shown according to Figures 22a to 22f The starting point of the embodiment of the sixth aspect shown. In particular, Figure 21aIn the embodiment, the input interface includes a microphone array composed of one or more microphones shown at 911, a user posture tracking system 914, and other sensors 919 that may be provided. The binaural processor in block 200 includes a direct sound processor and a reverberator for generating binaural signals, which are then aggregated by a signal aggregator 310 to obtain a two-channel binaural signal, which is then processed by the signal generator. The function of the signal aggregator is as follows: Figure 23b shown.
[0268] on the contrary, Figure 21b has a similar implementation, but the binaural filter parts are aggregated as shown in the filter aggregator blocks 250, 300, and the result of the filter aggregation is represented by Figure 21b The sound generator 300 or "auditizer" in Figure 23a The process is schematically shown in FIG.
[0269] According to the present invention defined in the sixth aspect, the reverberation processing is implemented in the second device and the direct sound processing is implemented in the first device 901. In addition, the functions of the input interface 100, the signal aggregator 310 and the signal generator 300 are also implemented in the second device. Figure 22a is implemented in the first device 901, Figure 23a treatment alternatives. Figure 22b Similar to Figure 22a , but with Figure 23b signal processing alternatives. Figure 22c Another embodiment is shown, which is Figure 22a and Figure 22b The embodiment of the present invention differs in that a second additional device 903 is provided. In particular, in this embodiment, the second device 903 typically processes Figure 2 The late reverberation processing of block 240, wherein the first additional device 902 performs Figure 2 The early reflection processing of block 230 and the direct sound processor in the first device perform Figure 2 Similarly, the signal aggregator 310 aggregates the individually convolved audio signals, such as Figure 23b The alternative is shown. Figure 22d Similar to Figure 22c But now the filter aggregation function is the same as Figure 23a treatment alternatives.
[0270] Figure 22eAnother implementation is shown in which even more than two additional devices are provided. For example, such an additional device 904 can be implemented to perform initialization tasks, such as the calculation of the image source and the image source position, thereby minimizing the use of the wearable device's battery. The additional device 904 then receives the microphone signal and the initial measurement data and uses the database, etc. to perform image source position processing and other initialization processes, such as determining the correct room impulse response, since these tasks are performed even less frequently than the calculation of the late reverberation part. Other ways of distributing processing tasks to even more additional devices are also useful. Figure 22e Once again have Figure 23b The processing alternatives Figure 22f have Figure 23a An alternative processing scheme for filter aggregation.
[0271] Figure 18 A preferred implementation of the process according to the sixth aspect is shown, but the process may also be applied to any of the other aspects. In block 801, the earlier single channel acoustic data or earlier raw representation obtained in block 150 of the fifth embodiment has been obtained.
[0272] In step 803, a new raw representation is acquired in response to control 802, which provides an activation signal to block 803 at regular intervals or upon a detected event (e.g., when a user moves from one room to another), thereby requiring an update of the entire room impulse response rather than the position of the user or listener. In block 804, the new raw representation is compared with an earlier raw representation, or the new single-channel acoustic data is compared with the earlier single-channel acoustic data, to determine whether an update is necessary.
[0273] If the deviation is determined to be approximately the threshold or an update condition in block 805, new single-channel acoustic data is determined 806. To gradually change from one RIR to the next, a blend of the earlier data and the new data is used, or alternatively, when the new data is not significantly different from the earlier data, the new data is used directly. In block 808, the earlier data in the memory is overwritten with the current data so that the current single-channel acoustic data or the current original representation is present during the next process in block 801.
[0274] Then, for the purpose of handling two-channel synthesis with different update rates, the Figure 20. In block 930, the currently used two-channel acoustic data for the earlier reflections and the later reverberation is stored. In block 931, it is assumed that the update of the two-channel acoustic data of the direct sound portion has been performed. In block 932, it is determined whether new data of the early reflection portion or the later reverberation portion is available. If this is confirmed, the new data is used together with the new data of the direct sound for sound generation. However, when it is determined in block 933 that the new data of the earlier reflection portion or the later reverberation portion is not available, the stored data of the earlier reflection portion or the later reverberation portion is used together with the new data of the direct sound portion. Therefore, due to the fact that the two-channel acoustic data of the portion with a reduced update rate is always stored, this data can be easily used together with the newly updated direct sound portion that requires a high update rate.
[0275] Subsequently, preferred embodiments of the present invention relating to a seventh aspect are shown, which relate to an improved separation of single-channel acoustic data and an improved combination of dual-channel acoustic data.
[0276] exist Figure 24a In block 600, the entire room impulse response is pre-processed, for example, from a database, or from measurements or from a synthesis process, so that the direct sound portion of the room impulse response is located at a predefined sample index. This pre-processed room impulse response is then forwarded from the input interface 100 to the two-channel synthesizer and, in particular, to the separation block 210. In block 601, the separation instant between the direct sound portion and the earlier reflected sound portion is determined, for example, midway between the maximum of the direct sound portion and the maximum of the first earlier reflection. In addition, or alternatively, a separation time is determined between the early reflection portion and the late reverberation portion, for example at the mixing time or, for the purpose of saving computing resources, some predetermined amount of time before the mixing time.
[0277] In block 602, at least one of the two adjacent parts is extended by a certain number of samples of the corresponding other part. For example, when a directional transfer function or a directional impulse response is used in the direct sound part, the direct sound part is removed and no extension or subsequent windowing by block 603 is required. However, when a directional transfer function is not used or the direct sound part of the RIR is used for some reason, the processing in blocks 602 and 603 is also applied to the part of the direct sound part at the first separation time. In block 603, at least the first earlier reflection part, the last earlier reflection part and the first late reverberation part are windowed using a window function that takes into account the extension (such as a Tukey window). Therefore, at the output of block 603, there is a windowed first early reflection part, a windowed last early reflection part and a windowed first late reverberation part.
[0278] Figure 24bshows the process of putting separately processed data together, as Figure 2 To this end, each portion is processed in a separate manner, as indicated by block 604, before being combined, and these manners are as described with respect to Figure 2 Then, an overlap-add is performed between the two-channel direct sound portion and the two-channel earlier reflection portion in block 250a, an overlap-add is performed between the two-channel earlier reflection portion and the two-channel late reverberation portion in block 250b, and finally, post-processing is performed in block 605 to obtain complete two-channel acoustic data for use in the present invention. Figure 1 The sound generator 300 is used.
[0279] In such Figure 25 In the preferred embodiment shown, for example, a directional impulse response or a directional transfer function plus an associated head-related impulse response is used to generate the direct sound portion. The result is a similar sound pattern to that previously discussed. Figure 24a A similar process as discussed is extended for n samples, and in block 603, windowing is performed, for example using a Tukey window.
[0280] Furthermore, as shown in block 605, each segment in the early reflection portion is overlapped and all sequences are overlap-added, and the initial time delay gap is preferably adjusted in block 610. After the initial time delay gap is adjusted, as shown in block 606, overlap-addition of each channel is performed to ultimately obtain aggregated two-channel data of the acoustic environment.
[0281] Figure 26 Shown Figure 25 In a preferred embodiment of the process performed in block 610 of FIG. , for initial time delay gap adjustment, an initial source-receiver distance or initial propagation time is determined using the initial source position and the initial receiver position, and is additionally placed at the position of the image source of the first reflection.
[0282] In block 612, the current source position and the current listener position are used together with the position of the image source of the first reflection to calculate the current distance or corresponding travel time. In block 613, the distance difference or the corresponding travel time difference is calculated, for example, the incremental ITDG is calculated in block 630, and in block 640, the ITDG is adjusted by shifting the earlier reflection portion (typically the already "connected late reverberation portion") more towards or away from the direct sound. For example, when the listener is closer to the sound source, the ITDG is greater than the initial time delay gap, thus shifting the earlier reflection portion away from the direct sound portion. As a result, the overlap no longer matches perfectly, and this can be compensated by padding some samples so that there is complete overlap at the beginning of the ER portion.
[0283] However, when the listener is further away from the sound source than in the initial measurement, the ITDG is smaller and Δ is negative. In this case, the earlier reflections are shifted closer to the direct sound portion, which is managed by simply chopping off the first few samples of the earlier reflections so that these samples are not overlap-added to the direct sound portion after the ITDG adjustment in block 610. Figure 25 In the direct portion of block 606.
[0284] Therefore, to maintain plausible distance perception, the Initial Time Delay Gap (ITDG) needs to be adapted to the listener's pose being synthesized. This acoustic feature describes the gap between the direct sound and the first reflection. Therefore, the temporal relationship between the DS and ER must be adapted. In basic binaural synthesizer implementations, this is achieved by temporally shifting the ER segments, as the system is designed to preserve the position of the DS components. Using an image source model, the ITDG can be calculated by taking the propagation time of the image source closest to the listener's position and subtracting it from the propagation time of the direct sound. This is done for the initial source-receiver constellation and the new constellation to be synthesized. The difference between the two ITDG values indicates how much the ER segments need to be shifted to represent the new situation. For example, when the listener is closer to the sound source, the ITDG will be larger compared to the initial constellation, and the ER segments will therefore be shifted slightly further away from the DS. In other implementations, this mechanism can be derived directly from the image source model by correlating individual reflections with the direct sound.
[0285] Subsequently, examples of the present invention related to the first aspect are summarized, wherein reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0286] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0287] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0288] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0289] a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0290] Wherein, the dual-channel synthesizer (200) is configured as
[0291] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part,
[0292] determining (222) directional information of the sound source for the listener position and the source position or orientation of the sound source, and
[0293] Directivity information is used in the calculation (220) of the two-channel acoustic data of the direct sound portion.
[0294] 2. An audio signal processor according to example 1, wherein the two-channel synthesizer (200) is configured to determine (227, 228) two channels of head-related data from a source position and a listener position or orientation in addition to directionality information, and to use (229) the two channels of head-related data and the directionality information in the calculation of the two-channel acoustic data of the direct sound part.
[0295] 3. An audio signal processor according to example 1 or 2, wherein the two-channel synthesizer (200) is configured to determine (222) emission direction information from a source position vector (423) of the sound source and a listener position vector (422) of the listener and a rotation of the sound source, and to derive the directionality information from a database of directionality information sets, wherein the directionality information sets are associated with specific source emission direction information.
[0296] 4. An apparatus according to example 2 or 3, wherein the two-channel synthesizer (200) is configured to use the source position vector (423) of the sound source and the listener position vector (422) of the listener and the rotation of the listener to derive the arrival direction (421) of the listener position or orientation.
[0297] 5. An apparatus according to any of the preceding examples, wherein the directional information is a directional impulse response or a directional transfer function, or wherein the two head-related data channels are a first head-related impulse response or a first head-related transfer function and a second head-related impulse response or a second head-related transfer function, or wherein the source emission direction information comprises an angle or an index into a database.
[0298] 6. The apparatus according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to determine a directional impulse response as the directional information,
[0299] determining a first head-related impulse response and a second head-related impulse response as two head-related data channels, and
[0300] The directional impulse response and the first head-related impulse response are combined by convolution in the time domain or using frequency-domain multiplication, and the directional impulse response and the second head-related impulse response are combined.
[0301] 7. The apparatus according to example 6, wherein the dual channel synthesizer (200) is configured to
[0302] Performing (261) a padding operation with the directional impulse response and the first head-related impulse response and the second head-related impulse response to obtain a padded function,
[0303] Transforming (262) the padded function into the frequency domain,
[0304] The frequency domain directionality information and the frequency domain head related data channel are multiplied (263) to obtain two frequency domain data channels, and
[0305] The two frequency domain data channels are transformed (264) into the time domain to obtain a time domain data portion of the direct sound portion of the dual channel acoustic data.
[0306] 8. An audio signal processor according to Example 6 or 7, wherein the two-channel synthesizer (200) is configured to adjust (265) the phase of the two-channel acoustic data by removing the phase shift introduced by the convolution, and truncate (266) the phase-adjusted two-channel acoustic data so that the length of the time domain representation of the two-channel acoustic data is equal to the length of the direct sound portion of the single-channel acoustic data describing the acoustic environment.
[0307] 9. The apparatus according to any of the preceding examples,
[0308] wherein the two-channel synthesizer (200) is configured to determine (221) an energy-related measure from the direct sound portion,
[0309] determining (223) an energy-related metric from raw directionality information determined for the listener position or orientation and the source position, and
[0310] The original directionality information is scaled (226) using a scaling value derived (224) from the energy-related metric to yield determined directionality information.
[0311] 10. The apparatus according to any of the preceding examples,
[0312] wherein the two-channel synthesizer (200) is configured to determine (225) distance scaling information from a distance between a source position and a listener position, and
[0313] The distance is taken into account (226) in the calculation of the two-channel acoustic data of the direct sound portion.
[0314] 11. The signal processor according to example 10,
[0315] The dual-channel synthesizer (200) is configured to generate amplified dual-channel acoustic data for the direct sound portion when the actual distance is lower than the distance in the initial case where the single-channel acoustic data has been determined, and to generate attenuated dual-channel acoustic data for the direct sound portion when the actual distance is greater than the distance in the initial case.
[0316] 12. The apparatus according to any of the preceding examples,
[0317] Therein, the dual-channel synthesizer (200) is configured to combine the directional information and the head-related impulse response as head-related channel data by padding (261) the two filters to an increased length, multiplying the two filters in the spectral domain (263), converting (264) the two multiplication results into the time domain, and removing (265) the introduced phase so that the center index of the result is similar to the center index of the direct sound portion of the single-channel acoustic data describing the acoustic environment.
[0318] 13. An audio signal processor according to any of examples 8 or 12, wherein the two-channel synthesizer is configured to apply the distance scaling information (226) to the result of the phase removal in the time domain.
[0319] 14. An audio signal processor according to any one of the preceding examples,
[0320] Wherein the two-channel synthesizer (200) is configured to update the calculation (222) of the direct sound portion more frequently than the calculation (232) of the early reflection portion or the calculation of the late reverberation portion.
[0321] 15. An audio signal processor according to any of the preceding examples, configured to include or access a memory comprising a set of directional information data for a plurality of angles relative to predetermined sound emission directions (430) of sound sources distributed on a cylinder or sphere around the location of the sound source, and
[0322] wherein the two-channel synthesizer is configured to derive (269, 270, 271) a directional information data set having a reference to a directional information data set closest to the sound emission direction from a sound emission direction determined (268) for a listener position and a sound source position and orientation, or to derive (271) two or more directional information data sets having reference information closest to the determined sound emission direction and to interpolate between the two or more directional information data sets to obtain directional information, or
[0323] Directional information is synthesized (272) using the determined sound emission direction and the directionality model of the sound source.
[0324] 16. A method for generating a dual-channel audio signal, comprising:
[0325] Provides single-channel acoustic data describing the acoustic environment;
[0326] synthesizing two-channel acoustic data from the single-channel acoustic data using the listener position or rotation; and
[0327] generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0328] The synthesis includes:
[0329] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part,
[0330] determining (222) directional information of the sound source for the listener position and the source position or orientation of the sound source, and
[0331] Directivity information is used in the calculation (220) of the two-channel acoustic data of the direct sound portion.
[0332] 17. A computer program for performing the method of example 16 when run on a computer or a processor.
[0333] Subsequently, examples of the present invention related to the second aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0334] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0335] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0336] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0337] A sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0338] The two-channel synthesizer (200) is configured to separate (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to process (220, 230, 240) the at least two parts separately to generate two-channel acoustic data for each part.
[0339] wherein the dual-channel synthesizer (200) is configured to split (231) the early reflection portion into a plurality of segments,
[0340] determining (232) a plurality of image source positions representing source positions of reflected sounds,
[0341] associating the image source locations with the segments using a matching operation, wherein the matching operation includes calculating a sound arrival time from each image source to the listener location and associating the image source locations with the corresponding segment having a time delay in the corresponding segment that best matches the sound arrival time of the corresponding image source location (234), and
[0342] Two-channel acoustic data for the direct sound is calculated using the image source positions associated with the segments.
[0343] 2. The audio signal processor according to Example 1,
[0344] The dual-channel synthesizer (200) is configured to determine (232) multiple image source positions using the initial measured initial source positions and initial receiving end positions to generate single-channel acoustic data of the acoustic environment and geometric data of the acoustic environment.
[0345] 3. An audio signal processor according to example 1 or 2, wherein the dual-channel synthesizer (200) is configured to determine the image source position using an image source method that models specular reflections in an acoustic environment.
[0346] 4. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to determine the image source position until a predetermined order is reached, and
[0347] For reflections in a segment that do not have an associated image source position or whose sound arrival times are not within a predetermined range that matches the time delay of the reflections in the segment, random or predetermined direction of arrival data or dual channel head related data is used (235).
[0348] 5. An audio signal processor according to any one of the preceding examples,
[0349] Wherein the two-channel synthesizer is configured to detect significant reflections in the early reflection portion and place a segment around each significant reflection, the segment having a predetermined length corresponding to the length of the head-related impulse response, or to divide the early reflection portion into a regular grid of reflection segments, each segment having a sample count and an overlap with an adjacent segment.
[0350] 6. An audio signal processor according to Example 5, wherein the dual-channel synthesizer is configured to detect significant reflections by comparing (283) a first average energy of each sample in a first window with a second average energy of each sample in a second window, wherein the sample count of the second window is greater than the sample count of the first window, and wherein a significant reflection is determined when the first average energy is greater than the second average by a predetermined amount.
[0351] 7. The audio signal processor according to example 6, wherein the predetermined amount is between 3 dB and 9 dB, or wherein the sample count of the first window is preferably at least 0.25 times smaller than the sample count of the second window.
[0352] 8. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to determine (284) arrival direction information for each segment from a listener position and an image source position associated with the respective segment, and to combine an early reflection portion in the segment with two head-related data channels associated with the arrival direction information to obtain at least a portion of the two-channel acoustic data for the segment.
[0353] 9. An audio signal processor according to any one of the preceding examples,
[0354] wherein the two-channel synthesizer (200) is configured to pad (285) the segments to the length of the two-channel acoustic data in the time domain,
[0355] converting the padded segments to the frequency domain and multiplying the frequency domain padded segments by each channel of the head-related two-channel data in the frequency domain to obtain segmented frequency domain two-channel acoustic data, and
[0356] Transform the segmented frequency domain dual-channel data into the time domain.
[0357] 10. The audio signal processor according to example 9, wherein the binaural synthesizer is configured to remove the introduced phase delay from the two-channel acoustic data in the time domain.
[0358] 11. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to generate two-channel acoustic data for each segment from a combination (239) of a specular portion derived using an image source position associated with the segment and a diffuse portion of the corresponding segment.
[0359] 12. An audio signal processor according to any one of the preceding examples,
[0360] The single-channel acoustic data describing the acoustic environment is a room impulse response or a room transfer function, or the two-channel acoustic data is a binaural two-channel head-related impulse response or a binaural two-channel head-related transfer function.
[0361] 13. An audio signal processor according to any one of Examples 2 to 11, wherein the dual-channel synthesizer is configured to maintain the image source position for a listener position at an initial receiving end position and a listener position different from the initial receiving end position, or for an initial source position or a source position different from the initial source position; or to maintain the association between the segment and the image source position for a source position at the starting source position or a source position different from the initial source position.
[0362] 14. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer is configured to determine directional information of the image sound source for a listener position and an image source position or orientation of the image sound source, and to use the directional information in the calculation (220) of the two-channel acoustic data of the earlier reflected sound portion.
[0363] 15. An audio signal processor according to example 14, wherein the directionality information for each image source is derived from the same set of directionality information determined for the direct sound portion, or wherein the orientation of the image sound source is determined by an image source model.
[0364] 16. An audio signal processor according to example 14 or 15, wherein directionality information is determined and used for a predetermined subset of segments in the early reflection portion.
[0365] 17. The audio signal processor according to example 14 or 15, wherein the predetermined subset of segments in the early reflection part comprises less than ten segments, preferably only 2 segments.
[0366] 18. A method for generating a two-channel audio signal, comprising:
[0367] Provides single-channel acoustic data describing the acoustic environment;
[0368] synthesizing binaural acoustic data from the monophonic acoustic data using the listener position or rotation; and
[0369] Generating a binaural audio signal from the audio signal and binaural acoustic data, wherein the generating comprises:
[0370] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part,
[0371] The early reflection portion is divided (231) into a plurality of segments,
[0372] determining (232) a plurality of image source locations representing source locations of reflected sounds, and
[0373] associating the image source positions with the segments using a matching operation, wherein the matching operation includes calculating a sound arrival time from each image source to the listener position and associating the image source position with a corresponding segment in the corresponding segment having a time delay that best matches the sound arrival time of the corresponding image source position (234), and
[0374] Two-channel acoustic data for the direct sound is calculated using the image source positions associated with the segments.
[0375] 19. A computer program for performing the method of example 18 when run on a computer or a processor.
[0376] Subsequently, examples of the present invention related to the third aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0377] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0378] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0379] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0380] a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0381] Wherein, the dual-channel synthesizer (200) is configured as
[0382] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part,
[0383] Therein, the two-channel synthesizer (200) is configured to calculate (230) two-channel acoustic data of the early reflection portion using a specular portion describing different early reflections and a diffuse portion describing the influence of diffuse reflections in the early reflection portion.
[0384] 2. An audio signal processor according to example 1, wherein the two-channel synthesizer is configured to compute (238) the diffuse reflection portion using a combination of the early reflection portion of the single-channel acoustic data and the two-channel noise sequence.
[0385] 3. An audio signal processor according to example 1 or 2, wherein the two-channel synthesizer (200) is configured to perform a weighted addition (239, 290) of the specular portion (292) and the diffuse portion (293), wherein the weight of the weighted addition is determined by a diffuse reflection coefficient indicating a degree of diffuse reflection of a segment of the early reflection portion of the single-channel acoustic data.
[0386] 4. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to determine the albedo from a ratio of a first energy average value per sample in a first window having a sample count n to a second energy average value per sample in a second window having a sample count m around the first window,
[0387] wherein the portion is considered to be completely specular when the ratio plus the first predetermined number divided by the second predetermined number is 1 or greater, or is considered to be completely diffuse when it is 0 or less, and wherein the second predetermined number is at least 3 dB greater than the first predetermined number, or the second predetermined number has a value within a range between 1.5 times and 2.5 times the value of the first predetermined number.
[0388] 5. The audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer is configured to split the early reflection portion into a plurality of segments and to calculate the specular portion and the diffuse portion for each segment.
[0389] 6. An audio signal processor according to any one of Examples 3 to 5, wherein the weight of the weighted addition is further determined by the position of the segment of the early reflection part relative to the direct sound part and the late reverberation part, thereby enhancing the weight of the specular part of the segment close to the direct sound part, and enhancing the weight of the diffuse reflection part close to the late reverberation part.
[0390] 7. An audio signal processor according to Example 6, wherein weights are determined so that the mirror portion of a segment that is temporally closer to the direct sound portion has a greater weight than the mirror data in a segment that is temporally closer to the late reverberation portion, or so that the diffuse reflection data of a segment that is temporally closer to the direct sequence portion has a lower weight than the mirror data of a segment that is temporally closer to the direct sound portion, or wherein the weights of the mirror data of the segment are determined using diffuse reflection measurements of the segment, and wherein the weights of the diffuse reflection data of the segment are determined using diffuse reflection measurements of the corresponding segment's mirror data.
[0391] 8. An audio signal processor according to any of the preceding examples, wherein the multi-channel synthesizer (200) is configured to
[0392] computing the specular portion in both channels using direction of arrival data of the early reflection portion that depends on the listener position or orientation and the source position and convolution of a head-related data channel associated with the direction of arrival data with the early reflection portion of the mono-channel acoustic data,
[0393] Computing the diffuse portion using a combination of the two-channel binaural noise data and the early reflection portion of the single-channel acoustic data, and
[0394] Combines the specular and diffuse parts.
[0395] 9. The audio signal processor according to example 8, wherein the two-channel synthesizer (200) is configured as
[0396] calculating specular components (292) in a plurality of segments of the early reflection portion to obtain first channel specular data for the plurality of segments and second channel specular data for the plurality of subsegments,
[0397] Calculating diffuse reflection parts in the same plurality of segments of the early reflection parts to obtain first-channel diffuse reflection segment data and second-channel diffuse reflection segment data,
[0398] For each segment, combining the segment's first channel specular data and the segment's first channel diffuse specular data to obtain a first channel of the segment's early reflection data, and
[0399] The segmented second channel specular reflection data and the segmented second channel diffuse reflection data are combined to obtain a second channel of segmented early reflection data.
[0400] 10. The audio signal processor of Example 9, wherein the multi-channel synthesizer (200) is configured to perform a linear combination using a first weighting coefficient and a second weighting coefficient, wherein the first weighting coefficient and the second weighting coefficient added together are substantially equal to one.
[0401] 11. The audio signal processing according to example 9 or 10, wherein the two-channel synthesizer (200) is configured to
[0402] When calculating the first channel specular segmentation data and the second channel specular segmentation data, the overlapping segments of the early reflection part of the single channel acoustic data are windowed using a window function.
[0403] Use a similar window function to window the overlapping segments of the first channel diffuse segmentation data,
[0404] Windowing the overlapping segments of the second channel diffuse segmentation data using a similar windowing function, and
[0405] Weighted addition is performed on the corresponding first-channel specular segmentation data and first-channel diffuse reflection segmentation data and second-channel specular segmentation data and second-channel diffuse reflection segmentation data to obtain two-channel acoustic data of the early reflection part.
[0406] 12. The audio signal processor according to example 11, wherein the two-channel synthesizer is configured to overlap and add the result data of the segmented sequence for each channel to obtain two-channel audio data of the early reflection part.
[0407] 13. An audio signal processor according to Example 12, wherein the two-channel synthesizer is configured to compensate (610) for an initial time delay gap that depends on the source position and the listener position by shifting the result of the segmented overlap-add operation in time relative to the direct sound portion to obtain an early reflection portion that has a timing relationship with the direct sound portion.
[0408] 14. A method for generating a two-channel audio signal, comprising:
[0409] Provides single-channel acoustic data describing the acoustic environment;
[0410] synthesizing binaural acoustic data from the monophonic acoustic data using the listener position or rotation; and
[0411] Generate a two-channel audio signal from an audio signal and two-channel acoustic data,
[0412] The synthesis includes:
[0413] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and processing (220, 230, 240) the at least two parts separately to generate dual-channel acoustic data for each part, and
[0414] Two-channel acoustic data for the early reflection portion is calculated (230) using a specular portion describing the different early reflections and a diffuse portion describing the contribution of the diffuse reflections in the early reflection portion.
[0415] 15. A computer program for performing the method of example 14 when run on a computer or a processor.
[0416] Subsequently, examples of the present invention related to the fourth aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0417] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0418] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0419] a binaural synthesizer (200) for synthesizing binaural acoustic data from monophonic acoustic data using a listener position or rotation; and
[0420] A sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0421] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and processing (220, 230, 240) the at least two parts separately to generate dual-channel acoustic data for each part, and
[0422] Wherein, the dual-channel synthesizer (200) is configured to use the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the first channel noise phase spectrum of the first channel for obtaining the dual-channel acoustic data, and use the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel sound data without the direct sound part or the amplitude spectrum of the late reverberation part and the second channel noise phase spectrum to calculate the dual-channel diffuse reflection part of the early reflection part or the dual-channel diffuse reflection part of the single-channel acoustic data without the direct sound part or the dual-channel diffuse reflection part of the late reverberation part.
[0423] 2. The audio signal processor according to example 1, wherein the first channel noise phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence.
[0424] 3. An audio signal processor according to Example 1 or 2, wherein the dual-channel synthesizer (200) is configured to calculate (530) a first spectrogram of an early reflection portion of single-channel acoustic data, or a first spectrogram of single-channel acoustic data without a direct sound portion, or a first spectrogram of a late reverberation portion of single-channel acoustic data, as well as a second spectrogram of a first channel noise phase spectrum and a third spectrogram of a second channel noise phase spectrum.
[0425] 4. An audio signal processor according to Example 3, wherein the dual-channel synthesizer is configured to calculate the first spectrogram as a first amplitude spectrum sequence, calculate the second spectrogram as a second phase spectrum sequence, calculate the third spectrogram as a third phase spectrum sequence, and combine the first amplitude spectrum sequence and the second phase spectrum sequence to obtain the first channel of the dual-channel diffuse reflection part, and combine the first amplitude spectrum sequence and the third phase spectrum sequence to obtain the second channel of the dual-channel diffuse reflection part.
[0426] 5. The audio signal processor according to example 3 or 4, wherein the second channel synthesizer is configured to use overlapping segments and a window function for each segment in the calculation of the first spectrogram, the second spectrogram and the third spectrogram.
[0427] 6. The audio signal processor according to example 4 or 5, wherein the first spectrogram, the second spectrogram, and the third spectrogram are calculated as complex spectra and converted into polar representations.
[0428] 7. An audio signal processor according to one of Examples 4 to 6, wherein the dual-channel synthesizer is configured to low-pass filter (448) the amplitude spectra of the amplitude spectrum sequence so that the first sequence of low-pass filtered amplitude spectra is combined with the second phase spectrum sequence and the third phase spectrum sequence.
[0429] 8. The audio signal processor according to example 7, wherein the low pass filter is a moving average filter.
[0430] 9. The audio signal processor according to example 8, wherein the moving average filter extends over a size between 0.1 and 0.75 octaves.
[0431] 10. An audio signal processor according to any one of Examples 3 to 9, wherein the multi-channel synthesizer is configured to perform (447) spectrum-by-spectrum low-pass filtering in a first spectral sequence of a first spectrogram, or in a spectrogram of a first channel or a second channel of a diffuse reflection part of a second channel, thereby low-pass filtering frequency intervals of adjacent spectra associated with the same frequency.
[0432] 11. The audio signal processor according to example 10, wherein the low-pass filter used for low-pass filtering is a moving average filter having 2 to 6 inputs.
[0433] 12. An audio signal processor according to Example 4, wherein the two-channel synthesizer (200) is configured to transform (450) the first channel of the two-channel diffuse reflection part and the second channel of the two-channel diffuse reflection part into the time domain to obtain overlapping blocks of the first channel and the second channel.
[0434] 13. An apparatus according to Example 12, wherein the dual-channel synthesizer is configured to perform an overlap and add operation (452) on the overlapping time-domain blocks of the first channel on the one hand, and an overlap and add operation (452) on the overlapping time-domain blocks of the second channel on the other hand to obtain the diffuse reflection part of the dual-channel representation.
[0435] 14. An apparatus according to any of the preceding examples, wherein the two-channel synthesizer is configured to use only the diffuse reflection part in the late reverberation part as the two-channel acoustic data, or to use a combination of the diffuse reflection part and the specular part as the two-channel acoustic data in the early reflection part.
[0436] 15. A method for generating a two-channel audio signal, comprising:
[0437] Provides single-channel acoustic data describing the acoustic environment;
[0438] synthesizing binaural acoustic data from the monophonic acoustic data using the listener position or rotation; and
[0439] generating a binaural audio signal from the audio signal and binaural acoustic data,
[0440] The synthesis includes:
[0441] separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and processing (220, 230, 240) the at least two parts separately to generate dual-channel acoustic data for each part, and
[0442] Using the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the first channel noise phase spectrum of the first channel for obtaining the dual-channel acoustic data, and using the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the second channel noise phase spectrum, calculate the dual-channel diffuse reflection part of the early reflection part or the dual-channel diffuse reflection part of the single-channel acoustic data without the direct sound part or the dual-channel diffuse reflection part of the late reverberation part.
[0443] 16. A computer program for performing the method of example 15 when run on a computer or a processor.
[0444] Subsequently, examples of the present invention related to the fifth aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0445] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0446] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0447] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0448] a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0449] The input interface (100) is configured to obtain (150) a raw representation associated with the single-channel acoustic data and to derive (151) the single-channel acoustic data using the raw representation and additional data stored in the audio signal processor or accessible by the audio information processor.
[0450] 2. The audio signal processor according to Example 1, wherein the input interface (100) is configured as
[0451] obtaining (150) an initial measurement of raw single-channel acoustic data as a raw representation,
[0452] deriving (101) a test fingerprint to access a pre-stored database having a set of associated reference fingerprints, wherein each reference fingerprint is associated with high resolution single channel acoustic data, wherein the high resolution single channel acoustic data has a higher resolution than the initial measurement, and
[0453] High-resolution single-channel acoustic data having a reference fingerprint that best matches the test fingerprint is retrieved (113) from a pre-stored database or synthesized (140) from the test fingerprint, from an initial measurement of raw single-channel acoustic data, or from geometric parameters.
[0454] 3. The audio signal processor according to example 1, wherein the input interface (100) is configured as
[0455] Get the initial measurements of the original single-channel acoustic data as a raw representation,
[0456] derive a test fingerprint, and
[0457] Single-channel acoustic data is synthesized (140) from a test fingerprint or from an initial measurement of raw single-channel acoustic data.
[0458] 4. An audio signal processor according to example 1, wherein the original representation is a geometric description of the acoustic environment, and wherein the input interface (100) is configured to perform an acoustic room simulation to derive single-channel acoustic data from the geometric description.
[0459] 5. The audio signal processor according to example 1, wherein the input interface (100) is configured to determine at least one of the following parameters RT60, EDC, DRR as a test fingerprint, and
[0460] The reference fingerprint includes at least one of the following parameters: RT60, EDC, and DRR.
[0461] 6. An audio signal processor according to any of the preceding examples, wherein the input interface (100) is configured to apply a psychoacoustic weighting function to the calculated fingerprint to obtain a fingerprint for accessing a pre-stored database (110) or for performing direct synthesis (140).
[0462] 7. An audio signal processor according to any one of examples 1 to 3, wherein the input interface (100) is configured to derive the fingerprint using a trained neural network, or to perform direct synthesis (140) from a raw representation associated with the single-channel acoustic data using a trained neural network.
[0463] 8. An audio signal processor according to Example 1, wherein the input interface (100) is configured to calculate a test fingerprint using a trained neural network, wherein the trained neural network is trained to classify single-channel acoustic data into categories of a single room, and wherein the input interface (100) is configured to synthesize (153) prototype single-channel acoustic data for a fingerprint indicating a matching room category, or retrieve (152) prototype single-channel acoustic data for a matching room type from a pre-stored database.
[0464] 9. The audio signal processor according to one of examples 1 to 5, wherein the input interface (100) is configured as
[0465] Deriving a test fingerprint such that the test fingerprint has a lower dimension than the original single-channel acoustic data,
[0466] deriving a lower dimensional reference fingerprint from a pre-stored database using the same process as used to derive the test fingerprint, and
[0467] Single-channel acoustic data with a reference fingerprint that minimizes the distance to the test fingerprint is selected.
[0468] 10. The audio signal processor according to any of the preceding examples, wherein the input interface (100) is configured to use natural sounds that a listener can produce in the initial measurement.
[0469] 11. The audio signal processor according to example 10, wherein the natural sound is applause, or speech, or a transient sound that a listener can produce.
[0470] 12. The audio signal processor according to example 1, wherein the input interface (100) is configured to
[0471] recording (150) a sound clip played by one or more speakers in an acoustic environment,
[0472] using a sound recognition process to determine (155, 156) the identification of the sound segment,
[0473] accessing (157) a database having at least an approximate representation of a sound clip played by one or more speakers without being affected by an acoustic environment, and
[0474] Single channel acoustic data is determined (159) using the recorded sound clips and the sound clips obtained from the database.
[0475] 13. The audio signal processor according to example 8, wherein the input interface (100) comprises a second trained neural network for generating single-channel acoustic data from a test fingerprint calculated by the first trained neural network.
[0476] 14. An audio signal processor according to any of the preceding examples, wherein the input interface (100) comprises a speaker and a microphone embedded in the mobile device, and wherein the input interface (100) is configured to perform initial measurements with the speaker and the microphone or with only the microphone embedded in the mobile device.
[0477] 15. An audio signal processor according to any of the preceding examples, wherein the input interface (100) is configured to receive new single-channel acoustic data at regular intervals or at specific events, compare the new single-channel acoustic data with the single-channel acoustic data, and replace the single-channel acoustic signal data with the new single-channel acoustic data when the deviation exceeds a deviation threshold, or compare the new initial measurement with the earlier initial measurement, or compare the new test fingerprint with the earlier test fingerprint, or compare the new original representation with the earlier original representation.
[0478] 16. An audio signal processor according to any of the preceding examples, wherein the input interface (100) is configured to store a history of earlier single channel acoustic data to allow mixing from the earlier single channel acoustic data to the new single channel acoustic data.
[0479] 17. The audio signal processor according to example 16, wherein the mixing comprises linear interpolation between the earlier single-channel acoustic data and the later single-channel acoustic data in the time domain or the frequency domain.
[0480] 18. A method for generating a dual-channel audio signal, comprising:
[0481] Provides single-channel acoustic data describing the acoustic environment;
[0482] synthesizing two-channel acoustic data from the single-channel acoustic data using the listener position or rotation; and
[0483] generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0484] Therein, the synthesis comprises obtaining (150) an original representation associated with the single-channel acoustic data and deriving (151) the single-channel acoustic data using the original representation and additional data stored in or accessible by the audio signal processor.
[0485] 19. A computer program for performing the method of example 18 when run on a computer or a processor.
[0486] Subsequently, examples of the present invention related to the sixth aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0487] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0488] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0489] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0490] a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0491] wherein the two-channel synthesizer (200) is configured to separate (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to process (220, 230, 240) the at least two parts separately to generate two-channel acoustic data for each part; and
[0492] The two-channel synthesizer comprises two physically separate devices (901, 902), wherein a first device (901) of the two physically separate devices is configured to process (220, 230) at least one of a direct sound portion and an early reflection portion, wherein a second device (903) of the two physically separate devices is configured to process (230, 240) at least one of an early reflection portion and a late reverberation portion, and wherein the first device (901) and the second device (902) are connected via a transmission interface (918, 925) and have separate power supplies (917, 924).
[0493] 2. An audio signal processor according to Example 1, wherein the first device (901) is configured to update the two-channel acoustic data of the direct sound part or the early reflection part more frequently than the second device (902) updates the two-channel acoustic data of at least one of the early reflection part and the late reverberation part.
[0494] 3. The audio signal processor according to example 1, wherein the transmission interface (918, 925) is configured to operate according to a wireless transmission protocol.
[0495] 4. An audio signal processor according to any of the preceding examples, wherein the first device (901) is a wearable device and further comprises an input interface (100) and a sound generator (300), and wherein the second device (102) is a mobile device or a fixed device separate from the wearable device.
[0496] 5. An audio signal processor according to any of the preceding items, wherein the wearable device (901) is an earbud device, a headphone device or an in-ear device, and wherein the mobile device or the fixed device is a mobile phone, a smartwatch, a tablet computer, a laptop computer or a fixed computer.
[0497] 6. An audio signal processor according to any of the preceding examples, wherein the first device (901) comprises a user tracking system (914) and is configured to transmit data of the user's position or direction to the second device (902).
[0498] 7. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer (200) is configured to separate (210) single-channel acoustic data into three parts: direct sound, early reflections and late reverberation,
[0499] wherein the two-channel acoustic data of the direct sound portion is generated by a first device (220), wherein the two-channel acoustic data of the early reflection portion is generated by a second device (902) (230), or wherein the two-channel acoustic data of the late reverberation portion is generated by a third device (903) (240), wherein the third device (903) is separated from the first device (901) and the second device (902).
[0500] 8. An audio signal processor according to Example 6, wherein the second device (902) is a mobile phone capable of accessing the Internet, and wherein the first device (903) is a remote computer connected to the mobile device via the Internet, and wherein the update frequency of the two-channel audio data of the late reverberation part is lower than the two-channel acoustic data of the early reverberation part.
[0501] 9. An audio signal processor according to any of the preceding examples, wherein the second device (902) is configured to receive a user position or orientation from the first device (901), provide two-channel acoustic data of the early reflection and / or late reverberation portion, and transmit the two-channel acoustic data of the early reflection portion and / or late reverberation portion to the first device.
[0502] 10. An audio signal processor according to any of the preceding examples, wherein the second device is configured to receive the user position or orientation and the audio signal from the first device and provide two-channel acoustic data of at least the early reflection portion, and
[0503] wherein the sound generator (300) is distributed to a first device (901) and a second device (902), wherein the first device is configured to generate a two-channel audio signal of a direct sound portion, wherein the second device is configured to generate a two-channel audio signal of at least an early reflection portion, and wherein the third device is configured to transmit the two-channel audio information of the early reflection portion to the first device.
[0504] 11. An audio signal processor according to any of the preceding examples, wherein the first device (901) is configured to delay the two-channel acoustic data of the direct sound portion with a delay value that covers the delay caused by transmission to and from the second device.
[0505] 12. An audio signal processor according to any one of the preceding examples,
[0506] The first device (901) has a memory for storing the second acoustic data of the early reflection part and / or the second acoustic data of the late reverberation part,
[0507] Wherein the two-channel synthesizer (200) or the sound generator (300) is configured to use the stored two-channel acoustic data in the calculation of the complete two-channel sound data when updated two-channel data of the direct sound part is available and updated two-channel acoustic data of the early reflection part or the late reverberation part is not available (933) due to different update rates of the first device (901) and the second device (902).
[0508] 13. The audio signal processor of any of the preceding examples,
[0509] The sound generator (300) is configured to aggregate the two-channel acoustic data of each part to obtain complete two-channel audio data, and combine the complete two-channel acoustic data with the multi-channel audio signal to obtain a two-channel audio signal, or combine the two-channel acoustic data of each part with the input audio signal to obtain a partial two-channel audio signal of each part, and aggregate the partial two-channel video signal to obtain a two-channel audio signal.
[0510] 14. An audio signal processor according to any of the preceding examples, wherein the two-channel analyzer is configured to update the two-channel acoustic data of each portion at a different rate, wherein the direct sound portion is updated more frequently than the remaining portions, or wherein the early reflection portion is updated less frequently than the direct sound portion and more frequently than the late reverberation portion, or wherein the late reverberation portion is updated less frequently than the remaining portions of the two-channel acoustic data of the acoustic environment.
[0511] 15. The audio signal processor of any of the preceding examples,
[0512] wherein the second device comprises a calculator or a reverberator network for generating or processing two-channel audio data of early reflections and / or late reverberation components, or wherein the update rate of the direct sound component is higher than 15 Hz, wherein the update rate of the early reflection component is greater than 5 Hz and lower than 15 Hz, or wherein the update rate of the late reverberation component is greater than 0.5 Hz and lower than 5 Hz.
[0513] 16. A method for generating a dual-channel audio signal, comprising:
[0514] Provides single-channel acoustic data describing the acoustic environment;
[0515] synthesizing two-channel acoustic data from the single-channel acoustic data using the listener position or rotation; and
[0516] generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0517] wherein the synthesizing comprises separating (210) the single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part, and
[0518] The synthesis comprises using two physically separate devices (901, 902), wherein a first device (901) of the two physically separate devices processes (220, 230) at least one of a direct sound portion and an early reflection portion, wherein a second device (903) of the two physically separate devices processes (230, 240) at least one of an early reflection portion and a late reverberation portion, and wherein the first device (901) and the second device (902) are connected via a transmission interface (918, 925) and have independent power supplies (917, 924).
[0519] 17. A computer program for performing the method of example 16 when run on a computer or a processor.
[0520] Subsequently, examples of the present invention related to the seventh aspect are summarized, wherein the reference numerals in parentheses should not be construed as limiting the scope of the examples.
[0521] 1. An audio signal processor for generating a two-channel audio signal, comprising:
[0522] An input interface (100) for providing single-channel acoustic data describing an acoustic environment;
[0523] a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; and
[0524] a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0525] Wherein, the dual-channel synthesizer (200) is configured as
[0526] separating (210) single-channel acoustic data into at least two parts including a direct sound part and at least one of an early reflection part and a late reverberation part, and processing (220, 230, 240) the at least two parts separately to generate dual-channel acoustic data for each part,
[0527] determining (601) a separation moment between a direct sound portion and an early reflection portion or between an early reflection portion and a late reverberation portion in single-channel acoustic data,
[0528] extending (602) at least one of the two parts of the separation instant by a specific number of samples to achieve an overlap at the separation instant, and
[0529] At least one extended portion is windowed using a specific window function that compensates for sample expansion (603).
[0530] 2. The audio signal processor of Example 1, wherein the number of overlapping samples is taken from the corresponding other part.
[0531] 3. An audio signal processor according to example 1 or example 2, wherein the window function is a Tukey window with lobes of width 2n, where n is a certain number of samples, and the two parts each extend by n samples.
[0532] 4. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer is configured to determine a separation instant between the direct sound portion and the early reflection portion such that the distance of the separation instant is substantially centered between a direct sound peak and a first early reflection peak, or to determine a separation instant between the early reflection portion and the late reverberation portion as a perceived mixing time of the acoustic environment, or at a predetermined amount of time before the perceived mixing time.
[0533] 5. The apparatus of any of the preceding examples, wherein the two-channel synthesizer is configured to perform an overlap-add operation on a first channel of the two-channel acoustic data of the direct sound portion, a first channel of the two-channel acoustic data of the early reflection portion, and a first channel of the two-channel acoustic data of the late reverberation portion after separate processing of the corresponding portions (220, 230, 240), and
[0534] The two-channel synthesizer is configured to perform an overlap-add operation on the second channel of the two-channel acoustic data of the direct sound part, the second channel of the two-channel acoustic data of the early reflection part, and the second channel of the two-channel acoustic data of the late reverberation part after separate processing (220, 230, 240) of the corresponding parts.
[0535] 6. An audio signal processor according to any of the preceding examples, wherein the dual-channel synthesizer (200) is configured to pre-process (600) the single-channel acoustic data by detecting a direct sound index in a time representation of the single-channel acoustic data, and clipping or extending a beginning portion of the time representation of the single-channel acoustic data by zero-valued samples so that the detected time index coincides with a predefined sample index offset from the beginning of the single-channel acoustic data.
[0536] 7. An audio signal processor according to any of the preceding examples, wherein the acoustic data describing the acoustic environment is a room impulse response or a room transfer function, or wherein the two-channel acoustic data is a binaural two-channel head-related impulse response or a binaural two-channel head-related transfer function.
[0537] 8. An audio signal processor according to any one of the preceding examples, wherein the sound generator (300) is configured to aggregate the two-channel acoustic data of each portion to obtain complete two-channel audio data, and combine the complete two-channel acoustic data with the input audio signal to obtain a two-channel audio signal, or combine the two-channel acoustic data of each portion with the input audio signal to obtain a partial two-channel audio signal of each portion, and aggregate the partial two-channel audio signals to obtain a two-channel audio signal.
[0538] 9. An audio signal processor according to any of the preceding examples, wherein the two-channel synthesizer is configured to determine (611) a current distance from a listener position to a source position relative to an initial distance of initial generation of the single-channel acoustic data, and to adjust (614) a time period between the two-channel acoustic data of the direct sound portion and the two-channel acoustic data of the early reflection portion so that the time period is extended when the current distance is lower than the initial distance, or shortened when the current distance is greater than the initial distance.
[0539] 10. The audio signal processor according to Example 9, wherein the multi-channel synthesizer is configured to add zero samples to the early reflection part before overlap-adding when the time period is expanded, or to remove excess samples when the time period is shortened.
[0540] 11. An audio signal processor according to example 9 or 10, wherein the multi-channel synthesizer (200) is configured to determine (611) a first initial time delay gap of an initial measurement of single-channel acoustic data provided by the input interface (100), determine (612) a second initial time delay gap of a current listener position and a current sound source position, calculate (630) a difference between the first initial time delay gap and the second initial time delay gap, and adjust (614) the first initial time delay gap by shifting the early reflection portion by the calculated difference.
[0541] 12. An audio signal processor according to example 11, wherein the multi-channel synthesizer is configured to determine (612) a second initial time delay gap by the difference between the propagation time of the first reflection from the image source position associated with the first segment of the early reflection portion to the current listener position and the propagation time from the current sound source position to the current listener position.
[0542] 13. An audio signal processor according to any one of the preceding examples,
[0543] The dual-channel analyzer is configured to store early generated dual-channel acoustic data of the early reflection part and the late reverberation part to which the window function is applied, and retrieve the stored dual-channel acoustic data to perform a channel overlap-add operation with the newly updated other parts of the dual-channel acoustic data (606).
[0544] 14. A method for generating a dual-channel audio signal, comprising:
[0545] Provides single-channel acoustic data describing the acoustic environment;
[0546] synthesizing two-channel acoustic data from the single-channel acoustic data using the listener position or rotation; and
[0547] generating a two-channel audio signal from an audio signal and two-channel acoustic data,
[0548] The synthesis includes:
[0549] separating (210) single-channel acoustic data into at least two parts including a direct sound part and at least one of an early reflection part and a late reverberation part, and separately processing (220, 230, 240) the at least two parts to generate dual-channel acoustic data for each part,
[0550] determining (601) a separation moment between a direct sound portion and an early reflection portion or between an early reflection portion and a late reverberation portion in single-channel acoustic data,
[0551] extending (602) at least one of the two parts of the separation instant by a specific number of samples to achieve an overlap at the separation instant, and
[0552] At least one extended portion is windowed using a specific window function that compensates for sample expansion (603).
[0553] 15. A computer program for performing the method of example 14 when run on a computer or a processor.
[0554] It is mentioned here that all the alternatives or aspects discussed above and all the aspects defined by the following claims or the independent claims in the preceding examples can be used alone, that is, there are no other alternatives or objects other than the alternative, object, example or independent claim considered. However, in other embodiments, two or more alternatives or aspects or examples or independent claims can be combined with each other, and in other implementation examples, all aspects or alternatives or all examples and all independent claims can be combined with each other.
[0555] Although some aspects are described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where the blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of the corresponding blocks or items or features of the corresponding apparatus.
[0556] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, on which electronically readable control signals are stored, which cooperate (or are capable of cooperating) with a programmable computer system to perform the corresponding method. Some embodiments according to the present invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein. Generally, embodiments of the present invention may be implemented as a computer program product having program code that, when executed on a computer, is executable to perform one of the methods. The program code may, for example, be stored on a machine-readable carrier. Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable carrier or non-transitory storage medium. In other words, an embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer. Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) that includes (has recorded thereon) a computer program for performing one of the methods described herein. Therefore, another embodiment of the inventive method is a data stream or signal sequence representing a computer program for executing one of the methods described herein. The data stream or signal sequence can, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet). Another embodiment includes a processing device, such as a computer or a programmable logic device, which is configured to or suitable for executing one of the methods described herein. Another embodiment includes a computer having a computer program for executing one of the methods described herein installed thereon. In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to execute some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can collaborate with a microprocessor to execute one of the methods described herein. Typically, these methods are preferably performed by any hardware device.
[0557] The above embodiments are intended to illustrate the principles of the present invention only. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended that the present invention be limited only by the scope of the claims to be filed, and not by the specific details set forth herein by way of description and explanation of the embodiments.
Claims
1. An audio signal processor for generating a two-channel audio signal, comprising: An input interface (100) for providing single-channel acoustic data describing an acoustic environment; a two-channel synthesizer (200) for synthesizing two-channel acoustic data from single-channel acoustic data using a listener position or rotation; as well as a sound generator (300) for generating a two-channel audio signal from an audio signal and two-channel acoustic data, wherein the two-channel synthesizer (200) is configured to separate (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and to process (220, 230, 240) the at least two parts separately to generate two-channel acoustic data for each part; and Wherein, the dual-channel synthesizer (200) is configured to use the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the first channel noise phase spectrum of the first channel for obtaining the dual-channel acoustic data, and use the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel sound data without the direct sound part or the amplitude spectrum of the late reverberation part and the second channel noise phase spectrum to calculate the dual-channel diffuse reflection part of the early reflection part or the dual-channel diffuse reflection part of the single-channel acoustic data without the direct sound part or the dual-channel diffuse reflection part of the late reverberation part.
2. The audio signal processor according to claim 1, wherein The first channel noise phase spectrum and the second channel noise phase spectrum are derived from a two-channel binaural noise sequence.
3. The audio signal processor according to claim 1 or 2, wherein: The two-channel synthesizer (200) is configured to calculate (530) a first spectrogram of an early reflection portion of single-channel acoustic data, or a first spectrogram of single-channel acoustic data without a direct sound portion, or a first spectrogram of a late reverberation portion of single-channel acoustic data, and a second spectrogram of a first channel noise phase spectrum and a third spectrogram of a second channel noise phase spectrum.
4. The audio signal processor according to claim 3, wherein the dual-channel synthesizer is configured to calculate the first spectrogram as a first amplitude spectrum sequence, calculate the second spectrogram as a second phase spectrum sequence, calculate the third spectrogram as a third phase spectrum sequence, and combine the first amplitude spectrum sequence and the second phase spectrum sequence to obtain the first channel of the dual-channel diffuse reflection part, and combine the first amplitude spectrum sequence and the third phase spectrum sequence to obtain the second channel of the dual-channel diffuse reflection part. 5 . The audio signal processor according to claim 3 , wherein the second channel synthesizer is configured to use overlapping segments and a window function for each segment in the calculation of the first spectrogram, the second spectrogram and the third spectrogram.
6. The audio signal processor according to claim 4 or 5, wherein: The first spectrogram, the second spectrogram, and the third spectrogram are calculated as complex spectra and converted into polar coordinate representations.
7. An audio signal processor according to any one of claims 4 to 6, wherein the dual-channel synthesizer is configured to low-pass filter (448) the amplitude spectrum of the amplitude spectrum sequence so that the first sequence of low-pass filtered amplitude spectra is combined with the second phase spectrum sequence and the third phase spectrum sequence. The audio signal processor according to claim 7 , wherein the low-pass filter is a moving average filter.
9. An audio signal processor according to claim 8, wherein the moving average filter extends over a size between 0.1 and 0.75 octaves.
10. An audio signal processor according to any one of claims 3 to 9, wherein the multi-channel synthesizer is configured to perform (447) spectrum-by-spectrum low-pass filtering in a first spectral sequence of a first spectrogram, or in a spectrogram of a first channel or a second channel of a diffuse reflection part of a second channel, thereby low-pass filtering frequency intervals of adjacent spectra associated with the same frequency. 11 . The audio signal processor according to claim 10 , wherein the low-pass filter used for low-pass filtering is a moving average filter having 2 to 6 inputs.
12. The audio signal processor according to claim 4, wherein The two-channel synthesizer (200) is configured to transform (450) a first channel of the two-channel diffuse portion and a second channel of the two-channel diffuse portion into the time domain to obtain overlapping blocks of the first channel and the second channel.
13. The device according to claim 12, wherein The two-channel synthesizer is configured to perform an overlap and add operation (452) on the overlapping time-domain blocks of the first channel on the one hand and the second channel on the other hand to obtain the diffuse portion of the two-channel representation.
14. The apparatus according to any one of the preceding claims, wherein the two-channel synthesizer is configured to use only the diffuse reflection part in the late reverberation part as the two-channel acoustic data, or to use a combination of the diffuse reflection part and the specular part as the two-channel acoustic data in the early reflection part.
15. A method for generating a dual-channel audio signal, comprising: Provides single-channel acoustic data describing the acoustic environment; synthesizing two-channel acoustic data from single-channel acoustic data using listener position or rotation; as well as generating a two-channel audio signal from an audio signal and two-channel acoustic data, The synthesis includes: separating (210) single-channel acoustic data into at least two parts consisting of a direct sound part and at least one of an early reflection part and a late reverberation part, and processing (220, 230, 240) the at least two parts separately to generate dual-channel acoustic data for each part, and Using the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the first channel noise phase spectrum of the first channel for obtaining the dual-channel acoustic data, and using the amplitude spectrum of the early reflection part or the amplitude spectrum of the single-channel acoustic data without the direct sound part or the amplitude spectrum of the late reverberation part and the second channel noise phase spectrum, calculate the dual-channel diffuse reflection part of the early reflection part or the dual-channel diffuse reflection part of the single-channel acoustic data without the direct sound part or the dual-channel diffuse reflection part of the late reverberation part.
16. A computer program for performing the method of claim 15 when run on a computer or processor.