System for measuring an impulse response
By outputting a repeating sinusoidal sweep with varying receiver locations and applying deconvolution and interpolation, the method efficiently generates full-range impulse responses, addressing the time-consuming nature of conventional IR measurement techniques.
Patent Information
- Application Number
- GB2024002868
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-03
AI Technical Summary
Existing methods for measuring impulse responses require numerous repetitions of outputting a sound source signal, waiting for reverberant reflections, and repositioning the sound source or receiver, which is cumbersome and time-consuming, especially when interpolating IRs in unmeasured locations.
A method involving a sound output device and receiver device that outputs a repeating sinusoidal sweep with varying relative locations of the receiver during signal recording, followed by deconvolution and interpolation to generate full-range impulse responses, reducing measurement time by recording only sub-ranges of the frequency sweep at each location.
This approach allows for efficient reconstruction of full-range impulse responses with reduced overall measurement time, suitable for applications like head-related transfer functions and reverb responses, by varying the relative location of the sound receiver device during the sinusoidal sweep recording process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD OF THE INVENTION The present invention relates to the field of 3D audio. In particular, the invention relates to methods and systems for recording a set of impulse responses. BACKGROUND Impulse Responses (IRs) represent, in the time domain, an acoustic transfer function in the frequency domain between two points in space. One way of measuring an IR of a sound source at a particular location relative to the sound receiver is to output, by the sound source, a Dirac impulse (or approximation thereof), and record the outputted audio with a sound receiver device (e.g. a microphone) positioned at the location to be measured. A perfect Dirac impulse contains an equal amount of all frequencies, all output at the same time. In this way, the signal that is received by the sound receiver device can be determined as representing a transfer function or change in audio between the sound source and the sound receiver device. Alternatively, another way of measuring an IR of a sound source at a particular location relative to the sound source is to output, by the sound source, a sinesweep signal. Unlike a Dirac impulse, which attempts to recreate all frequencies at the same time, this signal contains all frequencies spread over time. This may involve outputting a sinusoidal signal starting at a low frequency and increasing in frequency (typically logarithmically) until the desired frequency range is output. The resulting recording is then deconvolved with an inverse sine sweep which effectively delays and scales the amplitude of each frequency component individually such that the components in the processed signal occur at the same time, thus being equivalent to recording a Dirac impulse. This technique is generally robust to loudspeaker harmonic distortions and other noise sources. In situations which require the measurement of many different IRs, for example when calculating head related transfer functions (HRTFs), it is typical to measure a grid of IRs at locations which are close enough together such that it is possible to interpolate between measurements in post-processing to sufficiently approximate the IR at any position in the space. HRTFs describe the way in which a person hears sound in 3D and can change depending on the position of the sound source and the geometry of their head and upper body. HRTFs are used to calculate a received sound by combining an outputted audio signal with (e.g. multiplying or convolving with) a transfer function corresponding to a particular direction from which the sound source originates. Recording enough impulse responses at different grid locations to sufficiently characterize HRTFs for different sound source locations in three-dimensions usually requires setting up a measurement system that will position a sound source and sound receiver device, output a sine sweep signal (typically anywhere from less than 1 to more than 20 seconds in length) and then reposition either the source or receiver to measure the next location in the grid. Performing this technique results in a discrete set of ‘complete’ I Rs at a number of points in space, where the term complete refers to the fact that each IR represents the transfer function at that particular location over the entire desired frequency range. A problem with the techniques described above is that even when applying interpolation to approximate IRs in locations which are not directly measured, measuring an IR at a large number of different relative locations is still required to obtain accurate results. Each of these measurements involves outputting the source signal, waiting for any reverberant reflections, moving the sound source or receiver device, waiting for any equipment noise to settle, and repeating. Performing many repetitions of these steps can be cumbersome for the user. Accordingly, there exists a need for a solution to mitigate at least some of the problems described above associated with measuring impulse responses. SUMMARY OF INVENTION In a first aspect of the invention there is provided a method of recording an impulse response (IR) using a sound output device and sound receiver device, the method comprising: outputting, by the sound output device, an audio signal comprising a repeating sinusoidal sweep, wherein during each repetition, a frequency of the sinusoid varies progressively within a total frequency range; during each repetition of the sinusoidal sweep of the audio signal: varying a relative location of the sound receiver device with respect to the sound output device; recording, by the sound receiver device, a received audio signal such that, at any particular relative location, a response to a sub-range of the total frequency range of the sinusoidal sweep is recorded; deconvolving the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different relative locations of the sound output and sound receiver devices; and interpolating between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location. The inventors have found that, by varying a relative location of a sound receiver device with respect to a sound output device during the outputting and recording of a sinusoidal sweep signal, full range impulse responses for particular relative locations can be reconstructed from the recorded data. These reconstructed impulse responses approximate those that would have been recorded using conventional measurement techniques where the full range of the sine sweep signal is outputted at each measurement location. However, overall measurement time is reduced as at each measurement location, only a part or sub-range of the outputted sine sweep signal is recorded. Therefore, the invention provides a more efficient mechanism for recording impulse responses. The invention may be applied to a number of different application scenarios where the recording of multiple impulse responses is required. These may be, but are not limited to, recording a set of impulse responses characterizing, in three-dimensional space, head related transfer functions (HRTFs) for different sound source directions, recording a set of reverb responses at different locations within a room, or recording the response of an item of audio equipment (e.g. a speaker or a microphone) at varying locations. The sound output device may be any device suitable for outputting an audio signal, and typically comprises a speaker array with speakers disposed at different positions, thereby allowing a relative location to the sound receiver device to be adjusted by selecting which of the speakers in the array that the audio signal is output through. However, the sound receiver device may alternatively be a single speaker which is preferably movable so as to allow a relative location to the sound receiver device to be adjusted. The sound receiver device may be any device suitable for recording an audio signal, and typically comprises one or more microphones or arrays of microphones. The type of microphone or microphone array may be dependent on the application scenario. For example, for recording impulse responses characterizing HRTFs for different sound source directions, the sound receiver device may be a microphone or microphone array suitable for positioning in one or each ear of a human test subject, or may be a microphone built into one or each ear of a dummy head. The audio signal output by the sound output device may be a sine sweep signal that periodically repeats, each repetition being a frequency sweep typically increasing in frequency over time in a logarithmic fashion. The total frequency range associated with the audio signal may be a difference between a minimum frequency and maximum frequency associated with any sine sweep in the audio signal. Each repetition of the sine sweep in the audio signal may have a constant frequency range, or the frequency range of some repetitions may differ within the total frequency range. The audio signal may be continuously output by the sound output device or may be paused and played as necessary. As such, varying the relative location “during each repetition of the sinusoidal sweep of the audio signal” may refer to both varying the relative location simultaneously to outputting the audio signal, and outputting a portion of a repetition of the sine sweep at a first relative location, and another portion of the repetition of the sine sweep at a second relative location. The varying of the relative location of the sound receiver device with respect to the sound output device may involve displacing either one or both of the sound receiver device and the sound output device. Such a displacement may be translational, angular or be a combination thereof. The phrase “any particular relative location” in the context of recording a received audio signal may refer to any particular relative location amongst the relative locations traversed by varying the relative location of the sound receiver device with respect to the sound output device. The granularity of the relative locations may be determined by a motion path traversed in varying the relative location and a sampling rate at which the received audio signal is recorded. The deconvolving of the recorded audio signal may be interpreted conventionally as an operation which has the effect of delaying and amplitude scaling each frequency component of the recorded audio signal individually such that they ‘line up’, or are synchronised, in time, thus approximating a Dirac impulse. However, due to the variation of the relative location during each repetition of the sinusoidal sweep, the deconvolved signal associated with each sinusoidal sweep no longer represents an IR of a single location as is conventional, but instead comprises a multi-location response signal in which different frequency ranges (referred to as partial IRs) correspond to different relative locations. The interpolation between a combination of the generated multi-location response signals to generate an IR over the total frequency range may be understood as using frequency components corresponding to or closest to a particular relative location to reconstruct an impulse response at the particular location. In some embodiments, the sinusoidal sweep is divided into two or more parts, each part being output at a different relative location of the sound receiver device with respect to the sound output device. In this way, a measurement time to record the audio signal at each relative location is reduced and thus the overall measurement time is reduced, but it is still possible to generate a full range impulse response at a particular relative location utilizing the recorded audio signal as described herein. In other embodiments, the varying a relative location of the sound receiver device with respect to the sound output device is continuous during at least one repetition of the sinusoidal sweep. These represent a limiting case of the embodiments described in the paragraph above, whereby the sinusoidal sweep may be considered as divided into a large number of different parts (for example, 1024 or 2048 frequency bins), each having a short enough measurement time such that the varying a relative location may be performed continuously. In this way, a measurement time to record the audio signal at each relative location is further reduced, as is the overall measurement time. However, it is still possible to generate a full range impulse response at a particular relative location utilizing the recorded audio signal as described herein. In some embodiments, varying a relative location of the sound receiver device with respect to the sound output device comprises varying an angular displacement between the sound receiver device and sound output device. Such a displacement may be performed by rotating either one or both of the sound output and sound receiver devices. In many application scenarios where an audio transfer function between a sound output device and sound receiver device is dependent on one or more angles between them, it is typically most convenient to rotate the sound receiver device with respect to the sound output device. However, it is also possible to rotate the sound output device with respect to the sound receiver device. Typically, varying an angular displacement comprises rotating the sound receiver device about a vertical axis so as to adjust an azimuthal angle relative to the sound output device. Such examples may be suited to application scenarios where an elevation angle between the sound output and sound receiver devices may be changed by other means, for example by adjusting the elevation angle from which the audio signal is output. However, in other embodiments, varying an angular displacement between the sound receiver device and sound output device may comprise rotating the sound output device about a vertical axis so as to adjust an azimuthal angle relative to the sound output device, or rotating one or both of the sound output and receiver devices so as to adjust an elevation angle between them. In some embodiments, the rotating comprises rotating the sound receiver device at least one full rotation about the vertical axis. For example, when recording impulse responses characterizing HRTFs for different sound source directions and the sound receiver device comprises a microphone located at each ear of a real or dummy head, performing a full rotation allows the capture of responses for both ipsilateral and contralateral sound source directions. Typically, the rotation is of a constant angular velocity so as to ensure a consistent distribution of measurement points across the frequency range. However in some examples the speed of rotation may be varied during each repetition of the sinusoidal sweep to vary the distribution of measurement points across the frequency range. Preferably, the sinusoidal sweep repeats at least once every 5 degrees of rotation about the vertical axis. This rate of repetition of the sinusoidal sweep ensures that the recorded audio signal includes sufficient data at each angular displacement so as to be able to generate a full range impulse response at a particular angular displacement utilizing the recorded audio signal as described herein. The rate at which the sinusoidal sweep repeats with rotation, however, may vary with application scenario. For example, for measuring low frequency directivity of a speaker such as a subwoofer, it may be appropriate that the sinusoidal sweep repeats once every 45 degrees or 90 degrees of rotation. To record impulse responses at different relative locations in three-dimensional space, an elevation angle may be varied in addition to the azimuthal angle between the sound output and receiver devices. This applies, for example, when recording impulse responses which characterize HRTFs for different sound source directions. In some embodiments, the method further comprises, at a plurality of different azimuthal angular displacements of the sound receiver device: outputting the audio signal at different elevation angles; recording, by the sound receiver device, a received audio signal; deconvolving the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different azimuthal and elevation angles of the sound receiver device with respect to the sound output device; and interpolating between a plurality of different combinations of the generated multilocational response signals to generate a set of impulse responses characterizing head related transfer functions (HRTFs) for different azimuthal and elevation angles. By additionally outputting the audio signal at different elevations, for example by outputting the audio signal through speakers at different elevations or by adjusting an elevation angle of a speaker, audio signals are recorded such that impulse responses at different relative locations in three-dimensional space can be generated. In some embodiments, the sound output device comprises a plurality of speakers positioned in an arc in a vertical plane, the arc at least partially surrounding the sound receiver device such that each speaker is positioned at a different elevation angle to the sound receiver device. Such an arrangement of speakers provides an efficient way of adjusting an elevation angle by changing which speakers the audio signal is output through at a given time. In these examples, the outputting the audio signal may comprise, for preferably at least one full rotation of the sound receiver device about the vertical axis, one of: outputting the audio signal through only one of the plurality of speakers; and outputting the audio signal through more than one of the plurality of speakers, wherein the audio signal is temporally offset in each speaker. In the first option, the audio signal is output through only one of the speakers at a time such that all components of the signal received by the sound receiver device can be associated with a particular elevation angle. However, in other examples, more than one of the speakers may output the audio signal simultaneously, wherein the audio signals output by each different speaker are temporally offset such that components of the signal received by the sound receiver device can be attributed to different elevation angles. In some embodiments, each speaker may comprise multiple drivers which may be configured, for example by an internal crossover, to output different frequency bands of an outputted audio signal. In these embodiments, it may be considered that different frequency bands of the audio signal are output from different elevation angles according to the positioning of each driver with respect to the sound receiver device. In some embodiments, the method further comprises storing a plurality of grids of measurement points, each grid associated with a sub-range of the total frequency range, wherein each measurement point in a grid is a recorded response for the sub-range at a different relative location. A grid representation of recorded responses is typical in situations where response measurements are performed close enough together in space such that interpolation can be performed between the measurements in post-processing to sufficiently approximate the IR from any position in space. By varying the relative location of the sound output and receiver devices during each repetition of the sinusoidal sweep, discrete measurement grids may be formed associated with different frequency sub-ranges. To compute a full range IR at a particular location, one or more (typically three) measurement points from each grid closest to the particular location are interpolated together. Specifically, in some examples, interpolating between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location comprises interpolating between measurement points of each grid closest to the particular relative location to generate response signal portions corresponding to the frequency sub-ranges of each grid; and combining the generated response signal portions to generate the IR over the total frequency range. In some examples, one or more of the repetitions of the sinusoidal sweep is of a narrower frequency range than the other repetitions, preferably wherein the narrower frequency range has a higher minimum frequency. The inventors have found that, in certain circumstances, an audio transfer function varies between different relative locations of the sound output and sound receiver devices less for lower frequencies in the audible range than it does for higher frequencies. Consequently, measurement efficiency can be gained by outputting and recording response signals corresponding to higher frequencies more frequently than lower frequencies. In a specific example, for each pair of consecutive repetitions of the sinusoidal sweep, one of the repetitions is of a narrower frequency range than the other repetitions, preferably wherein the narrower frequency range has a higher minimum frequency. For human HRTF impulse responses, the minimum frequency in these examples may be 1 kHz or higher. However, in other application scenarios, a frequency below which an audio transfer function begins to vary less significantly between different relative locations of the sound output and sound receiver devices may be lower or higher than 1kHz. In some embodiments, where the method further comprises storing a plurality of grids of measurement points, a stored grid associated with a higher frequency range comprises a greater density of measurement points than a stored grid associated with a lower frequency range. Such a differing density of measurement points for lower and higher frequency ranges may be due to the different frequency at which the lower frequencies are output in repetitions of the sinusoidal sweep as described above, or may be obtained by discarding measurement points in one or more lower frequency range. In some embodiments, the method further comprises calculating, based on a time of output of each sub-range of the audio signal, a time of recording a corresponding response, and the variation of relative location of the sound receiver device with respect to the sound output device, an estimated relative location of the sound output at the time of recording; and associating each estimated relative location with the corresponding frequency sub-range in each multi-location response signal. Since there is a time difference associated with the output of a part of the audio signal by the sound output device and the receipt of that part of the audio signal at the sound receiver device, an accuracy of associating components of the outputted audio signal with a particular relative location can be improved by calculating and taking into account this time delay. This results in an improved accuracy of impulse responses generated from the recorded response signals. It is, however, considered that some margin of error is acceptable for the calculation of this time which is dependent based on the rate at which the relative location is varied. In some embodiments, the method further comprises transforming each multilocation response signal into the frequency domain, wherein the frequency domain representation comprises a plurality of frequency bins, each frequency bin being associated with a different relative location. In the frequency domain representation of the multi-location response signals, different frequency bands within the signal can be conveniently associated with different relative locations of the sound output and receiver devices. In some embodiments, the interpolation comprises a triangulation method, for example, Delaunay triangulation. In some embodiments, the sound output device and sound receiver device are in an anechoic environment. In application scenarios where it is advantageous to minimise any effect of the surrounding environment on the recorded response signals due to reflections, an anechoic environment such as an anechoic or semi anechoic chamber may be used to perform the method. In a second aspect of the invention there is provided a system for recording an impulse response (IR), the system comprising: a sound output device; a sound receiver device; and a processor configured to: control the sound output device to output an audio signal comprising a repeating sinusoidal sweep, wherein during each repetition, a frequency of the sinusoid varies progressively within a total frequency range; control the sound receiver device to, during each repetition of the sinusoidal sweep of the audio signal, record a received audio signal whilst a relative location of the sound receiver device with respect to the sound output device varied such that, at any particular relative location, a response to a subrange of the total frequency range of the sinusoidal sweep is recorded; deconvolve the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different relative locations of the sound output and sound receiver devices; and interpolate between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location. For advantages associated with the second aspect of the invention, these may be considered as in line with those described above in relation to the first aspect of the invention. In some embodiments, the sound output device comprises a plurality of speakers positioned, in use, in an arc in a vertical plane, the arc at least partially surrounding the sound receiver device in use such that each speaker is positioned at a different elevation angle to the sound receiver device. In some embodiments, the sound receiver device comprises at least one of: a dummy head microphone; and an in-ear microphone for positioning in the ear of a test subject in use. The dummy head microphone may include multiple microphones, each positioned in either a left ear or a right ear of the dummy head. The sound receiver device may, in some examples, comprise two in-ear microphones, one positioned in each ear of a test subject in use. In some embodiments, the system further comprises a rotating platform configured, in use, to rotate the sound receiver device about a vertical axis, wherein the processor is further configured to control the rotation of the rotating platform so as to adjust an azimuthal angle of the sound receiver device relative to the sound output device. The rotating platform may be motorised and / or comprise a rotary encoder, for example in the form of a stepper motor with an integrated encoder, which is controllable by the processor to rotate the sound receiver device to a chosen azimuth angle at a selected angular velocity. Although in the above-described embodiments the audio signal comprises a repeating sinusoidal sweep, in other examples, the audio signal may alternatively comprise a single sinusoidal sweep. This may be understood as the audio signal comprising a sinusoidal sweep that typically increases in frequency over time in a logarithmic fashion and terminates at a maximum frequency. An audio signal comprising a single sinusoidal sweep may be suitable in instances where the sound output device comprises an array of speakers positioned at different elevation angles with respect to the sound receiver device, and wherein each speaker is offset in along an azimuthal direction. In these examples, different portions of the sinusoidal sweep may be output through different speakers so as to adjust a relative location of the sound output device with respect to the sound receiver device. It will be appreciated that elements of the first aspect described above equally apply to the second aspect, along with their associated advantages. BRIEF DESCRIPTION OF DRAWINGS Embodiments of the invention are described below, by way of example only, with reference to the accompanying drawings, in which: Figure 1 schematically illustrates a side view of a test subject and system for measuring impulse responses; Figure 2A schematically illustrates a top-down view of a test subject and system for measuring impulse responses, wherein the test subject is positioned at an azimuth angle of zero with respect to a sound output device; Figure 2B schematically illustrates a top-down view of a test subject and system for measuring impulse responses, wherein the test subject is positioned an azimuth angle 0 with respect to a sound output device; Figure 3 schematically illustrates a method for recording an impulse response; and Figure 4 schematically illustrates a method for recording a set of impulse responses corresponding to HRTFs for different azimuthal and elevation angles. DETAILED DESCRIPTION The description of Figures 1, 2A and 2B provided below are explained with reference to a spherical coordinate system having an origin at the centre of the head of a test subject 20, and defined by a radial distance from the origin, an azimuth angle and an elevation angle. Figure 1 schematically illustrates a side view of the test subject 20 and a system for measuring impulse responses. The illustrated system may be employed for measuring a set of impulse responses which characterize HRTFs at different sound source locations. Typically, such a measurement system is set up in an environment which is anechoic such as an anechoic or semi anechoic chamber so as to minimise the influence of sound reflections on the measured responses. The measurement system shown in Figure 1 may be used to measure HRTFs specific to a person, for example when the test subject 20 is a real person, or may be used to measure generalised HRTFs, for example when the test subject 20 is a model or dummy head having an idealised geometry. In other (non-illustrated) examples of the invention, a system for measuring impulse responses may be provided suitable for other application scenarios where multiple impulse responses are to be recorded. These may be, but are not limited to, recording a set of reverb responses at different locations within a room, or recording the response of an item of audio equipment (e.g. a speaker or a microphone) at varying locations. As illustrated in Figure 1, the test subject 20 is positioned such that it is partially surrounded by a plurality of speakers 10, each of which form part of a sound output device for the measurement system. In this example, the speakers 10 are positioned in an arc shape in a vertical plane such that each speaker is positioned at a constant radial distance and at a different elevation angle with respect to the test subject 20. For example, the speaker 10 labelled in Figure 1 is positioned at a radial distance r and at an elevation angle <p from the test subject 20. The number of speakers 10 and / or the extent to which the speakers 10 surround the test subject 20 are not limited to those illustrated in the examples herein. For recording audio signals indicative of the response of the test subject to sound sources originating from different directions, the test subject 20 is provided with a sound receiver device 21 / 22, which is typically a microphone or array of microphones suitable for positioning in an ear of the test subject. When the test subject 20 is a real person, the sound receiver device 21 / 22 may be an in-ear microphone or microphone array suitable for positioning at a location proximal to an ear canal of the person. When the test subject 20 is a dummy head, the sound receiver device 21 / 22 may be a microphone or microphone array suitable for positioning at a location proximal to a modelled ear or ear canal of the dummy head, or may be a microphone or microphone array built into the dummy head at a location corresponding to an ear canal. In Figure 1, the sound receiver device 21 is shown as located at a position of a left ear of the test subject 20. However, as shown in Figures 2A and 2B, the sound receiver device may additionally comprise a sound receiver device 22 equivalent to the sound receiver device 21 but located in a right ear of the test subject 20. As is further illustrated in Figure 1, the test subject 20 is positioned such that it can be rotated about a vertical axis passing through the test subject 20, or in other words, such that the azimuth angle to a particular speaker 10 is varied with the rotation, but the elevation angle to a particular speaker 10 remains constant. Rotation of the test subject 20 about the vertical axis may be provided by way of a rotating platform (not illustrated). The rotating platform may be motorised and comprise a rotary encoder, for example in the form of a stepper motor with an integrated encoder, which is controllable to rotate the test subject 20 to a chosen azimuth angle and typically at a constant angular velocity. However, it is contemplated that in other examples of the invention, the sound receiver device may alternatively be rotatable about the vertical axis so as to provide an equivalent variation of the azimuthal angle. For recording a set of impulse responses which correspond to HRTFs, for example using the system illustrated in Figures 1, 2Aand 2B, a measurement process is described in the flow diagram illustrated in Figure 4. Elements of the measurement process may be controlled by a processor (not-illustrated) of the system illustrated in Figures 1,2Aand 2B. The process begins at step S201 by outputting, through at least one of the speakers 10, an audio signal comprising a repeating sinusoidal sweep. In some examples, the audio signal is output through only one of the speakers 10 at a time such that all components of the signal received by the sound receiver device 21 / 22 can be associated with a particular elevation angle. However, in other examples, more than one of the speakers 10 may output the audio signal simultaneously, wherein the audio signals output by each different speaker are temporally offset such that components of the signal received by the sound receiver device 21 / 22 can be attributed to different elevation angles. In all of the above examples, each speaker 10 may comprise multiple drivers which may be configured, for example by an internal crossover, to output different frequency bands of an outputted audio signal. In these examples, it may be considered that different frequency bands of the sinusoidal sweep are output from different elevation angles according to the positioning of each driver of the speaker with respect to the sound receiver device. For simplicity, the following examples assume that the audio signal is output through a single driver of only one of the speakers 10 at a given time. Each repetition of the sinusoidal sweep of the audio signal includes a sinusoid with a frequency that varies progressively within a total frequency range. The total frequency range may be, for example, an audible frequency range between 20 Hz and 20 kHz or a sub-range therein. Referring to Figures 2A and 2B, for a speaker 10 located at a fixed elevation angle <p, the audio signal received at the sound receiver device 21 at the left ear of the test subject can be modelled as a representation of the outputted audio signal modified by a frequency-dependent filter Similarly, the signal received at the sound receiver device 22 at the right ear of the test subject can be modelled as a representation of the outputted audio signal modified by a frequencydependent In an anechoic environment, the frequency dependent filters are dependent on geometric features of the test subject 20 such as the shape of the head, ears, ear canal, density of the head, size and shape of nasal and oral cavities. Moving to the next step of the flow diagram in Figure 4, during each repetition of the sinusoidal sweep of the audio signal, the (azimuthal) angular displacement between the speaker 10 outputting the audio signal and the test subject 20 is varied by virtue of rotating the test subject 20 about the vertical axis. During the repetitions of the sinusoidal sweep, the rotation of the test subject 20 may be continuous or may be intermittent such that defined parts of the sinusoidal sweep are output at different angular displacements between the sound receiver device 21 / 22 and speaker 10. The sound receiver device 21 / 22 then records, at step S203, a received audio signal such that, at any particular angular displacement, a response to a subrange of the total frequency range of the sinusoidal sweep is recorded. In one illustrative example, for an audio signal comprising a sinusoidal sweep between 20 Hz and 20 kHz, the sweep may be split into two halves: 20-10kHz and 10 kHz to 20 kHz, with the first half being recorded at a first azimuthal angle, and the second half being recorded at a second azimuthal angle. Alternatively, the sinusoidal sweep may be split into a large number of different parts such that each part is recorded by the sound receiver device 21 / 22 at a different azimuthal angle and the rotation of the sound receiver device 21 / 22 is continuous and typically at a constant angular velocity. This rotation of the sound receiver device 21 / 22 is illustrated, for example, in Figures 2A and 2B. Figure 2A schematically illustrates a top-down view of a test subject and system for measuring impulse responses, wherein the test subject 20 is positioned at a azimuth angle of zero with respect to the speaker 10. For simplicity, only a single speaker 10 at an elevation angle <p of the speakers illustrated in Figure 1 is shown. At a stage of the rotation where the azimuth angle is zero, the sound receiver device 21 / 22 receives a frequency filtered version of a portion of the sinusoidal sweep signal, wherein the filtering is represented by the functions / iL(0) at the left ear and hR(0) at the right ear. At a stage of the rotation where the azimuth angle is at a value 9, the sound receiver device 21 / 22 receives a frequency filtered version of a different portion (frequency sub-range) of the sinusoidal sweep signal, and the frequency filtering is represented by the functions hL(0) at the left ear and hR(0) at the right ear which are different to the functions hL(0) and hR(0). Typically, the audio signal and angular velocity of the rotation of the test subject 20 are chosen so as to ensure sufficient response data is recorded at each frequency sub-range such that a full range impulse response can be generated at any angular displacement through interpolation. This may involve ensuring that during the outputting of the audio signal, the test subject 20 rotates at least one full rotation about the vertical axis so as to capture both ipsilateral and contralateral data. Furthermore, the rate at which the sinusoidal sweep signal changes in frequency may be chosen in view of the intended angular velocity of the rotation so that the sinusoidal sweep signal repeats at least once every 5 degrees of rotation about the vertical axis. Once the required rotation and number of repetitions of the sinusoidal sweep in the audio signal are complete, the process moves to step S204 in which the recorded audio signal or signals are deconvolved to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different angular displacements of the speaker and sound receiver device 10. By performing steps S201-S204 as described above, the multi-location response signals correspond to different azimuthal angular displacements of the test subject 20, but to the same elevation angle as in this example, the audio signal was output from a single speaker 10. Consequently, to generate multi-location response signals associated with different elevation angles, step S205 involves repeating steps S201-S204, each time outputting the audio signal at a different azimuth angle with respect to the sound receiver device. In the illustrated example of Figure 1, this would typically involve changing the speaker 10 which outputs the audio signal so as to adjust the elevation angle. However, an equivalent adjustment to the elevation angle could be achieved by either moving the position of a single speaker 10 or rotating the test subject 20 about an axis perpendicular to the vertical axis. After obtaining multi-location response signals over a plurality of different azimuthal and elevation angles, the final step of the process S206 is to interpolate between a plurality of different combinations of the generated multilocational response signals to generate a set of impulse responses characterizing head related transfer functions (HRTFs) for different azimuthal and elevation angles. As described elsewhere herein, performing a process such as that illustrated in Figure 4 involves a reduced measurement time in comparison to the conventional mechanism of recording full range impulse responses at each desired combination of azimuthal and elevation angles. Figure 3 schematically illustrates a method for recording an impulse response according to the invention in a similar manner to that described above in relation to Figure 4. At step S101, a sound output device outputs an audio signal comprising a repeating sinusoidal sweep, wherein during each repetition, a frequency of the sinusoid varies progressively within a total frequency range. During each repetition of the sinusoidal sweep, steps S102 and S103 are performed, where S102 involves varying a relative location of a sound receiver device with respect to the sound output device and S103 involves recording, by the sound receiver device, a received audio signal such that, at any particular relative location, a response to a sub-range of the total frequency range of the sinusoidal sweep is recorded. At step S104, the recorded audio signal is deconvolved to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different relative locations of the sound output and sound receiver devices. The multi-location response signals recorded by the sound receiver device may, in some examples, be transformed into the frequency domain, wherein the frequency domain representation comprises a plurality of frequency bins, each frequency bin being associated with a different relative location. Finally, at step S105, interpolation is performed between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location. The interpolation may be performed using conventional techniques such as Delaunay triangulation. For both the processes illustrated in Figure 3 and in Figure 4, the multilocation response signals may be stored as a plurality of grids of measurement points, each grid associated with a sub-range of the total frequency range, wherein each measurement point in a grid is a recorded response for the sub-range at a different relative location. From this grid representation, the computation of a full range IR at a particular relative location involves interpolating between measurement points of each grid closest to the particular relative location. The density of a grid may be indicative of the amount of information stored for a given sub-range of frequencies within the total frequency range. In this way, the more measurement points which are stored in each grid, the higher the accuracy of a given full range IR interpolated from the measurement points. However, it has been found by the inventors that a reducing a density of measurement points in some grids may have a larger effect on the accuracy of a resultant computed IR than others. In some application scenarios, lower frequency portions of transfer functions may vary less with angular displacement than high frequency portions. Consequently, a density of measurement points in grids associated with lower frequencies (e.g. less than 1 kHz in some examples) may be lower than measurement points in grids associated with higher frequencies without observing a significant reduction in the quality of the resulting interpolated IRs, but at the same time reduces the amount of memory required to store the grids. Based on the observation above, efficient recording of the audio data can be performed to result in higher density grids at higher frequencies and lower density grids at lower frequencies. Specifically, one or more of the repetitions of the sinusoidal sweep may be output with a narrower frequency range than the other repetitions, preferably wherein the narrower frequency range has a higher minimum frequency. By outputting higher frequencies in the audio signal more frequently than lower frequencies, grids associated with higher frequencies are populated more densely with measurement points. One particular example of this is that every other repetition of the sinusoidal sweep is output having only a high frequency sub-range of the total frequency range. However, instead of modifying the output audio signal to output higher frequencies more frequently, memory may also be saved by discarding measurement points associated with lower frequency grids such that a stored grid associated with a higher frequency range has a greater density of measurement points than a stored grid associated with a lower frequency range. To further improve an accuracy of the computed impulse responses, it is advantageous to know accurately the exact relative location of the sound receiver device with respect to the sound output device at each part or frequency subrange of the outputted audio signal. However, because there is a delay between the sound output device outputting the signal and the sound receiver device receiving the signal, associating a part of the outputted audio signal with a relative location accurately may involve accounting for this delay. In some examples, based on a time of output of each sub-range of the audio signal, a time of recording a corresponding response, and the variation of relative location of the sound receiver device with respect to the sound output device, an estimated relative location of the sound output device at the time of recording may be calculated and associated with a corresponding frequency sub-range in each multi-location response signal. However, such an uncertainty in the relative location for a given frequency sub-range in the audio output signal may be minimised by lowering one or more of a rate at which the sinusoidal sweep varies in frequency, and an angular velocity of the rotation of the sound receiver device. Having described aspects of the disclosure in detail, it will be apparent that modifications and variations are possible without departing from the scope of aspects of the disclosure as defined in the appended claims. As various changes could be made in the above methods and products without departing from the scope of aspects of the disclosure, it is intended that all matter contained in the above description and shown in the accompanying drawings shall be interpreted as illustrative and not in a limiting sense.
Claims
1. A method of recording an impulse response (IR) using a sound output device and sound receiver device, the method comprising:outputting, by the sound output device, an audio signal comprising a repeating sinusoidal sweep, wherein during each repetition, a frequency of the sinusoid varies progressively within a total frequency range;during each repetition of the sinusoidal sweep of the audio signal:varying a relative location of the sound receiver device with respect to the sound output device;recording, by the sound receiver device, a received audio signal such that, at any particular relative location, a response to a sub-range of the total frequency range of the sinusoidal sweep is recorded;deconvolving the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different relative locations of the sound output and sound receiver devices; andinterpolating between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location.
2. The method according to claim 1, wherein the sinusoidal sweep is divided into two or more parts, each part being output at a different relative location of the sound receiver device with respect to the sound output device.
3. The method according to claim 1 or claim 2, wherein the varying a relative location of the sound receiver device with respect to the sound output device is continuous during at least one repetition of the sinusoidal sweep.
4. The method according to any preceding claim, wherein varying a relative location of the sound receiver device with respect to the sound output device comprises varying an angular displacement between the sound receiver device and sound output device.
5. The method according to claim 4, wherein varying an angular displacement between the sound receiver device and sound output device comprises rotating the sound receiver device about a vertical axis so as to adjust an azimuthal angle relative to the sound output device.
6. The method according to claim 5, wherein the rotating comprises rotating the sound receiver device at least one full rotation about the vertical axis.
7. The method according to claim 3, wherein the rotation is of a constant angular velocity.
8. The method according to any one of claims Error! Reference source notfound, to 7, wherein the sinusoidal sweep repeats at least once every 5 degrees of rotation about the vertical axis.
9. The method according to any one of claims Error! Reference source not found, to 8, further comprising:at a plurality of different azimuthal angular displacements of the sound receiver device:outputting the audio signal at different elevation angles;recording, by the sound receiver device, a received audio signal;deconvolving the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different azimuthal and elevation angles of the sound receiver device with respect to the sound output device; andinterpolating between a plurality of different combinations of the generated multilocational response signals to generate a set of impulse responses characterizing head related transfer functions (HRTFs) for different azimuthal and elevation angles.
10. The method according to any one of claims Error! Reference source not found, to 9, wherein the sound output device comprises a plurality of speakers positioned in an arc in a vertical plane, the arc at least partially surrounding thesound receiver device such that each speaker is positioned at a different elevation angle to the sound receiver device.
11. The method according to claim 10, wherein the outputting the audio signal comprises, for preferably at least one full rotation of the sound receiver device about the vertical axis, one of:outputting the audio signal through only one of the plurality of speakers; andoutputting the audio signal through more than one of the plurality of speakers, wherein the audio signal is temporally offset in each speaker.
12. The method according to any preceding claim, further comprising: storing a plurality of grids of measurement points, each grid associated with a sub-range of the total frequency range, wherein each measurement point in a grid is a recorded response for the sub-range at a different relative location.
13. The method according to claim 12, wherein interpolating between a combination of the generated multi-location response signals to generate an IR over the total frequency range for a particular relative location comprises interpolating between measurement points of each grid closest to the particular relative location to generate response signal portions corresponding to the frequency sub-ranges of each grid; and combining the generated response signal portions to generate the IR over the total frequency range.
14. The method according to claim 13, wherein one or more of the repetitions of the sinusoidal sweep is of a narrower frequency range than the other repetitions, preferably wherein the narrower frequency range has a higher minimum frequency.
15. The method according to claim 14, wherein for each pair of consecutive repetitions of the sinusoidal sweep, one of the repetitions is of a narrower frequency range than the other repetitions, preferably wherein the narrower frequency range has a higher minimum frequency.
16. The method according to any one of claims 12 to 15, wherein a stored grid associated with a higher frequency range comprises a greater density of measurement points than a stored grid associated with a lower frequency range.
17. The method according to any preceding claim, further comprising: calculating, based on a time of output of each sub-range of the audio signal, a time of recording a corresponding response, and the variation of relative location of the sound receiver device with respect to the sound output device, an estimated relative location of the sound output at the time of recording; and associating each estimated relative location with the corresponding frequency sub-range in each multi-location response signal.
18. The method according to any preceding claim, further comprising transforming each multi-location response signal into the frequency domain, wherein the frequency domain representation comprises a plurality of frequency bins, each frequency bin being associated with a different relative location.
19. The method according to any preceding claim, wherein each sinusoidal sweep in the audio signal increases in frequency logarithmically with time.
20. The method according to any preceding claim, wherein the interpolation comprises a triangulation method.
21. The method according to any preceding claim, wherein the sound output device and sound receiver device are in an anechoic environment.
22. A system for recording an impulse response (IR), the system comprising: a sound output device;a sound receiver device; anda processor configured to:control the sound output device to output an audio signal comprising a repeating sinusoidal sweep, wherein during each repetition, a frequency of the sinusoid varies progressively within a total frequency range;control the sound receiver device to, during each repetition of the sinusoidal sweep of the audio signal, record a received audio signal whilst a relative location of the sound receiver device with respect to the sound output device varied such that, at any particular relative location, a response to a sub-range of the total frequency range of the sinusoidal sweep is recorded;deconvolve the recorded audio signal to generate, for each repetition of the sinusoidal sweep, a multi-location response signal comprising a frequency response in which different frequency sub-ranges of the total frequency range are associated with different relative locations of the sound output and sound receiver devices; andinterpolate between a combination of the generated multilocation response signals to generate an IR over the total frequency range fora particular relative location.
23. The system according to claim 22, wherein the sound output device comprises a plurality of speakers positioned, in use, in an arc in a vertical plane, the arc at least partially surrounding the sound receiver device in use such that each speaker is positioned at a different elevation angle to the sound receiver device.
24. The system according to claim 22 or claim 23, wherein the sound receiver device comprises at least one of:a dummy head microphone;an in-ear microphone for positioning in the ear of a test subject in use.
25. The system according to any one of claims 22 to 24, further comprising a rotating platform configured, in use, to rotate the sound receiver device about a vertical axis, wherein the processor is further configured to control the rotation of the rotating platform so as to adjust an azimuthal angle of the sound receiver device relative to the sound output device.