How to get the location of a sound source
The method addresses the challenges of determining sound source positions in spatial audio recording by using advanced audio signal processing techniques, achieving accurate and flexible results across diverse applications without the need for expensive hardware.
Patent Information
- Application Number
- JP2024570354
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-31
- Filing Date
- 2023-05-31
- Publication Date
- 2025-06-05
AI Technical Summary
Existing spatial audio recording technologies, such as spatial audio microphones, face challenges in accurately determining the position of sound sources relative to a reference point, especially in noisy or reverberant environments, and are often expensive and hardware-dependent.
A method that uses a combination of audio signal processing techniques, including cross-correlation and phase transformation, to accurately determine the position of a sound source relative to a dedicated reference point, both in distance and angle, and is scalable and flexible across various hardware platforms.
The method achieves accurate and flexible determination of sound source positions, enabling high-quality spatial audio recordings across various applications, including podcasts, movies, live events, and virtual reality, without the need for expensive hardware.
Smart Images

Figure 2025517542000001_ABST
Abstract
Description
[Technical field]
[0001] This application claims priority to Danish patent application DK PA 2022 70280, filed May 31, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present invention relates to a method for obtaining the position of a sound source relative to a dedicated reference point.The present invention also relates to a computer system and a non-transitory computer readable storage medium.The present invention further relates to a sound recording device. [Background technology]
[0003] Sound field or spatial audio systems and formats (such as Ambisonics or Dolby Atmos) provide encoded sound information associated with a given sound scene. Such an approach allows for the assignment of positional information to sound sources within a sound scene. These techniques are already known in certain computer games, where recorded sounds are attached to the positions of game objects, but also in the live capture of events, for example the capture of large orchestras or sporting events. The number of possible applications is therefore enormous, ranging from the above-mentioned immersiveness, for example by giving the impression of participating in a sporting event, to virtual or augmented reality experiences.
[0004] Recording sound for such applications is often a challenge in itself using spatial audio microphones. Although these are useful for capturing live sound field information from specific points in space, they have some technical limitations as they are based on beamforming technology and are generally considered expensive. For example, sound quality may be degraded for people located at a large distance from the microphone. In noisier or more reverberant situations, or when multiple people are speaking, it is difficult to identify and separate individual sound sources for the purposes of equalization or other processing techniques.
[0005] On the other hand, audio content creators also recognize the need for high quality audio, including the use of spatial audio information, to improve the quality of audio recordings or to add additional sound effects that enhance the listener's immersion. Therefore, there is a need for a lower cost solution that achieves the benefits and advantages of advanced spatial audio microphones. This solution should preferably be hardware agnostic and flexible for use in a variety of scenarios. Summary of the Invention [Means for solving the problem]
[0006] The present disclosure with its proposed principles provides a method, computer system, and sound recording device that achieves some of the benefits and advantages mentioned above.
[0007] The inventors have found a method that provides for accurate determination of the position of a sound source relative to a dedicated reference point, both in distance and angle. The proposed method is scalable to various levels of quality, largely independent of the hardware used. However, the use of specific dedicated hardware significantly increases the functionality and resolution of the method. Furthermore, the method allows for offline and real-time processing. Thus, the proposed method can be included in a variety of applications, including, but not limited to, sound capture and processing for podcasts, movies, live or other events, audio and teleconferencing, virtual reality, video game applications, etc.
[0008] In one embodiment, the inventors propose a method for determining the position of a sound source relative to a dedicated reference point. In this regard, the expression "position" includes the distance from the sound source to the dedicated reference point, the angle based on one or two axes passing through the reference point, or a combination thereof.
[0009] The method obtains a first audio signal recorded by a microphone at a sound source or at a known position relative to the sound source. Similarly, a number of second audio signals are recorded at positions in known relationship to a dedicated reference point. This reference point may be defined, for example, by dedicated hardware having multiple microphones. The first audio signal and the number of second audio signals are synchronized in time.
[0010] Typically, it is assumed that the first audio signal is recorded close to the audio source. This means that the sound emitted by the audio source is recorded at a level higher than reflections, reverberation, and background noise due to the proximity between the microphone and the audio source, and that the distance is relatively small compared to the distance between the audio source and a dedicated reference point. However, the term "at the audio source" should not be understood in a very restrictive sense. Rather, this expression includes and allows that there is a certain distance between the actual audio source and the microphone. In other words, the location of the microphone relative to the actual audio source is well known. Similarly, the multiple second audio signals are recorded at different locations whose distances and angles relative to the reference point are known. Temporal synchronization is important for the method proposed in the following steps. Such temporal synchronization can be achieved in some examples by providing a common time base for any recorded audio signals. In some other examples, the recorded audio signals can be used to provide the time base, for example by timely correlating a dedicated start signal included and recorded in the first audio signal and the multiple second audio signals.
[0011] A modified cross-correlation, such as a generalized cross-correlation or a phase transform, is then calculated for each time frame between the first audio signal and at least one of the plurality of second audio signals to obtain at least one generalized cross-correlation or phase transform signal for each frame of the recorded audio signal. The length of the frame is typically adjustable and may be adjusted during estimation, for example when there is an indication that the sound source is moving.
[0012] The generalized cross-correlation or phase-transformed signals are then used to estimate the distance between the sound source and the dedicated reference point, the distance estimation being performed by estimating a time delay between the first audio signal and at least one of the plurality of second audio signals using the at least one phase-transformed signal.
[0013] The angle between the sound source and the dedicated reference point is estimated by a weighted least mean square evaluation of the time delay between each pair of the plurality of second audio signals, the weighted least mean square evaluation being dependent on the obtained phase shift signal between the first audio signal and each pair of the plurality of second audio signals.
[0014] The calculation of the phase transformation is performed in some aspects by correlating the first audio signal with at least one of the plurality of second audio signals in the frequency domain to obtain at least one correlation signal. The power spectrum is then normalized and the at least one correlation signal is transformed to the time domain. A Discrete Fourier Transformation (DFT) and an inverse DFT may be used, and in some examples, a Short-Time Fourier Transformation (STFT) and an inverse STFT or an ISTFT may be used, among others.
[0015] When a short-time Fourier transform (STFT) is used, the STFT is performed on at least one of the first audio signal and the plurality of second audio signals to obtain each spectrum. Then, a cross spectrum in each spectrum is obtained, and a spectral mask filter is applied to the obtained cross spectrum. After the application, an inverse short-time Fourier transform (ISTFT) is performed to obtain at least one phase-transformed signal. The above-mentioned mask filter can be estimated based on a signal-to-noise ratio in each frequency bin of the first audio signal. For example, a quantile filter, in particular a median filter, can be used to smooth the power spectrum for each time slice of the power spectrum derived from the first audio signal. The noise is estimated for each time slice depending on the previous time slice. Then, a filter parameter is set to 1 or 0 depending on whether the signal-to-noise ratio exceeds a predetermined threshold.
[0016] In some other aspects, the filter that acts on the signal-to-noise ratio in each frequency bin of the first audio signal can be estimated by using a residual signal from the noise reduction process as a noise estimate. The noise reduction process can optionally be based on machine learning.
[0017] In some embodiments, the time resolution may be increased by interpolation to improve the accuracy of the distance and angle. One approach is to upsample the first audio signal and at least one of the plurality of second audio signals before correlating the first audio signal and at least one of the plurality of second audio signals in the digital domain. Another approach is cubic interpolation. This can be done by upsampling the first audio signal and at least one of the plurality of second audio signals before a discrete or short-time Fourier transform, or by converting the first audio signal and at least one of the plurality of second audio signals from the frequency domain back to the time domain using a higher sampling frequency for the IDFT and ISTFT, respectively. Thus, in some examples, the frequency domain signal may be converted back to the time domain and to the digital domain at a conversion frequency higher than the conversion frequency used in the conversion step of the first audio signal and at least one of the plurality of second audio signals.
[0018] In some examples, each phase transformation signal can be calculated between the first audio signal and each of the multiple second audio signals, which can then estimate the distance from the audio source to each of the locations of the second audio signals, allowing further statistics and thereby improving accuracy.
[0019] Some further aspects relate to the step of estimating the time delay. For this purpose, it is proposed to search for a maximum in at least one phase-transformed signal, or alternatively to detect a first magnitude value that exceeds a given threshold and search for the maximum within a specified window centered on the first magnitude value. The specified window may be suitable when there is a possibility of crosstalk between different microphones that record various first audio signals, or when the microphones that record the audio signals are located further away from the actual sound source and sound reflections are present.
[0020] In other words, the specified window centered on the first intensity value provides a solution to suppress the recorded reflections from the audio signal, thereby reducing the risk of estimating distance or angle with a false positive result. The length of the specified window can be set for a limit to be inversely proportional to the signal bandwidth estimated from the highest frequency component of the first audio signal. Another limit can be the range of expected early reflections depending on the distance between the recording location of the first audio signal and the location of one or more sound sources. In some examples, the length of the specified window can be proportional to the maximum time of flight between the locations of the multiple second audio signals.
[0021] In some further examples, the dedicated reference point is approximately midway between recording locations of the multiple second audio signals, and the estimation of the distance between the sound source and the dedicated reference point uses an average value of a set of time delays between the first audio signal and each of at least one of the multiple second audio signals.
[0022] Some other aspects relate to estimating angles including azimuth and elevation angles. The weighted least mean squares is dependent on the acquired phase transform signal between a pair of the first audio signal and a plurality of second audio signals when the magnitude value of the acquired phase transform signal exceeds a given threshold and is within a specified window centered on the first intensity value.
[0023] The method proposed herein allows not only to estimate the angle and distance between a sound source and a reference point, but also to estimate the angle and distance between two microphones, for example two microphones attached to several speakers and spatially separated. Both microphones record an audio signal from the sound source. This audio signal is denoted as a first audio signal (recorded at the first microphone) and a further first audio signal (recorded at the location of the second microphone). In such a case, a phase transformation may be calculated between the first audio signal and the further first audio signal recorded at or associated with one or more sound sources to obtain a further phase transformation signal. Then, the distance between the position of the microphone (recording the first audio signal) and the position associated with the recording of the further first audio signal can be calculated by estimating the time delay between the first audio signal and the further first audio signal using the further phase transformation signal.
[0024] This proposed aspect provides a simple tool for calculating the distance between locations associated with two or more first audio signals. This is not only useful for estimating the possibility of crosstalk between two or more microphones (recording the first audio signals), including classifying audio signals as source signals or crosstalk based on whether the time delay is negative or positive, but also provides information about the relative distance between the microphones that can be used for post-processing to create location estimates. Thus, this approach can be used to obtain information about sound sources that are far away from the locations where the two (or more) first audio signals are recorded.
[0025] Other aspects relate to post-processing and the identification of movement of sound sources during processing. For fixed sound sources, the distance may not change between different frames (except for possible fluctuations due to estimation). However, if the sound source is moving slowly, the distance and angle will change over time. Such sound sources may be difficult to identify because moving sound sources affect the STFT by Doppler shift. Furthermore, the estimated noise can be identified as a moving source of two or more sound sources located at different positions (or vice versa).
[0026] To accommodate this observation, one embodiment proposes applying a noise reduction filter to the estimated distance and / or the estimated angle. Additionally or alternatively, a Kalman filter can be applied to the estimated distance and / or the estimated angle, respectively, to predict the likelihood of movement and its reduced result. In some examples, such filtering is performed by applying a gradient or divergence to the estimated distance and / or the estimated angle.
[0027] It has been found that a particular arrangement of the microphones recording the second audio signals may be suitable. Possible reflections or errors may be more easily identified and may cancel each other out. Thus, each position of a pair of the multiple second audio signals may be located at the same distance from the dedicated reference point and on a virtual line passing through the dedicated reference point.
[0028] It is useful to position the microphone that records the second sound signal in a dedicated location. For example, the multiple second sound signals may include four audio sound signals, two of which are recorded with a maximum spatial distance of a few centimeters. This distance is usually small enough to avoid simultaneous erroneous recording of direct and reflected sounds of the same sound source, while being large enough to provide sufficient difference when cross-correlating the first and second sound signals without using excessive upsampling.
[0029] The speed of sound through a material depends on its temperature. For accurate measurements, air temperature is measured, especially in the vicinity of multiple second sound sources. Such measurements can be repeated periodically to compensate for temperature changes during the recording session. Distances and angles can then be estimated depending on the derived air temperature, which changes the speed of sound in air.
[0030] In some further examples, a computer system is provided that includes one or more processors and a memory. The memory is coupled to the one or more processors and includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the proposed method and its various steps as described above. Similarly, a non-transitory computer-readable storage medium may be provided that includes computer-executable instructions for performing the method of any one of the preceding claims.
[0031] Another aspect relates to a recording device having a rectangular parallelepiped shape with a bottom surface, a top surface, and four sides. The recording device is adapted to be placed with its bottom on any substantially flat surface, such as a floor, a table, etc. The device may have a height slightly greater than its width or depth. In particular, the width and depth are similar or equal. The recording device also comprises a user interface accessible on the top surface. The user interface may comprise one or more buttons, displays, switches, etc., to provide information to the user and allow the user to interact with the device to realize the device's functionality. In this regard, the recording device may include a processor adapted to read and execute user instructions. Furthermore, in some examples, the processor is configured to process one or more audio signals, at least in part, using aspects of the principles proposed herein.
[0032] The recording device also comprises a number of microphones, in particular omnidirectional microphones, a pair of microphones arranged on each of the sides, a first microphone of the pair of microphones arranged at the top of each side and a second microphone of the pair of microphones arranged at the bottom of each side.
[0033] The distance between the first and second microphones of each pair of microphones is not set equal to the distance between the first microphones on the adjacent sides, in other words, the two adjacent microphones are spaced the same distance apart from each other.
[0034] In some embodiments, the distance from the first microphone to the top surface is greater than the distance from the second microphone to the bottom surface. In some other embodiments, the outer dimensions of the recording device can be slightly larger than the outer dimensions of the two opposing microphones, i.e., the microphones are slightly offset within the recording device. [Brief description of the drawings]
[0035] Further aspects and embodiments based on the proposed principles will become apparent in connection with the various embodiments and examples described in detail in conjunction with the accompanying drawings. [Figure 1] FIG. 1 illustrates an embodiment of the proposed method showing some processing steps for determining the location of a sound source. [Diagram 2] FIG. 1 illustrates the steps of a frequency-weighted phase transformation applying a spectral mask to obtain a filtered correlation signal. [Figure 3A] FIG. 1 illustrates a recording environment with multiple microphones for recording more complex sound field scenarios. [Figure 3B] FIG. 1 illustrates an embodiment of a sound field microphone embodying some aspects of the proposed principles. [Figure 4] FIG. 2 illustrates a process flow of a method according to some aspects of the proposed principles. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0036] The following embodiments and examples disclose different aspects and their combinations according to the proposed principles. The embodiments and examples are not necessarily drawn to scale. Similarly, different elements may be shown with increased or decreased size to highlight individual aspects. It goes without saying that the individual aspects of the embodiments and examples shown in the figures can be combined with each other without further ado without contradicting the principles of the present invention. Some aspects show conventional structures or shapes. It should be noted that in practice, slight differences or deviations from the ideal form may occur, but still do not contradict the idea of the present invention.
[0037] Additionally, the individual figures and aspects are not necessarily drawn to size, nor are the proportions between the individual elements necessarily correct in nature. Some aspects are emphasized by showing them enlarged. However, terms such as "upper", "lower", "larger", "smaller", etc. are properly expressed with respect to the elements in the drawings. Thus, by referring to the figures, the specific relationship between each element can be understood.
[0038] Fig. 3A shows an application of the method according to the proposed principle. This scenario corresponds to a typical audio recording session, where multiple audio signals are recorded to obtain the sound field of a scene. In this example, a speech recording of a natural person is used, but it can be understood that the method and principle disclosed herein are not limited to speech processing or locating the position of a natural person. Moreover, it can be used to localize any dedicated sound source relative to a reference point.
[0039] The scene includes two sound sources depicted as P1, P2, which in this embodiment are two people having a conversation in an at least partially enclosed space. Each person holds a microphone M1, M2, respectively, close to his or her body. Alternatively, the microphones M1, M2 are attached to each person's chest or body. Thus, it can be imagined that the microphones M1, M2 are at the location of each sound source. A number of second microphones M3, M4 are located at a position B1. The position B1 is also defined as a reference point. Thus, the people P1, P2 are each located at a certain distance and angle relative to the reference point B1, and spaced apart from each other. A wall W is located on one side and generates a reflection when each sound source P1, P2 speaks.
[0040] The microphones M1, M2, M3, and M4 are synchronized in time with each other, i.e., the sound recording in this scenario is done with a common time base. When recording a conversation, microphone M1 records the speech of person P1 and, with some delay, the speech of person P2. Similarly, microphones M3 and M4 record the speech of people P1 and P2 with some delay, depending on the speed of sound and the distance of person P1 from reference point B1. The delay varies depending on the distance, but in any case, the direct path from the sound source to one of microphones M3 and M4 is called the direct sound.
[0041] Now, assuming there is only a single sound source P1, the distance can be easily calculated using the direct sound, i.e. to the reference point B1, by measuring the time delay between the sound signal recorded by microphone M1 and one of the sound signals recorded by microphones M3 or M4, and multiplying by the speed of sound.
[0042] Since the speed of sound depends on temperature, a temperature sensor T1 is located near the microphones M3, M4 to measure the air temperature and compensate for the effects of temperature changes. The above scenario is very simple and not suitable for real-world scenarios. In this case, the wall W reflects a part of the speech, which is then recorded by microphone M1 at a relatively low value, but also by microphones M3, M4 with a slight delay, which may be at a relatively high level. Microphone M4 also records the speech. In some scenarios, reflected sound speech is superimposed on the ongoing speech. Due to possible constructive interference or other effects, the recording of the indirect reflected sound may be at a higher level than the direct sound. In more complex scenarios, a second sound source also provides a sound signal at the same time, resulting in the superposition of several different sound signals. Some of these different sound signals originate from sound sources P1, P2, and some of these different sound signals are reflections on the walls.
[0043] The present application aims to process the recorded signals in such a way that each sound source can be located and localized relative to a reference point.
[0044] Another application that also addresses the problem of associating specific location information with a sound source is in virtual reality (VR) applications, which typically include a 360° stereoscopic video signal with several objects in the virtual environment, some of which are associated with objects that correspond to sounds.
[0045] These objects (both visual and audio) are presented to the user, for example, via binocular and stereo headphones, respectively. The binocular headphones can track the position and orientation of the user's head (e.g., using an IMU / accelerometer) so that the video and audio played to the headphones and earphones, respectively, can be adjusted to maintain the illusion of virtual reality. For example, at any given time, only a portion of the 360° video signal is displayed to the user, which corresponds to the user's current field of view in the virtual environment. As the user moves or turns his head, the portion of the 360° signal displayed to the user changes to reflect how the user's view of the virtual world changes with that movement. Similarly, as the user moves, sounds emanating from different locations in the virtual scene may be subject to adaptive filtering of the left and right headphone channels, simulating the frequency-dependent phase and amplitude changes of sounds that occur in real life due to spatial offsets between the ears and the human head and scattering of the upper body.
[0046] Some VR works consist only of computer-generated images and separately pre-recorded or synthesized sound. However, it is becoming increasingly common to create "live-action" VR recordings using a camera with a 360° field of view and multiple microphones capturing the sound field. The recorded sound from the microphones is then processed in a manner according to the proposed principles and aligned with the video signal to produce a VR recording that can be played through headphones or earphones as described above.
[0047] Another application that addresses the problem of associating specific location information with an audio source is in Next Generation Audio (NGA) applications, which typically involve audio objects that contain metadata such as location.
[0048] These objects (both visual and audio) are presented to the user, for example, via head-tracked stereo headphones with binaural rendering. As a binocular headset, such headphones can track the orientation of the user's head (e.g., using an IMU / accelerometer), so that the audio played to the headphones can be adjusted to maintain the illusion of audio immersion. For example, as the user moves or turns his head, sounds emanating from different locations in the virtual scene, or in a scene recorded using this innovation, can be subject to adaptive filtering of the left and right headphone channels, simulating the frequency-dependent phase and amplitude changes of sounds that occur in real life due to spatial offsets between the ears and the human head, and scattering of the upper body.
[0049] Returning to Fig. 3A, the figure shows an embodiment of an audio recording device according to some aspects of the invention suitable for recording multiple audio signals used in the proposed method, in particular an Ambisonics microphone designed for Multiple Input and Multiple Output (MIMO) beamforming targeting directivity corresponding to spherical harmonics basis functions.
[0050] The audio recording device is formed as a rectangular parallelepiped, which has a certain dimension and is suitable for recording audio files. The cube shape also allows a display or user interface to be placed on the recording device so that the bottom of the device can be placed on a suitable surface and still be easily operated. The device can be placed on a stand by a screw at the bottom.
[0051] The eight microphones of the sound recorder are arranged in an octahedral configuration, i.e. in the centers of the faces of an octahedron. Beamforming (so-called ambisonics B-format conversion) comprises a weighted sum that depends on the spherical harmonic basis functions and the microphone configuration, and a set of filters applied to the beamformed signals adapted to the scattering of the sound recorder to achieve a flat frequency response. For wavelengths longer than the physical dimensions of the cube, the acoustic scattering can be approximated as a hard sphere. Therefore, the filters can be adapted to this approximation at low frequencies and simplified.
[0052] The surface of the recording device creates scattering that has the effect of preventing destructive interference when a signal of the same wavelength as the device is recorded by two microphones on opposite sides.
[0053] Thus, the recording device of Figure 3B has eight omnidirectional microphones, four of which, 1A, 1B, 1C, 1D, are located at the top, one microphone on each side, and likewise, four microphones, 2A, 2B, 2C, 2D, are located at the bottom, one microphone on each side. Since there are no microphones on the bottom or top, the cuboid can be placed on a surface, leaving space for a user interface.
[0054] The microphone tubes are placed in each recess so that they are slightly offset toward the center. The distance d between two adjacent microphones, e.g., 1A and 1B or 2A and 2B, is equal. In other words, adjacent microphones are equidistant from each other.
[0055]
number
[0056] is the upper frequency limit for spatial aliasing according to the Shannon criterion for any pair of adjacent microphones with a distance d.
number
[0057] Reference is now made to Fig. 1, which shows the various blocks of the method according to the proposed principle. For simplicity, the method is described using the above scenario of Fig. 3A and Fig. 3B. The method is suitable for post-processing of pre-recorded audio signals, but also for post-processing of real-time audio signals, for example during audio conferences, live events, etc. The method starts with providing one or more first audio signals and a number of second audio signals in blocks BM1, BM2, respectively. The recorded audio signals preferably have the same digital resolution, including the same sample frequency (for example 14 bits at 96 kHz). If different resolutions or sampling frequencies are used, it is appropriate to resample the various audio signals to obtain signals with the same resolution and sampling frequency.
[0058] The upper part of the figure including elements 3’, R1, 30A, and 31 relates to the identification of crosstalk that can occur between two or more first audio signals, i.e., audio signals recorded by microphones whose positions are to be determined. As described above, in block BM1, not only reflections but also direct sounds are recorded by two microphones. In order to determine which of the two or more microphones is actually positioned at each sound source, the signals recorded by the two microphones are filtered and cross-correlated to obtain the time difference in the cross-correlation.
[0059] For this purpose, both signals are processed using frequency-weighted generalized cross-correlation or phase transformation (3’). In the first step, each of the first signals is transformed into the frequency domain to obtain a time-frequency spectrum using the STFT. The filter of the spectral mask is first derived from the spectrum by generating a smooth power spectrum S(l,k) of the audio signal from the microphone, where l is the audio signal and k is each frame of the audio signal. For each frequency bin, a first-order filter estimates the noise n(l,k) in the current frame based on the previous frame. The overall noise n(l,k) is represented by the following equation.
[0060] n(l,k)=(1-α)log(S(l,k))+(n(l,k-1)) α Here, α takes different values depending on S(l,k)<log(n(l,k-1)). Therefore, the filter mask is 1 when the SNR exceeds a specific threshold and 0 otherwise. The result is different filter masks associated with each of the two first signals. In the next step, two pairs of the first signals are cross-correlated, and a cross-spectrum is generated by normalizing the result of the cross-correlation. Next, each estimated filter is applied to the normalized cross-spectrum, and an inverse STFT is performed to obtain a filtered correlation signal (see reference R1 for the sign). In this regard, the created cross-spectrum R xyFor , we use a filter Fx (for signal x) and the cross spectrum R yx It should be noted that for , a filter Fy (for signal y) is used. The filtered correlation signal is then used to estimate the signed time difference or delay of the direct sound at both microphones recording the first audio signal. The sign shown in block 31, i.e. dt>0 or dt<0, provides information which microphone is closer to the actual sound source. This microphone (and audio signal) is then associated with each sound source and the corresponding filter mask.
[0061] The above steps can be omitted if the association of the audio signals to each sound source is defined, i.e., only one first signal is recorded. Returning to FIG. 3, blocks 3, R2 to 35 shown at the bottom explain the various steps of estimating the distance and angle to the reference point. Block BM3 includes a number of second audio signals recorded by one or more second microphones whose location is fixed relative to the reference point. The second audio signals are recorded by a recording device, in this embodiment a total of eight second audio signals are provided. The locations of each of the second microphones are slightly different so that the angle can be obtained later, but are close enough so that effects such as reflections from walls can be determined and filtered. With regard to the method proposed herein, it is not necessary to fully evaluate all of the second audio signals in order to obtain the position. Rather, it has been shown that it is sufficient to evaluate the second audio signals recorded by the four microphones closest to the sound source. To obtain information about which of the audio signals is closest, various options can be used.
[0062] First, because all microphones are synchronized in time, the second audio signal can be used to identify the microphone that recorded the sound at the earliest time by estimating the correlation of the second audio signal between pairs of opposing microphones. Alternatively, blocks 3 and 30B can be used to evaluate the arrival of the sound at each microphone, as described further below.
[0063] The processing here is similar to that described for processing two or more first audio signals, except that in block 3, the first audio signal (the audio signal for which the distance and angle are to be determined) is now cross-correlated with at least one of the eight second audio signals. Block 3 can be performed for each of the second audio signals, providing eight filtered cross-correlated signals overall (see, for example, reference R2). Alternatively, as described above, the first audio signal is correlated with four second audio signals recorded by the microphones closest to the sound source.
[0064] Figure 2 shows the frequency-weighted phase transformation FW-PHAT in an exemplary embodiment. The two input signals are transformed into the frequency domain using the SFTF, and then the cross spectrum is derived therefrom. After normalizing the spectrum, the previously estimated filter is applied (in this case the filter of the spectral mask associated with the first audio signal). The result is then transformed back into the time domain using the inverse SFTF.
[0065] The time delay in blocks 30B, 30A is estimated by first identifying the maximum that the peak would have if the signals in the frequency-weighted PHAT were uncorrelated. For this, the variance of the noise is given by sigma=mean(mask) / framesize and the maximum of the noise is derived by sqrt(sigma*2*ln(framesize)). Then, in the frequency-weighted PHAT, we search for the first value above this maximum (which may include some headroom) and re-search for a local maximum close to the first value. The location of the maximum corresponds to the time-of-flight of the direct sound (n_max / sampling frequency). The distance is then given by the time-of-flight multiplied by the speed of sound, taking into account the temperature dependence of the speed of sound, which is assumed to be 20°C if the ambient temperature is not measured. The process in block 30B is repeated for each of the cross-spectra. In block 32, the various results are further processed using the average of the set of time-of-flight estimates. Then, in block 33, the distance is subtracted from this estimate.
[0066] In some embodiments, the time difference of arrival between a spot microphone and the center of the array for a distance (radius) to the sound source can be utilized, which is expressed as the vector m i and a l The time difference of arrival between the m i is the position of the first microphone that records the first audio signal, and a l is the location of microphone element l in the recording device for a total of L (L=8). The time difference can be obtained from a previous evaluation of the correlation signal, or it can be based on the time difference obtained by evaluating a sum of the correlation signals.
[0067] Blocks R3, 30C, and 34-36 are used to obtain the angle between the sound source and the reference point. To avoid the effects of room reflections, in block R2 a window function is used to truncate the FW-PHAT result of the first filtered correlation signal. As shown in R2 in FIG. 2, the window function has a width that depends on the distance between the second microphones. Since the second microphones that record the second sound signal are slightly spaced apart, the estimated distance between the sound source and each second microphone may also change. The width of the window function for truncating the first filtered correlation signal is approximately proportional to the maximum time-of-flight between the second microphones.
[0068] Here, a truncated set of filtered correlation signals can be upsampled to provide finer time resolution, and thus a more accurate estimate of the angle can be obtained.
[0069] The angle estimate is x i Location due to a sound source located at
number
number
number
[0070] This shows that the angle of arrival (both azimuth and elevation) can be estimated by trilateration when the source is far from the array compared to the array baseline (plane wave assumption).
number
number
[0071]
number
number
number
number
[0072] From the matrix, we can find a linear equation using the above matrix and proportional relationships, and the vector A (where
number
[0073] The array forms a tetrahedron with linearly independent rows, but other shapes, such as an octahedral shape, can also be used to provide orthogonal Cartesian coordinate axes.
number
number
number
number
number
number
[0074] Here, W is the component
number
number
number
[0075] This means:
number
number
[0076] The process flow of the method for determining distance and angle according to the proposed principle is shown in Figure 4. The method is suitable for real-time processing as well as for offline processing of multiple previously recorded sound signals forming a sound field.
[0077] The method includes, in step S1, obtaining a first audio signal recorded at a sound source whose distance and angle to a reference point needs to be determined. In step S2, a plurality of second audio signals are recorded in the immediate vicinity of the reference point or at least at a known location or position relative to the reference point. The first audio signal and the plurality of second audio signals are synchronized in time. Such time synchronization may be achieved by referencing all audio signals during the recording session to a common time base.
[0078] The various signals are then optionally pre-processed in step S3. For example, the recorded audio signals can be denoised or equalised to improve the results in the subsequent processing steps. However, care must be taken not to disturb the timing of the signals. Also, in some instances, it is useful to apply methods that preserve the phase information of the recorded signals during the pre-processing step S3. Furthermore, an STFT is performed on the first audio signal and each of the second audio signals. In step S3', the correlation between pairs of second audio signals is evaluated. Here, a pair of audio signals corresponds to signals recorded by two opposing microphones. The correlation determines the subset of microphones (at least four microphones) that are closest to the sound source. These microphones record each audio signal first. These second audio signals are marked to be used later in step S5.
[0079] In this embodiment, there is only a single first audio signal associated with a single sound source. The first audio signal is processed in step S4 by estimating a filter, in particular a filter of a spectral mask. The filter acts on the signal-to-noise ratio at each frequency of the first audio signal in the time domain. The resulting spectral mask contains a set of "1"s and "0"s for each frequency bin.
[0080] In step S5, the first audio signal is correlated in the frequency domain with each of the marked second audio signals of the plurality of second audio signals previously identified in step S3' to obtain at least one correlation signal. To obtain one or more filtered correlation signals, the cross-correlation may be normalized before applying the filter estimated in step S4.
[0081] So far, the steps for determining a distance or an angle are similar.
[0082] Here, the determination of the distance between the reference point and the sound source will be continued with the steps S6 to S8. Step S6 includes obtaining a first timing value in at least one filtered correlation signal that exceeds a dedicated threshold in the time domain. Then, in step S7, a second timing value corresponding to the threshold in at least one filtered correlation signal is obtained based on the first timing value. Both steps S6 and S7 may use the search for a maximum value mentioned above in the PHAT signal (i.e., the filtered correlation signal). The distance between the dedicated reference point and the sound source is based on the first and second timing values obtained, respectively, in step S8. It should be noted that the air temperature may be taken into account. In case of a pre-recorded signal, this information is stored and used in S9 to compensate for the effect of temperature affecting the speed of sound.
[0083] Step S9 is performed to derive and estimate the angle of the sound source from the reference point. More precisely, the angle between the sound source and the dedicated reference point can be determined by weighted least mean square evaluation of the time delay between each pair of the multiple second sound signals. The weighted least mean square evaluation depends on the acquired phase transformation signal between the first sound signal and the previously acquired subset of the second sound signal. For this reason, step S9 is performed multiple times. The angle of arrival (both azimuth and elevation) can be estimated by trilateration, in the usual case when the sound source is far from the array compared to the baseline of the array. The timing difference between the first sound signal and the different subsets of the second sound signals can be represented by a correlation result represented by PHAT. Alternatively, cross-correlation can be used before step S10.
[0084] In any case, the collection of timing differences (indicating the distance between the first microphone and each of the four second microphones in the recorder closest to the sound source) is processed to form an observation vector. The timing difference is somewhat proportional to the distance between two adjacent microphones. Considering the eight microphone recording apparatus presented herein, under the assumption of plane waves, a vector with 12 components can be formed. Thus, any subset of the vectors between pairs of microphones within the octahedron forms a vertical Cartesian coordinate axis. A sound source at a particular location is closest to four of those microphones.
[0085] Step S10 deals with one further aspect of processing audio signals that move over time. For example, if there are multiple first microphones, an active speaker detection algorithm may be used to identify the current active speaker and identify the first microphone associated with the current active speaker. For moving audio signals, a dynamic model and Kalman filtering can be utilized to estimate the location of the sound source at different times. The Kalman filter tracks the estimated state of the system and the variance or uncertainty of the estimate. The estimate is updated using the state transition model and measurements.
Claims
1. A method for obtaining the location of a sound source relative to a dedicated reference point, comprising the steps of: obtaining a first audio signal recorded by a microphone of or associated with one or more sound sources; obtaining a plurality of second audio signals, each recorded at a location having a known relationship to said dedicated reference point; the first audio signal and the plurality of second audio signals are synchronized in time; For the first audio signal: calculating a frequency weighted cross-correlation between the first audio signal and at least one of the plurality of second audio signals to obtain at least one frequency weighted cross-correlation signal; estimating a distance between the sound source and the dedicated reference point by estimating a time delay between the first audio signal and the at least one of the plurality of second audio signals using the at least one frequency weighted cross-correlation signal; and estimating an angle between the sound source and the dedicated reference point by a weighted least mean square evaluation of the time delay between each pair of the plurality of second audio signals, the weighted least mean square being dependent on the obtained frequency weighted cross-correlation signal between the first audio signal and each pair of the plurality of second audio signals. method.
2. Calculating the frequency weighted cross correlation comprises: correlating the first audio signal with the at least one of the plurality of second audio signals in a frequency domain to obtain at least one correlation signal; modifying the power spectrum by frequency weighting to obtain the at least one frequency weighted correlation signal; transforming the at least one correlation signal into the time domain; The method of claim 1 , comprising:
3. a respective frequency weighted cross-correlation signal is calculated between the first audio signal and each of the plurality of second audio signals; The method according to claim 1 or 2.
4. The step of transforming the at least one correlation signal into the time domain comprises: converting the first audio signal and the at least one of the plurality of second audio signals into the digital domain at a conversion frequency higher than a conversion frequency of the conversion step of the first audio signal and the at least one of the plurality of second audio signals; The method of claim 2 or 3, comprising:
5. The step of correlating the first audio signal further comprises: upsampling the first audio signal and the at least one of the plurality of second audio signals prior to correlating the first audio signal and the at least one of the plurality of second audio signals in the digital domain; The method according to any one of claims 2 to 4, comprising:
6. - calculating a phase transformation between the first audio signal and a further first audio signal recorded at or associated with one or more audio sources to obtain a further phase transformed signal; - estimating a distance between a position of the microphone and a position associated with the recording of the further first audio signal by estimating a time delay between the first audio signal and the further first audio signal using the further phase transformed signal; The method of claim 1 , further comprising:
7. Said step of calculating the frequency weighted cross correlation, in particular the phase transform, comprises: performing a Short-Time Fourier Transform (STFT) on the first audio signal and the at least one of the plurality of second audio signals to obtain a respective spectrum; obtaining a cross spectrum for each of the spectrograms; applying a spectral mask filter to the obtained composite cross spectrum; performing an Inverse Short-Time Fourier Transformation (ISTFT) to obtain at least one phase transformed signal; The method of any one of claims 1 to 6, comprising:
8. Said step of calculating the frequency weighted cross correlation, in particular the phase transform, comprises: estimating a filter acting on a signal-to-noise ratio in each frequency bin of the first audio signal, the estimating of the filter comprising: - applying a quantile filter, in particular a median filter, for each time slice (k) of the power spectrum derived from said one or more recorded first audio signals to smooth said power spectrum; estimating the noise for each time slice (k) according to the previous time slice; - evaluating, for a given frequency, whether the signal to noise ratio exceeds a predetermined threshold and setting a filter parameter for that frequency to 1 or 0 depending on said evaluation; The method of any one of claims 1 to 7, comprising:
9. estimating the filter acting on the signal-to-noise ratio in each frequency bin of the first audio signal comprises applying a residual signal from a denoising process as a noise estimate, the denoising process optionally being based on machine learning. The method according to claim 8.
10. The step of estimating a time delay comprises: Searching for a maximum in said at least one frequency-weighted cross-correlation, in particular in a phase-transformed signal, or Detecting a first intensity value that exceeds a given threshold and searching for a maximum value within a specified window centered on said first intensity value; 10. The method of claim 1 , comprising:
11. said search for a maximum in said at least one frequency weighted cross-correlation, in particular in a phase transformed signal, depends on time delay estimates in neighbouring time frames. The method of claim 10.
12. The window length of the specified window is inversely proportional to a signal bandwidth estimated from a highest frequency component of the first audio signal; expected early reflections as a function of the distance between a recording location of the first audio signal and the location of the one or more sound sources; proportional to a maximum time of flight between said positions of said plurality of second audio signals; Depends on at least one of The method of claim 10.
13. The dedicated reference point is approximately midway between recording locations of the plurality of second audio signals, and the estimation of the distance between the sound source and the dedicated reference point comprises: obtaining an average value of a set of time delays between the first audio signal and each of the at least one of the plurality of second audio signals; obtaining a time delay between the first audio signal and a signal formed by a sum of at least two of the plurality of second audio signals; The method includes the steps of:
13. The method according to any one of claims 1 to 12.
14. The weighted least squares mean is If the intensity value of the obtained phase-transformed signal exceeds a given threshold and is within a specified window centered on a first intensity value, the obtained frequency-weighted cross-correlation between the first audio signal and the pair of the plurality of second audio signals, in particular the phase-transformed signal, the time difference of arrival of the direct sound of a plurality of frequency weighted cross-correlation signals, the plurality of frequency weighted cross-correlation signals being provided by the first audio signal and the plurality of second audio signals; depends on one of 14. The method according to any one of claims 1 to 13.
15. applying a noise reduction filter to the estimated distance and / or the estimated angle; or applying a Kalman filter to the estimated range and / or the estimated angle; applying a gradient or divergence to the estimated distance and / or the estimated angle; 15. The method of claim 1, further comprising:
16. the respective positions of the pair of the plurality of second audio signals are located at the same distance from the dedicated reference point and on a virtual line passing through the dedicated reference point; 16. The method according to any one of claims 1 to 15.
17. the plurality of second audio signals includes at least four audio audio signals, two of the four audio signals being recorded with a maximum spatial distance of 15 cm; 17. The method according to any one of claims 1 to 16.
18. obtaining temperature information, in particular temperature information in the vicinity of the second sound sources; estimating the distance and / or angle in response to the acquired temperature information; 18. The method of any one of claims 1 to 17, further comprising:
19. one or more processors; a memory coupled to the one or more processors and containing instructions, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method of any one of claims 1 to 18; A computer system comprising:
20. 20. A method comprising: A non-transitory computer-readable storage medium.
21. A sound recording device having a rectangular parallelepiped shape with a bottom surface, a top surface, and four sides, the sound recording device adapted to be placed with the bottom surface on a surface; The recording device is a user interface accessible on said top surface; A plurality of microphones, in particular omnidirectional microphones; a pair of microphones is disposed on each of the sides, a first microphone of the pair of microphones is disposed on an upper portion of each of the sides, and a second microphone of the pair of microphones is disposed on a bottom portion of each of the sides; a distance between the first microphone and the second microphone of each pair of microphones is equal to a distance between the first microphones on adjacent sides; Recording device.
22. A distance from the first microphone to the top surface is greater than a distance from the second microphone to the bottom surface.
21. Sound recording device according to claim 20.
Citation Information
Cited By
Method for obtaining a position of a sound source
US12710502B2