Determining a spatial format representation of a sound scene

By employing a single high-quality microphone with affordable arrays to estimate acoustic parameters, the method addresses the cost and noise sensitivity issues of traditional spatial audio systems, achieving high-quality spatial audio capture and playback.

WO2025163235A1PCT designated stage Publication Date: 2025-08-07AALTO UNIV FOUND
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/FI2024/050406
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2024-08-06
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing spatial audio technologies require multiple microphones or microphone arrays for high spatial resolution, which can be costly and sensitive to noise, leading to poor quality and resolution in spatial audio recordings.

Method used

A method that uses a single high-quality monophonic microphone in combination with affordable and potentially noisy microphone arrays to estimate acoustic parameters, allowing for the determination of spatial format representation of a sound scene through binaural signal synthesis.

Benefits of technology

Enables high-quality spatial audio capture and playback at reduced costs by leveraging acoustic parameters to impose spatial characteristics on a single microphone signal, overcoming the limitations of traditional microphone arrays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FI2024050406_07082025_PF_FP_ABST
    Figure FI2024050406_07082025_PF_FP_ABST
Patent Text Reader

Abstract

According to an example aspect of the present invention, there is a method comprising taking as an input a single mono signal, capturing data about a sound scene using at least one sensor, estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources and determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.
Need to check novelty before this filing date? Find Prior Art

Description

DETERMINING A SPATIAL FORMAT REPRESENTATION OF A SOUND SCENE FIELD

[0001] The present invention relates to a field of spatial audio and more specifically, to a spatial format representation of a sound scene. BACKGROUND

[0002] Virtual Reality, VR, technology refers to a technology that provides a simulated experience of a real or an imaginary system to a user. Augmented Reality, AR, technology then refers to a type of VR technology, wherein a user may have an interactive experience with the simulated content. Mixed Reality, MR, technology then refers to a technology which combines a real-world environment with the simulated content. Such technologies provide fast-growing immersive platforms for communications, media, and education. The VR, AR and MR technologies involve overlaying, partially or completely, synthetic information over the user’s senses. Spatial audio is the audio platform for these technologies, involved in capturing, processing, and reproducing sound in 3 dimensions, as sound is heard and perceived in the real world. For spatial audio to be effective, it must provide sufficient auditory cues which convey the virtual representation of the sound scene and its plausibility. SUMMARY

[0003] According to some aspects, there is provided the subject-matter of the independent claims. Some embodiments are defined in the dependent claims.

[0004] According to a first aspect of the present disclosure, there is provided a method comprising taking as an input a single mono signal, capturing data about a sound scene using at least one sensor, estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources and determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

[0005] Example embodiments of the first aspect may comprise at least one feature from the following bulleted list or any combination of the following features: wherein the at least one sensor comprises an acoustic sensor, preferably a microphone array; wherein the at least one sensor comprises a vision sensor, preferably a camera; wherein taking the input comprises recording the single mono signal; recording the single mono signal using a single microphone; wherein the single microphone is of a higher quality than a microphone array used for capturing said data about the sound scene; the method further comprises estimating, based at least on the captured data, acoustic parameters comprising at least a parameter of diffuseness of the sound scene; the method further comprises estimating, based at least on the captured data, acoustic parameters comprising at least a reverberation time of a physical space of the sound scene; the method further comprises estimating, based at least on the captured data, acoustic parameters comprising at least a direct to reverberant ratio between the sound sources and an apparatus capturing said data.

[0006] According to a second aspect of the present disclosure, there is provided an apparatus comprising at least one processing core, at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processing core, cause the apparatus at least to, cause taking as an input a single mono signal, cause capturing data about a sound scene using at least one sensor, estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources and determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal. The at least one memory and the computer program code may be further configured to, with the at least one processing core, cause the apparatus to perform a method according to the first aspect.

[0007] According to a third aspect of the present disclosure, there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus at least to take as an input a single mono signal, capture data about a soundscene using at least one sensor, estimate, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources and determine a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal. The computer program may further comprise instructions which, when executed by the apparatus, cause the apparatus at least to a method according to the first aspect.

[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer readable medium having stored thereon a set of computer readable instructions that, when executed by at least one processor, cause an apparatus to at least to take as an input a single mono signal, capture data about a sound scene using at least one sensor, estimate, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources and determine a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal. The non-transitory computer readable medium may further have stored thereon a set of computer readable instructions that, when executed by at least one processor, cause the apparatus to perform a method according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG.1 illustrates a process in accordance with at least some embodiments of the present invention;

[0010] FIG. 2 illustrates acoustic parameter estimation in accordance with at least some embodiments of the present invention;

[0011] FIG. 3 illustrates scene matching in accordance with at least some embodiments of the present invention;

[0012] FIG. 4 illustrates an encoder in accordance with at least some embodiments of the present invention;

[0013] FIG. 5 illustrates an example apparatus capable of supporting at least some embodiments of the present invention;

[0014] FIG. 6 is a flow graph of a method in accordance with at least some embodiments of the present invention. EMBODIMENTS

[0015] Embodiments of the present invention relate to spatial audio and more specifically, to a spatial format representation of a sound scene. Spatial audio may be produced by relying on synthetic audio productions or recordings made with microphone arrays. For example, Ambisonics may be used to produce spatial audio from microphone arrays by relying on differences between levels of audio signals. To capture Ambisonics with high spatial resolution, microphone arrays may require 16 or more microphones in a strict geometric configuration. Furthermore, the encoding of Ambisonics may be highly sensitive to the microphone noise which can easily degrade its quality, and insufficient microphones may further produce poor spatial resolution.

[0016] Embodiments of the present disclosure therefore provide non-spatial mono recording which is combined with acoustic parameters, thereby leading to multichannel (2 or more) spatial audio format. Hence there is no need for 2 or more non-spatial audio signals, to generate multichannel (2 or more) spatial audio format.

[0017] For example, embodiments of the present invention provide a way for recording and binaurally reproducing spatial sound scenes using the audio from a single microphone, i.e., a single mono signal recorded by the single microphone. This may be realised by recording the sound scene using both a microphone array, which potentially comprises more affordable and lower quality capsules, and a monophonic microphone, possibly featuring a higher quality capsule. By analysing the array signals in the time- frequency domain and estimating perceptually meaningful spatial parameters, it is possible to determine what the appropriate binaural Spatial Covariance Matrices, SCMs, should be. The actual binaural signals may then be synthesised using a SCM matching renderer, which may take this higher-quality monophonic audio signal as input.

[0018] In some embodiments of the present invention, assumptions may be made on certain elements of sound fields. For example, an intermediate representation of the sound fields may be produced in the form of acoustic parameters. Such acoustic parameters may be exploited in many potential applications. For example, such acoustic parameters may beused to analyze and synthesize spatial sounds scenes with high resolution. In some example embodiments, a method that exploits such acoustic parameters may be referred to as a parametric spatial audio. From a recording perspective, the exploitation of such acoustic parameters parametric allows the use of fewer sensors and irregular microphone geometric arrangements and still achieve high-quality spatial recordings. That is, it is possible to determine spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

[0019] FIG.1 illustrates a process in accordance with at least some embodiments of the present invention. In FIG. 1, an apparatus is denoted by (10). The apparatus (10) may comprise an encoder (14), transmission / storage format block (16) and a decoder (18). Furthermore, in FIG. 1 mono signal is denoted by (18), spatial sound scene is denoted by (20), and spatial format is denoted by (22). The spatial format (22) may be a spatial format representation of the sound scene (20).

[0020] The encoder (12) may further comprise sensors, such as audio or visual sensors, Digital Signal Processing, DSP, and deep learning block, and analog-to-digital converter block, i.e., an audio interface. The transmission / storage format block (14) may further comprise acoustics parameters estimation block and digital audio signal block. The decoder (16) may further comprise a spatial mixer block and binauralizer block.

[0021] According to the process illustrated in FIG.1, the apparatus (10) may take as an input a mono signal (18) from the sound scene (20). That is, the apparatus (10) may record a single mono signal (18), possibly using a single microphone. For example, the apparatus (10) may comprise a monophonic microphone, possibly featuring a higher quality capsule, and record the single mono signal (18) using the monophonic microphone. Alternatively, the single mono signal may be speech or from an instrument which is monophonic. In general, the single mono signal is monophonic and taken as an input.

[0022] The single mono signal (18) may be fed to the analog-to-digital converter block. The single mono signal (18) may be processed in the analog-to-digital block and fed, after the processing, to the digital audio signal block of the transmission / storage format block (14). After processing at the digital audio signal block, the processed single mono signal (18) may be fed to the spatial mixer block of the decoder (16).

[0023] The sensors associated with the encoder (12) of the apparatus (10) may capture data about the spatial sound scene 20. Said data captured by the sensors may be fed to the DSP and deep learning block of the encoder (12). Said data captured by the sensor may depend on the sensors, i.e., whether an audio and / or vision sensors are used. In case of audio sensors, which may comprise at least one microphone array, said data may comprise audio signals. In case of vision sensors, which may comprise at least one camera, said data may comprise at least one image, or data about at least one image.

[0024] After the DSP and deep learning block, said data captured by the sensors may be fed to the acoustics parameters estimation block of the transmission / storage format block (14). After estimation of the acoustics parameters, the acoustics parameters may be fed to the spatial mixer block of the decoder (16). Alternatively, or in addition, the acoustics parameters may be used to generate spatial acoustics data.

[0025] At the spatial mixer of the decoder, the processed mono signal from the digital audio signal block may be mixed with the acoustics parameters from the acoustic parameters estimation block, to generate information about the spatial sound scene (20). Said information about the spatial sound scene (20) may be further fed to the binauralizer of the decoder (16) to determine the spatial format representation (22) of the sound scene (20) based on the acoustic parameters.

[0026] The apparatus (10) may be thus configured to produce a set of enhanced signals based at least one the single mono signal. The apparatus (10) may produce said enhanced signals such that said enhanced signals match the spatial acoustic properties of the spatial sound scene (20), i.e., the original recorded scene. The apparatus (10) may comprise an inputs, through which audio signals, such as the mono signal (18) can be recorded. The recorded audio signals (18) may comprise phantom power. The apparatus may comprise signal pre-amplifying modules to condition microphones and adjust input signal levels.

[0027] In some embodiments, the apparatus (10) may comprise acoustic and vision sensors. The apparatus (10) may estimate acoustic parameters from its surrounding environment through the acoustic and vision sensors. For example, an acoustic sensor may comprise an array of Micro-Electromechanical Systems, MEMS, microphones. A vision sensor may be a camera capable of recording in 360 degrees.

[0028] The process illustrated in FIG. 1 may be performed by the apparatus (10) andcomprise an encoding stage (performed by the encoder (12)), an intermediate storage ortransmission stage (performed by the transmission / storage format block (14)) and adecoding stage (performed by the decoder (16)). At the encoding stage, the apparatus (10)may use the data captured by its sensors and a combination of signal processing and deep learning.

[0029] In some embodiments, the apparatus (10) may be configured to estimate,based at least on the captured data, acoustic parameters comprising at least one of thefollowing: a number of sound sources in a sound scene (20) and a location of each of the sound sources; aparameter of diffuseness of the sound scene (20);a reverberation time of the physical space (RT60) of the sound scene (20);a direct to reverberate ratio between the sources and the apparatus (20).

[0030] The apparatus (10) may determine a spatial format representation of thesound scene (20) based at least on the acoustic parameters and the single mono signal. Insome embodiments, the process of estimating parameters might not interact with the recording process of the encoder (12).

[0031] FIG. 2 illustrates acoustic parameter estimation in accordance with at leastsome embodiments of the present invention. An audio signal may be captured with a recording microphone in the sound scene, that has no spatial properties is converted to a spatial multichannel audio signal of channels as follows.

[0032] The sound scene is captured in parallel from the analysis device. In oneembodiment, the apparatus (10), such as an analysis device, may have an integratedmicrophone array of channels, while in another the apparatus (10) may have anintegrated microphone array plus one or more cameras. The main monophonic microphone recording signal may be termed as a main signal, while the recordings from the apparatus (10) may be termed as analysis signals. The audio signals captured by the apparatus (10) may be denoted

[0033] Firstly, both the main signal and the analysis signals may be passed to thetime-frequency transform block, being transformed from the time-domain to the time-frequency domain through a short-time Fourier Transform or a filterbank suitable for audio coding, resulting in andrespectively. The denote the subsampled temporal index and frequency bin index of the time- frequency transformed signal.

[0034] Second, the analysis signals may be passed through a cross-spectral densityestimator. The SCM of the analysis signals may be computed by averaging multipleframes aswhich would introduce a latency of frames to the system. When lowlatency is required, a recursive averaging may be preferred aswherein is a smoothing constant.

[0035] Third, the analysis SCM may be passed to the source component counter. Inone embodiment, the number of dominant source components at each time-frequency point(t,f) may be determined by analysis of the eigenvalues of the analysis SCM. Morespecifically, an Eigenvalue Decomposition, EVD, may be applied to the SCM, returning an diagonal matrix of eigenvalues sorted in descending order, and an matrix of eigenvectors sorted according to theeigenvalues. The first K distinct high eigenvalues before a noticeable drop of eigenvaluemagnitude to a lower level may be taken as the number of dominant source components.One suitable method to find the first ditinct eigenvalues is SORTE (Second OrdersTatistic of Eigenvalues) as described in He, Z., Cichocki, A., Xie, S., & Choi, K. (2010).Detecting the number of clusters in n-way probabilistic clustering. IEEE Transactions onPattern Analysis and Machine Intelligence, 32(11), 2006-2021.

[0036] Based on that the dominant source number would be thenNote that to avoid source number overestimation, K may be limited to lessthan half the number of channels

[0037] In an alternative embodiment, the dominant source component number maybe estimated by a compact neural network block, such as You Only Look Once (YOLO).The neural network block may be a foreground object detector returning bounding boxesand class labels of the various objects captured in the video recording. The objects andlabels of the neural network block may then be passed to an object filter block, which mayretain objects based on pre-specified labels for potential sound sources.

[0038] Fourth, the analysis SCM and the source number may be passed to a sourceDirection-Of-Arrival, DoA, estimator. The method may thus benefit from independentsource components with distinct DoA estimates in order to model diverse scenarios with multiple spectrally non-overlapping sources in the scene, rather than broadband DoA estimates expressing a small number of dominant sources across all frequencies. Hence, narrowband estimators may be favoured. In one implementation, subspace methods may be used that receive an estimate of the source components K(t,f), the eigenvectors of the analysis SCM and the array steering vectorsfor a dense grid of Q predefined directions. The array steeringvector matrix is known either from a theoretical array response model, fromnumerical simulations, or from array calibration measurements. As an example, a subspacelocalization method suitable for this implementation is MUltiple SIgnal Classification(MUSIC), e.g., as in Schmidt, R., “Multiple emitter location and signal parameterestimation,” IEEE transactions on antennas and propagation, 34(3), pp.276–280, 1986.

[0039] MUSIC would return a spatial pseudo-spectrum of values for every grid pointand every time-frequency point as follows

[0040] The spatial pseudo-spectrum may be passed to a 1D peak finding algorithm ifDOA estimation is done only on horizontal angles, or a 2D peak finding algorithm if DOAestimation is done on the full sphere, to select the K most prominent peaks, returning KDoAs that belong to the dense grid of steering vectors

[0041] As an example, one such implementation for full sphere peak finding isdescribed in McCormack, L., Meyer-Kahlen, N. and Politis, A., 2023. Spatialreconstruction-based rendering of microphone array room impulse responses. AES: Journal of the Audio Engineering Society, 71(5), pp.267-280.

[0042] In an alternative embodiment, the DoA of the dominant sources may beestimated by using both parameters estimated by a microphone array and by a camera. Theobjects and labels retained by the object filter block may be passed to an object locationestimator block. The object location block may define the direction of the sources relativeto the center pixel of the object’s bounding boxes and the center of the captured image.Then, the signals of the microphone array may be passed through a spatial activititydetector block which returns probabilities for the activity of sound source coming from acertain direction. In one embodiment, this could be achieved by a spatial post-filter tomicrophone arrya signals as described by S. Delikaris-Manias and V. Pulkki, "CrossPattern Coherence Algorithm for Spatial Filtering Applications Utilizing MicrophoneArrays," in IEEE Transactions on Audio, Speech, and Language Processing, vol. 21, no.11, pp.2356-2367, Nov.2013, doi: 10.1109 / TASL.2013.2277928.

[0043] Then, the directions estimated from the camera feed may be combined withestimates acquired from the microphone array into an audiovisual DoA estimator blockwhich may output the narrowband DoA estimates from the microphone array biased by thedirections estimated from the camera feed. Hence, the DoAs might not be fully fixed basedon the camera feed and might fluctuate based on the acoustic paramaters. Regions may bealso defined to restrict the estimated audiovisual DoAs to visual region covered by the camera.

[0044] Fifth, analysis signals and estimated DoAs may be passed to the model powerparameter estimator. The vector of K source powersmay be estimated through a beamforming operation

[0045] where is a matrix of beamforming weights with each row focusing in oneof the estimated DoAs and extracts the diagonal values of a matrix. For this purpose, the matched-filter beamformer design could be selected for this task.,

[0046] whereis a small regularization parameter to stabilize the matrix inversionand an SxS sized identity matrix.is a matrix ofarray steering vectors selected from based on the analyzed DoAs .

[0047] An estimate of the ambience power may be obtained then by,

[0048] where stands for the trace operator, and is an ambiencebeamforming matrix. This can be perceived as an operation which spatially subtracts thematched-filter beamformer pattern from an unit sphere, and may be calculated as:. In this equation, is referring to the identity matrix.

[0049] The set of parameters would be the parameters required bythe rendering stage in order to upscale the monophonic main signal into a spatialmultichannel target signal which matches spatially the captured scene.

[0050] FIG. 3 illustrates scene matching in accordance with at least someembodiments of the present invention.

[0051] The set of parameters estimated at the recording stage may upscale therecorded monophonic signal into a spatial multichannel target signal. The spatial renderingtarget may switch between two modes: binaural mode rendering binaural signals, orambisonic mode rendering rendering signals in the Ambisonics spatial audio format. Inorder to produce these multichannel signals, appropriate target binaural SCMs andambisonic SCMs may be required. In such a case, these SCMs should include the spatial characteristics that binaural or ambisonic signals, at the end of the chain, should possess. It's important that these target SCMs are updated on a regular basis for every frequency bin and time window. For better understanding, the following matrices are presented for a single time window.

[0052] First, for a specific scene containing directional sounds, the estimatedsource powers and the DoAs may be passed to a target source SCM synthesizer,which may output a SCM, whose dimensions would depend on the targetplayback. The synthesizer may take as an input the source powers and the DoAsalong with a lookup table of Head-related Transfer Functions, HRTFs, values forbinaural rendering mode, or a lookup table of Spherical Harmonic, SH, values forambisonic rendering mode. The tables may be designed in such a way to include therespective HRTF or SH values for the DoA analysis grid . The dimensions of theHRTF lookup table may be 2xQxF, where F is the number of frequency bins or subbandsresulting from the time-frequency transform operation. The dimensions of the SH lookuptable may be (N+1)2xQ, where N is the target ambisonic order to be rendered.

[0053] In the case of a target binaural playback, may be defined using theequation:

[0054] where is a 2x1 HRTF vector for frequency f retrieved from the lookup tablefor DoA .

[0055] In the case of a target ambisonic playback, may be defined usingthe equation:where is a SH vector ( +1)2 1 retrieved from the lookup table forDoA .

[0056] Second, for a specific scene that includes ambient sounds, the estimatedambiance power may be passed to a target diffuse SCM synthesizer block whichoutputs a SCM , whose dimensions depend on the target playback.

[0057] For the case of a target binaural playback,may be computed usingthe formula:

[0058] Here, denotes a QxQ diagonal matrix containing grid weights normalizedto unity, , used to accommodate non-uniform simulation / measurement grids.

[0059] For the case of a target ambisonic playback, can be computedusing the formula:, where correspond to the matrix of spherical harmonics computed for the grid of directions .

[0060] Fourth, the main signal may be passed through a power spectraldensity estimator. A local time-frequency power estimate is computed forby averaging frames aswhich allows accurate alignment of the estimated powers with the previously estimated parameters. When low latency is required, a recursive averaging is preferred aswhich, however, would introduce some temporal misalignment between the estimated power and the estimated frames.

[0061] Fifth, the target SCMs and , and the power estimate may bepassed to a target SCM mixer and equalizer block. Initially, a total target binaural orambisonic SCM is produced via superposition

[0062] Then, the spectral color of the main signalis forced into the totaltarget in order for the final target spatial playback to retain the colouration of the original non-spatial recording. For that we modify the target SCM by equalizing its trace to that of the main signal power through

[0063] Sixth, the main signal may be passed to the prototype signalgenerator block. Prototype signals may be modified copies of the main signal that aremixed in order to generate that final spatial target signal such that it has the target spatialproperties modeled by the target SCM . According to this spatial covariancematching approach, a suitable number of prototyping signals in line with the targetplayback, may be required. In this setting, may comprise the main single-channelsignal and its copies acquired through a signal decorrelator . In one embodiment, the chosen signal decorrelatormay use a velvet noise based designed which achieves high decorrelation at high frequencies, but minimal decorrelation at lowfrequencies as described in Schlecht, S. J., Alary, B., Välimäki, V., Habets, E. A., et al.,“Optimized velvet-noise decorrelator,” in Proc. Int. Conf. Digital Audio Effects (DAFx- 18), Aveiro, Portugal, pp.87–94, 2018., based on the knowledge that the target SCMs may generally feature high coherence at low frequencies.

[0064] The prototyping signals for target binaural playback may then beand for the target ambisonic playback

[0065] Seventh, the prototype signals and the target binaural SCMs are passed to theoptimal mixing block, where the prototype signals may be adaptively mixed to produce thetarget format that have SCMs that match with the target . This adaptive mixing may be obtained by implementing the following equation: wherein, and may be optimal mixing matrices derived usinga spatial covariance matching framework, such as described by Vilkamo, J., Bäckström, T.,and Kuntz, A., “Optimized covariance domain framework for time–frequency processing ofspatial audio,” Journal of the Audio Engineering Society, 61(6), pp. 403– 411, 2013.Unlike the [.] decorrelator, it should be noted that the [.] decorrelator would produce significantly uncorrelated versions of signals at all frequencies.

[0066] Finally, the time-frequency domain signals may be passed to afrequency-time transform block which outputs the time domain signals via an inverseshort-term Fourier transform or synthesis filterbank suitable for audio coding.

[0067] FIG. 4 illustrates an encoder in accordance with at least some embodimentsof the present invention. The encoder (12) illustrated in FIG.2 may be the encoder (12) ofthe apparatus (10) illustrated in FIG. 4. The encoder (12) may have the following signalprocessing chain.

[0068] The encoder (12) may comprise a microphone array (24), spatial frequencytime domain transform block (26), camera (28), convolutional neural network (30), source estimation block (32), Direction of Arrival, DoA, estimation block 34, diffuseness estimation block (36) and acoustic parameters estimation block (38).

[0069] In the example of FIG. 2, the sensors illustrated in FIG. 1 comprise the microphone array (24). The microphone array (24) may capture audio data, such as audio signals, about the spatial scene (20). Said audio data may be fed to the frequency time domain transform block (26).

[0070] In the example of FIG. 4, the sensors illustrated in FIG. 1 further comprise the camera (28). The camera (26) may capture image data, such as video images, about the spatial scene (20). Said image data may be fed to the convolutional neural network block (30).

[0071] Said audio data and image data may be fed to the source estimation block (32). At the source estimation block (32), the strongest sources may be identified. Information about the strongest sources may be then fed to the DoA estimation block (34). At the DoA estimation block (34), the DoAs of the strongest sources may be determined. Said audio data from the spatial frequency time domain transform block (26) may be associated with the DoAs of the strongest sources and further fed to the diffuseness estimation block (36). After diffuseness estimation, acoustic parameters may be esstimated at block (38).

[0072] The parameters and signals may then be transmitted and stored. The decoder (12) may use the estimated parameters and the audio signals, such as the mono signal (18), to synthesize a spatial format representation of the sound scene (20), wherein the signals were originally recorded. The decoder (12) may generate various outputs such as Binaural Audio, Ambisonics and Spatial Room Impulse Responses.

[0073] Embodiments of the present invention therefore enable more stringent requirements on lowering cost, while retaining high-quality spatial audio capture and playback. In order to further lower costs, the necessary spatial parameters may be estimated using a more affordable lower-quality tetrahedral array, (featuring potentially noisier sensors), but to then use these spatial parameters to impose the spatial characteristics of the captured sound scene onto a single high quality / low noise omnidirectional microphone signal. Such an approach could lead to substantial cost reductions, and thus lower the barrier to entry for artists and audio engineers to enter the field of spatial audio recording.

[0074] The spatial format representation of the spatial sound scene (20) may be determined using a more affordable (and potentially noisy) microphone array, and the acoustic parameters employed to synthesise the playback signals by adaptively mixing a signal obtained from a single-channel (high quality) microphone that is suited near to the array during the recording.

[0075] FIGURE 5 illustrates an example apparatus capable of supporting at least some embodiments of the present invention. Illustrated is device 50, which may comprise, for example, the apparatus (10) of FIG. 1, or device 50 may be comprised in apparatus (10). Comprised in device 50 is processor 51, which may comprise, for example, a single- or multi-core processor wherein a single-core processor comprises one processing core and a multi-core processor comprises more than one processing core. Processor 51 may comprise, in general, a control device. Processor 51 may comprise more than one processor. When processor 51 comprises more than one processor, device 50 may be a distributed device wherein processing of tasks takes place in more than one physical unit. Processor 51 may be a control device. Processor 51 may comprise at least one application- specific integrated circuit, ASIC. Processor 51 may comprise at least one field- programmable gate array, FPGA. Processor 51, optionally together with memory and computer instructions, may be means for performing method steps in device 50, such as cause taking an input, cause capturing data, estimating and determining. Processor 51 may be configured, at least in part by computer instructions, to perform actions.

[0076] Device 50 may comprise memory 52. Memory 52 may comprise random- access memory and / or permanent memory. Memory 52 may comprise at least one RAM chip. Memory 52 may be a computer readable medium. Memory 52 may comprise solid- state, magnetic, optical and / or holographic memory, for example. Memory 52 may be at least in part accessible to processor 51. Memory 52 may be at least in part comprised in processor 51. Memory 52 may be means for storing information. Memory 52 may comprise computer instructions that processor 51 is configured to execute. When computer instructions configured to cause processor 51 to perform certain actions are stored in memory 52, and device 50 overall is configured to run under the direction of processor 51 using computer instructions from memory 52, processor 51 and / or its at least one processing core may be considered to be configured to perform said certain actions. Memory 52 may be at least in part external to device 50 but accessible to device 50. Memory 52 may be transitory or non-transitory. The term “non-transitory”, as used herein,is a limitation of the medium itself (that is, tangible, not a signal) as opposed to a limitation on data storage persistency (for example, RAM vs. ROM).

[0077] Device 50 may comprise a transmitter 53. Device 50 may comprise a receiver 54. Transmitter 53 may comprise more than one transmitter. Receiver 54 may comprise more than one receiver.

[0078] Device 50 may comprise a near-field communication, NFC, transceiver 55. NFC transceiver 55 may support at least one NFC technology, such as NFC, Bluetooth, Wibree or similar technologies.

[0079] Processor 51 may be furnished with a transmitter arranged to output information from processor 51, via electrical leads internal to device 50, to other devices comprised in device 50. Such a transmitter may comprise a serial bus transmitter arranged to, for example, output information via at least one electrical lead to memory 52 for storage therein. Alternatively to a serial bus, the transmitter may comprise a parallel bus transmitter. Likewise processor 51 may comprise a receiver arranged to receive information in processor 51, via electrical leads internal to device 50, from other devices comprised in device 50. Such a receiver may comprise a serial bus receiver arranged to, for example, receive information via at least one electrical lead from receiver 54 for processing in processor 51. Alternatively to a serial bus, the receiver may comprise a parallel bus receiver.

[0080] Device 50 may comprise further devices not illustrated in FIG. 5. In some embodiments, device 400 lacks at least one device described above.

[0081] Processor 51, memory 52, transmitter 53 and receiver 54 may be interconnected by electrical leads internal to device 50 in a multitude of different ways. For example, each of the aforementioned devices may be separately connected to a master bus internal to device 50, to allow for the devices to exchange information. However, as the skilled person will appreciate, this is only one example and depending on the embodiment various ways of interconnecting at least two of the aforementioned devices may be selected without departing from the scope of the present invention.

[0082] In some embodiments, device 50 may be the apparatus (10), or be comprised in the apparatus (10), and the apparatus (10) may further comprise the at least one sensor and / or the single microphone. Alternatively, device 50 may be the apparatus (10), or becomprised in the apparatus (10), and the apparatus (10) might not comprise the at least one sensor and / or the single microphone. In such a case, processor 51 may be configured to cause taking as an input a single mono signal and / or cause capturing data about a sound scene using at least one sensor, e.g., by transmitting a command to the at least one sensor and / or the single microphone, respectively. In response to the command, processor 51 may receive the input and / or said data. Processor 51 may be further configured to estimate acoustics parameters and determine a spatial representation of the sound scene (20). Thus, in some embodiments, the at least one sensor and / or the single microphone may be in a remote location from the apparatus (10) and the apparatus (10) may communicate with the at least one sensor and / or the single microphone via a communication network, using transmitter 53, receiver 54 and / or NFC transceiver 55.

[0083] FIG. 6 is a flow graph of a method in accordance with at least some embodiments of the present invention. The phases of the illustrated method may be performed by apparatus 10, an auxiliary device or a personal computer, for example, or in a control device configured to control the functioning thereof.

[0084] At step 61, the method may comprise taking as an input a single mono signal, preferably from a sound scene. At step 62, the method may comprise capturing data about the sound scene using at least one sensor. At step 63, the method may comprise estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources, and preferably a parameter of diffuseness of the sound scene. At step 64, the method may comprise determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

[0085] It is to be understood that the embodiments of the invention disclosed are not limited to the particular structures, process steps, or materials disclosed herein, but are extended to equivalents thereof as would be recognized by those ordinarily skilled in the relevant arts. It should also be understood that terminology employed herein is used for the purpose of describing particular embodiments only and is not intended to be limiting.

[0086] Reference throughout this specification to one embodiment or an embodiment means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment”in various places throughout this specification are not necessarily all referring to the same embodiment. Where reference is made to a numerical value using a term such as, for example, about or substantially, the exact numerical value is also disclosed.

[0087] As used herein, a plurality of items, structural elements, compositional elements, and / or materials may be presented in a common list for convenience. However, these lists should be construed as though each member of the list is individually identified as a separate and unique member. Thus, no individual member of such list should be construed as a de facto equivalent of any other member of the same list solely based on their presentation in a common group without indications to the contrary. In addition, various embodiments and example of the present invention may be referred to herein along with alternatives for the various components thereof. It is understood that such embodiments, examples, and alternatives are not to be construed as de facto equivalents of one another, but are to be considered as separate and autonomous representations of the present invention.

[0088] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the preceding description, numerous specific details are provided, such as examples of lengths, widths, shapes, etc., to provide a thorough understanding of embodiments of the invention. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the invention.

[0089] While the forgoing examples are illustrative of the principles of the present invention in one or more particular applications, it will be apparent to those of ordinary skill in the art that numerous modifications in form, usage and details of implementation can be made without the exercise of inventive faculty, and without departing from the principles and concepts of the invention. Accordingly, it is not intended that the invention be limited, except as by the claims set forth below.

[0090] The verbs “to comprise” and “to include” are used in this document as open limitations that neither exclude nor require the existence of also un-recited features. The features recited in depending claims are mutually freely combinable unless otherwise explicitly stated. Furthermore, it is to be understood that the use of "a" or "an", that is, asingular form, throughout this document does not exclude a plurality.

[0091] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. INDUSTRIAL APPLICABILITY

[0092] At least some embodiments of the present invention find industrial application in VR, AR and / or MR. ACRONYMS LIST AR Augmented Reality DoA Direction-Of-Arrival DSP Digital Signal Processing EVD Eigenvalue Decomposition HRTF Head-related Transfer Functions MEMS Micro-Electromechanical Systems MR Mixed Reality SCM Spatial Covariance Matrices SH Spherical Harmonic VR Virtual Reality REFERENCE SIGNS LIST 10 Apparatus 12 Encoder 14 Transmission / storage format 16 Decoder 18 Spatial formatMono signal Spatial sound scene Microphone array Spatial frequency time domain transform block Camera Convolutional neural network Source estimation block DoA estimation block Diffuseness estimation block Acoustic parameters block – 55 Structure of the device of FIG.5 – 64 Phases of the method of FIG.6

Claims

CLAIMS:

1. A method, comprising: taking as an input a single mono signal; capturing data about a sound scene using at least one sensor; estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources; and determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

2. A method according to claim 1, wherein the at least one sensor comprises an acoustic sensor, preferably a microphone array.

3. A method according to claim 1 or claim 2, wherein the at least one sensor comprises a vision sensor, preferably a camera.

4. A method according to any of the preceding claims, wherein taking the input comprises recording the single mono signal.

5. A method according to claim 4, further comprising recording the single mono signal using a single microphone.

6. A method according to claim 5, wherein the single microphone is of a higher quality than a microphone array used for capturing said data about the sound scene.

7. A method according to any of the preceding claims, further comprising: estimating, based at least on the captured data, acoustic parameters comprising at least a parameter of diffuseness of the sound scene.

8. A method according to any of the preceding claims, further comprising: estimating, based at least on the captured data, acoustic parameters comprising at least a reverberation time of a physical space of the sound scene.

9. A method according to any of the preceding claims, further comprising: estimating, based at least on the captured data, acoustic parameters comprising at least a direct to reverberant ratio between the sound sources and an apparatus capturing said data.

10. An apparatus comprising at least one processing core, at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processing core, cause the apparatus at least to: cause taking as an input a single mono signal; cause capturing data about a sound scene using at least one sensor; estimating, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources; and determining a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

11. An apparatus according to claim 10, wherein the at least one memory and the computer program code are further configured to, with the at least one processing core, cause the apparatus to perform a method according to any of claims 2 – 9.

12. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus at least to: take as an input a single mono signal; capture data about a sound scene using at least one sensor; estimate, based at least on the captured data, acoustic parameters comprising at least a number of sound sources in the sound scene and a location of each of the sound sources; and determine a spatial format representation of the sound scene based at least on the acoustic parameters and the single mono signal.

13. A computer program according to claim 12, further comprising instructions which, when executed by an apparatus, cause the apparatus at least to a method according to any of claims 2 – 9.

Citation Information

Patent Citations

  • Determination of spatialized virtual acoustic scenes from legacy audiovisual media

    US10721521B1

  • Methods and systems for recording mixed audio signal and reproducing directional audio

    US20210092514A1

  • Network-based processing and distribution of multimedia content of a live musical performance

    US20210204003A1