Devices and methods for capturing and processing audio signals

The handheld electronic device with a specialized microphone and processing circuitry enhances the quality of spatial audio capture and processing, improving computational efficiency and accuracy in localizing and separating sound sources in a 3D audio scene.

WO2025195578A1PCT designated stage Publication Date: 2025-09-25HUAWEI TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/057220
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing electronic consumer devices, such as smartphones, face limitations in capturing and processing spatial audio due to restricted microphone placement, acoustic scattering, and limited processing power, which affect the quality and accuracy of 3D sound recording and reproduction.

Method used

A handheld electronic consumer device with a specialized microphone array geometry, such as a right-angled tetrahedron or linear arrangement, and integrated processing circuitry performs pre-processing, localization, blind separation, disambiguation, and packaging of spatial audio signals to enhance sound source identification and ambient sound separation.

Benefits of technology

The solution allows for high-quality capture and processing of spatial audio, improving computational efficiency and accuracy in localizing and separating sound sources in a 3D audio scene, enabling offline recording and transmission of data for further audio processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024057220_25092025_PF_FP_ABST
    Figure EP2024057220_25092025_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing apparatus (120; 110) is disclosed comprising processing circuitry (121; 111) configured to implement the following processing stages for processing a plurality of spatial audio signals captured by a microphone array (116a-d): (a) pre- processing a mixture of the plurality of spatial audio signals captured by the microphone array (116a-d); (b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by the microphone array (116a-d); (c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatial audio signals captured by the microphone array (116a-d); (d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatial audio signals captured by the microphone array (116a-d); (e) performing a disambiguation of blind separation based on the localization results of processing stage c); (f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format; (g) performing an identification of ambient sound for separating the ambient sound from the mixture of the plurality of spatial audio signals; and (h) adding the separated ambient sound to the spatial multi-channel audio format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Devices and methods for capturing and processing audio signals TECHNICAL FIELDThe present disclosure relates to audio technology in general. More specifically, the disclosure relates to devices and methodsfor capturing and processing audio signals, in particular 3D spatial audio. BACKGROUND Smartphones, tablet computers and other types of electronic consumer devices have been produced in large quantities and havebecome prevalent across the world. Often, these devices are small, thin, portable, handheld, and have a small weight. Moreover,they are generally equipped with microphones and loudspeakers as well as processing resources usually in the form of severalmicroprocessors. Therefore, electronic consumer devices like smartphones and tablet computers are portable and accessibledevices that can perform data acquisition and processing of audio signals representing, for instance, a 3D spatial audio scene.Once recorded with an electronic consumer device like a smartphone, spatial audio data can then be transmitted to other audio processing devices for the purpose of 3D sound auralization. For example, a smartphone recording done by a person at a concertmay be transmitted (offline) to a remote computer in some other person`s room, and then rendered over headphones in orderto give the person in the room the impression of hearing spatial sound as if they were present at the concert (such a scenario isillustrated in figure 2a). Reproducing spatial audio in this manner can achieve better results (objectively and subjectively) ifthe original audio scene is recorded as an object-based audio format. Such a format may be used to separately transmit thesignals belonging to individual elements that make up the audio scene. Generally, these elements are referred to as objects.Alongside the signals of objects, metadata may also be transmitted, which may contain the position in 3D space of a staticobject, the trajectory in 3D space of a moving object, sound radiation models for an object, and / or a categorization label for anobject for identifying the type of an object.Typically, the objects in an audio scene are categorized into primary sources and ambience (as illustrated in the scenario shownin figure 2b, where, by way of example, the primary sound sources are a guitar, drums and a singer). The primary sources aredefined as the sound sources in the scene that are most significant based on a criterion such as the greatest generated soundlevel, direct sound from source as opposed to reflections from environment, and / or preferential (e.g. prioritizing human speech).The ambience is defined as all present sound that is not considered primary sources.Given a 3D audio scene, individual signals and individual metadata can be obtained for each primary source and for the ambience by processing the data captured from a microphone array. For example, certain array processing algorithms can identify the direction of arrival (DoA) of the loudest sound sources in a 3D audio scene, while some other algorithms can separate the individual signals of such sources from the scene recording. However, there are limitations when it comes to capturing and processing of spatial audio with an electronic consumerequipment like a smartphone. Firstly, recording of 3D sound requires several microphones with locations spread in space, buthandheld electronic consumer devices like smartphones have small dimensions so that the usable number of microphones andthe distances between them is restricted. Moreover, acoustic scattering due to smartphone geometry and material constructionmay contaminate the data recorded by the mounted microphones. Furthermore, the ability to process spatial audio is governedby the processing power of the chip of the electronic consumer device.SUMMARYIt is an objective to provide improved devices and methods for capturing and processing audio signals. The foregoing and other objectives are achieved by the subject matter of the independent claims. Further implementation formsare apparent from the dependent claims, the description, and the figures.According to a first aspect an audio processing apparatus is provided, wherein the audio processing apparatus comprisesprocessing circuitry, such as one or more processors, configured to implement the following processing stages for processinga plurality of spatial audio signals captured by a microphone array:a) pre-processing a mixture of the plurality of spatial audio signals captured by the microphone array;b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array;c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatialaudio signals captured by the microphone array;d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatialaudio signals captured by the microphone array;e) performing a disambiguation of blind separation based on the localization results of processing stage c);f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing an identification of ambient sound for separating the ambient sound from the mixture of the plurality ofspatial audio signals; andh) adding the separated ambient sound to the spatial multi-channel audio format.Thus, the audio processing apparatus according to the first aspect allows combining several audio signal processing algorithmshaving a beneficial performance when it comes to localizing or separating sound sources in a 3D audio scene recorded by amicrophone array. In this instance, beneficial performance refers to computational costs of the algorithms, the ability of thealgorithms to resolve a diversity of potential recorded audio signals, and quality of the outcome produced by the algorithms(i.e. accuracy of the localization or fidelity of the separated signals). In a further possible implementation form, the audio processing apparatus is an application server configured to receive theplurality of spatial audio signals captured by the microphone array via a data network.In a further possible implementation form, the plurality of spatial audio signals are captured by a microphone array comprisingat least four microphones, wherein each microphone captures a respective spatial audio signal component and wherein the fourmicrophones of the microphone array are arranged (a) at the corners of a notional right-angled tetrahedron located on a frontwall and a rear wall of a housing having substantially the shape of a rectangular cuboid or (b) along a straight line extendingalong a first side wall or a second side wall of the housing having substantially the shape of a rectangular cuboid.According to a second aspect a handheld electronic consumer device for capturing spatial audio is provided, wherein thehandheld electronic consumer device comprises a housing having substantially the shape of a rectangular cuboid with a primaryfront wall (also referred to as front face herein), a primary rear wall opposite to the front wall, a first side wall, a second sidewall opposite to the first side wall, a top wall and a bottom wall opposite to the top wall. Moreover, the handheld electronicconsumer device according to the second aspect comprises a microphone, i.e. transducer, array comprising at least fourmicrophones, wherein each microphone is configured to capture a respective spatial audio signal component and wherein the four microphones of the microphone array are arranged (a) at the corners of a notional right-angled tetrahedron located on thefront wall and the rear wall of the housing or (b) along a straight line extending along the first side wall or the second side wallof the housing. The handheld electronic consumer device according to the second aspect with its special microphone arraygeometry allows capturing spatial audio with high quality.In a further possible implementation form, the housing comprises a display screen providing at least a portion of the front wallof the housing and a casing defining the rear wall, the first side wall, the second side wall, the top wall and / or the bottom wallof the housing.In a further possible implementation form, a first microphone, a second microphone and a third microphone of the right-angledtetrahedron microphone array are arranged on the rear wall of the housing and a fourth microphone of the right-angledtetrahedron microphone array is arranged on the front wall of the housing.In a further possible implementation form, the right-angled tetrahedron microphone array is located in the vicinity of one ofthe corners of the housing.In a further possible implementation form, the edges of the right-angled tetrahedron have a length in the range from 10 mm, inparticular 20 mm to 30 mm. In a further possible implementation form, the microphones of the microphone array are arranged along a straight line extendingalong the first side wall or the second side wall of the housing with substantially uniform or non-uniform distances betweenadjacent microphones of the microphone array.In a further possible implementation form, the microphones of the microphone array are configured to be gain calibrated andphase-matched before capturing the respective spatial audio signal component.In a further possible implementation form, the handheld electronic consumer device comprises an audio processing apparatusaccording to the first aspect for processing the plurality of spatial audio signals captured by the microphone array.In a further possible implementation form, the handheld electronic consumer device comprises a communication interfaceconfigured to provide the plurality of spatial audio signals captured by the microphone array to a remote audio processingapparatus, such as an application server, for processing the plurality of spatial audio signals captured by the microphone array.In a further possible implementation form, the handheld electronic consumer device is a smartphone or tablet computer.According to a third aspect an audio processing method is provided, wherein the audio processing method comprises thefollowing processing steps for processing a plurality of spatial audio signals captured by a microphone array:a) pre-processing a mixture of the plurality of spatial audio signals captured by the microphone array;b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array;c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatialaudio signals captured by the microphone array;d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatialaudio signals captured by the microphone array;e) performing a disambiguation of blind separation based on the localization results of processing step c);f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing an identification of ambient sound for separating the ambient sound from the mixture; andh) adding the separated ambient sound to the spatial multi-channel audio format.The audio processing method according to the third aspect can be performed by the audio processing apparatus according tothe first aspect. Thus, further features of the audio processing method according to the third aspect result directly from thefunctionality of the audio processing apparatus according to the first aspect as well as its different implementation forms andembodiments described above and below.According to a fourth aspect a computer program product is provided, comprising a computer-readable storage medium forstoring program code which causes a computer or a processor to perform the method according to the third aspect, when theprogram code is executed by the computer or the processor. Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS In the following, embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which: Fig.1 is a schematic diagram illustrating an example of a handheld electronic consumer device according to an embodiment in the form of a smartphone with a microphone array for capturing audio signals, an audio processing apparatus according to an embodiment for processing the audio signals, and a headset for rendering the processed audio signals;Figs.2a and 2b are schematic diagrams illustrating exemplary application scenarios for an handheld electronic consumer deviceaccording to an embodiment in the form of a smartphone with a microphone array for capturing audio signals, and an audio processing apparatus according to an embodiment for processing the audio signals; Fig. 3 is a schematic diagram illustrating in more detail processing blocks and stages implemented by a handheld electronicdevice according to an embodiment with a microphone array for capturing audio signals and / or an audio processing apparatusaccording to an embodiment for processing the audio signals; Fig.4 shows different views of a housing of a handheld electronic device according to an embodiment; Figs.5a, 5b and 5c show a portion of a front wall, a portion of a rear wall as well as different uses cases for a handheld electronicdevice in the form of a smartphone according to a first embodiment;Figs.6a, 6b and 6c show a portion of a front wall, a portion of a rear wall as well as different uses cases for a handheld electronic device in the form of a smartphone according to a second embodiment;Fig. 7 shows for three microphones located at different locations of an exemplary smartphone housing the geometry (distanceaway of source acoustic centres not to scale), a plot of impulse responses for different DoAs, a plot of the magnitude of the impulse response frequency spectrum for different DoAs, and a plot of the phase of the impulse response frequency spectrum for different DoAs;Figs. 8a, 8b and 8c show more detailed views of some of the plots of figure 7, namely a plot of the magnitude of the frequencyspectrum of the impulse response for different DoAs on the left-hand side, and a plot of the phase of frequency spectrum of the impulse response for different DoAs on the right-hand side;Figs. 9a and 9b show differences in magnitude (figure 9a) and phase (figure 9b) of the impulse response frequency spectrumfor a reference electret capsule (referred to as “MicA”) versus 32 other electrets capsules, when measured with a two-portacoustic coupler device;Figs. 10a and 10b show diagrams illustrating the loudspeaker positions relative to a smartphone-like prototype, all placed inanechoic laboratory conditions and used to measure far-field impulse responses that describe the sound scattering from the smartphone shape;Fig. 11 shows a diagram illustrating the signal processing chain implemented by a data processing apparatus according to anembodiment for primary sound source and ambient sound identification using a microphone array mounted on a handheldelectronic consumer device according to an embodiment in the form of a smartphone; andFig. 12 is a flow diagram illustrating steps of an audio processing method according to an embodiment.In the following, identical reference signs refer to identical or at least functionally equivalent features. DETAILED DESCRIPTION OF THE EMBODIMENTS In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims. For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units eachperforming one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated inthe figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise. Before describing detailed embodiments in the following some technical background as well as terminology will be introduced making use of one or more of the following acronyms and abbreviations:DoA Direction of arrivalFFT Fast Fourier TransformMSCoh Magnitude Squared CoherenceSVD Singular Value DecompositionSTFT Short Term Fourier TransformLMS Steepest descent least means squaresLCMV Linearly Constrained Minimum VarianceSORTE Second Order Statistics of the Eigenvaluesp-MUSIC Pseudospectrum Multiple Signal Classificationaux-IVA Auxiliary Independent Vector AnalysisHOA High Order AmbisonicsPCB Printed circuit boardMEMS Micro-electromechanical systemsAs used herein, a “signal” may comprise data and information.As used herein, “noise” can refer to sound, but also can refer to non-audio signals. More specifically, “noise” refers to soundor behaviour of data in a signal that is considered unwanted or has no use. As used herein, a “spatial audio scene” is a scenario consisting of several sources of sound placed and radiating in a 3D environment.As used herein, “sound source identification” refers to the characterization of physical properties of a sound source within aspatial audio scene (e.g. position of source in space and source signal separated from scene).As used herein, a “primary sound source” is a significant sound source in an audio scene, based on a chosen criterion.As used herein, “ambient sound” refers to all the sound in an audio scene that is not a primary sound source, i.e. that is not ofmain interest.As used herein, an “audio object” is an individual element within a recorded audio scene (e.g. a sound source, an ensemble ofsound sources, environment reverberation), which typically consists of a signal (representing the variations of the sound solelydue to the source) and of some metadata.As used herein, “metadata” refers to information about an audio object that is not the audio signal itself, such as the position inspace, label of primary source or ambience.As used herein, “audio format” refers to a data structure or protocol for arranging audio data and any associated metadata suchthat it can be transmitted between devices. Audio data is typically segmented into independent entities called channels. As used herein, an “audio channel” refers to an independent entity within an audio format that represents an individual instance / stream of data. A channel is typically associated with the signal from one microphone, but it can also be associated with the signal of an audio object instead, or with the metadata from an audio scene.As used herein, “object-based audio” refers to audio formats where the signal of each object within an audio scene is transmittedas an individual audio channel, while all metadata is typically transmitted as one separate channel.As used herein, “audio mixture” refers to the recording of an audio scene from one or more microphones. At one microphone position, the recording contains the combination of sound from primary sources and from ambience.As used herein, “real-time” and “offline” refer to the transmission of data as it is created in the given application (typicallydone cyclically after a short time interval), or as one single chunk after some arbitrary amount of time has passed. As used herein, “free-field” refers to an idealized, theoretical 3D environment where no sound waves are reflected back frominfinite distance away. In other words, an infinitely large 3D open space filled with a medium in which sound can travel.As used herein, “far-field” refers to a situation where the distance between an acoustic source and an acoustic receiver hasbecome large enough such that the dimensions of the object would no longer affect the sound propagation between the two, at all wavelengths of interest, under idealized conditions. As used herein, a “filter” is a system used as a stage in a signal processing algorithm in order to alter the properties of a signal.As used herein, a “frequency spectrum” is the representation of a signal as the summation of multiple sinusoidal componentsof different contributions via Fourier analysis. Specifically, the term “frequency” is utilized when such representation is used to describe the variation of a physical quantity with time. As used herein, “digital” refers to a characteristic attributed to a system or device. This characteristic implies capturing, storing and manipulation of the data (signals) from the real world at specific moments in time called samples (e.g. a computer).As used herein, “auralization” refers to the process of recreating a recorded 3D audio scene via a system of loudspeakers orheadphones, in an environment different from that of the initial recording. Auralization typically utilizes psychoacoustic properties of binaural human listening to optimize the rendered illusion.As used herein, “informed” and “blind” in the context of a sound source identification algorithm denotes whether or not priorknowledge about the audio scene is utilized to generate outcomes.As used herein, “disambiguation” refers to the process of associating the signals generated by a blind sound source separationalgorithm with corresponding metadata, such as source location or source type (e.g. speech, sound instrument). As used herein, “sampling” refers to the discretization of values associated with the domain of a function representing thevariation of a physical quantity. Typically, sampling performed by digital systems and uniformly, maintaining the same distancebetween any two consecutive samples.As used herein, an “impulse response” refers to a function of time and space that describes the linear physics which govern thevariation of physical quantities within a given set of physical circumstances between a source of energy and a sensor (i.e. medium, boundary conditions, initial conditions, geometry).As used herein, “aliasing” refers to features that appear as error in the frequency spectrum of a sampled signal. Aliasing isstrictly related to how the signal is sampled and how much information is contained between samples in the original continuoussignal. This information is lost after sampling, which leads to the error in the frequency spectrum.As used herein, “convolution” refers to the mathematical process of combining the audio signal from a sound source with theassumed physics between the source and a known receiver in order to obtain the sound generated at that receiver.As used herein, “acoustically rigid” characterizes a surface to have a significantly large acoustic impedance in the normaldirection, i.e. the surface exhibits close to zero acoustic particle velocity such that sound waves cannot pass through it. This term can also be used to describe a sound propagation medium sharing a separating surface with another propagation medium.As used herein, “directivity” refers to a property of acoustic sensors and acoustic sources that describes how they acquire and,respectively, radiate sound as a function of position in space or of DoA around them. This property may also depend on thefrequency of the involved sound.Figure 1 is a schematic diagram illustrating a handheld electronic consumer device 110 according to an embodiment in theform of a smartphone 110 with a microphone array 116a-d for capturing audio signals, a remote audio processing apparatus120 according to an embodiment (for instance, a cloud application server 120) for processing the audio signals captured by thehandheld electronic consumer device 110, and an audio rendering apparatus 130 in the form of a headset or headphones 130for rendering the processed audio signals. Although in the embodiment shown in figure 1, the handheld electronic consumer device 110 is implemented as a smartphone with the usual formfactor of a smartphone, in other embodiments the handheldelectronic consumer device 110 may be, for instance, a tablet computer or another handheld electronic consumer device havinga similar formfactor as a smartphone. As illustrated in figure 1, in addition to the array of microphones 116a-d the smartphone 110 may comprise a processing circuitry 111, e.g. one or more processors 111 configured to process data for providing the functionality described in more detail further below. The processing circuitry 111 may be implemented in hardware and / or software and may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or general-purpose processors.As illustrated in figure 1, the smartphone 110 may further comprise a communication interface 113 configured to communicatewith the audio processing apparatus 120 (for instance a cellular and / or Wi-Fi interface 113 for communicating via a cellular orWi-Fi communication network 115a with the remote audio processing apparatus 120). The smartphone 110 may furthercomprise a memory 115 configured to store executable program code which, when executed by the processing circuitry 111,causes the smartphone to perform the functions and methods described herein. Similar to the smartphone 110 the audiorendering apparatus 130, in particular headphones 130 may comprise one or more processors 131, a communication interface133 for communication with the audio processing apparatus 120 via a cellular or Wi-Fi communication network 115b and amemory 135 as well.As further illustrated in figure 1, the audio processing apparatus 120 may be a cloud application server 120. The audioprocessing apparatus, in particular cloud server 120 may comprise processing circuitry 121, e.g. one or more processors 121 configured to process data for providing the functionality described in more detail further below. The processing circuitry 121 may be implemented in hardware and / or software and may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs), field-programmable gatearrays (FPGAs), digital signal processors (DSPs), or general-purpose processors. As illustrated in figure 1, the data processingapparatus 120 may further comprise a communication interface 123 configured to communicate with the smartphone 110, forinstance, via the Wi-Fi or cellular communication network 115a and / or with the audio rendering apparatus 130, for instance,via a Wi-Fi or cellular communication network 115b. The data processing apparatus, in particular cloud server 120 may furthercomprise a memory 125 configured to store executable program code which, when executed by the processing circuitry 121,causes the data processing apparatus, in particular cloud server 120 to perform the functions and methods described herein.In the following several detailed embodiments of the handheld electronic consumer device 110 with a smartphone formfactoras well as several detailed embodiments of the audio processing apparatus 120 will be described. Generally, according toembodiments disclosed herein the array of microphones 116a-d of the handheld electronic consumer device 110 with asmartphone formfactor comprises four microphones 116a-d, wherein each microphone 116a-d is configured to capture a spatialaudio signal and wherein the four microphones 116a-d of the microphone array 116a-d are arranged (i) at the corners of a right-angled tetrahedron located on a front wall 112a and a rear wall 112b of the housing 112 of the handheld electronic consumerdevice 110 (as illustrated in figures 5a-c) or (ii) along a straight line extending along a first side wall 112c or a second side wall112d of the housing 112 (as illustrated in figures 6a-c).The plurality of spatial audio signals captured by the microphone array 116a-d of the smartphone 110 may be processed by theprocessing circuitry 111 of the smartphone 110 itself or by the remote audio processing apparatus 120, e.g. application server120, which is configured to receive the plurality of spatial audio signals captured by the microphone array 116a-d from thesmartphone via the communication network 115a. For processing the plurality of spatial audio signals captured by themicrophone array 116a-d the smartphone 110 and / or the audio processing apparatus 120 may implement the followingprocessing stages, which will be explained in more detail below:a) pre-processing a mixture of the plurality of spatial audio signals captured by the microphone array 116a-d;b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array 116a-d;c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatialaudio signals captured by the microphone array 116a-d;d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatialaudio signals captured by the microphone array 116a-d;e) performing a disambiguation of blind separation based on the localization results of processing stage c);f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing an identification of ambient sound for separating the ambient sound from the mixture of the plurality ofspatial audio signals; andh) adding the separated ambient sound to the spatial multi-channel audio format.Thus, embodiments disclosed herein allow offline recording of a 3D audio mixture via the microphone array 116a-d mountedon the handheld electronic consumer device 110 with a smartphone formfactor, offline processing of the 3D audio mixture toidentify and describe the primary sound sources, offline processing of the 3D audio mixture to identify and describe the ambientsound, and offline packaging of the separated primary sound sources and separated ambient sound into a multi-channel audiodata format composed of signals and respective metadata for rendering by the audio rendering apparatus 130.Figures 2a and 2b are schematic diagrams illustrating exemplary application scenarios for the handheld electronic consumerdevice 110 according to an embodiment and the audio processing apparatus 120 according to an embodiment. Figure 3 is aschematic diagram illustrating in more detail processing blocks and stages implemented by the handheld electronic consumerdevice 110 according to an embodiment and / or the audio processing apparatus 120 according to an embodiment. As illustratedin figure 3, audio signals may be captured by the microphone array 116a-d of the smartphone 110 in different acoustic environments, such as an indoor auditorium, an open-air stage, household rooms, vehicle cabins, industrial spaces, publicgathering spaces, outdoor spaces and the like. The potential sources of sound may be, for instance, human speech, humansinging, musical instruments, machinery noise, environmental noise and the like. For all the scenarios illustrated in figure 3,the acoustic waves generated by the sound sources in air may be recorded by the handheld electronic consumer device 110, inparticular smartphone 110 with its microphone array 116a-d. For instance, the smartphone 110 may record sound with its microphones 116a-d when held by a user against one ear in a phone conversation position, when held in a video recordingposition with one or two hands, or when placed resting on a surface. With the special microphone array geometry of thesmartphone 110 according to embodiments disclosed herein, the spatial characteristics of the sources of sound in the givenenvironment may be inherently captured in the recorded audio signals. In this way, data describing a 3D audio scene may beacquired in digital form by means of the processing circuitry 111 of the smartphone 110. As already mentioned above, in an embodiment, the audio signals recorded in the manner described above are intended to be later transmitted and processed by the remote data processing apparatus 120. The data processing apparatus 120 is configured to apply sound source identification algorithms to the 3D audio scene captured as multiple channels of audio signals (i.e. one per microphone). The goal of the identification algorithms is to separate the primary sound sources signals from the 3D audio scene and identify their location in space, while also separating the ambient sound ina format that preserves its spatialcharacteristics. A criterion for classifying sound sources as primary versus ambient may be chosen for the previously outlinedidentification process. The signals from all sources of sound separated from the 3D audio scene may then be packaged by the data processing apparatus 120 into a multi-channel audio format alongside the relevant metadata (e.g. primary source position)intended to be further transmitted to the auralization device 130, e.g. headphones 130.The auralization device 130, e.g. headphones 130 receives the multi-channel audio format from the data processing apparatus120 and feeds it through rendering algorithms for 3D audio scenes (e.g. High order Ambisonics). The goal of the rendering algorithms is to recreate the same auditory perception as the original sound scene, either via loudspeakers placed in a new environment compared to the one in which the audio recordings, or via headphones. Some situations in which the former applies include living room audio, vehicle cabin audio, a recording studio, a movie theatre, and a virtual reality room. Headphones can be used during a number of activities, both indoor and outdoor, for example, while working at a desk, while taking care of the household, while doing sports, or while walking / commuting.For processing the plurality of spatial audio signals captured by the microphone array 116a-d in the way described above, thesmartphone 110 and / or the audio processing apparatus 120 may utilize, i.e. implement, one or more of the following algorithms(some of which will be described in more detail below): ^Magnitude Squared Coherence (MSCoh) – as disclosed in [MSCoh2023] (the list of references is provided at the endof the description); ^Singular Value Decomposition (SVD) – as disclosed in [SVD2023];^ Short Term Fourier Transform (STFT) – as disclosed in [HZ2019];^ Multi-channel steepest descent least means squares (LMS) error minimization – as disclosed in [SJE2001];^ algorithm for estimating the number of primary sources from a multi-channel microphone recording of a 3D audioscene, such as an algorithm for source enumeration from 2ndorder statistics of eigenvalues (SORTE); ^algorithm for localizing the positions of the primary sound sources from a multi-channel microphone recording of a3D audio scene, such as the pseudospectrum based multiple signal classification (p-MUSIC) algorithm; ^algorithm for separating the signals of the primary sound sources from a multi-channel microphone recording of a3D audio scene, such as linearly constrained minimum variance (LCMV) beamforming and / or auxiliary independentvector analysis (aux-IVA) algorithm. Algorithm for source enumeration from 2ndorder statistics of eigenvalues (SORTE), as its name suggests, is a general signal processing strategy that analyses the ordered set of eigenvalues associated with a signal`s auto-spectrum density, which can be obtained via SVD, for example. The algorithm processes the difference in magnitude (gap) between any two consecutiveeigenvalues in the ordered set obtained from SVD, clusters the gaps between each of the first ^ largest eigenvalues as a new setand computes the variance of the new set of gaps. Ultimately, the goal is to obtain how this variance scales from previousvalues as ^ is increased, i.e. as more and more of the gaps between the first few largest eigenvalues are used to compute thevariance. In the case of a sensor array that acquires signals generated by sources, the process described above applied to the auto-spectrum density matrix of the sensor data yields an estimate for the number of the uncorrelated sources with the largest contributions tothe sensor signals. The estimated number of sources is equal to the number of first eigenvalues ^ for which the variance of theset of gaps scales the least compared to the variance of the set of gaps created by the first ^ − 1 eigenvalues. A detailedexplanation of the theory behind SORTE is found in [ZhetAl2010].Let there be a coordinate system placed at the centre of an array of ^ microphones, such as the microphone array 116a-d ofthe smartphone 110, which are situated at the Cartesian coordinates ^^ = [^^ ^^ ^^]^ , ^ ∈ {1, 2 … ^}, and let the signalscaptured by these sensors be ^(^), where ^ is the sample index in the discretised time domain and ^(^) is a column vectorwith ^ elements, each one corresponding to a microphone channel. Also, let there be ^ < ^ primary acoustic sourcespositioned in 3D space that generate sound at the ^ microphones and let ^^ = ^ be a known, correct estimation of the numberof these sources. The SORTE algorithm can be implemented following the steps below: 1) Compute the STFT of ^(^) with chosen parameters for the windowing function, the overlap between windows, thenumber of frequency bins ^ used to compute the FFT in each window, and the sampling frequency.Return the frequency spectrum ^(^) as a ^ × ^ matrix, where ^ is the index of the frequency bin and ^ is the number of timedomain samples recorded by the microphones. 2) Compute the auto-spectrum density matrix ^^^(^) = ^{^^(^) ∙ ^(^)} =^ ^∙ ^^(^) ∙ ^(^) for each frequency bin ^,where ^{∙} represents the expected value operator. ^^^(^) has the size ^ × ^.3) Compute the SVD of the ^^^(^) = ^(^) ∙ ^(^) ∙ ^^(^) for each frequency bin ^, where:^ ^(^) = diag(^^ , ^^ ⋯ ^^^^ , ^^) is a square diagonal matrix containing the ^ eigenvalues ≥ ^^ ≥ ⋯ ≥of ^^^(^) on its diagonal, ordered from highest to lowest magnitude, where ^ ≥ ^^ because ^ <^, ^^(^) is an ^ × ^ matrix of the left eigenvectors,^ ^(^) is an ^ × ^ matrix of the right eigenvectors,The matrix ^^^(^) must be full rank for this to be possible.4) Compute ∇^^ = ^^ − ^^^^ for each ^ ∈ {1, 2 … ^ − 1}, i.e. the ^ − 1 magnitude differences (gaps) between each twoconsecutive eigenvalues in the ordered set. 5) For each value of ^ ∈ {1, 2 … ^ − 1}, compute the variance of the sequence {∇^^}^^^^^^, which is given by 6) For each value of ^ ∈ {1, 2 … ^ − 2}, compute the quantity SORTE(^), which is given by. 7) Compute the minimum value of SORTE(^) for ^ ∈ {1, 2 … ^ − 2}, which is the estimated number of sources capturedin the recordings of the microphones.The pseudospectrum based multiple signal classification (p-MUSIC) algorithm is an informed data processing technique forsensor arrays that localizes the uncorrelated sources with the largest contributions to the sensor signals. P-MUSIC is classified as a subspace method where the auto-spectral density matrix of the sensor data is analysed using the SVD and the obtained eigenvalues are ordered based on their magnitude. The eigenvalues with the largest magnitude that are all above an established threshold are assumed to each correspond to a source acquired by the sensors, while all the eigenvalues below the magnitude threshold form a subspace considered to be noise. There are several well-established variants of the MUSIC algorithm available in the literature. The p-MUSIC variant analyses the frequency spectra of the signals in narrow bands and requires prior knowledge of the underlying physics (i.e. impulse responses) between each sensor and a finite set of DoAs around the sensor array or a finite grid of 3D positions around the sensor array. Furthermore, a quantity denoted as the pseudospectrum is computed using the known physics for each DoA in the set or each location on the 3D grid. The pseudospectrum represents the result of an orthogonality check between the knownimpulse responses and the noise subspace eigenvectors obtained initially from the SVD. Under the assumptions of MUSIC, thepseudospectrum exhibits a maximum magnitude for the DoAs or 3D grid positions that correspond to sources because the known impulse response vectors are orthogonal to the eigenvectors of the noise subspace. The version of MUSIC first proposed in the literature is [RS1986].Let there be a coordinate system placed at the centre of an array of ^ microphones, such as the microphone array 116a-d ofthe smartphone 110, which are situated at the Cartesian coordinates ^^ = [^^ ^^ ^^]^ , ^ ∈ {1, 2 … ^}, and let the signalscaptured by these sensors be ^(^), where ^ is the sample index in the discretised time domain and ^(^) is a column vectorwith ^ elements, each one corresponding to a microphone channel. Also, let there be ^ < ^ primary acoustic sourcespositioned in 3D space that generate sound at the ^ microphones and let ^^ = ^ be a known, correct estimation of the numberof these sources. Lastly, let there be a set of ^ chosen directions ^^ , ^ ∈ {1, 2 … ^}, around the sensor array; these correspondto a 2D grid of co-elevation angles ^^and azimuth angles ^^over which the p-MUSIC algorithm performs its computation ofthe pseudospectrum ^(^, ^^, ^^), where ^ is the sample index in the frequency domain. The matrix ^(^) of size ^ × ^contains the known frequency spectra of the impulse responses that governed the physics between each chosen direction ^^and each microphone; each microphone corresponds to a row in the matrix while each chosen direction corresponds to a columnin the matrix. The p-MUSIC algorithm can be implemented following the steps below: 1) Compute the STFT of ^(^) with chosen parameters for the windowing function, the overlap between windows, thenumber of frequency bins ^ used to compute the FFT in each window, and the sampling frequency.Return the frequency spectrum ^(^) as a ^ × ^ matrix, where ^ is the index of the frequency bin and ^ is the number of timedomain samples recorded by the microphones. 2) Compute the auto-spectrum density matrix ^^^(^) = ^{^^(^) ∙ ^(^)} =^^^∙ ^ (^) ∙ ^(^) for each frequency bin ^,where ^{∙} represents the expected value operator. ^^^(^) has the size ^ × ^.3) Compute the SVD of the ^^^(^) = ^(^) ∙ ^(^) ∙ ^^(^) for each frequency bin ^, where:^ ^(^) = diag(^^ , ^^ ⋯ ^^^^ , ^^) is a square diagonal matrix containing the ^ eigenvalues ≥ ^^ ≥ ⋯ ≥of ^^^(^) on its diagonal, ordered from highest to lowest magnitude, where ^ ≥ ^^ because ^^, ^^(^) is an ^ × ^ matrix of the left eigenvectors,^ ^(^) is an ^ × ^ matrix of the right eigenvectors,The matrix ^^^(^) must be full rank for this to be possible.4) Classify the first ^^ eigenvalues as each belonging to one primary sound source. The remaining ^ − ^^ eigenvalues areconsidered to be part of the noise subspace. As such, let ^^^^^^(^) and ^^^^^^(^) be ^ × ^^ − ^^^ matrices containingonly the eigenvectors of the noise subspace. 5) For each frequency bin ^, compute the pseudospectrum ^(^, ^^ , ^^) for each DoA (^^, ^^), ^ ∈ ^ 1^^,^ ^^, which is givenby where ^^(^) represents the j-th column of the matrix ^ of impulse response, which corresponds to the underlying physicsbetween each of the ^ microphones and the j-th DoA.6) Average the pseudospectrum over all frequency bins ^ to obtain the frequency averaged pseudospectrum: 7) The magnitude ^^^^^^, ^^^^ of the frequency-averaged pseudospectrum can be interpreted as a finite and discretized3D surface given by the Cartesian coordinates (^^ , ^^, |^^|) in a new virtual coordinate system. This surface shouldexhibit exactly ^^ peaks corresponding to the DoAs of the ^ primary sources.Use a peak-finding algorithm on the surface described by ^^^^^^, ^^^^ to localize the primary source locations on the ^^^ , ^^^grid.8) (Optional) In order to improve localization accuracy, restrict the peak finding algorithm to not associate differentprimary sources to two peaks that are within a threshold of proximity to each other on the ^^^, ^^^ grid.Linearly constrained minimum variance (LCMV) beamforming is an informed data processing technique for sensor arrays that involves combining optimally filtered versions of each signal acquired by the sensors such that a desired directivity is obtained from the array, typically, a discrimination of signal captured in all but one chosen direction or for one chosen position in space. The technique relies on knowledge of the underlying physics between each sensor in the array and the chosen direction / position.The linearly constrained minimum variance (LCMV) beamformer is a special case of a minimum variance distortionlessresponse (MVDR) beamformer. The latter forces its output to be unity in a chosen direction or at a chosen position while also minimizing the variance of its output. In other words, any contributions that are not from the chosen direction or position are minimized. In the case of the LCMV, multiple directions or positions are chosen such that some are constrained via the MVDR (forced unity values) while others are constrained to have a null output (forced zero values). The theory behind LCMV and MVDR beamformers are detailed in [BR2015].Let there be a coordinate system placed at the centre of an array of ^ microphones, such as the microphone array 116a-d ofthe smartphone 110, which are situated at the Cartesian coordinates ^^ = [^^ ^^ ^^]^ , ^ ∈ {1, 2 … ^}, and let the signalscaptured by these sensors be ^(^), where ^ is the sample index in the discretised time domain and ^(^) is a column vectorwith ^ elements, each one corresponding to a microphone channel. Also, let there be a LCMV beamformer with ^ chosendirections ^^ , ^ ∈ {1, 2 … ^}, in which to force the beamformer output to either be unity or zero. The matrix ^(^) of size^ × ^, where ^ is the sample index in the frequency domain, contains the known frequency spectra of the impulse responsesthat governed the physics between each chosen direction ^^ and each microphone; each microphone corresponds to a row inthe matrix while each chosen direction corresponds to a column in the matrix. The LCMV beamformer can be implemented following the steps below: 1) Compute the STFT of ^(^) with chosen parameters for the windowing function, the overlap between windows, thenumber of frequency bins ^ used to compute the FFT in each window, and the sampling frequency.Return the frequency spectrum ^(^) as a ^ × ^ matrix, where ^ is the index of the frequency bin and ^ is the number of timedomain samples recorded by the microphones. 2) Compute the auto-spectrum density matrix ^^^(^) = ^{^^(^) ∙ ^(^)} =^^^∙ ^ (^) ∙ ^(^) for each frequency bin ^,where ^{∙} represents the expected value operator. ^^^(^) has the size ^ × ^.3) Compute the inverse of the ^^^(^) for each frequency bin ^ using SVD. The matrix ^^^(^) must be full rank for thisto be possible. 4) Compute the frequency spectrum of the optimal beamforming filters ^^(^) for which frequency bin k, from theequation: ^^(^) = ^^ ∙ [^^(^) ∙ ^^^^^(^) ∙ ^(^)]^^ ∙ ^^(^) ∙ ^^^^^(^), where: ^a superscript of “H” denotes the Hermitian transpose of a matrix,^ ^^(^) is a row vector of ^ elements representing the frequency spectrum of each optimal beamformingfilter at the frequency sample ^, ^^^ is a row vector containing ^ elements that are either 1 (beamformer maximum response) or 0(beamformer null response), where each element corresponds to a chosen direction ^^. 5) Compute the beamformer output ^(^) ∙ ^^(^) for each frequency bin ^ to obtain a column vector with ^ elements.6) Compute the inverse STFT of the beamformer output ^(^) ∙ ^^(^) with the previously used parameters for theanalysis windowing function, the synthesis windowing function, the overlap between windows, the number of frequency bins ^, and the sampling frequency.Return a column vector of ^ elements which represents the time domain of the beamformer output.7) Normalize the time domain beamformer output such that its amplitude values are contained in the range [−1,1].Auxiliary independent vector analysis (aux-IVA) algorithm is a blind data processing technique for sensor arrays that separatesthe signals generated by sources at the sensors based on maximizing the independence between the estimates of the extracted source signals. Said independence can be quantified in different ways, of which, a common method is maximizing non- Gaussianity (i.e. dissimilar information). The aux-IVA algorithm belongs to the family of iterative algorithms denoted as independent component analysis (ICA), in contrast to its counterparts, it processes the sensor signals in the frequency domain, while also accounting for inter-sensor delay differences rather than just inter-sensor magnitude differences. The aux-IVA algorithm considers the frequency spectra of the signals acquired by the sensors as the multiplication between an unmixing matrix and the frequency spectra of the original source signal vectors. The algorithm attempts to compute the unmixing matrix, i.e. the underlying physics that govern the physics between each source and each sensor. As such, theunmixing matrix can be considered as the frequency spectra of the impulse responses between each contributing source andeach sensor. However, the original source signal vectors are required to be non-Gaussian in order for the algorithm to achieve its goal. Furthermore, because the algorithm maximizes non-Gaussianity, the algorithm struggles to perform its goal as the sensor signals become a mixture of more and more source contributions. This is a direct result of the well-established central limit theorem, which states that “the distribution of sample means approximates a normal Gaussian distribution as the sample size gets larger, regardless of the population’s distribution”. Algorithms in the ICA family can be implemented in different ways, for example, when it comes to method of quantifying the non-Gaussianity, the method of solving the frequency bin permutation problem, and the method for optimizing the iterationstep size required to converge to a correct estimation. As will be appreciated, the aux-IVA algorithm uses negentropy (a basicconcept in information theory) as the measure of non-Gaussianity; solves the frequency bin permutation problem by computingthe unmixing matrix for the whole spectrogram of the mixture instead of unmixing each frequency individually; and minimizes an auxiliary function when computing the unmixing matrix which circumvents the need for optimizing the iteration step size. The theory behind using aux-IVA in the frequency domain is detailed in [PS1998], while the solution for the permutation problem is explained in [AH2006] and the solution for avoiding the step size optimization is explained in [NO2011]. As already described above, according to embodiments disclosed herein the array of microphones 116a-d of the handheldelectronic consumer device 110 with a smartphone formfactor comprises four microphones 116a-d, which are arranged (a) atthe corners of a right-angled tetrahedron located on a front wall 112a and a rear wall 112b of the housing 112 of the handheldelectronic consumer device 110 (as illustrated in figures 5a-c) or (b) along a straight line extending along a first side wall 112cor a second side wall 112d of the housing 112 (as illustrated in figures 6a-c).As illustrated in figure 4, the formfactor, i.e. the shape of a smartphone may be described as a thickened plate exhibiting lengthand width as well as rounded edges and rounded corners. Thus, the handheld electronic consumer device 110, in particularsmartphone 110 comprises a housing 112 having the shape of a rectangular cuboid (with rounded edges and corners) with a front wall (or front face) 112a, a rear wall (or rear face) 112b opposite to the front wall 112a, a first side wall 112c, a secondside wall 112d opposite to the first side wall 112c, a top wall 112e and a bottom wall 112f opposite to the top wall 112e. Asillustrated in figure 4, exemplary dimensions of the housing 112 are a length of 160 mm, a width of 80 mm, and a thickness of10 mm.Figures 5a, 5b and 5c show a portion of the front wall 112a, a portion of the rear wall 112b as well as different uses cases forthe handheld electronic consumer device 110, in particular smartphone for the first main embodiment, where the fourmicrophones 116a-d are arranged at the corners of a right-angled tetrahedron located on the front wall 112a and the rear wall112b of the housing 112 of the smartphone 110. As can be taken from figures 5a and 5b, in an embodiment, three microphones116b-c may be arranged flush with the surface on the side of the video camera 118a, 118b (i.e. the rear wall 112b), while thefourth microphone 116a is arranged flush with the surface on the side of the display screen (i.e. the front wall 112a). As willbe appreciated, the point-locations of the microphones 116a-d form a right-angled tetrahedron, where the three catheti (theedges that form the right angles) have, by way of example, a length of 20 mm (as illustrated in figures 5a and 5b). Of the threemicrophones 116b-d that share a plane, the microphone 116b closest to the corner may be placed, for instance, 20 mm awayfrom the two closest edges. Consequently, the fourth microphone 116a on the screen side, i.e. front wall 112a may also beplaced 20 mm away from the closest two edges.Figures 6a, 6b and 6c show a portion of the front wall 112a, a portion of the rear wall 112b as well as different uses cases forthe handheld electronic consumer device 110, in particular, the smartphone 110 for the second main embodiment where thefour microphones 116a-d are arranged along a straight line extending along the first side wall 112c or the opposite second sidewall 112d of the housing 112. In other words, for this second main embodiment the four microphones 116a-d may be equallyplaced in a line along the centre of the longest side surface 112c,d of the housing 112 of the smartphone 110. As will be appreciated, the length of the line formed by the positions of the four microphones 116a-d and the separation distance betweentwo adjacent microphones 116a-d dictate the low frequency limit and high frequency limit of the microphone array 116a-d, asexplained in [PAN2004]. Both these two quantities are governed by the length chosen for the smartphone 110. A longer lengthincreases the working frequency range of the microphone array 116a-d to include lower frequencies, but also increases the separation distance which lowers the high frequency limit. Even so, the geometry illustrated in figures 6a and 6b utilizes as much as possible of the available length to distribute the microphones 116a-d equally. Some small distance is required between the two microphone positions nearest to smartphone corners in order to account for the actual physical size of the acousticsensors. This small distance is left arbitrary in order to be adapted to the whichever construction components are available.As will be appreciated, both the tetrahedral array of microphones 116a-d illustrated in figures 5a and 5b and the line array ofmicrophones 116a-d illustrated in figures 6a and 6b may be easily integrated into currently available smartphones, reduce thenumber of utilized sensors while still enabling the capture of spatial characteristics in a 3D audio scene, be located at positionsof the smartphone 110 that are not covered by the smartphone user`s hand during common scenarios, and minimize the effecton the recorded signals generated by the acoustic scattering due to the smartphone 110 shape. Moreover, both the tetrahedralarray of microphones 116a-d illustrated in figures 5a and 5b and the line array of microphones 116a-d illustrated in figures 6aand 6b do not occupy any areas on the surface of the smartphone 110 that are typically used for other components.As will be appreciated, any three points in space form a plane, and a fourth out-of-plane point has a counterpart on the opposite side of the plane that has the same set of distances to the initial three points. In other words, the sound generated at a source location can have the same propagation distances relative to a microphone array with planar positions as those of another sound source location. In free-field, this translates to the two sources generating the same sound signal at each one of the planar microphones. Therefore, having a minimum of four microphones that are not all on a single plane guarantees that inter-mic differences are recorded for the sound generated by a given sound source in most possible scenarios. Such differences are necessary for the performance of some of the sound source identification algorithms utilized according to embodiments disclosed herein. In figure 5c and figure 6c the tetrahedral array of microphones 116a-d and, respectively, the line array of microphones 116a-dare depicted diagrammatically for multiple common use cases of the smartphone 110. These use cases are represented by single-hand video recording position, two-hand video recording position, phone call over-ear position, and when placed resting on a surface. As can be taken from figures 5c and 6c, the microphones 116a-d are not covered by the human hand in all cases forthe tetrahedral array shown in figure 5c and are partially covered only in a few cases for the line array shown in figure 6c.When the microphones 116a-d are covered by the human hand, the sound sources from the surrounding 3D audio scene may have their contributions altered in the recorded signals as the effect of the covering becomes stronger, as explained in [EBB2020]. For the tetrahedral array of microphones 116a-d illustrated in figures 5a and 5b and the line array of microphones 116a-dillustrated in figures 6a and 6b laboratory measurements with a prototype smartphone 110 have been performed, which showthat the acoustic scattering effect created by the shape on the audio signals recorded by the microphones 116a-d mounted onthe smartphone may be minimized for the most DoAs in comparison with other arrangements of the microphones, namely anarrangement where the microphones are located on the four thin side surfaces or an arrangement where the microphones arelocated on the two surfaces of largest area, near the side surfaces. The used prototype contained 32 off-the-shelf electretmicrophones of 6 mm diameter placed in key positions on the shape and mounted flush. The prototype body was constructed out of hardwood, which approximates an acoustically rigid surface in the surrounding air, and with the dimensions 320 mm x160 mm x 20 mm (L x W x T)). A total of 720 sound source positions on a virtual sphere were created in the far-field, at 2 maway from the centre of the prototype, by using an arrangement of 10 loudspeakers (Genelec model 8020C mounted on the semi-circular truss around the prototype and by rotating the prototype on a turntable (in 5°intervals). The sine sweep method [AF2000, AF2007] was used to capture data from the microphones. Further details of this experiment are described further below. The impulse responses obtained from the laboratory results, the magnitude of their frequency spectrum, and the phase of their frequency spectrum are illustrated in figure 7 for three key microphone positions. These are labelled as “Mic50”, “Mic49” and “Mic44”, and are placed, respectively, as far away as possible from all edges on one of the prototype`s surface of largest area, on one of the protype`s longest side surface, and on one of the prototype`s shortest side surface. In figure 7 the positions of the sound sources on the virtual sphere are drawn as different coloured rings of points around each respective microphone position depicted in the diagram. It is important to emphasize that the positions of the sources are shown relative to each microphone and differ from microphone to microphone, even though this is not noticeable in the presented diagram. This occurs because the microphones are situated some distance away off the axis of rotation of the turntable. Furthermore, the red ring in the diagram corresponds to the loudspeaker of co-elevation 0°. This ring has 72 source positions very close to each other for all 32 mics, which is also a result of the mics being situated away off the turntable axis of rotation. The sources are labelled with indices starting with 1, going from one ring to another, starting from the lowest co-elevation (top) on each ring. For a given microphone, the impulse responses, the magnitude of their frequency spectrum, and the phase of their frequency spectrum are plotted on the same graph as colormaps for all positions of sources, where the x-axis is time or frequency, the y-axis is the source index following the order described previously, and the colormap represents amplitude, magnitude, or phase. As will be appreciated, figure 7 illustrates the directivity of the microphone and how this is affected by the presence of the scattering shape. The three colormaps for the impulse responses clearly show where the first largest peak occurs on the x-axis (bright yellow). If all sources around a microphone would be at the same distance away from a free-field microphone, the bright yellow strip on the colormap would be straight and parallel with the y-axis. However, the sources are not all at the same distance away from any given mic as some mic positions are not on the rotation axis of the turntable; these variations are on the scale of millimetres. Furthermore, the scattering shape introduces additional delay to the propagation path due to its dimensions,depending on the microphone location on the scatterer relative to the source. Because of this, the bright yellow strips for“Mic50” and “Mic49”, which are located on larger features of the scattering body, are not very straight compared to the onefor “Mic44”, which introduces minimal delays to the propagation path for the majority of source DoAs due to being situatedon the thinnest feature of the shape. Moreover, from the colormaps for the magnitude of the frequency spectrum, it can be takenfrom figure 7 that, at low frequencies, the response does not vary as much with source DoAs as at high frequencies, for all threemicrophones. The range of magnitude variation with DoAs is 11 - 12 dB in the low frequency regime for all three microphones.However, at high frequencies up to 7500 Hz, the range of variation of magnitude with DoAs is 71 dB for “Mic50” and only 22- 23 dB for “Mic49” and “Mic44”. The regions on the colormap for “Mic50” where the magnitude drops significantly can beclearly distinguished. From the data of the colormaps for the unwrapped phase of the impulse response frequency spectrum, it can be observed that there is some minimal variation with DoAs across the frequency range from 200 Hz to 7500 Hz for “Mic49” and “Mic44”. However, this is not the case for “Mic50”, which showcases significant, distinguishable increases in the descending slope of the unwrapped phase at high frequencies (up to 7500 Hz). These increases signify that, for those regions of the colormap, the scattering shape is introducing delay in addition to the free-field travel path between source and microphone. In figures 8a-c the magnitude and the phase plots for all three microphones are magnified to show the colormap only for the magenta ring of sources, which is situated at 90°co-elevation, i.e. the same elevation as the centre of mass of the scattering shape and same elevation as the acoustic centre of “Mic50”. In figures 8a-c, the colormaps can be considered circular on the y-axis, as only one ring is shown. The sources that illuminate the microphone positions and those that shadow it are marked as regions on the graphs. It can be observed from the plots that “Mic44” which is shadowed by all source positions on this ring, shows minimal variation with DoA in the magnitude and phase of the impulse response frequency spectrum. “Mic49” shows some slight variations with source DoA at high frequencies, which are the most significant for the shadowing sources. “Mic50” also shows variations with source DoA at high frequencies, but these are a lot more significant compared to the other cases. The shadowing positions exhibit significant drops in the magnitude of the frequency spectrum, which start at a much lower frequency than the slight variations observed for “Mic49”. Likewise, the unwrapped phase for the shadowing positions of “Mic50” suffer significant increases in steepness of the descending slope at high frequencies. This indicates that additional delays are introduced to the free-field travel path between source and microphone. Overall, “Mic44”, which is on a surface of 20 mm x 160 mm, shows the least amount of variation with DoAs, followed “Mic49”, which is on a surface of 20 mm x 320 mm, and “Mic50”, which is in the centre of a surface of 160 mm x 320. This suggests than the thinner features of the scatteringobject are more desirable for mounting a microphone if the goal is to have the inter-microphone differences altered as little aspossible by the scattering.The experimental findings shown in figures 7 and 8a-c illustrate solely the sound scattering produced by the smartphone-likeshape. In practice, smartphone devices have an outer shell that engulfs a cavity filled with tightly packed electronic components,all of which can be made from materials with elastic properties. The sound scatterings produced by sound waves impinging onan elastic structure manifest as a contribution from rigid behaviour combined with a contribution form elastic behaviour. This is explained in detail throughout [MJ&DF1986]. However, the elastic contribution of the behaviour is minimal when the acoustic impedance of the structure is significantly greater than that of the surrounding propagation medium, i.e. when the propagation medium changes from one that offers minimum opposition to traveling sound waves to one that offers significant opposition. This is the case for air as a propagation medium and a structure constructed from a dense solid material that lacks cavities filled with air or other light fluids. As such, the sound scattering generated in air solely due to the rigid smartphone shape is a sufficient approximation for the sound scattering generated when the smartphone is constructed out of all the different components commonly utilized in practice. According to an embodiment, the microphones 116a-d of the microphone array are configured to be gain calibrated and phase-matched before capturing the spatial audio signal. As will be appreciated, calibration and phase-matching of the microphones116a-d used for sound source identification is important because both informed and blind algorithms typically rely on differentacoustic sources placed in 3D space generating different behaviours in the signals at one or more microphone positions away.These different behaviours generally depend on the wavelength of the involved sound waves and can manifest in both themagnitude and the phase of the frequency spectrum of the microphone signals.According to an embodiment, a standard method for calibrating and phase-matching electret capsules may be used, asrepresented by the acoustic coupling chamber method. This is explained for magnitude calibration in [GW&TE1892]. The goal is to place two microphone capsules such that their diaphragms are exposed to the inside of a small, sealed chamber. Uniform sound pressure of high enough level is created in the chamber for an intended working frequency range via a compression driver unit, and impulse responses are measured, for example, via the sine sweep method [AF2000, AF2007]. In this way, it is guaranteed that both diaphragms experience the same frequency-dependent magnitude and phase over the working range, regardless of their different position. For an ideally constructed chamber, any differences encountered between the signals measured at the microphones is then a result of construction differences between the two units. These differences can be accounted for via inverse filtering designed base on the results of the measurement. If no significant phase-differences are identified across the working frequency range, the microphones can be calibrated by manipulating only the magnitude of each microphone via real-valued gains. The calibration and phase-matching of two microphones 116a-d can be done either as absolute or as relative. In the former case, one of the two microphones must already be calibrated to output known values of a specific measuring unit. For the latter, the goal is to bring the magnitude and the phase of the frequency spectrum as close as possible between the two microphones. However, in this case, the absolute values in measurement units of the recorded results may be offset from the true value by a scaling factor. When more than two microphones require calibration and phase-matching, one microphone that is chosen as areference, and all other microphones are matched to it respectively. For absolute calibration, the reference microphone may becalibrated beforehand.Exemplary results for the calibration and phase-matching of microphones captured with an exemplary coupling chamber areshown in figures 9a and 9b for 32 electret microphones as differences in magnitude and, respectively, phase of the impulse response frequency spectrum. One reference electret microphone “MicA” was used to compute the differences with all other 32 used electret microphones and the performed calibration was relative. Three of the electret capsules exhibited outlier behaviour different from the rest and, thus, were disregarded. Ignoring the three disregarded microphones, the magnitude plots show that, in the range between 200 Hz and 4000 Hz, the 29 lines are relatively flat and are, approximately, the same behaviour shifted up and down the y-axis. Furthermore, in this range, the maximum gain offset between any two lines is approx.6 dB. This value represents the maximum span of the range of values for the difference in magnitude with the reference mic. Above 4000 Hz, the first resonant peak due to the design of the used exemplary acoustic coupling chamber. In terms of the phase plot, the 29 lines are relatively flat and parallel between 200 Hz and 1000 Hz, and then, some lines start to slightly slope up to 4000 Hz. At 200 Hz, the maximum span of the range of values for the difference in phase with the reference mic is equivalent to an alignment offset distance of 3.1 cm, which represents 1.8% of the wavelength. At 4200 Hz, the maximum span of the range of values for the difference in phase with the reference mic is equivalent to an alignment offset distance of 0.38 cm, which represents 4.6% of the wavelength. Between 200 Hz and 4200 Hz, the alignment offset does not become smaller than 1.8% or greater than 4.6% for the 29 lines. All the findings described above suggest that the electret microphones may be calibrated only by adjusting the real-valued gains for each unit. This is due to phase-matching via inverse filtering being a tedious processwhere ideal results are difficult to achieve. As such, it may not be worth employing to address the small phase offsetsencountered for the measured electret capsules. For gain calibration, the data between 200 Hz and 4000 Hz for each of the 29usable lines in the magnitude plot from figures 9a and 9b was fitted with a first-degree polynomial, i.e. a linear fit. The real-valued coefficient of the zeroth order polynomial was then used as a scaling factor to shift each line on the y-axis in order tobring it as close as possible to the reference mic, i.e. the x-axis. This process preserved the phase to what is shown in figures9a and 9b, and created a new maximum gain offset of approx. 1.7 dB for the magnitude in the range between 200 Hz and 4000Hz. In conclusion, the amount of magnitude and phase differences exhibited after calibrating each mic with a real-valued gain are adequate for using the calibrated data in the sound source identification algorithms. The informed sound source identification algorithms implemented by the data processing apparatus 120 according to an embodiment may operate on the basis that the underlying physics governing the 3D audio scene recorded by the microphonesare known. These physics include: the sound radiation of present sources of sound; the sound propagation path from each soundsource to each microphone; the sound scattering effects from obstacles present in the 3D audio scene; and / or the recording capabilities of each microphone 116a-d as a sensor of acoustic waves. As will be appreciated, the goal of sound source identification with the microphones 116a-d mounted on the handheld electronicconsumer device 110, in particular smartphone 110 should be achieved for a variety of potential 3D audio scenes in order tobe offered as a commercial option to a user base. The physics governing a 3D audio scene depend on many parameters (e.g. position of each sound source and sound receiver relative to each other, frequency spectrum generated by each sound source and its relation to those of all other sound source, presence of sound scattering obstacles and their position relative to sourcesand receivers) and there are too many possible variations to measure individually ahead of time to later inform the algorithms.As such, a representative set of governing physics may be utilized in the informed sound source identification algorithms inorder to resolve a variety of possible 3D audio scene. This representative set may be measured in a controlled setting such as alaboratory, then stored on the smartphone 110 and / or the data processing apparatus 120 to be later used as metadata in the audio signal processing chain. For obtaining the necessary assumed physics the sound scattering properties of a rigid smartphone shape placed in the free-field have been captured by using a finite set of many sound sources on a virtual surface in the far-field and a large number ofmicrophone positions mounted on the smartphone 110 in key locations. As already described above, the rigid smartphone shapeis a sufficient approximation for the sound scattering in air generated when the smartphone 110 is constructed out of all thedifferent components commonly utilized in practice. This approach is effective when all underlying processes that generate thesound in the 3D audio scene are linear. This is because, under linearity any sound scattering obstacles or boundaries of roomspresent in the scene can be considered as equivalent virtual sound sources placed in the free-field around the scattering shape and as generating sound together with the actual acoustic sources. In practice, certain types of sources such as noise generated by large moving machinery or engines and high-level sound generated by large loudspeakers at concerts may not exhibit linear behaviour. Furthermore, the sound scattering properties of the smartphone shape measured only for far-field sources may notbe a good representation of the underlying physics when the sound sources are positioned close to the scattering shape. Theproximity required for the representation to become inaccurate depends on the size of the smartphone shape relative to the frequency of sound waves generated by the sources. An exhaustive measurement of the sound scattering physics requires thatthe finite set of many sound sources be used at multiple distances away from the scattering shape, i.e. a sampling of the volumerather than of a spherical surface.The measurement approach described above was realized in the anechoic conditions of the Doak Laboratory at the Institute ofSound and Vibration research, in Southampton. A prototype smartphone-like device was constructed for this purpose and was designed to house 32 off-the-shelf electret microphones of 6 mm diameter placed in key positions on the shape and mounted flush. The prototype body was constructed out of hardwood, which approximates an acoustically rigid surface in thesurrounding air, and with the dimensions 320 mm x 160 mm x 20 mm (L x W x T)). A total of 720 sound source positions on a virtual sphere were created in the far-field, at 2 m away from the centre of the prototype, by using an arrangement of 10loudspeakers (Genelec model 8020C) mounted on the semi-circular truss around the prototype and by rotating the prototypeon a turntable (in 5°intervals). A diagram of the experiment geometry is presented in figures 10a and 10b. The sine sweep method [AF2000, AF2007] was used to capture data from the mics and then convert this to impulse responses. All the equipment used in the experiment was controlled from MATLAB on a Mac Mini computer via a Dante network. Dante is aprotocol that allows communication over Ethernet cables for certain multi-channel audio equipment, such as specializedsoundcards, ADC / DAC units, and signal acquisition systems, where any signal in the chain can be transmitted to any device while maintaining a constant low latency across all equipment. During the experiment described above, this latency was set in software to the value of 5 ms. As already described above, for processing the plurality of spatial audio signals captured by the tetrahedral array of microphones116a-d illustrated in figures 5a and 5b or the line array of microphones 116a-d illustrated in figures 6a and 6b the smartphone110 and / or the audio processing apparatus 120 may implement the following processing stages:a) pre-processing a mixture of the plurality of spatial audio signals captured by the microphone array 116a-d;b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array 116a-d;c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatialaudio signals captured by the microphone array 116a-d;d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatialaudio signals captured by the microphone array 116a-d;e) performing a disambiguation of blind separation based on the localization results of processing stage c);f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing an identification of ambient sound for separating the ambient sound from the mixture of the plurality ofspatial audio signals; andh) adding the separated ambient sound to the spatial multi-channel audio format.Thus, as will be appreciated, this offline processing chain of the 3D audio mixture captured by the microphones 116a-d maybe used for obtaining: an estimation of the number of primary sound sources present in the audio scene; an estimation of theDoAs of the primary sound sources present in the audio scene; an estimation of the separated signals of the primary soundsources present in the scene; and / or an estimation of the separated signals at each microphone representing the ambient soundin the scene. A detailed block diagram of this processing chain implemented by the smartphone 110 itself and / or the dataprocessing apparatus 120 is shown in figure 11, where: a solid, thin and single-coloured line denotes a single channel audiosignal (such as the exemplary single channel audio signal 1101); a solid, thick, shaded line denotes a multi-channel audio signal(such as the exemplary multi-channel audio signal 1103); a dashed, single-coloured line denotes metadata (such as theexemplary metadata 1105); and the pair (^, ^) refers to co-elevation and, respectively, azimuth about the center of thesmartphone 110. By way of example, the audio signals captured by the microphones 116a-d and processed by the processingchain illustrated in figure 11 are based on three primary sound sources represented by a guitar, a human singer, and a drum, as well as ambient sound consisting of human babble and some environmental noise.Thus, in the scenario illustrated in figure 11, the 3D audio scene consists of ^ = 3 primary sound sources together with ambientsound sources, all recorded with ^ = 4 microphones 116a-d mounted on the smartphone 110 with the tetrahedral geometryillustrated in figures 5a and 5b or the line geometry illustrated in figures 6a and 6b.Let there be a coordinate system placed at the centre of the smartphone 110, where spherical coordinates are presented as co-elevation angle ^ and azimuth angle ^. Let the angular positions of the primary sources be (^^, ^^), (^^ , ^^), and (^^ , ^^) for^ ^In the pre-processing stage (a) the signal processing chain implemented by the smartphone 110 and / or the data processingapparatus 120 prepares the raw multi-channel mixture ^(^) = [^^(^) ^^(^) ^^(^) ^^(^)]^and turns it into the pre-processedmixture ^(^) = [^^(^) ^^(^) ^^(^) ^^(^)]^. This transformation consists of two actions: calibration (see processingblock 1111), and pre-filtering (see processing block 1113). The former involves the adjustment of the microphone signals basedon the calibration data gathered in the way described above. If the microphone sensors 116a-d used in the device 110 are alreadyphase-matched, the calibration 1111 reduces to applying real-valued amplitude gains to each channel. Otherwise, morecomplicated inverse filtering may be required to ensure phase-matching. A second aspect of the pre-processing stage (a)implemented by the smartphone 110 and / or the data processing apparatus 120 may comprise pre-filtering 1113, which typicallyinvolves anti spatial aliasing filters to account for the lower and upper frequency limits imposed by the maximum and,respectively, minimum separation distance between the microphones 116a-d, as described in [PAN2004]. The calibration andprocessing filters should preserve, as best as possible, the inter-microphone differences in magnitude, but more importantly the phase of the frequency spectrum of the original recorded audio signals. This is because the performance of some sound source identification algorithms rely on these differences. After the pre-processing stage (a) in the pre-processing stage (b) of the signal processing chain implemented by the smartphone110 and / or the data processing apparatus 120 an algorithm for estimation of the number of the primary sound sources presentin the recorded 3D audio scene may be used on the mixture ^(^) (see processing block 1115). This yields the estimatednumber of primary sources ^^, which may then be utilized in some of the following stages of the signal processing chain, namelythe informed primary sound source localization (see processing block 1117) in processing stage (c), and the separation of theambient sound (see processing block 1125). The estimation of the number of primary sound sources based on a highest soundlevel criterion may be achieved by the processing block 1115 using the SORTE algorithm already described above. chain illustrated in figure 11 does not perform any action that classifies the primary sound sources contained in the mixture based on their timbre, i.e. whether they are specific musical instruments, human speech, machinery noise, or something else. The resulting signals for the separated ambient sound at each microphone location can be arranged as ^^=[^^(^, ^^) ^^(^, ^^) ^^(^, ^^) ^^(^, ^^)]^.In the approach described above for separating the ambient sound signals (i.e. the processing stage (g) implemented by theprocessing block 1125), the estimated number of primary sound source ^^ is not directly involved. However, this number maybe required when implementing the approach at code level in actions such as pre-allocating the correct amount of memory andrunning loops over fixed number of steps. Such actions help with optimizing runtime of the code and reduce computational times.Ultimately, the separated ambient sound ^^ is integrated as part of the spatial audio format alongside the separated primarysound source in processing stage (h) (which may also be implemented by the processing block 1127). For the basic spatialaudio format described previously, four new channels, one for each ambient signal ^^(^, ^^) separated at the microphoneposition ^^, require to be added to the previous three corresponding to the separated primary source signals. Furthermore, the positions ^^of the four microphones 116a-d in space may be added as an additional metadata channel, such that the spatialproperties of the ambient sound can be later auralized by the audio rendering device 130.Figure 12 is a flow diagram illustrating an audio processing method 1200, which is a summary of the different processingstages (a)-(h) described in the context of figure 11. The audio processing method 1200 illustrated in figure 12 comprises thefollowing processing steps for processing a plurality of spatial audio signals captured by the microphone array 116a-d:a) pre-processing 1201 a mixture of the plurality of spatial audio signals captured by the microphone array 116a-d;b) estimating 1203 the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array 116a-d;c) performing 1205 an informed localization of one or more primary sound sources in the mixture of the plurality ofspatial audio signals captured by the microphone array 116a-d;d) performing 1207 a blind separation of one or more primary sound source signals in the mixture of the plurality ofspatial audio signals captured by the microphone array 116a-d;e) performing 1209 a disambiguation of blind separation based on the localization results of processing step c);f) packaging 1211 one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing 1213 an identification of ambient sound for separating the ambient sound from the mixture; andh) adding 1215 the separated ambient sound to the spatial multi-channel audio format.The method 1200 can be performed by the data processing apparatus 120, e.g. application server 120 according to anembodiment or the electronic device 110, in particular smart phone 110 according to an embodiment. Thus, further features ofthe method 1200 result directly from the functionality of the data processing apparatus 120 and the electronic device 110 aswell as their different embodiments described above and below.The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the present disclosure (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step). In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described embodiment of an apparatus is merely exemplary. For example, the unit division is merely logical function division and may be another division in an actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not bephysical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the unitsmay be selected according to actual needs to achieve the objectives of the solutions of the embodiments.In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

[0002] List of references cited:[AF2000] A. Farina. “Simultaneous Measurement of Impulse Response and Distortion with a Swept-SineTechnique”. In: Audio Engineering Society Convention 108, Paris, 2000.[AF2007] A. Farina. “Advancements in Impulse Response Measurements by Sine Sweeps”. In: AudioEngineering Society Convention 122, 2007.[AH2006] A. Hiroe. “Solution of permutation problem in frequency domain ICA, using multivariateprobability density functions”. In: International Conference on Independent Component Analysis and Signal Separation. Springer.2006, pp.601–608.[AKetAl2019] A. Küçük et al. “Real-time convolutional neural network-based speech source localization onsmartphone”. In: IEEE Access 7, 2019, pp.169969–169978.[AmpMe2021] AmpMe. AmpMe,Play Music louder. 2021. URL:https: / / www.ampme.com / ?locale=en_US.[APetAl2018] A. Politis, S. Tervo, and V. Pulkki. “Compass: Coding and multidirectional parameterization ofambisonic sound scenes,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP),(Calgary, Canada), pp.6802–6806, Apr.2018.[BR2015] B. Rafaely. “Chapter 7: Beamforming with noise minimization”. In: “Fundamentals of SphericalArray Processing”.1st edition, Springer-Verlag Berlin Heidelberg, 2015.[EBB2020] E. B. Brixen. “The effect of hand position on handheld microphones' frequency response anddirectivity”. In: Audio Engineering Society Convention 148, May 2020.[GW&TE1892] G. S. K. Wong and T. F. W. Embleton. “Three-port two-microphone cavity for acousticalcalibrations”. In: The Journal of the Acoustical Society of America 71, 5, May 1982, pp.1276– 1277.[JLetAl2015] J. Liu et al. “Snooping keystrokes with mm-level audio ranging on a single phone”. In:Proceedings of the 21st Annual International Conference on Mobile Computing and Networking. 2015, pp.142–154.[HZ2019] H. Zhivomirov. “On the Development of STFT-analysis and ISTFT-synthesis Routines and theirPractical Implementation”. In: TEM Journal. Volume 8, Issue 1, February 2019, pp.56-64.[LG&DE2001] L. Girod and D. Estrin. “Robust range estimation using acoustic and multimodal sensing”. In:Proceedings 2001 IEEE / RSJ International Conference on Intelligent Robots and Systems. Expanding the Societal Role of Robotics in the Next Millennium (Cat. No.01CH37180). Vol.3. IEEE.2001, pp.1312–1320.[LM&AP2022] L. McCormack and A. Politis, “Estimating and reproducing ambience in ambisonic recordings,”in Proc.30th Eur. Signal Process. Conf. (EUSIPCO), (Belgrade, Serbia), pp.314–318, Aug.2022.[MH&GF2011] M. H. Hennecke and G. A. Fink. “Towards acoustic self-localization of ad hoc smartphonearrays”. In: 2011 Joint Workshop on Hands-free Speech Communication and Microphone Arrays. IEEE.2011, pp.127–132.[MJ&DF1986] M. Junger and D. Feit. “Chapter 11: Elastic scatterers and waveguides”. In: “Sound, Structures,and Their Interaction”, 2ndedition. Cambridge, MA: MIT Press, 1986.[MSCoh2023] The MathWorks, Inc. “mscohere – Magnitude squared coherence”. In: “Spectral AnalysisDocumentation”, mathworks.com. Accessed: December, 2023. [Online]. Available: https: / / uk.mathworks.com / help / signal / ref / mscohere.html.[NO2011] N. Ono. “Stable and fast update rules for independent vector analysis based on auxiliary functiontechnique”. In: 2011 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE.2011, pp.189–192.[NSetAl2018] N. Shankar et al. “Influence of MVDR beamformer on a speech enhancement based smartphoneapplication for hearing aids”. In: 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE.2018, pp.417–420.[NSetAl2019] N. Shankar, G. S. Bhat, and I. Panahi. “Comparison and real-time implementation of fixed andadaptive beamformers for speech enhancement on smartphones for hearing study”. In: Proceedings of Meetings on Acoustics 178ASA. Vol.39.1. Acoustical Society of America.2019, p.055009.[PAN2004] P.A. Nelson. “Chapter 3: Source identification and location”. In: “Advanced Applications inAcoustics, Noise and Vibration”, 1st edition. Edited by K. F. Fahy and J. Walker. CRC Press, 2004.[PLetAl2015] P. Lazik et al. “Ultrasonic time synchronization and ranging on smartphones”. In: 21st IEEE Real-Time and Embedded Technology and Applications Symposium. IEEE.2015, pp.108–118.[PS1998] P. Smaragdis. “Blind separation of convolved mixtures in the frequency domain”. In:Neurocomputing 22.1-3, 1998, pp.21–34.[RS1986] R. Schmidt. “Multiple emitter location and signal parameter estimation”. In: IEEE transactionson antennas and propagation 34, 3, 1986, pp.276–280.[SJE2001] S.J. Elliott. “Chapter 4: Multichannel Control of Tonal Disturbances”, In: “Signal Processing andits Applications”. Academic Press, 2001, pp.186-191.[SLetAl2017] S. Li et al. “Auto++ detecting cars using embedded microphones in realtime”. In: Proceedings ofthe ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1.3, 2017, pp.1–20.[SSetAl2014] S. Sur, T. Wei, and X. Zhang. “Autodirective audio capturing through a synchronized smartphonearray”. In: Proceedings of the 12th annual international conference on Mobile systems, applications, and services.2014, pp.28–41.[SSSetAl2019] S. S. Sandha et al. “Exploiting smartphone peripherals for precise time synchronization”. In: 2019IEEE Global Conference on Signal and Information Processing (GlobalSIP). IEEE.2019, pp.1– 6.[STetAl2007] S. Takada et al. “Sound source separation using null-beamforming and spectral subtraction formobile devices”. In: 2007 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. IEEE.2007, pp.30–33.[SVD2023] The MathWorks, Inc. “svd – Singular value decomposition”. In: “Linear AlgebraDocumentation”, mathworks.com. Accessed: December, 2023. [Online]. Available: https: / / uk.mathworks.com / help / matlab / ref / double.svd.html.[VP2007] V. Pulkki, “Spatial sound reproduction with directional audio coding,” J. Audio Eng. Soc.(JAES), vol.55, pp.503–516, June 2007.[VSPetAl2023] V.S. Paul, H. Nara and & J. Hollebon. “Extraction of ambience sound from microphone arrayrecordings for spatialisation”. In: Immersive and 3D Audio: from Architecture to Automotive (I3DA). IEEE.2023.[ZHetAl2010] Z. He, A. Cichocki, S. Xie and K. Choi. “Detecting the Number of Clusters in n-Way ProbabilisticClustering”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.32, no.11, November 2010, pp.2006-2021.

Claims

CLAIMS1. An audio processing apparatus (120; 110) comprising processing circuitry (121; 111) configured to implement thefollowing processing stages for processing a plurality of spatial audio signals captured by a microphone array (116a-d):a) pre-processing a mixture of the plurality of spatial audio signals captured by the microphone array (116a-d);b) estimating the number of primary sources in the mixture of the plurality of spatial audio signals captured by themicrophone array (116a-d);c) performing an informed localization of one or more primary sound sources in the mixture of the plurality of spatialaudio signals captured by the microphone array (116a-d);d) performing a blind separation of one or more primary sound source signals in the mixture of the plurality of spatialaudio signals captured by the microphone array (116a-d);e) performing a disambiguation of blind separation based on the localization results of processing stage c);f) packaging one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing an identification of ambient sound for separating the ambient sound from the mixture of the plurality ofspatial audio signals; andh) adding the separated ambient sound to the spatial multi-channel audio format.

2. The audio processing apparatus (120) of claim 1, wherein the audio processing apparatus (120) is an applicationserver (120) configured to receive the plurality of spatial audio signals captured by the microphone array (116a-d) via a datanetwork.

3. The audio processing apparatus (120; 110) of claim 1 or claim 2, wherein the plurality of spatial audio signals arecaptured by a microphone array (116a-d) comprising four microphones (116a-d), wherein each microphone (116a-d) capturesa spatial audio signal and wherein the four microphones (116a-d) of the microphone array (116a-d) are arranged (a) at thecorners of a right-angled tetrahedron located on a front wall and a rear wall of a housing of an electronic device (110), whereinthe housing has the shape of a rectangular cuboid or (b) along a straight line extending along a first side wall or a second sidewall of the housing.

4. An electronic device (110), wherein the electronic device (110) comprises:a housing (112) having the shape of a rectangular cuboid with a front wall (112a), a rear wall (112b) opposite to the front wall(112a), a first side wall (112c), a second side wall (112d) opposite to the first side wall (112c), a top wall (112e) and a bottomwall (112f) opposite to the top wall (112e); anda microphone array (116a-d) comprising four microphones (116a-d), wherein each microphone (116a-d) is configured tocapture a spatial audio signal and wherein the four microphones (116a-d) of the microphone array (116a-d) are arranged (a) at the corners of a right-angled tetrahedron located on the front wall (112a) and the rear wall (112b) of the housing or (b) along astraight line extending along the first side wall (112c) or the second side wall (112d) of the housing (112).

5. The electronic device (110) of claim 4, wherein the housing (112) comprises a display screen (114) providing at leasta portion of the front wall (112a) of the housing (112) and a casing defining the rear wall (112b), the first side wall (112c), thesecond side wall (112d), the top wall (112e) and / or the bottom wall (112f) of the housing (112).

6. The electronic device (110) of claim 4 or 5, wherein a first, a second and a third microphone of the right-angledtetrahedron microphone array (116a-d) are arranged on the rear wall (112b) of the housing (112) and a fourth microphone ofthe right-angled tetrahedron microphone array (116a-d) is arranged on the front wall (112a) of the housing (112).

7. The electronic device (110) of claim 6, wherein the right-angled tetrahedron microphone array (116a-d) is located inthe vicinity of one of the corners of the housing (112).

8. The electronic device (110) of claim 6 or 7, wherein the edges of the right-angled tetrahedron have a length in therange from 10 mm to 30 mm.

9. The electronic device (110) of claim 4 or 5, wherein the microphones of the microphone array (116a-d) are arrangedalong a straight line extending along the first side wall (112c) or the second side wall (112d) of the housing (112) with uniformor non-uniform distances between adjacent microphones of the microphone array (116a-d).

10. The electronic device (110) of any one of claims 4 to 9, wherein the microphones of the microphone array (116a-d)are configured to be gain calibrated and phase-matched before capturing the spatial audio signal.

11. The electronic device (110) of any one of claims 4 to 10, wherein the electronic device (110) comprises an audioprocessing apparatus (110) according to claim 1 for processing the plurality of spatial audio signals captured by the microphonearray (116a-d).

12. The electronic device (110) of any one of claims 4 to 10, wherein the electronic device (110) comprises acommunication interface (113) configured to provide the plurality of spatial audio signals captured by the microphone array(116a-d) to an audio processing apparatus (120) for processing the plurality of spatial audio signals captured by the microphonearray (116a-d).

13. The electronic device (110) of any one of claims 4 to 12, wherein the electronic device (110) is a smartphone (110)or tablet computer (110).

14. An audio processing method (1200) comprising the following processing stages for processing a plurality of spatialaudio signals captured by a microphone array (116a-d):a) pre-processing (1201) a mixture of the plurality of spatial audio signals captured by the microphone array (116a-d);b) estimating (1203) the number of primary sources in the mixture of the plurality of spatial audio signals captured bythe microphone array (116a-d);c) performing (1205) an informed localization of one or more primary sound sources in the mixture of the plurality ofspatial audio signals captured by the microphone array (116a-d);d) performing (1207) a blind separation of one or more primary sound source signals in the mixture of the plurality ofspatial audio signals captured by the microphone array (116a-d);e) performing (1209) a disambiguation of blind separation based on the localization results of processing stage c);f) packaging (1211) one or more localized and separated primary sources into a spatial multi-channel audio format;g) performing (1213) an identification of ambient sound for separating the ambient sound from the mixture; andh) adding (1215) the separated ambient sound to the spatial multi-channel audio format.

15. A computer program product comprising a computer-readable storage medium for storing program code which causesa computer or a processor to perform the method (1200) of claim 14 when the program code is executed by the computer orthe processor.

Citation Information

Patent Citations

  • Ambisonic signal generation for microphone arrays

    US20190069083A1

  • Separating and rendering voice and ambience signals

    US20220059123A1

  • Multi-source audio processing systems and methods

    US20230115674A1

  • Spatial sound characterization apparatuses, methods and systems

    US9955277B1