Method for processing audio and audio processing system

CA3320419A1Pending Publication Date: 2025-09-11NOMONO AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CA3320419
Authority / Receiving Office
CA · CA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2025-03-03
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing audio processing methods for podcasts and other media types are cumbersome and time-consuming, requiring significant manual effort to separate and process sound components, and are not optimized for playback on various devices, especially in environments with background noise and multiple sound sources.

Method used

A method for processing audio signals that automatically separates voice and non-voice portions from multiple sound sources, synchronizes them, and stores them in an audio-based object format, enabling flexible playback and immersive experiences across different devices.

Benefits of technology

The method reduces manual effort by automatically separating and processing audio components, allowing for optimized playback on various devices and enhancing the immersive experience by preserving spatial information.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention concerns a method for processing an audio signal containing a plurality of in particularly timely synchronized audio signals, wherein the plurality of audio signals comprises a first voice sound portion from a first sound source and at least one second voice sound portion from a second sound source spatially distanced from the first sound source; and at least a non-voice portion and each audio signal of the plurality of audio signals recorded by a microphone of a plurality of microphones, wherein the microphones of the plurality of microphones are spatially separated from each other. The plurality of recorded audio signals are processed to obtain voice tracks and the position of the sound sources and generate one or more audio objects therefrom, these audio objects configured to be played back on a variety of sound systems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR PROCESSING AUDIO AND AUDIO PROCESSING SYSTEM

[0002] The present application claim priority from DK patent application PA 2024 70074 dated March 06 , 2024 the disclosure of which is incorporated herein in its entirety by reference . The present invention concerns a method for processing a plurality of audio signals and an audio processing system.

[0003] Podcast and some other media types require high level audio content to provide the necessary quality the listener and / or viewer of a certain media content expects and requires . In some environment audio 15 recording does not meet those expectations . While in studio recordings , any background noise is usually suppressed efficiently, other environments may contain different background noise , including quiet talks from other persons , noise from machinery nearby, natural sounds and so forth .

[0004] Furthermore , podcasts and other media types are played back on different devices , ranging from simple headphones to home entertainment systems using either several channels , like 5 . 1 or 7 . 2 and surround sound technologies like Dolby Atmos® . There is also the general trend to provide podcasts and other sound driven media with an increased immersive experience .

[0005] As a result , recorded audio is usually processed by the content producer before being published . Such processing usually includes denoising and gaining or loudness correction to improve the listening experience . In modern record and playback systems , which are used to generate spatial audio content , sound including speech is recorded using a plurality of microphones arranged in various locations . The microphones do not only record the desired sound ( e . g . the speech of a person) , but also crosstalk, whereas the speech of a person is recorded with different microphones , as well as noise and all kinds of ambient sound .

[0006] While recording spatial audio provides various benefits increasing a listener' s experience , the post-processing also becomes more complicated . This aspect becomes even more relevant , if artificial sound effects shall be provided . Those may range from interaction with an ( artificial audience ) to environmental sounds , like cars driving by, leaves rustling, water splashing and the like .

[0007] For these reasons , various post-processing tools and software have been proposed to improve the quality of the recorded audio . Still , the user of such tools needs to provide a lot of manual input from such recordings , including but not limited to separating the various sound components , e . g . voice , noise , effects from each other, process them and mixing them back together . The created media are then often stored in a proprietary format . The overall process is cumbersome and requires a significant time effort . It is assumed that one hour of recording requires at least the double time in post-processing with the various formats requiring a lot of storage on a computer readable medium .

[0008] It is therefore an obj ect of the present invention to provide a method for processing audio , that offers not only a flexible way of handling recorded audio content , but also a general storage to be able to play back such content on a variety of different devices .

[0009] SUMMARY OF THE INVENTION

[0010] This and other obj ects are addressed by the subj ect matter of the independent claims . Features and further aspects of the proposed principles are outlined in the dependent claims .

[0011] The inventors propose a combination of processing a plurality of recorded sound signals representing a sound scene with at least two spatially separated sound sources with storing them in an audio-based obj ect format . The plurality of recorded sound signals are synchronized in time or at least its respective time correlation is known . The audio-based obj ect format enables playback on various devices , thus preserving not only the spatial information, but enabling amending spatial and other information generating an immersive audio experience .

[0012] The plurality of recorded sound signals is automatically processed to separate the portions of one or more voices and the non-voice portions and subsequently process them as separate stems . The non voice portions may contain ambient , noise and non-voice event as described above . The benefit of processing them as stems lies in the possibility to change several characteristics of the different potions separately without affecting the respective other ones .

[0013] In addition, the process automatically identifies and extracts various features of the respective portions , including but not limited to language detection, speaker detection, relevance or importance of the respective portions in regard to the overall plurality of recorded signals and so forth . Those may be used later when compiling audio obj ects into a sound scene . In addition, a user may change and vary the extracted features to optimize the immersive experience if needed .

[0014] In some aspects , the inventor proposes a method for processing a plurality of audio signals . The method comprises providing a plurality of timely synchronized audio signals , wherein the plurality of audio signals comprises a first voice sound portion from a first sound source and at least one second voice sound portion from a second sound source spatially distanced from the first sound source . The expression timely synchronized or generally synchronized means that their sampling rate as well as its respective recording time , i . e . its start and stop in relation to each other are known . This can be achieved either by a master-slave configuration of microphones recording the sound signals , a common time clock or any other precise time trigger . Thus , the recorded signals are not shifted in time with respect to each other , or at least any time shift is known . The signals content however can be shifted because speech or any other sound originating from a sound source may travel different distances towards the respective microphone and thus will be recorded at different times .

[0015] The plurality of sound signals also comprises at least a non-voice portion . Each audio signal of the plurality of audio signals is recorded by a microphone of a plurality of microphones , wherein the microphones of the plurality of microphones are spatially separated from each other .

[0016] While in some instances , the sound sources and at least some of the microphones are associated with each other, this is not necessary . Rather , the position of the sound sources and the position of the microphones can be different . Furthermore , the proposed principle is able to detect movement of the sound sources through the sound space , whereas it is assumed that the position of the plurality of microphones are fixed .

[0017] The plurality of recorded sound signals are then processed by for instance obtaining a first voice track from the plurality of sound signals , said first voice track comprising at least a part of the first voice sound portion and optionally a part of non-voice portion . Furthermore , at least one second voice track from the plurality of sound signals is obtained, said at least one second voice track comprising at least a part of the at least one second voice sound portion and optionally a part of the non-voice portion .

[0018] A position of the first sound source and the at least one second sound source in relation to each other and / or in relation to a predetermined reference point is determined from the recorded plurality of sound signals . This determination can be repeated periodically or continuously such that movement and / or position vectors are obtained for each sound source in relation to the predetermined reference point .

[0019] As a result , each voice portion and / or voice track is associated with the respective sound source , with the position of the sound source being determined throughout the entire sound scene . The above steps are performed based on the recorded sound signals as such, without the need for a user interaction . In this regard, the expression user interaction refers to any action taken by the user to associate ( not necessary confirm) a sound voice to a sound source , a non-voice portion to a sound source and / or voice portion, to set the position of a sound source ( not change a determined position ) , to manually separate voice portions from the recorded audio signals etc . In other words , the inventor proposes a method, that performs the necessary generation of the various voice tracks , and the positional information with the voices automatically without a manual and time-consuming human workflow . Then, the method comprises the step of defining a rendering type , said rendering type associated with one or more rendering channels . The rendering type can be set automatically based on the number of voices , sound sources their respective positional information in the sound scene , but also determined by the user . The rendering type usually comprises information about the desired output format (usually a loudspeaker layout ) . Typical rendering types include references to the channel-based or obj ect based formats , as explained further below in greater detail . In this regard, the one or more channels to which the rendering type is associated does not necessarily refer to channels in the channel-based format . Rather , the rendering type defines the channels which may be used in a subsequent render to produce the output sound signals . If the rendering type is a channel based format , the channels defined therein may correspond to the channels in the channel based format . If the rendering type is an obj ect based format , the channels associated with the rendering type may contain information about the audio obj ect .

[0020] In a subsequent step , sound rendering information for the one or more channels associated with the determined position of the first sound source in relation to the reference point are generated . This process may be repeated for the one or more channels associated with the determined position of the second sound source in relation to the reference point , as well as to any other non-voice portion .

[0021] In a next step, at least one audio obj ect is created . The audio obj ect contains a unique ID identifying one or more stored tracks corresponding to at least the first voice track, an identifier to the rendering type , said rendering type referring to the one or more rendering channels and the generated sound rendering information for the one or more channels associated with the determined position of the first sound source .

[0022] This step is succeeded by generating at least one audio content obj ect , said at least one audio content obj ect containing a reference to the at least one audio obj ect and at least one data packet containing characteristics of one of the plurality of audio signals , the first voice track and the at least one second voice track .

[0023] Examples of the characteristics of one of the plurality of audio signals , the first voice track and the at least one second voice track may include , but are not limited to language of a voice , importance of a voice or a soundtrack, either as absolute importance or in relation to the other voices .

[0024] Finally, one or more files are stored on a computer readable storage medium, said one or more files comprising the generated at least one audio content obj ect , the generated at least one at least one audio obj ect and the one or more stored tracks identified by the unique ID .

[0025] The proposed method may also generate another audio obj ect , said audio obj ect containing a unique ID identifying one or more stored tracks corresponding to the at least one second voice track . The audio obj ect may further include the identifier of the rendering type , said rendering type referring to the one or more rendering channels and the generated sound rendering information for the one or more channels associated with the determined position of the second sound source .

[0026] In other words , the proposed method separates the individual voices originating from respective sound sources from each other . This is done that voice portions belonging to the same voice are associated with a sound source . Voices may contain interruptions and pauses . However , the proposed method separates the voices from each other , thus resulting in voice portions belonging to the same voice and the same sound source .

[0027] In some instances , the step of obtaining a first voice track from the plurality of sound signals may comprise the step of include a speaker detection, as to identify the pauses and interruptions in each voice . In this regard, pauses and interruption below a certain time threshold e . g . less than a second or less than 750 ms are not considered a pause . However, when generating the voice track, those areas may be attenuated to improve the listener' s experience . The step of speaker detection may provide a plurality of triggers or markers over time indicating a pause in the voice of a sound source . Those triggers or markers are used in the step of obtaining the respective voice track for data compression purposes . The information of the speaker detection may also be used in generation of audio obj ects and when generating the sound rendering information .

[0028] In some further aspects , the step of generating at least one audio obj ect further comprises the step of generating at least one more audio obj ect , said at last one more audio obj ect further containing a unique ID identifying one or more stored tracks corresponding to the at least one second voice track; and the identifier of the rendering type , said rendering type referring to the one or more rendering channels . The at least one more audio obj ect also contains the generated sound rendering information for the one or more channels associated with the determined position of the second sound source .

[0029] Hence , the present method proposes to generate -based on the plurality of sound signals and the extracted portions thereon- a plurality of audio obj ects , in which each audio obj ect contains the voice associated with a sound source .

[0030] In some further aspects , the proposed method comprises the step of obtaining a non-voice portion from the plurality of sound signals , said non-voice portion comprising at least a part of all non-voice portions in the plurality of sound signals . Based on the rendering type , sound rendering information is generated for the non-voice portion . Such non- voice portion may comprise a sound bed or ambient portion for the sound scene , but can also include time constraint effects , like clapping of hands , a car driving by and the like . Such effects are usually different from background noise or ambient sounds and consequently separated therefrom and extracted from the plurality of recorded sound signals .

[0031] The step of generating at least one audio obj ect comprises the step of generating at least one non-voice containing audio obj ect , said at least one non-voice containing audio obj ect containing a unique ID identifying one or more stored tracks corresponding to the part of the non-voice portion; and the identifier of the rendering type , said rendering type referring to the one or more rendering channels . The at least one non-voice containing audio obj ect further comprises the generated sound rendering information for the non-voice portion .

[0032] In some aspects , the step of generating at least one audio obj ect in the proposed method comprises the step of generating an ambient sound scene from the plurality of recorded sound signals , particularly in the B-format . Hence , the proposed method results in some aspects in a plurality of different audio obj ects , which can either be combined in audio content obj ects or summarized into an audio program.

[0033] Some aspects concern the processing of the various audio portions and the generation of voice tracks . As mentioned, audio portions are associated with a sound source , for example a speaker . However, the sound source may not generate a continuous speech or voice , for example a person being the sound source may pause for a certain period of time . Hence , there may be different voice portions belonging to the same sound source or there is one voice portion associated to a sound source that contains longer period of silence .

[0034] It is therefore advisable to detect such interruptions . This can be achieved either for the plurality of all recorded signals , e . g . by detecting speech or the absence of speech in the various voice portions for each sound source . The approach simply detects the speech irrespectively of the voice portions .

[0035] In another aspects , the method proposes to detect speech or the absence of speech dependent on the sound source , i . e . only in those voice portions that are associated with the respective sound source for this purpose , the voice portions are separated in the plurality of recorded sound signals and associated with a respective sound source . A voice track may be generated from the voice portions , and the speech detection may be performed on the voice track . Detecting the speech or the absence of speech in the voice track or the voice portions used to generate said voice track has benefits because speech and pauses in said voice track are independent on any other voice track . Consequently, a pause one sound source can be correctly detected, even if another sound source is continuing to produce speech or another non-voice portion .

[0036] The information about the presence and absence of speech is then used in the audio obj ect . More particularly, the audio obj ect contains information derived from the detected speech or the absence of speech associated with one of stored tracks corresponding to said voice track and the unique ID . The information may include a time stamp as well as a reference to the respective portion of the voice tack, thereby indicating which portion of the voice track shall be played at which time . This reduces the overall storage necessary for the respective track, as pauses and silence in the actual voice track can be omitted . Instead, the audio obj ect provides the correct time information which portion shall be played back at which time .

[0037] In some aspects , the steps of generating one or more files comprises the step of associating at least some of the one or more stored tracks with the identifier of the rendering type . This may be useful for certain channel based rendering types , in which each track corresponds to a channel .

[0038] Some aspects as mentioned above concern the rendering type . As stated above , the rendering type may contain a definition as to the output format , i . e . whether the output format shall be channel based or obj ect based . In channel-based audio , a soundtrack is created by recording a separate audio track ( channel ) for each speaker . Common speaker arrangements for channel-based surround sound systems include the so called 5 . 1 and 7 . 2 systems , which utilize five and seven surround channels respectively, and one or more low-frequency channel .

[0039] By selecting such rendering type , each soundtrack must be created for a specific speaker configuration, hence the development of industrystandard configurations such as 2 . 1 ( stereo ) , 5 . 1 and 7 . 1 . In addition, there are certain headphones configurations like stereo and ambisonics formats offering an immersive sound experience . A maj or audio format for higher order ambisonics is B-format . This is a multichannel audio format where the individual channels do not correspond directly to speaker feeds . There is no "front left" channel as in classical systems . Instead, the channels contain components of the sound field that are combined during a later decoding step .

[0040] Obj ect-based audio addresses this by representing a sound scene as multiple separate audio obj ects , each of which comprises one or more audio signals and associated metadata . Each audio obj ect is associated with metadata that defines a location and traj ectory of that obj ect in the scene . Obj ect-based audio rendering involves the rendering of audio obj ects into loudspeaker signals to reproduce the authored sound scene as well as specifying the location and movement of an obj ect , the metadata can also define the type of obj ect and the class of renderer that should be used to render the obj ect . For example , an obj ect may be identified as being a diffuse obj ect or a point source obj ect .

[0041] Example for an obj ect-based rendering type is Dolby Atmos®, in which the various sound obj ects are embedded into an ambient sound bed and the renderer takes care of the positional and other meta information during rendering .

[0042] In some aspects , the rendering type includes a definition to one or more of the above-mentioned formats .

[0043] Some aspects concern the step of obtaining the first voice track and / or the second voice track from the plurality of sound signals . It may comprise separating from each of the at least two audio signals a clean voice stem, a crosstalk stem and a noise stem . These stems are defined to be the types of sound potentially contained in each of the recorded audio signals . For the purpose of this application, the following stems are defined as clean voice stem, crosstalk stem and noise stem .

[0044] The clean voice stem comprises substantially the speech component , particularly from the sound source to which the microphone is associated with but is otherwise deprived of other sound portions . In a high-quality studio environment with a single microphone and no sound reflections or other ( noise ) sources , the recorded audio signal would most likely represent the clean voice stem . However, in the presently proposed method, the plurality of timely synchronized audio signals can each comprise voice , crosstalk and noise .

[0045] The crosstalk stem comprises a crosstalk portion recorded by the microphone associated with the respective sound source . If the microphone is not associated with the respective sound source , then there may be not crosstalk for such sound source or in the respected microphone , but rather j ust recorded voice and possible noise .

[0046] In some instances the term crosstalk refers to an audio component that is recorded by a microphone , which is not associated with the sound source causing the audio component . As an alternative , the term crosstalk refers to an audio component that is recorded by a microphone , which is not associated with the sound source causing the audio component whereas the sound source is associated with another microphone . Hence , in the latter case a voice component coming from a sound source with no microphone associated with may then be considered a voice component and not crosstalk . In addition, crosstalk refers to an audio component reflected at a wall or an obstacle but coming from the audio source associated with the microphone . Simply speaking crosstalk often refers to recorded background speech or recorded voices from one or more persons that are not current or desired speakers as well as any reflected portion of the speaker' s voice .

[0047] It may be that a crosstalk portion is comprehensible to a listener but j ust degraded, whereas a noise portion usually does not provide useful information to a listener . While the former audio component is usually different from the direct recorded sound of the sound source , a reflection may have a similar frequency and time distribution, but is delayed, attenuated and phase shifted . Consequently, the latter can also be handled as part of the noise stem. The noise stem comprises any noise portion . It is defined as a portion of the recorded signal that does not provide useful or derivable information to a listener . It is not related to the sound source but originates from a different source ( one can consider reflection as originating from the wall or an obstacle ) . Typical sound sources are machinery, circuits , the clapping of hands in an audience sounds of nature like a stream or river flowing . One s killed in the art can identify many more sound sources that produce or contribute to the noise stem in a recorded audio signal . In some instances , one may distinguish between different types of noise sources , like event-noise and background noise . The event noise is usually shorter and may be louder than background noise . Clapping of hand of an audience , a car driving by a horn and the like can be considered event noise .

[0048] After separation of the recorded audio signals into the respective stems , each stem of each audio signal can be processed individually . Alternatively, two or more stems from the same audio signal can be processed using the same processing function and parameter set . Here all variations and combinations are possible providing a significant improvement over conventional techniques , in which a certain processing function is applied to the audio signal as a whole , but not to the separate and individual stems . In particular , the separated stems of each audio signal can be stored in a non-volatile memory for later use .

[0049] The method according to the proposed principle proposes to apply a first processing function to the clean voice stem of at least one of the at least two audio signals to obtain a first processed clean data stem . Particularly, independently thereof , the first processing function is also applied to at least one of the crosstalk stem and the noise stem of said at least one audio signal to obtain a first processed crosstalk stem and / or first processed noise stem . In this regard, applying a function to a certain stem or more generally an audio signal or portion thereof means that the respective stem or portion is used as input for said function, such the content of said respective stem is changed or amended accordingly . The first processed clean voice stem may in some aspects combined with at least a portion of the first processed crosstalk stem and / or first processed noise stem to obtain a processed combined stem . The processed combined stems of each of the at least two audio signals are subsequently mixed to provide an output signal .

[0050] Separation of the recorded audio signals into the different stems provide several advantages over conventional techniques . For one , the processing function subsequently applied to the various stems can be optimized for the respective stem . Furthermore , some processing function will provide better results if applied to only a single stem and not a combination . For example , de-esser and breath detection should preferably only be applied to the clean voice stem to avoid identifying noise portions as part of the voice or speech . Likewise , an automated gain control should avoid increasing crosstalk, while the speaker is silent . The proposed principle allows developing functionality specifically for one of the separated stems that add to the content provider' s flexibility when processing audio content . In some more instances , the step of applying a first processing function to the clean voice stem further comprises applying a second processing function to at least one of the clean voice stems to obtain at least one processed clean voice stem, said at least one processed clean voice stem being the input for the first processing function .

[0051] Consequently, one can either apply various processing functions to the individual stems , or the same processing function to two or more of the stems prior to combining them back together . It is also possible that an output of a processing function acts as an input for the next processing function . Stems may therefore be processed using a plurality of processing function, wherein some of those processing function are applied to different stems , while other processing functions are used solely on a single stem. Furthermore , depending on the functionality, one may also consider applying certain functions only to stems of dedicated audio signals , and not to each of the recorded audio signals .

[0052] The following sections refer to the functionality of certain processing functions . The implementation of those can be different , including for example open-source libraries , but also proprietary solutions . Furthermore , some of the processing functions are static , i . e . the use of a fixed algorithm or approach, while others include a trained deep learning network optimized to provide the respective functionality . In some instances , the second processing function comprises an adaptive equalizing function, configured to f requency-selectively gain or attenuate portions of the clean voice stem of at least one of the at least two audio signals . The gain or attenuation can be done on each of the data stems with different parameter sets . This will allow adj usting the volume of different sound sources to a common level without increasing crosstalk and noise .

[0053] In addition, it may be useful in some instances , to mute portions of the clean voice stem of at least one of the at least two audio signals , particular in those portions , in which no speech by the sound source associated with the at least one audio signal is present . This function is of particular usefulness in cases , in which the different sound sources are persons talking or having a discussion . Muting or at least significantly reducing the volume of those signals not associated with the current speaking person may improve the overall quality, as undesired signal portions from other sound sources are suppressed . This function may be applied to all clean voice stems , as it also required to identify speech and pauses in the respective data stems , to be able to mute them selectively .

[0054] Consequently, in some aspects , the second function applied to a data stem of one of the at least two audio signals also utilize the other data stems to derive various parameters therefrom including , but not limited to markers associated with voice or speech . Some other function concern improving the actual speech or voice included in the data stem . In some aspects , the second function is configured to reduce or eliminate the excessive prominence of sibilant consonants within the speech component of at least one of the clean voice stems of at least one of the at least two audio signals . This is also referred to as de- essing . In some instances , 25 the second processing function comprises a breath detection functionality, configured to evaluate each of the clean voice stems of at least one of the at least two audio signals , particularly evaluate those separately from each other , to mark portion of the respective stems , in which a breath by the associated sound source is detected .

[0055] In some instances , the breath detection, can be subsequently used to adj ust the gain of the respective stems , thereby preventing that breathing sounds are increased . Alternatively, the detected breath of a sound source is adj usted automatically, e . g . by attenuating the respective portion . Breath detection is also used later in the generation of the audio obj ect .

[0056] Some further aspects concern the first processing function, which is a processing function which is applied to one of the data stems or processed data stems (with one or more of the above mentioned second processing functions ) and one of the noise stem and the crosstalk stem . The function can be applied with the same parameters to various stems or with different parameter sets . However , the first processing function is applied to those stems separately, which means that the data stems and the noise stem or crosstalk stem are used separately as input to the first processing function . This will provide flexibility, because the results of the processed data , noise and crosstalk stems can be optimized separately by applying the respective function individually .

[0057] The first processing function comprises for example a high-pass filtering functionality configured to filter frequency portions below a threshold frequency . The threshold frequency can be set to different values depending on the type of stem ( e . g . data, noise and / or crosstalk) . The threshold frequency can be in particular below 100Hz and more particularly below 70Hz . The high-pass filtering function is applied in particular to one of the clean voice stems and one of the noise stems .

[0058] In some further instances , the first processing function may comprise an adaptive levelling functionality, configured to frequency- selectively gain or attenuate portions of at least one of the clean data stem of at least one of the at least two audio signals and in particular the noise stem associated with the at least one of the at least two audio signals . Such function may also use markers associated with the voice or speech in the data stem . For instance , by evaluating those markers , one may selectively attenuate or gain the crosstalk and noise stems during time periods , in which no voice or speech is present in the data stem . Such time and stem selective gain may contribute to a more natural listening experience .

[0059] In some instances , the step of separating the stems from each f the plurality of recorded audio signals or a subset thereof comprises the step of separating a crosstalk portion as crosstalk stem from the respective recorded audio signal using those audio signals . Additionally, a noise portion is separated for each of the at least two audio signals , in which in particular the crosstalk portion has already been separated, wherein optionally the two separation steps are subsequently executed . Consequently, the separation is stems is done in two different steps , whereas a crosstalk portion is separated in a first step and a noise portion is separated from the audio signal in a subsequent step . If all three stems are summed up together , one derives again at the recorded audio signals .

[0060] Another aspect concerns the determination of the position of the respective sound sources in regard to each other within the sound scene from the plurality of record audio signals . In this regard, the expression "position" does include the distance from the sound source to the dedicated reference point , an angle based on one or two axes through the reference point or a combination thereof . It is assumed that the plurality of recorded sound signals is synchronized in time . The method itself can act on the original recorded sound signals , but also on the already separated voice stems .

[0061] In some instances , some sound signals are recorded by microphones close to the respective sound source , meaning that the distance is between the microphone to the sound source relatively low compared to the distance between the sound source and the dedicated reference point . However, the term "at the sound source" is not to be understood in a very limited sense . Rather, the expression shall include and allow for a certain distance between the actual sound source and a microphone . Similarly, other sound signals of the plurality of sound signals ( e . g . ambient sound signals ) are recorded at different locations , for which the distance and angle to the reference point is known . Time synchronization is important for the proposed method in subsequent steps . Such time synchronization can be achieved in some instances by providing a common time base for any sound signal recorded . In some other instances , the recorded sound signals can be used to provide the time base , e . g . by timely correlating a dedicated start signal that is recorded and included in the first and the plurality of second sound signals .

[0062] A filter is now estimated for the first sound signal acting on the signal-to-noise ratio in each frequency bin of the signals recorded by microphones close to the sound sources ( referred to as first sound signals ) in a time-frequency domain . Then, the first sound signal is correlated with at least one of the other ones of the plurality of recorded sound signals ( e . g . the recorded ambient sound signals referred to as second sound signals ) in the frequency domain to obtain at least one correlated signal . In some instances , the first sound signal is correlated with each of the plurality of second sound signals to obtain a plurality of correlated signals .

[0063] The previously estimated filter is applied to the at least one correlated signal to obtain at least one filtered and correlated signal .

[0064] In the next step , one can estimate the distance between the dedicated reference point and the sound source . For this purpose , a first timing value in the at least one filtered and correlated signal exceeding a dedicated threshold in the time domain is obtained . A second timing value corresponding to a threshold value in the at least one filtered and correlated signals based on the first timing value is also obtained . The distance between the dedicated reference point and the sound source is now derived based on the respective obtained first timing value and second timing value . If more than a single filtered and correlated signals are derived in the previous step, a plurality of distances is obtainable enabling to improve the distance determination, ( e . g . by calculating a mean value including error margins and the like ) .

[0065] Alternatively or also additionally, an angle of a sound source relative to an axis through the dedicated reference point can be calculated . For this purpose at least two filtered and correlated signals and an optional a priori estimate or knowledge of microphone locations providing the plurality of second sound signals are utilized . In a subsequent step after applying the above-mentioned filter and correlating the filtered first signal with at least two of the plurality of second sound signals in the frequency domain to obtain at least two correlated sound signals , the at least two filtered and correlated sound signals are truncated around a specific time period . Then, a cross correlation between pairs of truncated filtered and correlated sound signals are obtained . The angle of arrival of the filtered first sound signal is derived by proj ecting the obtained cross correlation in a spherical spatial space based on the a priori estimate or knowledge of microphone locations providing the plurality of second sound signals .

[0066] With the proposed method, it is possible to obtain distance and angle independently of each other . Correlating the two different signals can be done off-line on recorded and stored sound signals as well as in real-time if necessary, enabling the method to be used in a variety of applications . Further, it is possible to exchange first and second sound signals as well as deriving slowly moving sound sources . By correlating different first sound signals , one may further improve the accuracy of the proposed method . The method is robust against reflections of sound, which is useful during recording sessions in closed space .

[0067] SHORT DESCRIPTION OF THE DRAWINGS

[0068] Further aspects and embodiments in accordance with the proposed principle will become apparent in relation to the various embodiments and examples described in detail in connection with the accompanying drawings in which Figure 1A shows an illustration of a sound scene with several sound sources providing voice and non-voice portions in accordance with some aspects for the proposed principle ;

[0069] Figure IB illustrates various recorded voice and non-voice portions over time to illustrate some aspects of the proposed principle ;

[0070] Figures 2A and 2B illustrate to embodiments of a method for processing a plurality of a timely synchronized audio signal in accordance with some aspects of the proposed principle ;

[0071] Figure 3 shows an exemplary method of a process for separating voice and noise of the different sound sources and generates voice tracks therefrom in accordance with some aspects of the proposed principle ;

[0072] Figures 4A and 4B illustrates an exemplary embodiment of a process for obtaining the position of sound sources using the plurality of recorded sound signals in accordance with some aspects of the proposed principle ;

[0073] Figure 5 shows a schematic overview over the structure of an audio program including one or more audio obj ects in accordance with the proposed principle ;

[0074] Figure 6 illustrates an example of an audio program containing one or more audio obj ects in the ADM format that can be processed using the proposed method and system .

[0075] DETAILED DESCRIPTION

[0076] The following embodiments and examples disclose various aspects and their combinations according to the proposed principle . The embodiments and examples are not always to scale . Likewise , different elements can be displayed enlarged or reduced in size to emphasize individual aspects . It goes without saying that the individual aspects of the embodiments and examples shown in the figures can be combined with each other without further ado , without this contradicting the principle according to the invention . Some aspects show a regular structure or form . It should be noted that in practice slight differences and deviations from the ideal form may occur without , however , contradicting the inventive idea .

[0077] In addition, the individual figures and aspects are not necessarily shown in the correct size , nor do the proportions between individual elements have to be essentially correct . Some aspects are highlighted by showing them enlarged . However , terms such as "above" , "over" , "below" , "under" "larger" , "smaller" and the like are correctly represented with regard to the elements in the figures . So it is possible to deduce such relations between the elements based on the figures .

[0078] Figure 1A illustrates a typical environment , in which an audio session is recorded creating a sound scene . The environment depicted in Figure la includes three speakers also identified as sound sources SSI , SS2 and SS3 .

[0079] The respective speakers are arranged in an environment with one or more walls restricting the sound space in one direction . The speakers are arranged around a central reference point , in which a microphone array MA is positioned . Microphone array MA comprises a plurality of microphones being set spatially slightly apart and looking into different directions . The microphones are configured to record directionally dependent ambient sound . Furthermore , the second speaker or sound sources SS2 comprises a portable microphone M2 directly attached to him. One can assume that M2 is associated with him in this regard . Likewise , the third speaker SS3 comprises another portable microphone MS , which is directly attached to his body, for example close to his mouth . In contrast , thereto speaker SSI does not have an additional microphone attached to his body close by or generally associated with him. Furthermore , possible events like a car driving by creating an event El can occur over time , but are usually random and / or timely limited .

[0080] The three speakers are talking to each other , for example in a podcast an interview and the like . Their voices as well as the background noise and the events are recorded by the microphone array MA as well as the two microphones , M2 and M3 .

[0081] During the recording session, several events and situations can occur , which are explained in more detail . If speaker SS2 or SS3 are speaking , the respective microphones M2 and M3 will record their voices substantially without delay, as the microphones are located close to the respective body and mouth . In addition, the voice of speaker SS2 will reach the microphone M3 and vice versa after some delay and subsequently recorded as well . Said voice component is considered crosstalk for the purpose of this example . Likewise , the speech and voice of speaker SS3 will generate crosstalk in microphone M2 . Additional crosstalk portions may occur due to reflections of the respective voices on the wall and reaching the corresponding microphones . However, in this regard, a reflection of the voice from speaker SS3 being recorded at microphone M3 is also considered across talk in its recorded signal . Consequently, microphones M2 and M3 may record crosstalk of two different types , one being the voice and speech of the respective other sound source , and one being the own voice and speech but recorded with a certain delay due to reflections on obstacles and the like .

[0082] In contrast , thereto speaker SSI does not contain its own microphone associated with him . As a result , any of his / her voice components and speech is recorded by microphones M2 , M3 with a certain delay as a so- called direct sound DS . The delay between of those portions in the two recordings at microphones M2 and M3 allow determining the position of SSI in regard to the two other sound sources SS2 and SS3 . Similar as for the voices and speech of the sound sources SS2 and SS3 , the speech and voice of sound sources SS2 may be reflected at the wall and further delayed before being recorded by one of the microphones M2 and M3 . That reflected portion is also considered crosstalk for the purpose of this example .

[0083] The microphone array MA being located at the reference point for this sound scene also records the various voice portions and speech portions of the respective sound sources and speakers SSI , SS2 and SS3 . More particularly, such voices are recorded either as direct sounds DS with various delays and also as crosstalk when being reflected at the wall or other obstacles . In addition, it may also record any noise either from nose events like El or as background noise ( again with possible reflections thereof at the wall ) . As a result , the microphones of microphone array MA record a plurality of sound signals , which contain the voice and speech portions of the various sound sources with a slight delay to each other . Furthermore , the plurality of recorded signals also contains the recorded signals of microphones M2 , and M3 including the voice and speech portions of the sound sources SSI to SS3 as well as the respective cross talks .

[0084] Apart from the voice and speech components as well as the cross talks thereof , noise and other events can occur in the sound scene during the recording session . An example is given in figure 1A, in which event El is created by a car driving by . This additional non-voice portion is recorded by the two microphones , M2 , M3 as well as the microphone array M3 with different delays . As those events usually happen randomly and are only present during a short period of time , it is considered a noise event in contrast to other noise portions , which are inherently and continuously present . Those other noise portions may include lower or high-frequency noise of machinery being present close to the environment , or natural sounds like wind whistling and the like . Consequently, the overall noise being recorded by the microphone array and the microphones can be differentiated in noise events as well as background noise . These two aspects are processed separately in the proposed method .

[0085] Figure IB illustrates a plurality of audio components being present in a sound scene over a certain period of time being recorded with the various microphones . The time axis in Figure IB is arranged in the center , with two voice portions VP1 and VP2 being arranged above the time axis and various noise portions being arranged below . Each voice portion belongs to a speaker and both speakers are spatially distanced from each other . In the present example , both voice portions Vpl and Vp2 ( associated with sound source SSI and SS2 , respectively as depicted in f Figure 1A) include various voice components , that is the voice portions is not an uninterrupted speech but due to the nature of this example , there are pauses by the two speakers , both are talking at the same time and so forth . Particularly a voice component of the first voice portions is present between the times the TO and T-l . Then, after a short pause , the speaker SS2 answers creating a voice component of the voice portion Vp2 . The speaker makes a pause between time T3 and T4 and then continues . After another pause , the first speaker answers creating another voice component of the voice portion VI at time T6 till T9 . The discussion may continue throughout the session . Furthermore , speaker SS2 starts talking at T8 , then between T8 and T9 both sound sources are producing a voice component ( i . e . talking at the same time ) , thereby creating overlapping components of voice portions Vpl and Vp2 during that timeframe .

[0086] In addition, the sound scene includes a constant background noise indicated by a constant noise stream NS1 . Furthermore , between times T5 and T6 an individual noise event is taking place , which ends slightly after the first speaker starts talking again .

[0087] The microphones including the microphone array will record the respective voice portions VP1 and VP2 , as well as the noise portions with various delay towards each other . Particularly, microphone M2 will record the voice portion VP1 without substantial delay as it is closely arranged to the sound source itself , while recording the sound portion VP2 as well as probably the noise event NE1 with a slight delay . Likewise , the microphone M3 will record the voice portion VP1 with a slight time delay due to its distance to the sound source generating VP1 , while recording the voice portion VP2 without any substantial delay . In contrast , thereto the microphone array MA, being distanced and spaced apart from the respective sound sources will record both voice portions with a slight delay . The time correlation between the voice portions VP1 and VP2 being recorded at the respective microphones provides the possibility to determine a position of the sound sources with respect to the reference point both in the distance and angles , thereby creating a virtual sound scene .

[0088] Furthermore , the recorded signals may all include components of all voice portions , VP1 , VP2 as well as the respective noise portions . By combining the individual signals properly, one can extract the pure voice portions as so-called voice stem as well as the noise portions being separated in an event stem as well as background noise stem . Moreover , the recorded noise by the microphone array MA can be presented by an ambient sound bed, which can be used later during processing to create an audio obj ect embedding the actual voice and speeches of the respective speaker in a more general sound bed .

[0089] Figure 2 illustrates an exemplary embodiment of a method for processing a plurality of recorded signals in accordance with the proposed principle .

[0090] In step SI , the plurality of signals is recorded, for example , in an environment depicted in Figures 1A and IB . The recording session will generate a plurality of recorded and timely synchronized audio signals . The expression "timely synchronized" shall be understood in this regard that the respective microphones are using the same time base in regard to each other . This can be achieved for example in a master-slave configuration, in which the microphone array MA provides a master clock, to which the other microphones M2 and M3 are synchronized to . Another possibility would be to create an audio signal by the various microphones or a central system, which are recorded by the microphones and subsequently synchronized to . Likewise , one can create an RF beacon at periodic time , which are received by the microphones and subsequently synchronized . The time synchronization will allow the determination of the position of the microphones in regard to each other Even further, if the sound sources are moving through the sound space , the time delay between the recorded signal varies to , which can be determined . Hence , it is possible to follow the sound sources through the sound space . Following the recording session, the step S2 determines the respective sound sources by separating the component of the various sound portions and noise portions as well as cross talk top portions from each other . This step called separation of the portions into respective stems is explained in greater detail below .

[0091] For this purpose , components of voice portions of belonging to the same sound source are identified . The signals are processed such as to extract one or more voice stems (preferably as many as there are speakers ) , crosstalk stems associated with the voice stems and one or more noise stems . The voice stems include substantially the pure voice of the respective sound sources , while the crosstalk stem provides the cross talks between the different microphones . The stem separation can be performed for each of the microphones M2 and Ml , as well as each of the microphones in the microphone array .

[0092] The separated voice stems from the different microphones but belonging to the same source can then be combined to create a voice track in step S3 . Alternatively, one can also create a combined voice stem from all the recorded signals at once . Furthermore , during separate of the step , one can process the respective stem, i . e . adj usting the gain, detect actual speaking portions and the like . Noise and events can generate a separate noise and event track accordingly .

[0093] After determination of the sound sources and its separation into the voice stem in step S2 , the various stems can also be used to determine the position of the sound source with regard to the reference point . While this step S4 is illustrated to be performed after the generation of the respective voice stems in step S2 , the step of determining the position of the sound sources can also be performed directly on the recorded signals . However, it might be suitable to perform this step after the separation, as for example , the timing information in the separated cross talks and voice stems can be used for this purpose .

[0094] In Step S3 , one or more voice tracks are generated . Those track contain the actual audio signal i . e . the voice of one speaker as a PCM data file and the like . In some aspects , the pauses between the component of a voice stems are removed to compress the data , however in this regard such information is preserved so that during playback, the correct audio including the pauses are played .

[0095] In step S5 , the rendering type is determined based on the obtained voice tracks as well as the determined position . A user may interact by choosing a specific render type based on his desires or the intended purpose for the recorded sound session . N some aspects , the render type is set automatically based on the recording session . The render type may be channel-based or obj ect-based . Still the proposed method will generate audio obj ects irrespective of the selected render type . The render type may also specify the necessary render information to be added to the sound obj ect . For example , if the render type is selected to be stereo , two channels are subsequently created for the audio obj ect , one containing the information and content for the right channel and one of the left channel .

[0096] However, the previously retrieved information, like position, speaker detection, loudness , etc . is used to generate the tracks for the two channels . Likewise , noise and / or crosstalk is mixed back in to provide a user with the impression that a certain voice or noise is coming from the left side or from the right side .

[0097] In more complex situations like 7 . 2 or 5 . 1 as well as in obj ect-based information, the immersive experience for a listener can be increased using more complex scenarios . However, the audio signal for the respective channels may be generated from the audio and noise / cross talk tracks and then added as part of the audio obj ects . This may also include processing the time information in regard to position of the sound source to generate moving sound sources , for example . As a result , when the sound source moves through the sound space , thereby changing the voice stem, this information can be added into the rendering type and rendering information to create the audio obj ect thereof . As a result , the user can get the experience that a certain sound flows associated with a voice track in step S3 is moving from left to right or vice versa . In obj ect-based systems , the tracks can be stored with the additional rendering information provided in step S 6 in the audio obj ect , as for such approaches , the external render will generate the output based on the information in the audio obj ect and an external speaker configuration .

[0098] The determined render type and provides render information ins Steps S5 , S 6 together with the respective position information of step S4 , the obtained voice tracks in step S3 as well as other information is then used in step S7 to generate one or more audio obj ects . Further information being extracted from the plurality of recorded sound signals , for example , language of the voices , gender , importance of a voice in regard to other voices etc . is used to generate an audio content obj ect in step S3 . Finally, the obtained voice tracks associated with the render type and render information as well as the audio obj ect are stored in one or more files in step S 9 .

[0099] The format thereof may follow standards like ADM or Dolby Atmos®, which are obj ect-based formats . They store the information of render type and information of audio obj ects in a file separate from the actual audio tracks . Examples of an ADM format and the processing of such file is given in Figures 5 and 6 . The audio tracks are associated with the audio obj ects as well as further information thereof to provide the render with the correct information when outputting the audio over a speaker system .

[0100] Figure 2B illustrates a further embodiment of the method of processing a plurality of recorded audio signals in accordance with the proposed principle . In this example , a plurality of sound signals is recorded in a time synchronous manner . The plurality of recorded audio signals contains at least two voices and a non-voice portion . The two voices are coming from to spatially distanced sound sources . The non-voice portions may include ambient sound, noise events and general background noise . In this particular example , the two voice portions are extracted step S30 , thereby generating two separate voice stems . In this example , stems are extracted for each of the plurality of recorded sound signals and subsequently combined such that the voice components ( i . e . the components belonging to a certain source ) are combined into the above- mentioned voice stem. The noise as part of the non-voice portion is extracted in step S31 , the noise may include noise events , e . g . clapping of hand and so forth . Background noise that is generally present as well as other sound portions are considered ambient sound that is extracted in step S32 .

[0101] In contrast to the previous exemplary embodiment , the position of sound sources within the sound space are directly extracted from the plurality of recorded sound signals in step S40 .

[0102] After extraction of the two voice portions , also referred as pure voice stems , one or more functions are applied to each voice stem to obtain further information therefrom . In particular the voice stems are analyzed to detect the speaker in step S200 . This will generate a time vector indicating when the speaker in the respective voice stem is actually talking or silent . In some aspects , pauses in the speech are determined and compared to a threshold distinguishing between breath and an actual pause . Furthermore , one may determine the language during the speech, the gender of the speaker , whether the speaker changes his / her intonation, speed of his talking , the overall time he is speaking in comparison to the total length of the voice stems and so forth . The latter may be connected to the voice stem importance determined in step S203 . For example , the more a speaker talks , the more important he is considered . The importance can be used in the rendering information and the audio obj ect later on .

[0103] For the extracted non-voice portion, and event detection function is applied in step S201 , detecting random occurring and time limited events in the non-voiced stem. Typical event may include a car driving by, clapping of hand in case of audience being present , a machine noise starting or stopping, horn blowing and the like . All those events occur only during a short period of time with different volume . In addition to this , the events themselves may be ordered based on their classification and importance in step S202 . For example , the clapping of hands and an audience may be considered more relevant compared to another event or the background noise itself . In step S35 , the extracted voice stems are used together with the speaker detection to generate the voice tracks . More particularly, the voice tracks for each speaker and / or voices from the same source are generated separately, whereas the speaker detection information and / or the positional information ( to confirm that the sound in the voice stems are from the same source ) is used to provide a continuous speech component in the track . This means the track does not comprise a significant interruption, e . g . a pause by the speaker for instance . While this reduced the overall size of the voice tracks , it is necessary to provide the speaker detection information as a time vector to ensure the correct parts of the voice track are played back correctly in time . This information is part of the generated audio obj ect indicating which part of a voice track shall be output as well as the output time thereof .

[0104] The reduction in the voice stem using the speaker detection information will reduce the file size while maintaining the information, when and for which duration a certain portion of the voice track shall be played back .

[0105] Likewise , an event track may be generated in step S204 using the extracted non-voice stem as well as the detected events in step S201 and S202 .

[0106] The generated voice tracks and the generated event track are then used as an input to generate respective audio obj ects . Each audio obj ect contains a voice track, the speaker detection information as well as the voice importance and other parameters . For example , the speaker detection will provide the time information ( e . g . start time and duration ) , when a certain voice component in the generated voice track shall be output . Likewise , the events are being output based on the time information provided during step S201 and S202 .

[0107] As a result , a plurality of audio obj ects are generated, for example , in this case four audio obj ects . The first and second audio obj ect contain the voice tracks including their respective time information and other information as well as one event track and the overall ambient sound be . The generated audio obj ects can then be combined with an audio content information and stored in one or more data files .

[0108] Figure 3 illustrates an exemplary embodiment of the method for processing a plurality of recorded audio signals in accordance with some aspects of the proposed principle . The processed recorded audio signals result in separated stems for the voice portions , the noise portions and the corresponding crosstalk for each sound source . In this regard it is notes that as explained in Figure 3 of the present application, the plurality of recorded sound signals contains the voice portions of each sound source , and the non-voice portions . If a microphone is associated with a sound source ( e . g . close by or attached to it ) , one can also associate the voice portions of said particular sound source with the respective microphone . In this instance , the voice of the other sound source will be recorded as crosstalk and vice versa .

[0109] In a possible example , the sound scene comprises three speakers set physically apart . Two speakers have a microphone attached to it . An ambisonics microphone array is also arranged to record the scene and provide a position reference . Hence , the plurality of audio signals comprises two signals recorded by individual microphones and another set recorded signals recorded by the ambisonics microphone array . In view of the above , the voices of two speakers recorded by a microphone not associated with a speaker are considered crosstalk with regard to said microphone . In addition, all reflections of voices ( irrespective of the sound source ) recorded by said microphones can be considered either crosstalk or noise . For the third speaker, the situation may slightly vary, as no microphone is directly attached to him or her . Here , voice reflections can be considered crosstalk, but not necessarily the voices of the other speakers .

[0110] Consequently, a recorded signal comprises one or more voice portions , wherein all other portions of the signal , i . e . non-voice portions can be referred to as crosstalk or different kinds of noise . The plurality of recorded signals is recorded by various microphones set spatially apart from each other to be able to identify sound sources based on associated voice and / or other portions to it . The plurality of sound signals is stored in a substantially lossless format in a storage device SD . Typical recording formats include the wav format , aiff , alac , PCM, WavPack and the like . For the ambisonics microphone array, one can alternatively store its recorded signals in a format that allows to include the directional recording information . Suitable formats include an ambisonic format and more particular ambisonic B- format containing the speaker-independent representation of the sound field and the sound scene . The various stored files representing the recorded sound scene can be organized in folders , such that processing of the individual signals is usually performed in a single loop creating processed audio content . Offline processing is usually performed, that is recording is done separately and independently of subsequent processing .

[0111] The proposed method enables to split the individual steps into tasks that are performed automatically and do not require a high user interaction, thus reducing effort for the workflow . Tasks include the above-mentioned recording, the separation of the plurality of audio signals into different voice , crosstalk and noise stems , the generation of audio tracks and its subsequent combination into audio obj ects . For example a producer may record the audio scene , upload it to a storage device SD as illustrated in Figure 3 , from which the various recorded audio signals are processed in accordance with a proposed method .

[0112] In a first step SI , some pre-processing of the stored plurality of audio signals is performed . This includes but is not limited to resampling of the respective signals to a common sampling rate . In particular , the common sampling rate is higher than the sampling rate at which the audio signals are stored . Re-sampling to a higher sampling rate may increase computational effort later on, but also produces better results for the individual channels and allow additional functionality like position estimation with higher accuracy . Typical sampling rate may include , but are not limited to 48 kHz , 96kHz and 192 kHz . Furthermore , as all channels and signals are timely synchronized, one may consider offset removal or initial cutting prior to processing to avoid processing portions of the audio content , that is either uninteresting or will not be used in the final audio product .

[0113] In a subsequent step S2 , a stem separation is performed to separate the above-mentioned three different signal portions in each audio signal from each other , enabling to identify the voice portions , the noise and a possible crosstalk . This step is performed depending on the nature of the signal portion . For example , crosstalk portion may be separated first using either fixed algorithm or trained deep learning networks for identifying the crosstalk and separating it from the channel .

[0114] For this purpose , it is suitable to also evaluate the other channels . In the present embodiment for example , the signals of the first two recording devices attached to the two speakers are evaluated together with support from the recorded signals of ambisonics microphone array, as devices Ml and M2 are usually recording the crosstalk . In other words , to separate the crosstalk in each stem, the proposed method utilizes the audio signals recorded by the other microphones . The identified crosstalk for each recorded signal is separated from the original signal , such that the new signals now only comprise the noise and the voice stems . The identified crosstalk is stored separately for each associated voice portion .

[0115] In a next step and / or parallel to the identification and separation of the crosstalk stem, the noise is identified and subsequently separated from the remaining signal in each channel . Similar to the crosstalk stem, the identified noise stem is subtracted from each signal , leaving only the voice stem . The identification and separation can be changed, i . e . the noise is identified first and separated from the original signal . In any case , each channel ( i . e . each audio signal ) may finally comprise a noise stem including all types of background noise , a crosstalk stem including the identified crosstalk portion associated with a voice portions and a voice or speech portion including the substantially pure speech portion as included in the original signal . With two recording devices and a microphone array having six directional microphones , one obtains up to several tens of different stems , although not each and every stem is required for the following process steps .

[0116] Following the next steps in the method in accordance with the proposed principle , each stem can now be processed separately . For this purpose , the respective stem is used as input to a function, tweaking the stem or portions thereof and providing a processed output stem . This is also referred to as applying a function to one or more stems . Examples of possible functions are presented herein further below . Functions can be mixed and its parameter , if any altered depending on the desired results . Some stems concerning the same part , e . g . referring to the same voice portion and such can be combined .

[0117] It is noted that for the present application, simply removing or attenuating a noise and / or crosstalk stems is not considered a function . Rather , applying a function to a stem will alter the stems or portions of it , but also provide an output that is processed further in subsequent steps .

[0118] As illustrated in Figure 3 , function Fl is applied to the separated voice or speech stem and produces a processed speech stem in step S3 . The nature of function Fl is such that it is suitable mainly for the voice and speech stem, but not for the other two types of stems . The crosstalk stem is input into function F3 in step S5 , altering the crosstalk stem and providing a processed crosstalk stem. Some functions expect an input from two or more stems . For example , function F2 can process the noise stem and the voice stem or processed voice stem as an input as illustrated in step S4 . The function F2 processes both of them separately, that is without any interference of any of the two stems to the respective other . Consequently, a function is applied to the processed voice stem and separately thereof to the noise stem . Examples for functions that are suitable to process the noise or crosstalk stem and the voice stems are provided further below . Several of such functions can be applied to the respective stem to alter portions of it as deemed necessary . In contrast to conventional processing methods , the proposed principle provides a greater flexibility and adj ustment possibilities , as the stems for each channel are processed separately and optionally independent of each other . In some instances , common parameters can be used for the respective function in some of the channels and / or stems . This will allow a fast and efficient workflow for processing the audio content . The functions can also be applied automatically from each other without the need from the user .

[0119] The processed stems or portions thereof are then combined back together in step S 6 . It has been found that processing the voice alone and having it as final audio content often sounds artificial and not natural . Hence , it is suitable to combine some processed noise or also crosstalk and other undesirable signal portions with the processed voice stem to create a more natural hearing experience . The level of the process stem is usually smaller than the original noise level to improve the overall audio content , but also prevent the impression of an "artificial speech" or "artificial voice" . Likewise , a portion of the processed crosstalk stem is then combined into the already combined noise and voice stems in step S7 . The result is an even more natural sound experience . The processed channels with the adj usted and processed voice stems are then mixed together . Additional processing can be performed including but not limited to spatial audio like binaural mixing, ambisonics , stereo mixing and the like . In addition, the result can be down-sampled or finally stored in a lossy format like AAC or mp3 .

[0120] Referring now to Figure 4A, illustrating various blocks of the method in accordance with the proposed principle . In this regard, Figure 4 and the following explanation is based on the international publication WO2023 / 118382 , whose additional context is incorporated herein by reference in its entirety . The application describes in more details the aspects of determining the distance and angle from a given location as a reference point . The proposed method herein and in the international publication WO2023 / 118382 can be used on the recorded sound signals but also on the extracted and separated voice stems.

[0121] For the purpose of simplicity, the method is explained using the abovedescribed scenario of Figure 1. The method is suitable for postprocessing of pre-recorded sound signals but also for real-time sound signals e.g. during an audio conference, a live event, and the like. The method starts with providing one or more first sound signals and a plurality of second sound signals in blocks BM1 and BM2, respectively. The recorded sound signals preferably comprise the same digital resolution including the same sample frequency (e.g. 14bit at 96kHz. In case different resolutions or sampling frequencies are used, it is advisable to re-sample the various sound signals to obtain signals with the same resolution and sampling frequency.

[0122] The upper portion of the depicted method including elements 3' , Rl, 30A and 31 concerns the identification of possible crosstalk between two or more sound signals, that is sound signals, which are recorded by microphones, for which the position is to be determined. As mentioned previously, reflections, but also direct sound are recorded by the two microphones in block BM1. To determine, which of the two or more microphone is actually positioned at the respective sound source, the signals recorded by the two microphones are to be processed filtered and cross correlated to obtain a time difference in the cross correlation .

[0123] For this purpose, both signals are processed using a frequency weighted phase transformation as indicated in 3' and 3. In a first step, each of the first signals are transformed into the frequency domain to using an STFT to obtain a time-frequency spectrum. A spectrum mask filter is derived from the spectrum by first generating a smooth power spectrum S(l,k) , with 1 being the sound signal from the microphone and k the respective frame of the sound signal. For each frequency bin a first order filter estimates the noise n(l,k) in the current frame based on previous frame. The overall noise n(l,k) is given by n(l,k) = (l-o) log(S(l,k) ) + (n(l,k-l) )o with different a depending on S ( 1 , k ) <log ( n ( 1 , k-1 ) ) . Hence , the filter mask is 1 when the SNR is above a certain threshold and otherwise 0 . The results are different filter mas ks , associated with each of the signals to be processed .

[0124] In a next step, the cross spectrum is generated by cross correlating two pairs of the first signals and normalizing the result of the cross correlation . Then, the respective estimated filter is applied to the normalized cross spectrum and an inverse STFT is performed to obtain a filtered and correlated signal .

[0125] In this regard, one should note that for the cross spectrum made Rxy one should use the filter Fx ( for the signal x ) and for the cross spectrum Ryx the filter Fy ( for signal y) . The filtered and correlated signals in 30A are then used to estimate the signed time difference or delay of the direct sound in both microphones recording the first sound signals , see reference 31 . The sign, i . e . dt>0 or dt<0 depicted in block 31 provides information, which microphone is closer to the actual sound source . Consequently, this microphone ( and sound signal ) is then associated with the respective sound source and the corresponding filter mas k .

[0126] The above-mentioned steps can be omitted if the association of sound signals to the respective sound source is defined, i . e . if only one first signal is recorded . Referring back to Figure 4A, the blocks 3 , R2 to 35 illustrated in the lower part describe the various steps of estimating the distance to the reference points and the angle .

[0127] Block BM3 contains a plurality of sound signals recorded by one or more second microphones whose location is fixed in regard to the reference point . The location of each of the second microphones is slightly different to be able to obtain the angle later on, but close enough that effects like reflections from the wall and the like can be determined and filtered . In the present example four different second sound signals are present each recorded by a different second microphone . The process now is similar as described with the processing of the two or more first sound signals . However , in block 3 , the sound signal for which distance and angle shall be determined is now cross correlated with at least one of the four second sound signals . Block 3 can be performed with each of the second sound signals to provide overall four filtered and cross correlated signals , see reference R2 for an example .

[0128] Figure 4B shows the frequency weighted phase transformation in an exemplary embodiment . The two input signals are transformed into the frequency domain using an SFTF and then the cross spectrum is derived from it . After normalizing the spectrum, the previously estimated filter , in this case a spectrum mas k filter associated with the first sound signal is applied . The result is then transformed back into the time domain using an inverse SFTF .

[0129] The time delay in blocks 30B and 30A are estimated by first identifying the maximum value a peak would have if the signals in the frequency weighted PHAT would be uncorrelated . For this purpose , the noise variance is given by sigma=mean (mas k) / frame size and the maximum value of the noise derived by sqrt ( sigma*2 *ln ( frame size ) . Then, a search is performed for the first value in the frequency weighted PHAT that exceeds this maximum (possible including a scale for some headroom) and the search refined for a local maximum close to that first value . The location of the maximum corresponds to the time of flight for the direct sound ( n max / sampling frequency) . The distance is then given by the time of flight multiplied by the speed of sound under consideration of the temperature dependency of the speed of sound .

[0130] The process in block 30B is repeated for each of the cross-spectrum . The various results are further processed in block 31 by using the mean of the set of time of flights estimates . The distance is then deducted in block 33 from this estimate . To obtain the angle between the sound source and the reference points , blocks R3 , 30C and 34 to 36 are used .

[0131] To avoid any influence of room reflections a window function is used to truncate the FW-PHAT results of the first filtered and correlated signal in block R2 . The window function comprises a width, which is dependent on the distance between the second microphones . As the second microphones recording the second sound signals are spaced apart slightly, the estimated distances between the sound source and the respective second microphone may also vary . The width of the window function for truncating the first filtered and correlated signals is substantially proportional to the maximum of the time of flight between the second microphones .

[0132] The now truncated set of filtered and correlated signals are up-sampled to provide a finer time resolution, resulting in a more precise estimate for the angle . The cross correlation between pairs of up-sampled truncated first signals is subsequently calculated . Consequently, one will receive a total of 6 results ( 4 truncated filtered signals results in 6 different pairs ) . The location of the maximum of the cross correlation of a pair of the up-sampled truncated first signals corresponds to the time difference of arrival of the first signal to the respective second microphones . The time difference of arrival is mapped to the angle of incidence making use of the knowledge about the location of the second microphones to the reference point . This means that the cross correlation can be proj ected in the spherical spatial space instead of the time domain . The approach depicted in block 34 is similar to the steered response step in an SRP-PHAT approach, with the 6 pairs of cross-correlations corresponding to the PHAT . The proj ected estimates are then simply summed together and a search for the maximum is conducted in block 35 . The location of the maximum corresponds to the angle of arrival .

[0133] The overall diagram of the ADM model is given in Fig . 5 . The model contains the main elements of various aspects as mentioned above , they are filled with the respective parameters obtained during processing the plurality of audio signals . The Audio Obj ect as generated by the proposed method based on the rendering type and the information obtained during processing generally corresponds to the various boxes in the Format section and the audio obj ect in the content section .

[0134] The ADM model is divided into a content part and the respective format . It also shows the chunk of a file with the actual stored tracks , which have been generated from the plurality of recorded audio signals . As shown, each track contains a unique ID that is also stored in the audioObj ect . Each track also connects to the audioPackFormat and indirectly also to the audioChannelFormat . The render type connects the audio track with the respective audioPackformat , the audioChannelFormat and where applicable also the audioTrackFormat . Further, the render type also fills the respective format definitions .

[0135] Hence , the BWF file contains a number of audio tracks with a list of numbers corresponding to each track in the file . However, the list can actually be longer than the number of tracks in the list , because a single track may have different definitions at different times so will require multiple audioTrackUIDs and references . An example for such is an audio event , e . g . hand clapping and such, which are played more than one time .

[0136] Likewise , the audioTrackFormat may not be unique for each track, and thus as in the present example the audioTrackFormatID is not unique . For example audioTrackFormat may include to definitions for stereo playback ( as defined by the render type ) , that is a definition for the left and one for right channel . Consequently, several tracks can include the same audioTrackFormatIDs . In this example , only two different audioTrackFormatIDs will need to be defined with several tracks sharing the same audioTrackFormatID .

[0137] The audioStreamFormat and audioStreamFormatID is contained in the audioTrackFormat to define and describe a decodable signal . Their respective ID basically answer the question what format the respective track has and how to decode it . Inside audioStreamFormat there will be a reference to either an audioChannelFormat or audioPackFormat that will describe the audio stream, i . e . what type f audio stream the track is , i . e . channel based or obj ect based, a higher order ambisonics or a group of channels . audioChannelFormat defines the typeDef inition attribute , which is used to define what the type of channel is . The typeDef inition attribute can be set for example but not limited to one of ' DirectSpeakers ' , 'HOA' , 'Matrix' 'Obj ects ' or ' Binaural ' set during selection of the render type . For each of those types, there is a different set of sub-elements to specify the static parameters associated with that type of audioChannelFormat . To allow audioChannelFormat to describe dynamic channels (i.e. channels that change in some way over time) , it uses audioBlockFormat to divide the channel along the time axis. The audioBlockFormat element will contain a start time (relative to the start time of the parent audioObject) and duration. Within audioBlockFormat there are time-dependent parameters that describe the channel which depend upon the audioChannelFormat type. For example, this set of parameters allow moving a sound source though the sound scene by properly defining the 'azimuth' , 'elevation' and 'distance'

[0138] At least one audioBlockFormat is required and so static channels will have one audioBlockFormat containing the channel's parameters.

[0139] If audioStreamFormat refers to an audioPackFormat, it describes a group of channels. An audioPackFormat element groups together one or more audioChannelFormats that belong together (e.g. a stereo pair) . This is important when rendering the audio, as channels within the group may need to interact with each other.

[0140] The reference to an audioPackFormat containing multiple audioChannelFormats from an audioStreamFormat usually occurs when the audioStreamFormat contains non-PCM audio which carries several channels encoded together. For example, 'stereo' , '5.1' , '1st order Ambisonics' would all be examples of an audioPackFormat. Note that audioPackFormat just describes the format of the audio. For example, a file containing 5 stereo pairs will contain only one audioPackFormat to describe 'stereo' . It is possible to nest audioPackFormats ; a '2nd order HOA' could contain a '1st order HOA' audioPackFormat alongside audioChannelFormats for the R, S, T, U & V components.

[0141] The formats only defined how the actual tracks are played back, but not which ones belong together, nor what is represented in them. For this purpose the present method also generates an audio object. AudioObject is used to determine which tracks belong together and where they are in the file. This element links the actual audio data with the format, and this is where audioTrackUID comes in. For example, an audioObject may contain references to two audioTrackUIDs representing stereo audio tracks, one for the left and one for the right channel. It will also contain a reference to audioPackFormat, which defines the format of those two tracks as a stereo pair. The audioObject element also contains start and duration attributes. This start time is the time when the signal for the object starts in a file or recording. Thus, if start is "00:00:10.00000", the signal for the object will start 10 seconds into the track in the audio file .

[0142] AudioObjects can be nested, for example, to contain not only references to the two audioTrackUIDs carrying the stream, but also references to two audioOb ects. This allows to generate playback for various channelbased systems, i.e. creating an object for 5.1. and one for 2.0 and nest both together into one audio object.

[0143] AudioObject is referred to by audioContent, which gives a description of the content of the audio; it may contain parameters such as language (if there is dialogue) as detected earlier by the proposed method, loudness parameters, importance of a certain audio object in regard to other and other information. These values are either generated during processing of the plurality of recorded audio signals, but may be adapted and changed by a user later on.

[0144] AudioProgramme brings all the audioContent together; it combines them to make the complete 'mix' . In other standards, like Dolby ATMOS® bed, records and instance records are used as organizational standards to identify content and object type present in the program.

[0145] Figure 6 shows an exemplary structure of an audio program having an audio object including several tracks included. The diagram shows how the defined elements relate to each other. The top half of the diagram covers the elements that describe the 4 channels of the 1st order HOA (N3D method) . The chunk in the middle shows how the four tracks are connected to the format definitions. The content definition elements are at the bottom of the diagram, with the audioObj ect element containing the track UID references to the UIDs in the chunk .

Claims

CLAIMS1 . Method for processing a plurality of audio signals , the method comprising :Providing a plurality of in particularly timely synchronized audio signals , wherein the plurality of audio signals comprises a first voice sound portion from a first sound source and at least one second voice sound portion from a second sound source spatially distanced from the first sound source ; and at least a non-voice portion; and each audio signal of the plurality of audio signals recorded by a microphone of a plurality of microphones , wherein the microphones of the plurality of microphones are spatially separated from each other ;Obtaining a first voice track from the plurality of sound signals , said first voice track comprising at least a part of the first voice sound portion and optionally a part of non-voice portion;Obtaining at least one second voice track from the plurality of sound signals , said at least one second voice track comprising at least a part of the at least one second voice sound portion and optionally a part of the non-voice portion;Determining from the plurality of sound signals a position of the first sound source and the at least one second sound source in relation to each other and / or in relation to a predetermined reference point ;Defining a rendering type , said rendering type optionally associated with one or more rendering channels ;Based on the rendering type , generating sound rendering information for the one or more channels associated with the determined position of the first sound source in relation to the reference point ;Based on the rendering type generating sound rendering information for the one or more channels associated with the determined position of the second sound source in relation to the reference point ;Generating at least one audio obj ect , said audio obj ect containing o a unique ID identifying one or more stored tracks corresponding to at least the first voice track; and o an identifier to the rendering type , said rendering type referring to the one or more rendering channels ; o the generated sound rendering information for the one or more channels associated with the determined position of the first sound source ;Generating at least one audio content obj ect , said at least one audio content obj ect containing : o a reference to the at least one audio obj ect ; o at least one data packet containing characteristics of one of the plurality of audio signals , the first voice track and the at least one second voice track .Generating one or more files on a computer readable storage medium, said one or more files comprising the generated at least one audio content obj ect , the generated at least one at least one audio obj ect and the one or more stored tracks identified by the unique ID .2 . The method according to claim 1 , further comprising the steps of :Generating another audio obj ect , said audio obj ect containing o a unique ID identifying one or more stored tracks corresponding to the at least one second voice track; and o the identifier of the rendering type , said rendering type referring to the one or more rendering channels ; o the generated sound rendering information for the one or more channels associated with the determined position of the second sound source .3 . The method according to any of the preceding claims , wherein the step of generating at least one audio obj ect further comprises the step of generating at least one audio obj ect , said audio obj ect further containing : o a unique ID identifying one or more stored tracks corresponding to the at least one second voice track; ando the identifier of the rendering type , said rendering type referring to the one or more rendering channels ; o the generated sound rendering information for the one or more channels associated with the determined position of the second sound source .4 . The method according to any of the preceding claims , further comprising the steps of :Obtaining a non-voice portion from the plurality of recorded sound signals , said non-voice portion comprising at least a part of non-voice portion; andBased on the rendering type generating sound rendering information for the non-voice portion; wherein the step of generating at least one audio obj ect comprises the step ofGenerating at least one audio obj ect , said at least one audio obj ect containing o a unique ID identifying one or more stored tracks corresponding to the part of the non-voice portion; and o the identifier of the rendering type , said rendering type referring to the one or more rendering channels ; o the generated sound rendering information for the non-voice portion .5 . The method according to any of the preceding claims , wherein the step of generating at least one audio obj ect comprises the step of :Generating an ambient sound scene from the plurality of recorded sound signals , particularly in the B-format .6 . The method according to any of the preceding claims , wherein the step of obtaining one of the first voice track and the second voice track comprises the step ofDetecting speech or the absence of speech in the voice track or the voice portions used for generating said voice track, in particular using a trained time convolutional network; wherein the audio obj ect contains information derived from the detected speech or the absence of speech associated with one ofstored tracks corresponding to said voice track and the unique ID .7 . The method according to any of the preceding claims , wherein the steps of generating one or more files comprises the step of :Associating at least some of the one or more stored tracks with the identifier of the rendering type .8 . The method according to any of the preceding claims , wherein the step of obtaining the first voice track and / or the second voice track from the plurality of sound signals comprises the steps of :- Separating from the plurality of recorded audio signals a clean voice stem, a crosstalk stem and a noise stem, whereas o the clean voice stem comprises substantially only a speech component , o the crosstalk stem comprises a crosstalk portion, and o the noise stem comprises a noise portion of the plurality of recorded signals ;- Applying one or more processing function to the clean voice stem to obtain a processed clean voice stem and- Generating the respective voice track from the processed clean voice stem.9 . The method according to claim 8 , further comprising the steps of- Applying -particularly independently- one or more processing functions to at least one of the crosstalk stem and the noise stem to obtain a processed crosstalk stem and / or processed noise stem;- Combining at least a portion of the processed crosstalk stem and / or processed noise stem into the processed clean voice stem to provide a processed combined stem for generating the respective voice track therefrom .10 . Method according to claim 8 , wherein the one or more processing function comprises at least one of :- a speech detection;- a breath detection function, configured to evaluate the clean voice stem, particularly evaluate those separately from each other ,to mark the respective stem, in which a breath by the associated sound source is detected; an adaptive equalizing function, configured to frequency- selectively gain or attenuate portions of the clean voice stem;- a muting function configured to mute portions of the clean voice stem, in which no speech by the associated sound source is detected;- de-essing function configured to reduce or eliminate the excessive prominence of sibilant consonants within the speech component of the clean voice stem.11 . Method according to any of claims 8 to 10 , wherein the one or more processing function comprises at least one of :- a high-pass filtering function configured to filter frequency portions below a threshold frequency, said threshold frequency in particular below 100Hz and more particularly below 70Hz , wherein the high-pass filtering function is applied in particular to one of the clean voice stem and / or the noise stem; an adaptive levelling function, configured to frequency- selectively gain or attenuate portions of at least one of the clean voice stem and the noise stem .12 . Method according to any of claims 8 to 11 , wherein combining at least a portion of the first processed crosstalk stem and / or first processed noise stem into the first processed clean voice stem comprises :- Combining a portion of the separated crosstalk stem of one of the at least two audio signals into least one processed clean voice stem of said one of the at least two audio signals , in particular prior to applying the first processing function .13 . The method according to any of the preceding claims , wherein the step of obtaining the first voice track and / or the second voice track from the plurality of sound signals comprises the steps of- Separating a crosstalk portion as crosstalk stem from at least one sound signal of the plurality of recorded sound signals , said crosstalk portion associated with the respective voice portion- Separating a noise portion as noise stem from the at least one sound signal .14 . Method according to claim 13 , wherein the step of separating a crosstalk portion comprises Inputting at least a subset of the plurality of recorded sound signals into an artificial network, said artificial network having been trained to identify crosstalk portion in the at least one sound signal , wherein the artificial network utilizes the respective other recorded audio signals to identify crosstalk portions in the at least one sound signal .15 . Method according to any of the preceding claims , further comprises prior to separating from each of the at least two audio signals at least one of :- Re-sampling each of the two audio signals to a common sampling rate , in particular one of 38 kHz , 88 . 2 kHz , 96 kHz and 192 kHz .- Normalising each of the at least two audio signals .16 . Audio processing system comprisingAt least one memory storing at least a plurality of in particularly timely synchronized audio signals , wherein the plurality of audio signals comprises a first voice sound portion from a first sound source and at least one second voice sound portion from a second sound source spatially distanced from the first sound source ; and at least a non-voice portion; and each audio signal of the plurality of audio signals recorded by a microphone of a plurality of microphones , wherein the microphones of the plurality of microphones are spatially separated from each other ;One or more processors coupled to the memory;One or more computer readable programs stored in the at least one memory and configured to cause the one or more processors when executed the method according to any of claims 1 to 15 .17 . Computer readable medium containing a program that when executed causes a processor to execute the method according to any of claims 1 to 15 .