Method for processing audio content and system
The method processes synchronized audio signals to separate and adjust loudness of speech, crosstalk, and noise components, addressing manual intervention issues and enhancing audio quality.
Patent Information
- Application Number
- PCT/EP2025/057239
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-25
AI Technical Summary
Existing audio recording and processing technologies struggle to effectively separate and adjust the loudness of speech, crosstalk, and noise components in audio content, often requiring manual intervention and leading to unnatural sound levels and artifacts.
A method that processes multiple synchronized audio signals from different microphones to separate speech, crosstalk, and noise components, adjusting their loudness independently to create a more natural sound experience without manual input.
Automatically separates and adjusts loudness of speech, crosstalk, and noise components, reducing manual effort and enhancing the naturalness of the audio experience.
Smart Images

Figure EP2025057239_25092025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR PROCESSING AUDIO CONTENT AND SYSTEM
[0002] The present application claims priority of Danish patent application DK PA 2024 70084 dated March 20 , 2024 , the disclosure of which is incorporated herein by reference in its entirety . The present invention concerns a method for processing audio content , said audio content having at least two timely synchronized audio signals recorded by a respective microphone , said microphone associated with a respective sound source , wherein each of the at least two timely synchronized audio signals comprise a speech component associated with the respective sound source and a secondary component unrelated to said respective sound source . The present invention also to a computer system .
[0003] BACKGROUND
[0004] Podcast and some other media types require high level audio content to provide the necessary quality the listener and / or viewer of a certain media content expects and requires . In some environments , audio recording does not meet those expectations . While in studio recordings , any background noise is usually suppressed efficiently, other environments may contain different background noise , including quiet talks from other persons , noise from machinery nearby, natural sounds and so forth . Furthermore , individual speakers are changing the voices , for example peaking louder or calmer depending on the environment .
[0005] In modern record and playback systems , which are used to generate spatial audio content , sound including speech is recorded using a plurality of microphones arranged in various locations . The microphones do not only record the desired sound ( e . g . the speech of a person) , but also crosstalk, whereas the speech of a person is recorded with different microphones , as well as noise and all kinds of ambient sound .
[0006] While recording spatial audio provides various benefits increasing a listener' s experience , the post-processing also becomes more complicated . This aspect becomes even more relevant , if artificial sound effects shall be provided . Those may range from interaction with an (artificial audience) to environmental sounds, like cars driving by, leaves rustling, water splashing and the like.
[0007] For these reasons, various post-processing tools and software have been proposed to improve the quality of the recorded audio. Such processing usually includes denoising and gaining or loudness correction to improve the listening experience. Still, the user of such tools needs to provide a lot of manual input from such recordings . One aspect of such processing already mentioned above includes equalization of the various sound levels of the speakers using loudness adjustment to provide a smoother play back without significant jumps in sound levels.
[0008] The expression "loudness" describes the subjective perception of sound pressure. More formally, it is defined as the "attribute of auditory sensation in terms of which sounds can be ordered on a scale extending from quiet to loud" . The relation of physical attributes of sound to perceived loudness includes of physical, physiological and psychological components. The industry has standardized the loudness, however partially with different meanings and measurement standards . Some definitions, such as ITU-R BS.1770 refer to the relative loudness of different segments of electronically reproduced sounds, such as for broadcasting and cinema. Others, such as ISO 532A (Stevens loudness, measured in sones) , ISO 532B (Zwicker loudness) , DIN 45631 and ASA / ANSI S3.4, have a more general scope and are often used to characterize loudness of environmental noise. More modern standards, such as Nordtest AC0U112 and ISO / AWI 532-3 (in progress) take into account other components of loudness, such as onset rate, time variation and spectral masking.
[0009] Content producers are usually processing the recorded audio content prior to publishing it. The above-mentioned loudness adjustment usually acts on the recorded audio signal, which can lead to undesired enhancement of noise, crosstalk or other artifacts. This may reduce the listener' s experience or cause unnatural variations of the sound level. Avoiding those issues often requires manual input increasing time and effort. It is therefore an obj ect of the present invention to provide a method for processing audio content , that offers not only a flexible way of handling recorded content but reduces the necessary manual effort without sacrificing the listener' s experience .
[0010] SUMMARY OF THE INVENTION
[0011] This and other obj ects are addressed by the subj ect matter of the independent claims . Features and further aspects of the proposed principles are outlined in the dependent claims .
[0012] The inventor proposes a combination of processing a plurality of recorded sound signals having at least two timely synchronized audio signals recorded by a respective microphone , said microphone associated with a respective sound source . Each of the at least two timely synchronized audio signals comprise a speech component that is associated with the respective sound source . In a non-limiting example , the sound source is a speaker . The recorded signal at ach sound source also comprises a secondary component unrelated to said respective sound source .
[0013] For example , such unrelated secondary component may comprise speech of another sound source , referred to as cross talk, noise from a third source and the like . As the at least two audio signals are synchronized in time , one can use the correlation of a speech component recorded at one source and recorded as cross talk at the other component for various purposes . These include but are not limited to acquiring position for the sound sources relative to each other and to a reference point , and separating the respective component associated with the sound sources and the noise from each other .
[0014] Consequently, the method proposes to process the at least two audio signals to separate the portions of one or more voices and the nonvoice portions and subsequently process them as separate stems . The non voice portions may contain ambient , noise and non-voice event as described above . The benefit of processing them as stems lies in the possibility to change several characteristics of the different potions separately without affecting the respective other ones . In addition, the process automatically identifies and extracts various features of the respective portions , including but not limited to language detection, speaker detection, relevance or importance of the respective portions in regard to the overall plurality of recorded signals and so forth . Those may be used later when compiling audio obj ects into a sound scene . In addition, a user may change and vary the extracted features to optimize the immersive experience if needed .
[0015] In some aspects , the inventor proposes a method for processing recorded audio content , said audio content having at least two timely synchronized audio signals recorded by a respective microphone , said microphone associated with a respective sound source , wherein each of the at least two timely synchronized audio signals comprise a speech component associated with the respective sound source and a secondary component unrelated to said respective sound source .
[0016] The expression "timely synchronized" , "synchronized in time" or generally synchronized means that their sampling rate as well as its respective recording time , i . e . its start and stop in relation to each other are known . This can be achieved either by a master-slave configuration of microphones recording the sound signals , a common time clock or any other precise time trigger . Thus , the recorded signals are either not shifted in time with respect to each other , or any time shift between the recorded individual audio signals is known . The signals content however can be shifted because speech or any other sound originating from a sound source may travel different distances towards the respective microphone and thus will be recorded at different times .
[0017] The at least two audio signals are then processed separating from each of the at least two timely synchronized audio signals at least one clean data stem, at least one crosstalk stem and a noise stem . While the at least one clean data stem comprises substantially only the speech component associated with one sound source , the at least one crosstalk stem comprises a crosstalk portion of the secondary component of the other ones of the at least two sound sources . The noise stem comprises a noise portion of the secondary component . The separation may in some instances be applied to each of the recorded audio signals associated with a sound source , such that each signal may in some aspects separated into a speech component as clean data stem, a noise stem and a crosstalk stem.
[0018] A median loudness in the at least one clean data stem is determined in a subsequent step and a target loudness is acquired based on the determined median loudness . The target loudness can be understood as the preferred loudness after processing the audio content . The target loudness may be acquired particularly for the at least one clean data stem and in particular based on the determined median loudness . Similar target loudness can be acquired for the noise stem as well based on the determined median loudness . This will allow adj usting both stems separately depending on the user' s preferences .
[0019] Obtaining the median loudness may be beneficial over the average loudness , in case the speech or clean data stem has significant changes in loudness level . Other benefits become apparent further down below . In a subsequent step, at least one of the noise stem and the crosstalk stem is processed . Particularly, the loudness of the at least one of the noise stem and the crosstalk stem are adj usted based on the acquired target loudness . In some further aspects , the loudness of the data stem is also adj usted based on the target loudness .
[0020] Finally, at least a portion of the at least one of processed noise stem and processed crosstalk stem is combined with the at least one clean data stem to provide a processed combined stem . The present invention therefore proposes to adj ust the loudness of speech, noise and / or crosstalk separately from each other . The proposed approach avoids the generation of sudden loudness changes or enhancement of undesired components as they may occur in conventional solutions . This is achieved by first separating the audio signals into three different components and processing those separately .
[0021] The above steps are performed based on the recorded audio content as such, without the need for a user interaction apart from setting the loudness target . In this regard, the expression user interaction refers to any action taken by the user to associate ( not necessary confirm) a speech or voice to a sound source , a non-voice portion to a sound source and / or voice portion, to set the position of a sound source ( not change a determined position ) , to manually separate speech portions from the recorded audio signals etc . In other words , the inventor proposes a method that is able to perform the necessary generation of the various stems and the subsequent loudness adj ustment without a manual and time-consuming human workflow .
[0022] The proposed method can be used to process stored recorded audio signals , which is audio signals which are stored for an unspecified period of time and then processed in accordance with the proposed principle . Alternatively, the proposed method is used for streamed audio content , which processes the signal as they are recorded . Consequently, the method is not only applicable for pre-recorded audio content but also for streaming events like phone or video conferences , live podcasts and the like .
[0023] In some aspects , the proposed method further comprises the steps of mixing the processed combined stems of each of the at least two audio signals to provide an output signal . The output signal can either be played back, stored or reamed to a third site . It includes the adj usted clean speech component with mixed in noise , noise event and the like to offer a sound environment as desired by the content creator . In some further aspects , the step of determining a loudness comprises the step of detecting speech in the at least one clean data stem . The loudness is determined in the at least one clean data stem upon detection of speech in the at least one clean data stem. The detection of speech might be beneficial , as it prevents from a skewed determination of the loudness . For example , the median loudness is preferred over the average loudness , particularly if the audio content includes several longer quiet passages interrupted by noisy passages .
[0024] Furthermore the detection of speech is used or other parameters as well and reduces the overall effort . Particular, section of the speech stem not containing any voice do not need to be adj usted . By detecting the speech and thus also the quiet or silent portions in between, one can choose in some aspects to determine the loudness for each speech portion individually or determine the overall median loudness .
[0025] In some further aspects the step of determining a median loudness comprises the step of detecting speech in the at least one clean data stem and comparing one of the level of the at least one clean data stem and the average level of the at least one clean data stem with a predefined threshold . Upon detection of speech in the at least one clean data stem and upon detection that one of the amplitude of the at least one clean data stem and the average amplitude of the at least one clean data stem exceeds the threshold, the median loudness is determined in the at least one clean data stem .
[0026] In this regard, some aspects concern the step of detecting speech . Usually, speech comprises short interruption, for example by a speaker taking breath . Those interruptions are not only short in time , but also comprise a certain pattern of sound . Hence , it is possible to include breath detection when detecting the speech and thus differentiate between real pauses in the speech and breathing . The step of determining the median loudness in the at least one clean data stem may then be dependent on the detected breath .
[0027] In some aspects , the step of determining a median loudness is conducted for each clean data stem . Those determined loudness can then be combined or otherwise processed, for example to define a common loudness or to differentiate between the different speakers using different loudness ' s .
[0028] Some further aspects concern the step of acquiring a target loudness . This step may also include acquiring a target for each of the clean data stems . The target loudness may be different and can for example depend on either the content creator' s preference , the "importance" of the respective speaker or sound source , the overall time a specific sound source is active and so forth . These parameters can be determined during processing, and particularly during the stem separation as well as the speech detection . After the target loudness is determined, changes for reaching the target for each of the clean data stem and in particular for each channel in each of the clean data stems can be determined . Likewise in some aspects , changes for reaching the target loudness in the noise and crosstalk stems can be determined . For this purpose , the loudness in those stems is to be obtained first , and the necessary changes can be calculated therefrom and the desired target loudness .
[0029] Consequently, the method according to the proposed principle comprises in the step of processing at least one of the noise stem and the crosstalk stem adj usting the loudness in the at least one clean data stem based on the acquired target loudness .
[0030] The step of acquiring a target loudness comprises in some aspects applying a rolling adj ustment window with a defined window size with a defined overlap between two consecutive windows , whereas the loudness of the signal in each rolling adj ustment window is adj usted based on the target loudness and the loudness in said rolling window .
[0031] In some further aspects , the step of processing at least one of the noise stem and the crosstalk stem comprises applying a slow noise reducer based on the acquired loudness target , in particular by reducing one of the loudness and the level of noise portions above the acquired loudness target ; and applying a noise compressor to reduce fast changing noise portions . The two approaches of processing noise enable to adj ust the loudness of base noise as well as so-called noise event , like hand clapping, whistling and the like .
[0032] Some other aspects concern the step of separating from each of the at least two timely synchronized audio signals . In some aspects thereof a crosstalk portion is separated as crosstalk stem from each of the at least two timely synchronized audio signals using the at least two timely synchronized audio signals . For this purpose a trained neural network can be used, into which the at least two timely synchronized audio signals are input . This improves the separation as the neural network can use the speech component from the respective other audio signal . Furthermore , a noise portion can be separated from each of the at least two timely synchronized audio signals . The two separation steps are subsequently executed or in parallel .
[0033] In some further aspect , the step of separating a crosstalk portion comprises inputting the at least two timely synchronized audio signals into an artificial network, said artificial network having been trained to identify crosstalk portion in one of the at least two timely synchronized audio signals . As stated previously, the artificial network may utilize the respective other audio signals or signals to identify crosstalk portions in the one of the at least two timely synchronized audio signals .
[0034] Some further aspects concern the step of combining at least a portion of the at least one of processed noise stem and processed crosstalk stem with the at least one clean data stem. For this purpose a portion of the processed crosstalk stem of one of the at least two timely synchronized audio signals may be combined with at least one processed clean data stem of said one of the at least two timely synchronized audio signals , in particular prior to adj usting the loudness in the at least one clean data stem. Hence , the loudness would then be adj usted on the combined stems and not separately . This may result in a more natural sound experience .
[0035] The step of combining at least a portion of the at least one of processed noise stem and processed crosstalk stem with the at least one clean data stem comprises a combination of a portion of the separated noise stem into a stem that comprises the clean data stem adj usted by the acquired target loudness and a portion of the crosstalk stem, the latter either adj usted by the acquired target loudness or with a different loudness as stated above .
[0036] The audio content can be recorded in a variety of environment and is not limited to a studio and the like . Particularly, it only requires a number of microphones associated with the sound sources or more particularly, arranged at dedicated locations , such that the position of the sound sources with reference to the position of the microphones can be determined . Consequently, while it is useful to associate a microphone with a sound source and bring such close together , such approach is not required . In some aspects however , the number of microphones configured to record the sound environment is at least as large as the number of sound sources or speakers .
[0037] In some aspects , a first microphone is associated with a first sound source and at least one second microphone is associated with a second sound source . Furthermore , an ambisonics microphone is arranged in a sound environment containing the first and second sound source . The environment is recorded with each of the microphone , such that a plurality a respective audio signal is obtained, each audio signal being recorded by a microphone . The recorded signals can then be processed in real time in accordance with the proposed principle or optionally storing the recorded audio signals , in particularly in a lossless format .
[0038] In some aspects , a mobile phone , a tablet , a smartwatch or any other wearable can be used . Hence , at least one audio signal may be recorded in some aspects by a mobile microphone and four audio signals are recorded by stationary microphones that are arranged in a fixed position towards each other and timely synchronized with the mobile microphone .
[0039] Some aspects concern a computer system . The computer system can be or comprise distributed components and can also be implemented as a virtual computer system. This allows scaling of the necessary components and computational power adapting to the needs of the content creator . The computer system comprises in some aspects at least two timely synchronized microphones , at least one of said microphones associated with a sound source , in particular a mobile sound source .
[0040] It is understood that the at least two timely synchronized microphones can be spatially separated from the remaining parts of the computer system . Particularly the at least two microphones may be mobile in some aspects , while the remaining parts of the system can include hardware that is fixed at a location . The location of the remaining hardware may be distributed and can even cross countries . Hence , it is possible to provide several microphones at different locations , with those microphones configured to transmit the recorded content to a central arrangement having the remaining parts of the system . This enables hardware sharing .
[0041] The computer system in accordance with the proposed principle comprises one or more processors and an optional storage device for storing audio signals recorded by said timely synchronized microphones . The storage device may temporarily store the unprocessed audio content and may also be used for storing the processed audio content . In real time application the storage device can be omitted . The computer system comprises a memory having a program stored therein, said program comprising instructions , which when executed on the one or more processors perform the method according to the proposed principle .
[0042] As mentioned before , the microphones may be configured in some instances to transmit recorded audio signals via a network to the storage . This can be achieved either directly via a wireless communication interface or via a relay, which stored the audio signals temporarily . In such instances , the communication between the at least two microphones and the relay can be different compared to the communication between the relay and the storage . For example , the communication between the microphones and the relay may utilize a Bluetooth protocol , while the communication between the relay and the storage comprises a WIFI connection .
[0043] In some aspects the system according to the proposed principle may further comprise a microphone array including two or more directional microphones for recording a third timely synchronized audio signal . The microphone array can act as the above-mentioned relay . Further , the microphone array is configured in some aspects to time synchronize the at least two microphones . In some further aspects , the microphone array is configured to wirelessly connect to an access point for transmitting the audio signals recorded by said timely synchronized microphones and the third timely synchronized audio signal to the storage device or the memory for real time processing in accordance with some aspects of the proposed principle .
[0044] SHORT DESCRIPTION OF THE DRAWINGS Further aspects and embodiments in accordance with the proposed principle will become apparent in relation to the various embodiments and examples described in detail in connection with the accompanying drawings in which
[0045] Figure 1 shows an embodiment of a method for processing recorded audio content with a plurality of timely synchronized audio signals in accordance with some aspects of the proposed principle ;
[0046] Figure 2 illustrates a time stem diagram illustrating several voice and other portions originating from different sound sources showing some aspects of the proposed principle ;
[0047] Figure 3 shows time-level diagram with several adj ustment windows to illustrate some aspects of the proposed principle ;
[0048] Figure 4 illustrates an embodiment for a processing step applying a certain function to different stem in accordance with some aspects of the proposed principle ;
[0049] Figure 5 illustrates an exemplary sound environment , in which audio content is recorded for being processed with a method in accordance with some aspects of the proposed principle .
[0050] DETAILED DESCRIPTION
[0051] The following embodiments and examples disclose various aspects and their combinations according to the proposed principle . The embodiments and examples are not always to scale . Likewise , different elements can be displayed enlarged or reduced in size to emphasize individual aspects . It goes without saying that the individual aspects of the embodiments and examples shown in the figures can be combined with each other without further ado , without this contradicting the principle according to the invention . Some aspects show a regular structure or form. It should be noted that in practice slight differences and deviations from the ideal form may occur without, however, contradicting the inventive idea.
[0052] In addition, the individual figures and aspects are not necessarily shown in the correct size, nor do the proportions between individual elements have to be essentially correct. Some aspects are highlighted by showing them enlarged. However, terms such as "above", "over", "below", "under" "larger", "smaller" and the like are correctly represented with regard to the elements in the figures. So it is possible to deduce such relations between the elements based on the figures .
[0053] Referring first to Figure 5, which illustrates an exemplary sound scene, in which audio content is recorded. The sound scene also referred to as sound environment contains two speakers, which are having a debate for example. The speakers are referred to as sound sources SSI and SS2, respectively. Generally, for the purpose of this application, the term sound source refers to an entity, -most often a speaker-, whose sound, i.e. utterance, speech or voice is to be recorded. In contrast thereto, other sound or noise sources, like noise coming from machinery or some background people talking but not intended to be recorded, are not considered a sound source. In the sound environment of Figure 5, two recording devices are present. The recording device Ml and M2 are associated with and closely arranged to one of the respective sound source. The recording device Ml and M2 are microphones recording the sound environment and storing the recordings each as a sound signal. In some instances, the microphones Ml and M2 are so- called omnidirectional microphones, i.e. they usually do not have a preferred recording directions . The microphones Ml and M2 are often attached to the sound source's body and carried around with them in case the person or sound source is moving.
[0054] In addition, the sound environment comprises a third recording device M3, for example in the form of a microphone array. The microphone array comprises one or more directional microphones, i.e. figure of eight microphones. The recording device M3 is usually fixed at a dedicated location during the recorded session . All microphones Ml to M3 are timely synchronized, that is , their internal clocks are synchronized . They may also record with the same sampling rate . Time synchronization can be achieved in a master slave fashion, for example , in which the recording device M3 triggers the time sync using a wireless transmission . The wireless transmission uses a low energy protocol like Bluetooth or one of its derivates . Consequently, the start time of the recordings for each microphone is either synchronized as well or even equal , that is , each microphone starts recording at the same time .
[0055] When recording , the microphones in the present embodiment record the sound signals and store it in a memory for later processing . The plurality of all recorded sound signal may also be referred to as audio content .
[0056] Each recorded signal may contain several components , which are outlined in greater detail with respect to recording devices Ml to M3 . Let' s assume sound source SSI (being a person) is talking . In such instance , the recording device Ml is recording the voice or speech with almost a negligible delay due to its close location to sound source SSI . After a short delay, the speech is also recorded at recording device M3 using the directional microphones . After a further delay, the speech is recorded at the recording device M2 located at sound source SS2 . The recorded speech at recording device Ml is referred to as direct speech or direct voice , while the recorded signal at microphone M2 is referred to as crosstalk, due to the fact that it does not originate from sound source SS2 . Likewise , when person SS2 is speaking, the voice recorded at recording device Ml is referred to as direct speech, while its recording at microphone M2 is referred to as crosstalk . Usually crosstalk has a smaller amplitude than the original voice and because of the time synchronization or the microphones , one can distinguish between direct speech and crosstalk .
[0057] Apart from those direct voice and speech portions being direct voice or crosstalk, the voice can also be reflected at a wall or any other obstacle . Furthermore , any machinery, artificial noises from electric or mechanical devices and even other persons talking may be present in the sound environment and recorded at the respective microphones usually with different delays . The latter are usually incomprehensible . Any noise from such devices as well as the incomprehensible utterance from the background is referred to as background noise BN, while the reflected portion are either also identified as noise or as crosstalk ( the identification partially depends on their physical parameters like delay, attenuation, phase shift and so forth ) .
[0058] Figure 2 in this regard illustrates a plurality of speech and voice portions Vpl and Vp2 as well as noise portions NS1 and a noise event portion NE1 being present in the sound scene over a certain period of time being recorded with the various microphones . Of course the various microphones also record cross talk portions , but those are associate with one of the two voice potions . After stem separation as outlined in greater detail below, those four portions may be present . If needed crosstalk i . e . coming from reflection and used to create a more natural environment may be mixed back in as part of the noise .
[0059] The time axis in Figure IB is arranged in the center , with two voice portions Vpl and Vp2 being arranged above the time axis and the two noise portions NS1 and NE1 being arranged below . Each speech portion belongs to a speaker and both speakers are spatially distanced from each other .
[0060] In the present example , both speech portions Vpl and Vp2 ( associated with sound source SSI and SS2 , respectively as depicted in figure 1A) include various speech components , that is the speech portions are not an uninterrupted speech but due to the nature of this example , there are pauses by the two speakers , both are talking at the same time and so forth . Particularly, a speech component of the first voice portions is present between the times the TO and Tl . Then, after a short pause , the speaker SS2 answers creating a speech component of the voice portion VP2 . The speaker makes a pause between time T3 and T4 and then continues . After another pause , the first speaker answers creating another voice component of the voice portion VI at time T6 till T9 . The discussion may continue throughout the session . Furthermore , speaker SS2 starts talking at T8 , then between T8 and T9 both sound sources are producing a speech component ( i . e . talking at the same time ) , thereby creating overlapping components of voice portions Vpl and Vp2 during that timeframe .
[0061] In addition, the sound scene includes a constant background noise indicated by a constant noise stream NS1 . Furthermore , between times T5 and T6 an individual noise event is taking place , which ends slightly after the first speaker starts talking again .
[0062] The microphones including the microphone array will record the respective speech portions Vpl and Vp2 , as well as the noise portions with various delay towards each other . Particularly, microphone M2 will record the voice portion Vpl without substantial delay as it is closely arranged to the sound source itself , while recording the sound portion Vp2 as delayed crosstalk as well as probably the noise event NE1 with a slight delay . Likewise , the microphone M3 will record the voice portion Vpl as crosstalk with a slight time delay due to its distance to the sound source generating Vpl , while recording the voice portion Vp2 without any substantial delay . In contrast , thereto the microphone array MA, being distanced and spaced apart from the respective sound sources will record both voice portions with a slight delay . The time correlation between the voice portions Vpl and Vp2 being recorded at the respective microphones provides the possibility to determine a position of the sound sources with respect to the reference point both in the distance and angles , thereby creating a virtual sound scene .
[0063] Hence , the recorded signals may include components of all voice portions , Vpl , Vp2 as well as the respective noise portions . By combining the individual signals properly, one can extract the pure voice portions as so-called voice or speech stem as well as the noise portions being separated in an event stem as well as background noise stem . Moreover, the recorded audio signals by the microphone array MA contribute to the voice or speech stem but can also be presented by an ambient sound bed, which can be used later during processing to create an audio obj ect embedding the actual voice and speeches of the respective speaker in a more general sound bed . Referring back to Figure 5 , each recorded audio signal may include several superimposed portions of the above types , whereas the ratio of each portion may vary over time . In accordance with the proposed principle these signal portions are categorized into one of the three categories or stems , namely a voice stem, a noise stem and a crosstalk stem for each recorded signal , also referred to a channel . Each stem has its own characteristics making them distinct from each other but is also characterized by the type of its content .
[0064] During recording , the recorded sound is either stored in each of the microphones . In such embodiments each microphone Ml to M3 comprise a respective memory . After recording, the recorded data is transferred to a cloud service CS that stores the recorded data in a structured folder , database or any other suitable location . For data transfer , the microphones Ml and M2 may transfer their recordings first to recording device M3 using a first transmission protocol like for example Bluetooth and the like . The recording device M3 then transmits all audio signals to the cloud service .
[0065] In a further embodiment , the recording devices Ml and M2 already transmit their recordings during the session to device M3 , thereby reducing the amount of their memory needed . Devices Ml and M2 store their recordings only temporary when the wireless connection deteriorates . Array M3 stores the recordings and then transmits them via an access point ( not shown in Figure 5 ) to the cloud service CS . The latter approach offers a higher flexibility as the recordings can be stored in device M3 till a wireless connection is established to an access point and from there to the cloud service .
[0066] In yet a further embodiment , suitable for example for real time applications , the array M3 collects the recorded audio form the microphones Ml and M2 and transmits the audio signals over a wireless or wired communication link . In yet another embodiment , all microphones use a wireless or wired transmission and transfer the respective recorded signals in real time for storage or further processing . The present invention is not limited in this regard to a certain approach . In the present non-limiting embodiment , the recorded audio signals are stored in a storage device SD as channels , wherein each channel is associated with a recording device and / or microphone thereof . The totality of all channels is referred to as audio content . While the channel contains the recorded audio signal , preferably in a lossless format , it may also comprise certain metadata . The audio content is transferred to a processing system, that is configured to perform the method according to the proposed principle .
[0067] Figure 1 illustrates an exemplary embodiment of the method for processing audio content in accordance with some aspects of the proposed principle . As stated above , each channel of the audio content contains a signal recorded by a single microphone or a microphone set .
[0068] For example , some channels comprise an omnidirectional sound signal , whereas the sound signals are stored in a substantially lossless format . Typical recording formats include the wav format , aiff , alac , PCM, WavPack and the like . Further, some additional microphones are arranged in a predetermined configuration providing directional recording that is stored an ambisonic format . Such directional information, whether those are stored separately or in contained in the audio signal itself can subsequently be used during processing of the signals . Ambisonics and more particular ambisonic B-format can be used to store the signal containing a speaker-independent representation of a sound field .
[0069] The various stored files representing the recorded audio content can be organized in folders , such that processing of the individual signals is usually performed in a single loop creating processed audio content . Offline processing is usually performed, that is recording is done separately and independently of subsequent processing .
[0070] The separation of recording and processing enables a work split , whereas a producer may record the audio content , upload it to a storage device SD as illustrated in Figure 1 , from which the various recorded audio signals are processed in accordance with a proposed method . In a first step SI , some pre-processing of the stored audio signals is performed . This includes but is not limited to re-sampling of the respective signals to a common sampling rate . In particular, the common sampling rate is higher than the sampling rate at which the audio signals are stored . Re-sampling to a higher sampling rate may increase computational effort later on, but also produces better results for the individual channels and allow additional functionality like position estimation with higher accuracy . Typical sampling rate may include , but are not limited to 48 kHz , 96kHz and 192 kHz . Furthermore , as all channels and signals are timely synchronized, one may consider offset removal or initial cutting prior to processing to avoid processing portions of the audio content , that is either uninteresting or will not be used in the final audio product .
[0071] In a subsequent step S2 , a stem separation is performed to separate the above-mentioned three different signal portions in each audio signal from each other , that is in each channel from each other . This step is performed depending on the nature of the signal portion . For example , crosstalk portion may be separated first using either fixed algorithm or trained deep learning networks for identifying the crosstalk and separating it from the channel .
[0072] For this purpose , it is suitable to also evaluate the other channels . In the present embodiment for example , the channels of the first two recording devices Ml and M2 are evaluated together with support from the recorded signal of microphone M3 , as devices Ml and M2 are usually recording the actual speech and the crosstalk from the respective other speaker . In other words , to separate the crosstalk in each stem, the proposed method utilizes the audio signal on the other channels . The identified crosstalk for each channel is separated from the original signal , such that the respective signal in each channel now only comprise the noise and the voice stems . The identified crosstalk is stored separately for each channel .
[0073] In a next step and / or parallel to the identification and separation of the crosstalk stem, the noise is identified and subsequently separated from the remaining signal in each channel . Similar to the crosstalk stem, the identified noise stem is subtracted from each signal , leaving only the voice stem . The identification and separation can be changed, i . e . the noise is identified first and separated from the original signal . In any case , each channel ( i . e . each audio signal ) may finally comprise a noise stem including all types of background noise , a crosstalk stem including the identified crosstalk portion and a voice or speech portion including the substantially pure speech portion included in the original signal . The noise stem may comprise a constant noise stem and an event noise stem .
[0074] With two recording devices and a microphone array having six directional microphones , one obtains up to 24 different stems , although not each and every stem is required for the following process steps .
[0075] Following the next steps in the method in accordance with the proposed principle , each stem can now be processed separately . For this purpose , the respective stem is used as input to a function, tweaking the stem or portions thereof and providing a processed output stem . This is also referred to as applying a function to one or more stems . Functions can be mixed and its parameter , if any, altered depending on the desired results .
[0076] As illustrated in Figure 1 , function Fl is applied to the separated voice or speech stem and produces a processed speech stem in step S3 . Function Fl determined the median loudness in the speech stem . Referring back to Figure 2 and speech stem Vpl . The method determines the median loudness for the speech stem Vpl either for the two portions separately or for the two portions combined . The benefit of obtaining the median loudness for each portion of speech stem Vpl separately lies in the fact that the overall loudness can be adj usted for each portion of the speech stem separately based on the median loudness of the respective portion of the speech stem and a desired target loudness . This approach ensures a substantial equal loudness for each portion of the speech stem.
[0077] For the purpose of determining the median loudness , a speech detection function is also applied in step S3 . The speech detection detects the timing , in which the speaker is actually speaking . Referring to Figure 2 , the speech detection function provides two-time vectors indicating speech between times TO and T1 as well as between times T 6 and T9 . The obtained vector, e . g . a probability vector indicating the probability of speech during speech stem Vpl is digitized ( i . e . being a binary vector ) and used to determine the median loudness during the time when speech S detected .
[0078] The median loudness is different from the average loudness , which can be obtained as well . This is beneficial , particularly in certain environment . For example , if the first portion of speech stem Vpl comprises a relatively quiet but long part followed by a relative loud but short part , the average loudness tends to overemphasize the loud part . A median loudness in this regard is closer to the natural environment and can determination of the median loudness copes better with short loud sections than a determination of the average loudness .
[0079] The determined median loudness and the desired target loudness is obtained for each speech stem . The desired target loudness acquired in step S3 can be the same but can also be different . The latter is suitable for example , if the main speaker is to be emphasized . The main speaker can be set manually by the content creator, but also determined using the speech detection function, for example by selecting the speech stem with the longest speech as main speaker .
[0080] The determined loudness and acquired target loudness are then used in the following step to adj ust the loudness using functions F2 and F3 in step S4 and step 4 , respectively .
[0081] Figure 3 shows a possible adj ustment of the loudness to the target level . Figure 3 shows a diagram of loudness or level of a sound signal over time . In this regard, one can adj ust the sound level or the loudness or both of them . Both adj ustments are possible for the proposed method . In the following, the proposed principle is used for loudness or sound level adj ustments . It is understood that the s killed person knows the difference and can adj ust at least one of the parameters properly to achieve the desired effect . The target loudness is given by two threshold values UT for the upper threshold and LT for the lower threshold around the median value M as illustrated . As illustrated the upper threshold is different from the lower threshold in this exemplary embodiment , although both limits can also have the same distance to the median . In the present example this means that loudness portions below the median are enhanced earlier, as the enhancement shall occur at a limit closer to the median .
[0082] If the sound level or the loudness is above the upper threshold UT , the level or loudness is to be reduced, if the sound level or loudness is below the lower threshold level , the level or loudness is to be enhanced . Reduction of loudness occurs for example in the time window between T2 and T3 , between T 6 and T7 and between T9 and T10 . The loudness is enhanced T4 and T5 for example . The adj ustment is done based on the target loudness and the median loudness , e . g . the enhancement or reduction is done stepwise . Furthermore , as indicated herein, the sound signal is divided into certain time windows and the level or loudness is adj usted for each window separately . The time windows are shifted, such that there is an overlap between two consecutive time windows . For example , if each window comprises a size of 200ms , then the window is shifted by 100ms for example to achieve a 50% overlap .
[0083] The overlap with the previous and following time window avoids sudden j umps in level or loudness , which deteriorates the listener' s experience . In addition, one can include a "looking ahead" or "looking past" functionality, such that for the loudness adj ustment not only the actual window is used but also the signal in the past few and probably next few windows . The smoothing will further reduce sudden j ump in the level or loudness and ensures that the level or loudness adj ustment follows the actual sound level , particular longer level changes and ignores short level peaks . The look ahead option is available for recorded sound signals or sound signal which are streamed with a small delay, i . e . quasi-real time audio content .
[0084] As a result , the loudness still varies for the respective speech, but is adj usted with regard to the original sound level . The adj ustment equalizes the sound level without affecting the listener' s experience . The loudness adj ustment is performed to one of the above-mentioned loudness standards like such as ITU-R BS . 1770 , ISO 532A or Nordtest AC0U112 .
[0085] The function F2 is acting slightly differently upon the noise stem in step S4 . The target loudness is usually significantly lower for the noise stem. For the noise stem one can also distinguish between the basic noise NS1 and sudden short noise events NE1 , see Figure 2 . In case of basic noise , the adj ustment is again based on the desired target loudness for noise , e . g . a fraction of the overall target loudness , and the determined loudness for the noise stem and / or the speech stems . The adj ustment is usually performed similar by comparing the level or loudness within a rolling overlapping window with a threshold . However , in contrast to the speech stem adj ustment , the noise level is not enhanced but either kept as is ( in case the noise loudness or noise level is below the threshold) or reduced ( in case the noise loudness or noise level is above the threshold) . Noise events Nel in contrast may also be adj usted to emphasize or reduce the loudness in such events . For example hand clapping may be enhanced, while a car driving by may be reduced .
[0086] With regard to step S5 , a similar function F3 is applied to the crosstalk stem adj usting its loudness as well . The approach can be similar as with function F2 applied to the noise stem. In addition, function F3 may include further functions to be applied to the crosstalk .
[0087] Several other functions can be applied to the respective stems in this approach as functions F2 and F3 , respectively to alter portions of the respective stems as deemed necessary . These include for example a de- esser functionality that reduces hissing sounds in the speech stems . In contrast to conventional processing methods , the proposed principle provides a greater flexibility and adj ustment possibilities , as the stems for each channel are processed separately and optionally independent of each other . In some instances , common parameters can be used for the respective function in some of the channels and / or stems . This will allow a fast and efficient workflow for processing the audio content .
[0088] The processed stems or portions thereof are then combined back together in step S6 . It has been found that processing the speech alone and having it as final audio content often sounds artificial and not natural . Hence , it is suitable to combine some processed noise or also crosstalk and other undesirable signal portions with the processed voice stem to create a more natural hearing experience . The level of the process stem is usually smaller than the original noise level to improve the overall audio content , but also prevent the impression of an "artificial speech" or "artificial voice" .
[0089] Likewise , a portion of the processed crosstalk stem is then combined into the already combined noise and voice stems in step S7 . The result is an even more natural sound experience . The processed channels with the adj usted and processed voice stems are then mixed together in step S8 . Additional processing can be performed including but not limited to spatial audio like binaural mixing, ambisonics , stereo mixing and the like . In addition, the result can be down-sampled or finally stored in a lossy format like AAC or mp3 .
[0090] Figure 4 again illustrates some aspects of the proposed method . Some of the mentioned functions are optional , adj ustable and can be interchanged with other functions applied to the speech of the noise / crosstalk stems . However, it may be useful to perform any analysis prior to altering the respective stems .
[0091] The combined speech stem or alternatively the speech and voice stem separately ( after being analyzed for example ) are filtered using a high pass filter with a cut-off frequency at 70Hz . The filter will reduce any left-over noise portion in the speech stem, but also low frequency components of the speech . It has been found that its presence or absence will not deteriorate the listener' s experience . This high-pass filter is also applied to the noise and / or crosstalk stems using the same cutoff frequency . The adaptive levelling function according to the proposed principle is applied to the speech stem ( or the combined speech stem) as well as to the other stem ( s ) . The adaptive levelling follows the outlined principle , adj usting the loudness or level as explained above . In the specific example of Figure 4 , the noise stem and the combined speech stem are altered . The respective target level for the adaptive levelling can be pre-set or individually adj usted for each stem . Similar to the other function, it is possible to use previously or additional evaluated metadata for the adj ustment . Further, the pre-set parameter may be set globally that is for each channel and not only individually for each stem .
[0092] Consequently, the present method proposes to apply certain functions either separately to one or more stems or combine stems together and then apply a function to it . Depending on the function, the results will be improved in comparison to conventional techniques , in which the audio signal as a whole is processed .
[0093] Further illustrated in Figure 4 are two more functions that are applied to the combined speech stem or alternatively to the pure speech stem . These functions are a de-esser that identifies and attenuates sibilants in a speaker' s voice and a mute of inactive speakers . The latter can re-use the results of the active speaker detection or identify silent portions in a speech combined speech stem and mute them. In more complex situations , this function improves the comprehensibility . Both functions improve the listener' s experience in more complex sound environments , in which a plurality of speakers are present and partially talking at once .
[0094] It has been found that complete removal of background noise is often perceived as artificial , particular in environments , in which some background noise is expected . For example a studio recording is different compared to an outdoor interview, or a recording in front of an audience . To create a more natural sound experience , a portion of the processed noise stem is mixed together with speech stem and / or the processed speech stem . This may lead to situation, in which an audio signal contains the speech portion, that is gained or processed in a first way, while noise and crosstalk portions are processed differently . Depending on the functions applied on the individual stems , the end result after combining the processed stems may not be possible , when the audio signal is processed as a whole as in conventional systems .
[0095] The proposed method is implemented in a computer system in the cloud of example . Referring back to figure 8 sows an example of a recorded audio content being transferred to a cloud server CS . The cloud server contains a storage SD as well as a memory and one or more processors . The proposed method allows for parallel processing , i . e . separating the stem on different processors or processing those on different processors . The cloud server contains a computer readable medium that if executed on the respective one or more processors perform the proposed method . It should be noted that the adaptive equalization can be applied to real time audio , thus the one or more processors enhances and processes the audio content and transmits it backs to a listener . This approach enables enhancement of speech during real time applications like phone conferences , online seminars , life podcasts and the like .
Claims
CLAIMS1 . Method for processing recorded audio content , said audio content having at least two timely synchronized audio signals recorded by a respective microphone , said microphone associated with a respective sound source , wherein each of the at least two timely synchronized audio signals comprise a speech component associated with the respective sound source and a secondary component unrelated to said respective sound source ; the method comprising the steps of :- separating from each of the at least two timely synchronized audio signals at least one clean data stem, at least one crosstalk stem and a noise stem, whereas o the at least one clean data stem comprises substantially only the speech component , o the at least one crosstalk stem comprises a crosstalk portion of the secondary component of the other ones of the at least two sound sources , and o the noise stem comprises a noise portion of the secondary component ;- determining a median loudness in the at least one clean data stem;- acquiring a target loudness , in particular for the at least one clean data stem and in particular based on the determined median loudness ;- processing at least one of the noise stem and the crosstalk stem by adj usting its loudness based on the target loudness ;- combining at least a portion of the at least one of processed noise stem and processed crosstalk stem with the at least one clean data stem to provide a processed combined stem .2 . The method according to claim 1 , further comprising the steps of : mixing the processed combined stems of each of the at least two audio signals to provide an output signal .3 . The method according to any of the preceding claims , wherein the step of determining a median loudness comprises the step of :Detecting speech in the at least one clean data stem;Upon detection of speech in the at least one clean data stem :o Determining the median loudness in the at least one clean data stem .4 . The method according to any of claims 1 to 2 , wherein the step of determining a median loudness comprises the step of :Detecting speech in the at least one clean data stem;Comparing one of the amplitude of the at least one clean data stem and the average amplitude of the at least one clean data stem with a threshold;Upon detection of speech in the at least one clean data stem and upon detection that one of the amplitude of the at least one clean data stem and the average amplitude of the at least one clean data stem exceeds the threshold : o Determining the median loudness in the at least one clean data stem .5 . The method according to claim 4 or 3 , wherein the step of detecting speech also comprises : Detecting breath in the at least one clean data stem; wherein the step of Determining the median loudness in the at least one clean data stem is dependent on the detected breath .6 . The method according to any of the preceding claims , wherein the step of determining a median loudness is conducted for each clean data stem .7 . The method according to any of the preceding claims , wherein the step of acquiring a target loudness comprises the step of : Acquiring a target for each of the clean data stems ;Determining changes for reaching the target for each of the clean data stem and in particular for each channel in each of the clean data stems .8 . The method according to any of the preceding claims , wherein the step of processing at least one of the noise stem and the crosstalk stem comprises :adj usting the loudness in the at least one clean data stem based on the acquired target loudness .9 . The method according to claim 7 , wherein the step of acquiring a target loudness comprises applying a rolling adj ustment windows with a defined window size with a defined overlap between two consecutive windows , whereas the loudness of the signal or the level of the signal in each rolling adj ustment window is adj usted based on the target loudness and the loudness or level in said rolling window .10 . The method according to any of the preceding claims , wherein the step of processing at least one of the noise stem and the crosstalk stem comprises : applying a slow noise reducer based on the acquired loudness target , in particular by reducing one of the loudness and the level of noise portions above the acquired loudness target ; and applying a noise compressor to reduce fast changing noise portions .11 . The method according to any of the preceding claims , wherein the step of separating from each of the at least two timely synchronized audio signals comprises :- separating a crosstalk portion as crosstalk stem from each of the at least two timely synchronized audio signals using the at least two timely synchronized audio signals ;- separating for each of the at least two timely synchronized audio signals , in which in particular the crosstalk portion has been separated, a noise portion from each of the at least two timely synchronized audio signals ; wherein optionally the two separation steps are subsequently executed .12 . The method according to claim 11 , wherein the step of separating a crosstalk portion comprises : inputting the at least two timely synchronized audio signals into an artificial network, said artificial network having been trained to identify crosstalk portion in one of the at least two timely synchronized audio signals ,whereas optionally the artificial network utilizes the respective other audio signals to identify crosstalk portions in the one of the at least two timely synchronized audio signals .13 . Method according to any of the preceding claims , wherein the step of combining at least a portion of the at least one of processed noise stem and processed crosstalk stem with the at least one clean data stem comprises : combining a portion of the processed crosstalk stem of one of the at least two timely synchronized audio signals with at least one processed clean data stem of said one of the at least two timely synchronized audio signals , in particular prior to adj usting the loudness in the at least one clean data stem.14 . Method according to any of the preceding claims , wherein the step of combining at least a portion of the at least one of processed noise stem and processed crosstalk stem with the at least one clean data stem comprises :- combining a portion of the separated noise stem into a stem that comprises the clean data stem adj usted by the acquired target loudness and a portion of the crosstalk stem .15 . Method according to any of the preceding claims further comprising :- associating a first microphone to a first sound source and at least one second microphone to a second sound source ;- arranging an ambisonics microphone in a sound environment containing the first and second sound source ;- recording with each of the microphone a respective audio signal ;- optionally storing the recorded audio signals , in particularly in a lossless format .16 . Method according to any of the preceding claims , wherein at least one audio signal is recorded by a mobile microphone and four audio signals are recorded by stationary microphones that are arranged in a fixed position towards each other and timely synchronized with the mobile microphone .17 . System comprising :- at least two timely synchronized microphones , at least one of said microphones associated with a sound source , in particular a mobile sound source ;- one or more processors ;- a storage device configured to store audio signals recorded by said timely synchronized microphones ;- a memory having a program stored therein, said program comprising instructions , which when executed on the one or more processors perform the method according to the preceding claims .18 . System according to claim 17 , further configured to transmit recorded audio signals from the at least two timely synchronized microphones via a network to the storage .19 . System according to claim 17 or 18 , further comprising- microphone array including two or more directional microphones for recording a third timely synchronized audio signal , wherein the microphone array is configured to time synchronize the at least two microphones .20 . System according to claim 19 , wherein the microphone array is configured to wirelessly connect to an access point for transmitting the audio signals recorded by said timely synchronized microphones and the third timely synchronized audio signal to the storage device .
Citation Information
Patent Citations
Enhancing intelligibility of speech content in an audio signal
US10096329B2
Methods and systems for processing and mixing signals using signal decomposition
US20150317983A1