How to process recorded audio content
By separating audio signals into clean data, crosstalk, and noise stems and processing them individually, the method addresses the uncanny valley effect in spatial recordings, enhancing the listening experience and reducing complexity.
Patent Information
- Application Number
- JP2025537641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-02
- Filing Date
- 2023-12-20
- Publication Date
- 2026-01-14
Smart Images

Figure 2026501359000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to DK Patent Application No. PA202370001, filed January 2, 2023, the disclosure of which is incorporated herein by reference in its entirety. The present invention relates to a method for processing recorded audio content and a system arranged to carry out the method. [Background technology]
[0002] Podcasts and some other media types require a high level of audio content to deliver the necessary quality that listeners and / or viewers of that particular media content expect and desire. In some environments, audio recordings may not meet these expectations. While studio recordings typically efficiently suppress all background noise, other environments may contain a variety of background noises, such as other people speaking quietly, noise from nearby machinery, or sounds of nature.
[0003] As a result, recorded audio is typically processed by content creators before being released. Such processing typically includes noise reduction and amplification or loudness correction to enhance the listening experience. In modern recording systems used to record spatial audio content, sounds, including speech, are recorded using multiple microphones placed in various locations. The microphones not only record the desired sound (e.g., a person's speech), but also record crosstalk if the person's speech is recorded with different microphones. Recording spatial audio offers various benefits that enhance the listener's experience, but also makes post-processing more complex.
[0004] For this reason, various post-processing tools and software have been proposed to improve the quality of recorded audio. However, it has been observed that in certain recordings, audio processing can cause an uncanny valley effect, meaning that the reproduced audio content sounds artificial and unnatural to the listener. To avoid this effect, audio content producers must invest a great deal of effort and time.
[0005] The purpose of this application is to improve the listener's experience without placing an additional burden on the producers of audio content. Summary of the Invention
[0006] These and other objects are solved by the subject matter of the independent claims. Features and further aspects of the proposed principle are outlined in the dependent claims.
[0007] It has been found that audio processing of multiple recorded audio tracks for noise reduction or crosstalk suppression often operates on individual tracks to generate audio content. While this is generally useful, it has also been found that such processing is often the cause of the uncanny valley, i.e., an acoustic experience that is perceived as artificial and unnatural. In spatial recordings, where much effort is put into generating more natural audio content, adjusting noise reduction, crosstalk suppression, gains of various channels, etc., can be rather complex and require several iterations.
[0008] Therefore, the present inventors propose an improved technique for processing recorded audio content, including separately recorded but time-synchronized audio signals. The concept proposes separating the various components in each audio signal based on their respective types (hereafter referred to as stems) and processing each type separately for each signal. In this way, the various parts (or stems) are processed separately, and then at least some of them are recombined. This allows processing for different parts (called stems) to be optimized separately. Furthermore, because specific processing functions are often related to specific types, potential interference from other components is minimized.
[0009] Based on the generated audio content, this approach allows for the flexible creation of natural sound environments that provide a pleasant listener experience and avoid the uncanny valley. At the same time, separating each audio signal into various stems reduces the subsequent computational load and allows for more flexible and individual handling in post-processing.
[0010] In some aspects, a method for processing recorded audio content is proposed, the audio content including at least two time-synchronized audio signals recorded by respective microphones, at least one microphone associated with a respective audio source, and each of the at least two audio signals including a speech component associated with the respective audio source and a secondary component unrelated to the respective audio source.
[0011] The proposed method therefore operates on at least two time-synchronized audio signals, at least one of which is associated with a sound source, such as a person, and at least one other of the at least two microphones is associated with a second sound source or can be located in a nearby position, in the latter case the at least one other microphone is typically fixed.
[0012] The proposed method comprises the steps of separating clean data stems, crosstalk stems, and noise stems from each of at least two audio signals. These stems are defined as types of sounds contained in the audio signals. For the purposes of this application, the following stems are defined as clean data stems, crosstalk stems, and noise stems.
[0013] A clean data stem substantially contains speech components, particularly speech components from the sound source with which the microphone is associated, but is deprived of other audio portions. In a high-quality studio environment with a single microphone and no sound reflections or other (noise) sources, the recorded audio signal is most likely to represent a clean data stem.
[0014] A crosstalk stem contains a crosstalk portion recorded by a microphone associated with each sound source. The term crosstalk refers to sound components recorded by microphones not associated with the sound source causing the sound component. Crosstalk also refers to sound components reflected off walls or obstacles but originating from the sound source associated with the microphone. Simply put, crosstalk often refers to the recorded background speech or recorded voices of one or more people who are not the current or desired speaker, as well as the reflected portions of the speaker's voice.
[0015] To the listener, the crosstalk portion may be understandable but simply degraded, whereas the noise portion typically does not provide any useful information to the listener.
[0016] Although the former audio components are usually different from the direct recording of the sound source, the reflections may have a similar frequency and time distribution, but are delayed, attenuated, and phase shifted, and therefore can also be treated as part of the noise stem.
[0017] Noise stems include noise parts of secondary components. A noise stem is defined as a part of a recorded signal that does not provide any useful or derivable information to the listener. Noise stems are not related to the sound source but originate from different sources (reflections can be attributed to walls or obstacles). Typical sound sources are machines, circuits, and natural sounds such as flowing streams or rivers. In a recorded audio signal, one skilled in the art can identify many more sound sources that generate or contribute to noise stems.
[0018] After separating a recorded audio signal into individual stems, each stem of each audio signal can be processed separately, or two or more stems from the same audio signal can be processed using the same processing function and parameter set, which offers a significant improvement over prior art techniques that apply a processing function to the entire audio signal rather than to each individual separated stem, and many variations and combinations are possible.
[0019] In particular, the separated stems of each audio signal can be stored in non-volatile memory for later use.
[0020] A method according to the proposed principle applies a first processing function to a clean data stem of at least one of the at least two audio signals to obtain a first processed clean data stem. In particular, independently of the clean data stem, the first processing function is applied to at least one of the crosstalk stems and noise stems of the at least one audio signal to obtain a first processed crosstalk stem and / or a first processed noise stem.
[0021] In this regard, applying a function to a stem, or more generally to an audio signal or part thereof, means that the respective stem or part is used as input for the function, and the content of the respective stem is changed or modified accordingly.
[0022] The first processed clean data stem is then combined with at least a portion of the first processed crosstalk stem and / or the first processed noise stem to obtain a processed combined stem, and the processed combined stems of each of the at least two audio signals are then mixed to provide an output signal.
[0023] Separating a recorded audio signal into different stems offers several advantages over conventional techniques. For one, processing functions subsequently applied to the various stems can be optimized for each stem. Furthermore, some processing functions produce better results when applied only to a single stem rather than in combination. For example, de-esser and breath detection are preferably applied only to the clean data stem to avoid identifying noise portions as voice or speech. Similarly, automatic amplification control should avoid increasing crosstalk during speaker silence. The proposed principles allow for the development of specialized functions for one of the separated stems, providing content providers with greater flexibility when processing audio content.
[0024] In some further examples, applying a first processing function to the clean data stems further includes applying a second processing function to the at least one clean data stem to obtain at least one processed clean data stem, the at least one processed clean data stem being an input to the first processing function.
[0025] Thus, different processing functions may be applied to individual stems, or the same processing function may be applied to two or more stems before being recombined. It is also possible for the output of a processing function to serve as the input for the next processing function. Thus, a stem may be processed using multiple processing functions. Some of these processing functions may be applied to different stems, while others may only be used on a single stem. Furthermore, some functions may only be applied to stems of a dedicated audio signal, rather than to each recorded audio signal.
[0026] The following sections describe the functionality of specific processing functions. These implementations include, for example, open-source libraries, but also proprietary solutions. Furthermore, some processing functions are static, i.e., use fixed algorithms or approaches, while other processing functions include pre-trained deep learning networks optimized to provide the respective function.
[0027] In some examples, the second processing function includes an adaptive equalization function configured to frequency-selectively amplify or attenuate a portion of at least one clean data stem of the at least two audio signals, where the amplification or attenuation can be performed for each data stem having a different parameter set, thereby adjusting the volume of different audio sources to a common level without increasing crosstalk and noise.
[0028] It may also be useful to mute portions of the clean data stem of at least one of the at least two audio signals, particularly portions where there is no speech from the audio source associated with at least one of the audio signals. This feature is particularly useful when the different audio sources are people talking or arguing. Muting or at least significantly reducing the volume of these signals not associated with the current speaker may improve overall quality by suppressing unwanted signal portions from other audio sources. This feature may be applied to all clean data stems so that they can be selectively muted, given the need to identify speech and pauses in each data stem.
[0029] Thus, in some embodiments, the second function applied to one data stem of the at least two audio signals also utilizes the other data stem to derive various parameters, including, but not limited to, markers associated with voice or speech.
[0030] Some other functions relate to improving the actual voice or speech contained in the data stems. In some embodiments, the second function is configured to reduce or remove excessive prominence of sibilants in at least one speech component of at least one clean data stem of the at least two audio signals. This is also referred to as sibilance suppression. In some examples, the second processing function includes a breath detection function configured to evaluate each of the at least one clean data stem of the at least two audio signals, particularly evaluating them separately from each other, and marking portions of each stem in which breathing by an associated audio source is detected. In some examples, the breath detection is then used to adjust the gain of each stem, thereby preventing increased breathing sounds. Alternatively, detected breath of a certain audio source is automatically adjusted, for example, by attenuating the corresponding portion.
[0031] Some further aspects relate to a first processing function that is applied to one of the data stems or processed data stems (using one or more of the second processing functions described above) and one of the noise stems and crosstalk stems. This function can be applied to the various stems with the same parameters or with different sets of parameters. However, the first processing function is applied to the stems individually, so that the data stems and the noise or crosstalk stems are used separately as inputs to the first processing function. This allows for increased flexibility, as the results of each of the processed data, noise, and crosstalk stems can be optimized separately by applying each function separately.
[0032] The first processing function may include, for example, a high-pass filter function configured to filter frequency portions below a threshold frequency. The threshold frequency may be set to different values depending on the type of stem (e.g., data, noise, and / or crosstalk). The threshold frequency may be specifically below 100 Hz, more specifically below 70 Hz. The high-pass filter function may be applied specifically to one of the clean data stems and one of the noise stems.
[0033] In further examples, the first processing function may include an adaptive leveling function configured to frequency-selectively amplify or attenuate at least a portion of a clean data stem of at least one of the at least two audio signals, and in particular, a portion of a noise stem associated with at least one of the at least two audio signals. Such a function may use markers associated with voice or speech in the data stem. For example, by evaluating these markers, crosstalk and noise stems may be selectively attenuated or amplified during periods when no voice or speech is present in the data stem. Such time-selective and stem-selective amplification may contribute to a more natural listening experience.
[0034] In some examples, the step of separating stems from each of the at least two audio signals includes using the at least two audio signals to separate crosstalk portions from each of the at least two audio signals as crosstalk stems. Also, in particular, for each of the at least two audio signals from which the crosstalk portions have already been separated, a noise portion is separated. Optionally, these two separation steps are performed sequentially. Thus, the separation of stems is performed in two distinct steps, where the crosstalk portion is separated in a first step and the noise portion is separated from the audio signal in a subsequent step. Adding all three stems together again derives the recorded audio signal.
[0035] It may be useful to obtain information from other audio signals to identify crosstalk portions, thereby improving the performance of the crosstalk separation step.
[0036] Some aspects relate to performing the separating step. In some examples, separating the crosstalk portions includes inputting the at least two audio signals to an artificial neural network, the neural network being trained to identify the crosstalk portions in one of the at least two audio signals, and optionally, the artificial neural network utilizing each of the other audio signals to identify the crosstalk portions in one of the at least two audio signals.
[0037] In some examples, the audio signals are preprocessed to simplify subsequent processing steps or to adjust the characteristics of the audio signals to match them. In some embodiments, the at least two audio signals are preprocessed before separating at least one stem from each of the at least two audio signals. Such preprocessing may include, among other things, resampling each of the two audio signals to a common sampling rate of one of 38 kHz, 88.2 kHz, 96 kHz, and 192 kHz. Additionally or alternatively, each of the at least two audio signals may be normalized.
[0038] The step of combining the processed stems can be performed in a variety of ways, depending on the desired effect, audio content, or listener preference. It has generally been found that adding a small amount of crosstalk and / or noise to the listener's experience is preferable, as it is perceived as more natural. Thus, the step of combining at least a portion of the first processed crosstalk stem and / or the first processed noise stem with the first processed clean data stem may include integrating a portion of the separated crosstalk stem of one of the at least two audio signals with at least one processed clean data stem of one of the at least two audio signals.
[0039] This combining can occur at various stages, but typically occurs after at least one processing function has been applied to the data stem. Thus, before integrating portions of the noise and / or crosstalk stems into the data stem, the data stems are processed and adjusted according to selected processing functions. For example, a data stem can be subjected to one or more second processing functions before being combined with portions of one of the other stems. The noise and crosstalk stems can also be processed before being integrated with the other stems. Separating the audio signal into three different stem types and processing them independently of each other provides all permutations of possible combinations, thus providing a high degree of freedom of choice.
[0040] In some further examples, combining at least a portion of the first processed crosstalk stem and / or the first processed noise stem with the first processed clean data stem to provide a processed combined stem includes integrating a portion of the separated noise stem into a stem including the clean data stem and a portion of the crosstalk stem processed with the second function. In such embodiments, a portion of the processed noise stem is combined with a data stem having a portion of the crosstalk stem already integrated. The data stem and / or the crosstalk stem may be processed to the extent that it is applied to one of the above functions.
[0041] After recombining the various stems, or portions thereof, loudness correction and normalization can be performed on each of the processed combined stems. The result is a plurality of processed audio signals, at least one of which is associated with a sound source and has noise and crosstalk reduced by individually adjustable portions. These processed audio signals can be mixed together according to a desired overall output.
[0042] Such mixing can include the evaluation of spatial information, allowing the processed audio signals to be placed in a virtual audio space. This includes binaural formats, Ambisonics B-format, and other formats. Alternatively, audio signals can be mixed as a stereo signal.
[0043] In some embodiments, the proposed principle further includes associating a first microphone with a first sound source and at least one second microphone with a second sound source. An Ambisonics microphone array is also arranged in a sound environment including the first and second sound sources. For purposes of the method, one of the microphones records a respective audio signal, which is then stored in a non-volatile memory. The format used should preferably be lossless to avoid losing information. The use of microphones, each associated with a dedicated audio source, in addition to the Ambisonics microphone array offers several advantages. Processing of the audio signals, particularly data and crosstalk stem processing, allows for the identification and location of sound sources in the sound environment.
[0044] In one embodiment, at least one audio signal is recorded by a mobile microphone and four audio signals are recorded by fixed microphones positioned at fixed positions facing each other and time-synchronized with the mobile microphones. Such audio signals offer the possibility to generate and manipulate virtual sound spaces. The proposed principle supports such applications by generating clean data stems with separate information on crosstalk and noise.
[0045] Some other aspects relate to a system comprising at least two time-synchronized microphones, at least one of which is associated with a sound source, in particular a moving sound source. The system comprises one or more processors and a memory for storing audio signals recorded by the time-synchronized microphones. The memory stores a program, the program including instructions for performing the method according to the claims when executed on the one or more processors.
[0046] The microphones may be located in different locations and may be far away from the processor and memory or storage. Furthermore, the microphones may record audio signals, which are then uploaded to a storage cloud for subsequent processing. This facilitates workload distribution, with one person responsible for the initial recording and a second person responsible for processing the recorded audio content. Therefore, the system may be configured to transmit audio signals recorded from at least two time-synchronized microphones over a network to a storage device.
[0047] In some embodiments, the system further includes a microphone array for recording a third time-synchronized audio signal, the microphone array configured to time-synchronize at least two microphones. The microphone array may include one or more directional microphones having a figure-of-eight directivity characteristic. The microphone array may operate as a master, and at least two microphones associated with each audio source may operate as slaves. The microphones may be configured as lavalier microphones.
[0048] In this regard, the microphone array comprises an interface for wireless connection to an access point using, for example, a Wi-Fi protocol. Communication between the at least two microphones and the microphone array may use a different protocol, such as Bluetooth. The microphone array is configured to transmit the audio signals recorded by the time-synchronized microphones and the third time-synchronized audio signal to a storage device via the wireless interface. [Brief explanation of the drawings]
[0049] Further aspects and embodiments in accordance with the proposed principles will become apparent in connection with the various embodiments and examples that are described in detail with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates an embodiment of a method for processing recorded audio content having multiple time-synchronized audio signals according to some aspects of the proposed principles. [Figure 2] 1 illustrates an example of a data object that can be used to facilitate a method for processing audio content in accordance with some aspects of the proposed principles. [Figure 3] FIG. 1 illustrates an embodiment of a pre-processing step of a method for processing audio content according to some aspects of the proposed principles. [Figure 4] 10 illustrates exemplary method steps for separating each stem according to some aspects of the proposed principles. [Figure 5] 1 illustrates some exemplary method steps for processing an audio stem according to some aspects of the proposed principles. [Figure 6] 1 illustrates one embodiment of processing steps for applying specific functions to different stems according to some aspects of the proposed principles. [Figure 7] 1 illustrates some exemplary method steps for processing the recombined stems and storing them in a suitable format according to some aspects of the proposed principles. [Figure 8] 1 illustrates an exemplary sound environment in which audio content is recorded to be processed in a manner according to some aspects of the proposed principles; DETAILED DESCRIPTION OF THE INVENTION
[0050] The following embodiments and examples disclose various aspects and their combinations according to the proposed principles. The embodiments and examples are not necessarily drawn to scale. Aspects of the embodiments and examples shown in the drawings may be combined with each other unless otherwise specified, as long as they do not violate these principles. In practice, slight differences and deviations from the specific embodiments may occur, but these do not violate the spirit of the invention. Some steps may be reversed or interchanged with other steps, but the relationships between elements can be inferred based on the figures.
[0051] First, referring to FIG. 8, FIG. 8 illustrates an exemplary sound environment in which audio content is recorded. The sound environment includes, for example, two speakers having a discussion. The speakers are referred to as sound sources SS1 and SS2, respectively. In general, in this application, the term sound source refers to the entity (in most cases, the speaker) whose sound, i.e., speech, voice, or voice, is to be recorded. In contrast, other sound or noise sources, such as noise from machinery or background people talking that are not intended to be recorded, are not considered sound sources. In the sound environment of FIG. 8, each sound source includes a recording device M1, M2 associated with and positioned proximate to each sound source. The recording devices M1 and M2 are microphones that record the sound environment and store the recording as an acoustic signal. In some examples, the microphones M1 and M2 are so-called omnidirectional microphones, which typically do not have a preferred recording direction. The microphones M1, M2 are attached to the sound source body and often move with the sound source body when the sound source moves.
[0052] The sound environment also includes a third recording device M3, for example in the form of a microphone array. The microphone array includes one or more directional microphones, i.e., figure-of-eight directional microphones. The recording device M3 is typically fixed in a dedicated position during the recording session. All microphones M1 to M3 are time-synchronized, i.e., their internal clocks are synchronized. Each microphone may also record at the same sampling rate. Time synchronization can be achieved, for example, in a master-slave manner, in which the recording device M3 triggers time synchronization using wireless transmission. The wireless transmission uses a low-energy protocol such as Bluetooth or one of its derivative protocols. Therefore, the recording start times of each microphone are synchronized or equal, i.e., each microphone starts recording at the same time.
[0053] During recording, the microphones capture the sound environment and store it in memory for post-processing. The recorded signal may contain several components, which will be described in more detail with respect to recording devices M1-M3. Assume that sound source SS1 (person) is speaking. In this case, recording device M1 is located close to sound source SS1, so it records the speech with negligible delay. After a short delay, the speech is also recorded by recording device M3 using a directional microphone. After a further delay, the speech is recorded by recording device M2, located at sound source SS2. The speech recorded by recording device M1 is called direct speech or direct audio, and the recording signal at microphone M2 is called crosstalk because it is not attributable to sound source SS2. Similarly, when person SS2 is speaking, the audio recorded by recording device M1 is called direct speech, and the recording at microphone M2 is called crosstalk. Crosstalk typically has a smaller amplitude than the original audio, and time synchronization or microphones can distinguish between direct speech and crosstalk.
[0054] Apart from these direct sound and direct speech portions being direct sound or crosstalk, the sound may also be reflected off walls or any other obstacles. Furthermore, the sound environment may include mechanical noise, artificial noise from electrical or mechanical devices, or the sounds of other people speaking, typically recorded with different delays by each microphone. The latter is usually unintelligible. The noise from such devices and the unintelligible speech from the background is referred to as background noise (BN), and the reflected portions are identified as either noise or crosstalk (the identification depends in part on physical parameters such as delay, attenuation, and phase shift).
[0055] Thus, each recorded audio signal may contain several overlapping segments of the above types, while the proportion of each segment may vary over time. According to the proposed principle, these signal segments are classified into one of three categories or stems: speech stems, noise stems, and crosstalk stems. Each stem has its own unique characteristics that distinguish it from the others, but it is also characterized by its content type.
[0056] During recording, the recorded sound is stored in each microphone. In such an embodiment, each microphone M1-M3 has its own memory. After recording, the recorded data is transferred to the cloud service CS, which stores the recorded data in a structured folder, database, or any other suitable location. For data transfer, microphones M1 and M2 may first transfer their recordings to the recording device M3 using a first transmission protocol, such as Bluetooth. The recording device M3 then transmits all audio signals to the cloud service. In a further embodiment, the recording devices M1 and M2 reduce the amount of memory required by transmitting their recordings to device M3 during the session. Devices M1 and M2 only record temporarily if the wireless connection deteriorates. Array M3 stores the recordings and then transmits them to the cloud service CS via an access point, not shown in FIG. 8. This latter approach provides greater flexibility, as the recordings can be stored on device M3 until a wireless connection to the access point is established and the cloud service is connected from there.
[0057] The recorded audio signals are stored in a storage device SD as channels, each channel associated with a recording device and / or microphone. The collective set of all channels is referred to as audio content. A channel contains the recorded audio signal, preferably in a lossless format, and may include certain metadata, as described below. The audio content is transferred to a processing system configured to execute a method according to the proposed principles.
[0058] Figure 1 shows an example embodiment of a method for processing audio content according to some aspects of the proposed principles. As mentioned above, each channel of audio content comprises a signal recorded by one microphone or set of microphones.
[0059] For example, some channels may contain omnidirectional audio signals, but the audio signals are stored in a substantially lossless format. Typical recording formats include wav, aiff, alac, PCM, WavPack, etc.
[0060] Additionally, several additional microphones may be placed in predetermined locations to provide directional recordings that are stored in Ambisonic format. Such directional information, whether stored separately or included in the audio signal itself, can be used during signal processing. Ambisonics, and more specifically the Ambisonic B format, is used to store signals that contain speaker-independent representations of the sound field.
[0061] The various stored files representing the recorded audio content are organized into folders, and the processing of the individual signals is typically performed in a single loop to create the processed audio content. Typically, offline processing is performed, i.e., the recording is made separate and independent of any subsequent processing.
[0062] The separation of recording and processing allows for a division of labor: on the one hand, producers may record audio content, upload the recorded audio content to a storage device SD, and process the various recorded audio signals according to the proposed method, as shown in Figure 1 .
[0063] The first step, S1, involves preprocessing a portion of the stored audio signals. This includes, but is not limited to, resampling each signal to a common sampling rate. In particular, the common sampling rate is higher than the sampling rate at which the audio signals are stored. Resampling to a higher sampling rate may increase subsequent computational effort, but it provides better results for individual channels and enables additional functions, such as more accurate position estimation. Typical sampling rates include, but are not limited to, 48 kHz, 96 kHz, and 192 kHz. Furthermore, because all channels and signals are time-synchronized, offset removal or initial cuts can be considered before processing, avoiding processing portions of audio content that are not used or interesting in the final audio product.
[0064] In the next step S2, stem separation is performed to separate the three different signal parts in each audio signal, i.e., in each channel, from each other. This step is performed depending on the nature of the signal parts. For example, the crosstalk part is first separated using either a fixed algorithm or a trained deep learning network to identify and separate the crosstalk from the channels.
[0065] For this reason, it is preferable to evaluate other channels as well. For example, in this embodiment, the channels of the first two recording devices M1 and M2 are evaluated with support from the recording signal of microphone M3, since devices M1 and M2 typically record crosstalk. In other words, to separate the crosstalk in each stem, the proposed method utilizes the audio signals in the other channels. The identified crosstalk for each channel is separated from the original signal, so that each signal for each channel contains only noise and the audio stem. The identified crosstalk is stored separately for each channel.
[0066] In the next step and / or in parallel with the identification and separation of the crosstalk stems, noise is identified and then separated from the remaining signal of each channel. Similar to the crosstalk stems, the identified noise stems are subtracted from each signal, leaving only the speech stems. This identification and separation can be modified so that the noise is identified first and then separated from the original signal. In any case, each channel (i.e., each audio signal) may ultimately include noise stems containing all types of background noise, crosstalk stems containing identified crosstalk portions, and a speech or speech portion containing substantially pure speech portions contained in the original signal.
[0067] With two recording devices and a microphone array with six directional microphones, up to 24 different stems are acquired, although not all stems are necessarily required for the following processing steps.
[0068] By following the next steps of the method according to the proposed principles, each stem can be processed individually. To this end, each stem is used as input to a function that adjusts the stem or a portion thereof and provides a processed output stem. This is also referred to as applying a function to one or more stems. Examples of possible functions are further provided below. Functions can be mixed and, where appropriate, their parameters can be modified depending on the desired result. Note that, for the purposes of this application, simply removing or attenuating noise and / or crosstalk stems is not considered a function. Rather, applying a function to a stem modifies the stem or a portion thereof, but also provides an output that is further processed in subsequent steps.
[0069] As shown in Figure 1, function F1 is applied to the separated audio or speech stems to generate processed speech stems in step S3. The nature of function F1 makes it suitable primarily for audio and speech stems, but not for the other two types of stems. In step S5, the crosstalk stems are input to function F3, which modifies the crosstalk stems and provides processed crosstalk stems.
[0070] Some functions expect input from more than one stem. For example, as shown in step S4, function F2 can process a noise stem and a speech stem or a processed speech stem as input. Function F2 processes both separately, i.e., processes one of the stems without interfering with the other corresponding stem. Therefore, the function is applied to the noise stem separately from the processed speech stem. Further examples of functions suitable for processing noise or crosstalk stems and speech stems are provided below.
[0071] Some of these functions can be applied to each stem, modifying parts of it, as needed. In contrast to traditional processing methods, the proposed principle offers greater flexibility and adjustability, since stems for each channel are processed separately, optionally independent of each other. In some instances, common parameters for each function can be used across several channels and / or stems. This allows for a fast and efficient workflow for processing audio content.
[0072] The processed stems or portions thereof are then recombined in step S6. Processing only the audio to produce the final audio content has been found to often sound artificial and unnatural. Therefore, it is suitable to combine some of the processed noise, or crosstalk and other undesired signal portions, with the processed audio stem to create a more natural listening experience. The level of the processed stem is typically lower than the original noise level, improving the overall audio content but also preventing the impression of "artificial speech" or "artificial voice."
[0073] Similarly, a portion of the processed crosstalk stem is combined with the already combined noise and speech stem in step S7, resulting in a more natural sound experience. In this manner, the processed channels with the adjusted and processed speech stem are mixed together. Additional processing can include, but is not limited to, spatial audio such as binaural mixing, ambisonics, and stereo mixing. The result can also be downsampled or ultimately saved in a lossy compression format such as AAC or mp3.
[0074] Figure 2 shows in more detail one embodiment of a data object for audio content. The data object contains the original audio content, as well as individual stems, parameters for each applied function, and the overall result. Some parameters or measurement results are also stored here. Thus, the data object provided with the method according to the proposed principles provides a non-destructive process that allows the user to supervise, undo, or modify the processed stems.
[0075] The data object contains the path and name of the recorded audio signal on a separate recording device associated with the audio source. These are called "lavalier_file_path". The data object also contains the path to the recorded audio signal in Ambisonics file format, called "ambisonics_file_path". This file contains the recorded audio signal from a microphone array that provides directional information for the audio signal itself. With these two pieces of information, the data object contains the original audio content.
[0076] The data object also provides metadata, for example, each feature applied to the audio content, which is useful because it allows different features and algorithms to be compared and the best one selected.
[0077] Additional fields in the data object are mostly empty and are filled in as the original audio content is processed. For example, fields such as "crosstalk_stems", "noise_stems", and "clean_stems" contain a list of separate files containing the crosstalk, noise, and audio stems for each audio signal listed in the path above. Further information such as the rms SNR and rms noise values may also be included. These parameters and values are included if they are reused at various stages of the audio content processing, reducing computational effort.
[0078] Figure 3 provides an example of the various steps during the pre-processing of each audio signal. The audio signals are synchronized in time. In an optional initial step, Ambisonics microphone recordings are converted to B format. To speed up the processing process, sound files are loaded into cache memory. This step is useful during subsequent processing steps, such as identifying and separating crosstalk in various audio signals. The sampling rate of each sound file is extracted and identified. If necessary, resampling may be performed to sample all sound files mentioned in the embodiment to a common sampling rate of 48 kHz as data tracks. To avoid signal clipping during processing, possible offsets are identified and removed, and peak normalization is performed to reduce peaks to 0.9.
[0079] Finally, an optional first notch EQ filter is applied to identify and reduce notches in each audio signal. The notch filter acts on the overlapping portion of each signal, including audio, crosstalk, and noise. Typically, notch EQ filtering is used to suppress or reduce peak noise that occurs at a specific frequency with a relatively large amplitude.
[0080] After pre-processing of the various audio signals, each individual channel is separated into its respective speech, noise, and background stems. The separation may be performed independently of the signal processing. However, in some cases, it may be suitable to process the audio signal of each channel by applying a function to identify portions of the respective signal. Figure 4 shows an embodiment in which crosstalk and noise are identified and stored in additional stems that are then subtracted from the input signal to obtain pure audio stems.
[0081] The first step of the separation process is to identify crosstalk using a trained deep learning model. The trained model uses one channel (or preprocessed audio signal) as input and each of the other acoustic signals to identify crosstalk portions, taking into account delays, reflections, and different levels. In the example sound environment of FIG. 8, the signals from recording devices M1, M2, and M3, along with the other signals, can be used as inputs to identify crosstalk portions. Typically, crosstalk is delayed compared to direct speech and may also be attenuated. The crosstalk portions identified in this way are stored in a separate stem and then subtracted from the original audio stem.
[0082] The modified audio signal is fed to a noise suppressor (denoiser), which identifies noise portions in the remaining signal. The noise can have various origins and be part of the frequency range of the human voice. A trained deep learning model is used as the noise suppressor. In some examples, a series of different noise suppressors can be used, including, for example, a trained deep learning model for identifying dynamic noise and a static noise suppressor. The training model and / or algorithm often depends on the type of noise to be removed from the signal. Similar to a crosstalk reducer, this noise suppressor includes settings that allow the user to identify noise and / or crosstalk, but also to select the amount of each portion to be removed.
[0083] As with the crosstalk reducer, the identified noise is used to fill in the noise stems, which are then subtracted from the audio signal, resulting in a crosstalk- and noise-free residue known as the audio stem or speech stem. This process is repeated until each stem in each channel is isolated.
[0084] In this way, the various stems are processed separately and different features are applied to each. Figure 5 shows some example features applied to speech stems. Some of these features are optional and / or interchangeable with other features. This approach allows the selection of each feature applied to a speech stem depending on the desired result. Additionally, because the original speech stem is stored in the data object, new features can be implemented. This also allows different processed speech stems to be compared to select the adjustments and parameters that provide the best performance.
[0085] The speech and audio stem are input to a first function, the output of which is then fed to a further function that triggers a chain of processes. In particular, the speech stem is evaluated to detect the speaker of the utterance and mark them accordingly. This function runs separately and adds some meta-information to the stem. This function does not modify the actual audio stem, but simply analyzes the audio stem and generates metadata including one or more timestamps.
[0086] Similarly, the audio stem is input to a function that detects the speaker's breath. Recorded breaths typically contain specific frequencies with specific levels and waveforms. They are typically audible and can create an unpleasant experience for the listener. The audio stem is analyzed using a dedicated trained deep learning network. Sections where breaths are detected are marked accordingly. In a post-processing stage, these sections may be slightly attenuated rather than emphasized or amplified. In an alternative embodiment, the breath detection function may attenuate sections where breaths are detected. Breath detection is applied to at least the direct speech audio stem, and, if necessary, to the channels of the Ambisonic file.
[0087] Additionally, the active portions of the audio stems for all channels (e.g., identified during active speaker detection) are used to identify and locate the speaker's position in the sound environment. In this regard, it may be advantageous to also use information in the Ambisonics file. To locate the speaker, the time delay and phase shift of various signals can be analyzed with respect to the location of a reference point. In the example of FIG. 8, the reference point is the location of recording device M3. Furthermore, while this function is applied to each audio stem, it may be appropriate to use the crosstalk portion of each signal as an input parameter. The location information can then be used to selectively amplify or attenuate each speaker, as well as to adjust the sound environment. For example, a virtual listener can be created and sound sources can be positioned at a desired distance and angle from the virtual listener.
[0088] The audio stems are also analyzed for signal-to-noise ratio. The above functions are primarily used to analyze the audio stems, and the results are stored in a data object for later use. However, audio stem processing is not limited to these steps. Rather, audio stems can be modified. For example, a level gain EQ can be applied to adapt the individual audio stems to a common level.
[0089] In contrast to the processed and analyzed speech and audio stems, the crosstalk stems remain unchanged. A portion of each crosstalk stem is then mixed with each analyzed and processed speech stem. A portion of the crosstalk stem, for example, the attenuated crosstalk stem, is added to the processed analyzed audio stem for each channel. The result includes the processed audio and speech portion and a portion of the crosstalk. This combined stem is then equalized using a trained deep learning model or an algorithm with adjustable parameters. Adaptive EQ is applied to the combination of the speech stem and the portion of the crosstalk stem, not to each stem individually.
[0090] The following steps in the process are followed, as shown in Figure 6. These functions are optional, adjustable, and interchangeable with the analysis functions shown in Figure 5. However, it may be useful to perform some analysis before modifying each stem.
[0091] The combined speech stem, or the speech stem and audio stem individually (e.g., after analysis), is filtered with a high-pass filter having a cutoff frequency of 70 Hz. The filter reduces the noise portion remaining in the speech stem, but also reduces the low-frequency components of the speech. It has been found that the presence or absence of the filter does not detract from the listener's experience. This high-pass filter is also applied to the noise and / or crosstalk stems using the same cutoff frequency. Similarly, the adaptive leveling function is applied to other stems in addition to the speech stem (or combined speech stem). In the example shown in Figure 6, the noise stem and combined speech stem are modified. The target levels for adaptive leveling can be preset or adjusted individually for each stem. As with other functions, adjustments can use pre-evaluated metadata. Furthermore, the preset parameters are for each channel and can be set individually for each stem, as well as globally.
[0092] Therefore, the method proposes applying a function to one or more stems individually, or to combine the stems and then apply a function, which, in some cases, improves the results compared to prior art techniques rather than processing the entire speech signal.
[0093] Figure 6 further shows two features that are applied to combined or pure speech stems: a sibilance suppressor that identifies and attenuates sibilants in a speaker's voice, and a silent speaker mute. The latter can reuse the results of active speaker detection or identify and mute silent portions of combined speech stems. In more complex situations, this feature improves intelligibility. Both features improve the listener's experience in more complex sound environments where multiple speakers are present and partially speaking simultaneously.
[0094] It has been found that complete removal of background noise is often perceived as artificial, especially in environments where some background noise is expected. For example, a studio recording is different from an outdoor interview or a recording in front of an audience. To create a more natural acoustic experience, a portion of the processed noise stem is mixed with the speech stem and / or the processed speech stem. This can result in a situation where an audio signal, including the speech portion, is gain adjusted or processed in a first way, while the noise and crosstalk portions are processed differently. Depending on the functions applied to the individual stems, processing the entire audio signal as in conventional systems may not produce the final result after combining the processed stems.
[0095] Finally, post-processing functions are applied to the entire stem, including the combined speech, noise, and crosstalk stems, as shown in Figure 7. This may include, for example, loudness normalization. Loudness normalization is the process of modifying the gain to achieve a target average amplitude throughout the entire recording. This average may be a simple measure of average power, such as an RMS value, or a measure of human perceived loudness, such as those provided by ReplayGain, Sound Check, and EBU R128. Standard loudness normalization reference levels vary by location and application. Loudness normalization is applied to at least the channel for the main talker or direct voice.
[0096] And finally, the audio signals are stored again according to the path specified in the data object, which then contains the processed audio signals of the recording devices associated with each sound source.
[0097] The data object can then be used to mix the channels according to the desired output. Several options are possible, as shown in Figure 7. In some examples, the stored audio signals are remixed into a stereo signal. However, thanks to the position information obtained by the proposed method, an optional smart stereo positioning can be applied, which can change or adjust the virtual position of the sound source on the two stereo channels. Optionally, a further loudness normalization is applied to the mixed audio, and the result is written in the desired format.
[0098] Alternatively, the processed audio signal can be spatially processed with different presets to generate binaural mixed sounds that are stored in memory. Sound files can also be output in Ambisonics B format. [Explanation of symbols]
[0099] M1, M2 recording devices M3 Ambisonics Recorder SS1, SS2 sound source BN background noise SD storage device
Claims
1. 1. A method for processing audio content in which at least two time-synchronized audio signals are recorded by respective microphones, the microphones being associated with respective sound sources, each of the at least two audio signals having a speech component associated with the respective sound source and a secondary component unrelated to the respective sound source, the method comprising: - separating a clean data stem, a crosstalk stem and a noise stem from each of said at least two audio signals, the clean data stem includes substantially only the speech component; said crosstalk stems comprise crosstalk portions of said secondary components of other sound sources of said at least two sound sources, - separating the noise stems, which contain the noise parts of the secondary components; - applying a first processing function to the clean data stem of at least one of the at least two audio signals to obtain a first processed clean data stem; - applying said first processing function, in particular separately, to at least one of said crosstalk stems and said noise stems of said at least one audio signal to obtain a first processed crosstalk stem and / or a first processed noise stem; - combining at least a portion of said first processed crosstalk stem and / or said first processed noise stem with said first processed clean data stem to provide a processed combined stem; - mixing said processed combined stems of each of said at least two audio signals to provide an output signal; A method comprising:
2. applying a first processing function to the clean data stem; - applying a second processing function to at least one of said clean data stems to obtain at least one processed clean data stem that is an input of said first processing function; The method of claim 1.
3. The second processing function is a breath detection function configured to evaluate each of the clean data stems of at least one of the at least two audio signals, in particular to evaluate the clean data stems independently of each other, and to mark stems in which a breath is detected with an associated sound; an adaptive equalization function configured to frequency-selectively amplify or attenuate portions of said clean data stem of at least one of said at least two audio signals; a muting function configured to mute portions of the clean data stem of at least one of the at least two audio signals that are not speech-generated by a sound source associated with the at least one audio signal; - a sibilance suppression function configured to reduce or remove excessive prominence of sibilance in at least one speech component of the clean data stem of at least one of the at least two audio signals, The method of claim 2.
4. The first processing function is a high-pass filter function adapted to filter frequency portions below a threshold frequency, in particular below 100 Hz, more particularly below 70 Hz, and applied in particular to one of said clean data stems and to one of said noise stems; an adaptive leveling function configured to frequency-selectively amplify or attenuate portions of at least one of the clean data stems of at least one of the at least two audio signals and in particular the noise stems associated with said at least one of the at least two audio signals, The method according to any one of claims 1 to 3.
5. The step of separating from each of the at least two audio signals comprises: - using said at least two audio signals, separating crosstalk portions from each of the at least two audio signals as crosstalk stems; - separating a noise portion from each of said at least two audio signals, in particular from each of said at least two audio signals from which said crosstalk portion has been separated, optionally said two separation steps being performed sequentially. The method according to any one of claims 1 to 4.
6. The step of separating the crosstalk portion includes: inputting the at least two audio signals to an artificial network trained to identify portions of crosstalk in one of the at least two audio signals, optionally wherein the artificial network utilizes each other audio signal to identify portions of crosstalk in the one of the at least two audio signals. The method of claim 5.
7. before the step of separating from each of the at least two audio signals, - resampling each of said at least two audio signals to a common sampling rate, in particular one of 38 kHz, 88.2 kHz, 96 kHz and 192 kHz; - normalizing each of said at least two audio signals, The method according to any one of claims 1 to 6.
8. The step of combining at least a portion of the first processed crosstalk stem and / or the first processed noise stem with the first processed clean data stem comprises: - combining a portion of the separated crosstalk stem of one of the at least two audio signals with at least one processed clean data stem of said one of the at least two audio signals, in particular before applying said first processing function, The method according to any one of claims 1 to 7.
9. Combining at least a portion of the first processed crosstalk stem and / or the first processed noise stem with the first processed clean data stem to provide a processed combined stem includes: - combining a portion of the separated noise stem with a stem that includes a portion of the clean data stem and the crosstalk stem that has been processed with a second function, The method according to any one of claims 1 to 8.
10. The step of mixing the treated bonded stems comprises: - performing loudness normalization of each of said processed combined stems, The method according to any one of claims 1 to 9.
11. - associating a first microphone with a first sound source and at least one second microphone with a second sound source; - placing an Ambisonics microphone in a sound environment containing said first and second sound sources; - recording respective audio signals with said first and second microphones; - storing the prerecorded audio signal, in particular in a lossless format; The method of any one of claims 1 to 10, further comprising:
12. At least one audio signal is recorded by a mobile microphone; Four audio signals are recorded by fixed microphones placed at fixed positions facing each other and time-synchronized with the mobile microphones; The method according to any one of claims 1 to 11.
13. at least two time-synchronized microphones, at least one of which is associated with a sound source, in particular a moving sound source; one or more processors; a storage device for storing the audio signals recorded by said time-synchronized microphones; a memory on which a program containing instructions is stored, said instructions being adapted to perform the method according to claims 1 to 12 when executed on said one or more processors; A system comprising:
14. further configured to transmit the recorded audio signals from the at least two time-synchronized microphones to the storage device over a network. The system of claim 13.
15. a microphone array including two or more directional microphones for recording a third time-synchronized audio signal; the microphone array is configured to time synchronize at least two microphones; 15. A system according to claim 13 or 14.
16. the microphone array is configured to wirelessly connect to an access point for transmitting the audio signals recorded by the time-synchronized microphones and the third time-synchronized audio signal to the storage device; 16. The system of claim 15.