Music synthesizer with spatial metadata output

By designing a device for a synthesizer, the device can directly generate audio signals and their associated spatial metadata based on the control signal, solving the problem of difficulty in integrating spatialization processes in the prior art, and improving the flexibility and efficiency of sound design.

CN117897765BActive Publication Date: 2025-06-17DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280059728.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-03
Filing Date
2022-08-24
Publication Date
2025-06-17
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Existing synthesizers are difficult to effectively integrate spatialization processes in object-based production, limiting creative options for sound design, and requiring manual editing of spatial metadata, which is inefficient.

Method used

An apparatus is designed including a first stage for obtaining an audio signal, a second stage for modifying the audio signal based on the control signal, and a third stage for generating spatial metadata based on the control signal. The device allows direct generation of spatial metadata related to the audio signal, avoiding intermediate rendering steps and enhancing flexibility in sound design.

Benefits of technology

The creative options in object-based production are broadened, allowing spatialization to be processed as an integrated part of the sound design process, avoiding laborious editing of spatial metadata, and improving processing efficiency and creativity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117897765B_ABST
    Figure CN117897765B_ABST
Patent Text Reader

Abstract

A device for generating and / or processing an audio signal is described. A device includes: a first stage for obtaining an audio signal; a second stage for modifying the audio signal based on one or more control signals for shaping the sound represented by the audio signal; a third stage for generating spatial metadata associated with the modified audio signal based at least in part on the one or more control signals; and an output stage for outputting the modified audio signal together with the generated spatial metadata. A corresponding method, as well as a corresponding program and a computer-readable storage medium, are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 240,383, filed on September 03, 2021 (Reference No.: D21053USP1) and European Application No. 21194849.2, filed on September 03, 2021 (Reference No.: D21053EP), which are hereby incorporated by reference. Technical field

[0003] The present disclosure relates to methods and apparatuses for generating and / or processing audio signals. The present disclosure further describes techniques for synthesizing or processing sounds with spatial metadata output. These techniques can be applied to, for example, music synthesizers and audio processors. Background art

[0004] A synthesizer is an electronic musical instrument that generates audio signals. Generally, electronic musical instruments use electronic circuits to produce sounds, usually in response to user input. In the broadest sense, this can include analog instruments, DSP - based instruments running as "virtual" instruments on dedicated hardware or computers, or sample - based instruments.

[0005] Existing synthesizers generate mono, stereo, or multi - channel audio signals. This means that when used for object - based production (e.g., Dolby production), a spatialization process separate from the sound generation process is applied to represent the synthesizer audio signal as an object. Currently, this is achieved by first configuring the output of the rendering synthesizer for a specific channel, importing the rendered audio into a digital audio workstation (DAW) (such as Pro Tool), and then using a panner to generate the associated spatial metadata.

[0006] Generally, there is a need for improved techniques for object - based production of sound signals generated by synthesizers and / or for object - based production of sound signals processed by audio processors. Summary of the invention

[0007] In view of the above, the present disclosure provides an apparatus for generating and / or processing audio signals, as well as a corresponding method, computer program, and computer - readable storage medium, which have the features of the respective independent claims.

[0008] According to one aspect of the present disclosure, there is provided an apparatus for generating or processing an audio signal. The apparatus may include a first stage for obtaining an audio signal. The apparatus may further include a second stage for modifying the audio signal based on one or more control signals for shaping (e.g., modifying, altering) the sound represented by the audio signal. The control signals may be generated by one or more modulators that affect how the audio signal is modified. For example, one or more modulators may involve or include a low-frequency oscillator (LFO) and / or an envelope. The apparatus may further include a third stage for generating spatial metadata related to the (modified) audio signal based at least in part on the one or more control signals. The spatial metadata may be metadata for indicating to an external device how to render the modified audio signal. For example, the spatial metadata may be metadata for object-based rendering. For example, it may include an indication of the position (e.g., Cartesian position in 3D space) and / or size of an audio object associated with (e.g., represented by) the audio signal. The apparatus may further include an output stage for outputting the modified audio signal together with the generated spatial metadata. In some embodiments, the apparatus may process more than one audio signal in parallel (i.e., multiple audio signals). In some other embodiments, more than one audio signal (i.e., multiple audio signals) may be processed serially. For example, in some such embodiments, the processing of a given audio signal may depend on an earlier audio signal and / or its metadata. As another example, in some embodiments, the processing may involve modifying existing spatial metadata of the audio signal obtained by the first stage.

[0009] Configured as above, the techniques implemented by the proposed apparatus significantly broaden the range of creative options available in object-based production and allow spatialization to be treated as an integral part of the sound design process. In addition, laborious editing of spatial metadata can be avoided, and creative intentions that cause specific shaping operations on the audio signal can be implemented directly and efficiently in the spatial metadata.

[0010] In some embodiments, the second stage may be adapted to apply time-dependent modifications to the audio signal. Wherein, the time-dependence of the modification may depend on one or more control signals. The one or more control signals may be time-dependent. Thereby, the potential time-dependence of the expected behavior of the sound source can be easily used to describe the spatial characteristics of the audio object that describes / represents the sound source.

[0011] In some embodiments, the second stage may include at least one of the following to modify the audio signal: a filter; an amplifier; a low-frequency oscillator; an audio delay; a driver; and a flanger effect.

[0012] In some embodiments, the second stage may be adapted to apply a filter to the audio signal. It should be understood that the second stage may include a filter. For example, the filter may be any one of a high-pass filter, a low-pass filter, a band-pass filter, or a notch filter. The characteristic frequency of the filter may be controlled by one or more control signals (e.g., based on one or more control signals). Thus, the characteristic frequency may be time-dependent. For example, the characteristic frequency of the filter may be a cut-off frequency. For example, one or more (time-dependent) control signals may be generated by an LFO. Thereby, in some embodiments, the characteristic frequency may be changed periodically.

[0013] In some embodiments, the second stage may be adapted to apply an amplifier to the audio signal. It should be understood that the second stage may include an amplifier. The gain of the amplifier may be controlled by one or more control signals (e.g., based on one or more control signals). Thus, the gain may be time-dependent. For example, one or more control signals may be generated by an LFO. Thereby, in some embodiments, the gain may be changed periodically.

[0014] In some embodiments, the second stage may be adapted to apply an envelope to the audio signal by using the amplifier. Then, the third stage may be adapted to generate the spatial metadata at least in part based on the shape of the envelope. For example, the spatial metadata may indicate the time-dependent position of an audio object, wherein the position changes according to the shape of the envelope (e.g., undergoes a linear translation when the envelope indicates a non-zero gain).

[0015] In some embodiments, obtaining an audio signal may include generating an audio signal by using one or more oscillators. It should be understood that the first stage may include one or more oscillators. For example, such a device may relate to a synthesizer, such as a music synthesizer.

[0016] In some embodiments, an audio signal may be generated by one or more oscillators at least in part based on one or more control signals. For example, at least one of the frequency, pulse width, and phase of one or more oscillators may be controlled by one or more control signals (e.g., based on one or more control signals). Thereby, the control signals that affect the generation of the audio signal may be used as a basis for generating the spatial metadata. This provides an additional function for capturing artistic intent when generating spatial metadata.

[0017] Alternatively, obtaining an audio signal may include receiving an audio signal. For example, the audio signal may be received from an external source such as a sound database. For example, such a device may relate to an audio processor, such as an effects audio processor.

[0018] In some embodiments, one or more control signals may be at least partially based on user input.

[0019] In some embodiments, the output stage may be adapted to output one or more audio streams based on the modified audio signal together with the generated spatial metadata. For example, each output audio stream may have a spatial metadata stream. Further, when the apparatus processes more than one audio signal in parallel, each (modified) audio signal may have an output audio stream. Alternatively, at least one output audio stream may be generated by mixing two or more (modified) audio signals.

[0020] According to another aspect of the present disclosure, a method of generating or processing an audio signal is provided. The method may include obtaining an audio signal. The method may further include modifying the audio signal based on one or more control signals for shaping the sound represented by the audio signal. The method may further include generating spatial metadata associated with the (modified) audio signal at least partially based on the one or more control signals. The method may further include outputting the modified audio signal together with the generated spatial metadata.

[0021] In some embodiments, the one or more control signals may be time-dependent.

[0022] In some embodiments, the spatial metadata may be metadata for instructing an external device how to render the modified audio signal.

[0023] In some embodiments, modifying the audio signal may include applying a time-dependent modification to the audio signal. Wherein, the time-dependency of the modification may depend on one or more control signals.

[0024] In some embodiments, the method may further include modifying the audio signal by at least one of the following: a filter; an amplifier; a low-frequency oscillator; an audio delay; a driver; and a flanger effect.

[0025] In some embodiments, modifying the audio signal may include applying a filter to the audio signal. Wherein, the characteristic frequency of the filter may be controlled by one or more control signals.

[0026] In some embodiments, the characteristic frequency of the filter may be a cut-off frequency.

[0027] In some embodiments, modifying the audio signal may include applying an amplifier to the audio signal. Wherein, the gain of the amplifier may be controlled by one or more control signals.

[0028] In some embodiments, modifying the audio signal may include applying a clipping wave to the audio signal by using an amplifier. Then, generating the spatial metadata may be at least partially based on the shape of the clipping wave.

[0029] In some embodiments, obtaining the audio signal may include generating the audio signal by using one or more oscillators.

[0030] In some embodiments, the audio signal may be generated by one or more oscillators at least partially based on one or more control signals.

[0031] Alternatively, obtaining the audio signal may include receiving the audio signal.

[0032] In some embodiments, one or more control signals may be at least partially based on user input.

[0033] In some embodiments, outputting the modified audio signal may include outputting one or more audio streams based on the modified audio signal together with the generated spatial metadata.

[0034] According to another aspect, a computer program is provided. The computer program may include instructions that, when executed by a processor (e.g., a computer processor, a server processor, etc.), cause the processor to perform all steps of the methods described throughout this disclosure.

[0035] According to another aspect, a computer-readable storage medium is provided. The computer-readable storage medium may store the aforementioned computer program.

[0036] According to yet another aspect, a device is provided that includes a processor and a memory coupled to the processor. The processor may be adapted to perform all steps of the methods described throughout this disclosure. For example, the device may relate to a computer system, a server (e.g., a cloud-based server), or a system of servers (e.g., a system of cloud-based servers).

[0037] It will be understood that device features and method steps may be interchanged in various ways. In particular, as will be understood by those skilled in the art, the details of the disclosed methods may be implemented by the corresponding device, and vice versa. Furthermore, any statement above regarding the method (and, for example, its steps) should be understood to equally apply to the corresponding device (and, for example, its blocks, stages, units, etc.), and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which

[0039] Figure 1 is a block diagram schematically illustrating an example of a synthesizer.

[0040] Figure 2 is a block diagram schematically illustrating an audio processing chain including a synthesizer, a spatialization module, and an object-based rendering module,

[0041] Figure 3 is a block diagram schematically illustrating an example of a synthesizer according to an embodiment of the present disclosure,

[0042] Figure 4 is a flowchart schematically illustrating an example of a method for generating and / or processing an audio signal according to an embodiment of the present disclosure, and

[0043] Figure 5 is a block diagram of an apparatus for performing a method according to an embodiment of the present disclosure. Detailed Description

[0044] The drawings and the following description are illustrative only and relate to preferred embodiments. It should be noted that, based on the following discussion, alternative embodiments of the structures and methods disclosed herein will readily be recognized as viable alternatives that may be employed without departing from the principles claimed herein.

[0045] Reference will now be made in detail to several embodiments, examples of which are illustrated in the drawings. It should be noted that, where feasible, like or similar reference numerals may be used in the drawings and the reference numerals may indicate like or similar functionality. The drawings depict embodiments of the disclosed apparatus (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

[0046] The present disclosure generally relates to a music synthesizer (and audio processor) having a spatial metadata output. The generation of the metadata may be driven by control signals related to direct / literal user input and / or internal modulation signals. Thus, the control signals are generally time-dependent.

[0047] Broadly speaking, a synthesizer may be considered a collection of connected sound generation / shaping elements and modulation elements. The sound generation / shaping elements create or act directly on an audio signal. For example, an oscillator may generate a raw tone / waveform, then a filter may shape its timbre, and an amplifier may introduce level dynamics. To make the generated sound lively, these elements should vary over time, which is accomplished by applying modulation elements that do not themselves emit sound but affect the behavior of the sound generation / shaping elements.

[0048] Figure 1A simple example of a synthesizer 100 is shown. The synthesizer 100 includes a sound generation block (or stage, module, etc.) 110 and a sound modulation block (or stage, module, etc.) 120. For example, the sound generation block 110 further includes one or more oscillators (oscillator elements) 112, one or more filters (e.g., voltage controlled filters (VCFs)) 114, and one or more amplifiers (e.g., voltage controlled amplifiers (VCAs)) 116. The sound modulation block 120 can modulate the actual generation of the audio signal by the oscillator 112 and / or modulate any subsequent shaping of the sound signal by, for example, the filter 114 or the amplifier 116. Here and hereinafter, it should be understood that an audio signal can be an electronic representation of a sound signal. It should further be understood that the shaping of the sound corresponds to the modification / change of the audio signal representing the sound.

[0049] For example, the sound modulation block 120 may include a low frequency oscillator (LFO) 124 capable of producing a cyclic sub-audio (e.g., below 20 Hz) frequency output. The LFO 124 may be used to modulate the cutoff frequency of a filter (e.g., filter 114), or any other characteristic frequency of the filter. Another common element is an envelope 126, which (e.g., via an amplifier 116) generates a one-time modulation pattern for the sound signal generated by the sound generation block 110. These modulation elements are typically paired with a control surface 122 such as a keyboard or sequencer, which is operated by a musician, for example. Thus, modulation may depend on user input and may typically be time-dependent.

[0050] These aforementioned elements of the sound generation block 110 and / or the sound modulation block 120 may be implemented using a DSP, analog circuits, digital circuits, or any combination thereof.

[0051] The output stage of the synthesizer 100 ( Figure 1 The synthesizer 100 may include a mixing circuit that combines internal audio signals (and optionally, effects) that are output as one or more audio signals. When outputting multiple audio signals, the synthesizer 100 may be configured to output a set of signals having some notion of spatiality, such as a stereo signal or even a multi-channel signal. Further, for example, an internal modulation signal such as a signal generated by an LFO may be used to influence spatial qualities, such as sound image position. In this case, the output is effectively rendered. The internal signals mixed together in the output stage may be the outputs of the oscillator element 112, which are mixed to form a speech signal or (for example, in a multi-speech synthesizer) multiple speech signals, which may be based on oscillators, sample playback, or any other suitable tone generation technique. These signals may also be output from the synthesizer 100 individually, rather than being mixed together.

[0052] Thus, the synthesizer 100 is based on the rendered channel output, such as mono or stereo, or even multi-channel output (e.g., 3.1.2 multi-channel output, 5.1 multi-channel output, 5.1.2 multi-channel output, 7.1.4 multi-channel output, etc.).

[0053] There are various available methods for generating a stereo output in a synthesizer. One method for generating stereo from mono is to use effects such as chorus, flanging, delay, or reverb. The synthesizer architecture can be mainly mono, where the final effect processor has a mono input and then applies different processing to create left and right signals to create a stereo output. In some cases, the final effect processor can involve a chorus effect output stage.

[0054] Another method is to allow the user to pan voices in the stereo field. This panning can be completely manual, or can be related to certain characteristics of the sound, such as pitch, or can be modulated over time using, for example, an LFO.

[0055] Furthermore, a multi-timbral synthesizer is capable of playing back more than one program simultaneously at a time. This allows one program to be assigned to a Lower Patch and one program to be assigned to an Upper Patch. Then, these different programs or patches can be positioned in the stereo field.

[0056] For example, the channel rendering output of the synthesizer 100 can be used for an object-based workflow, such as Dolby panner, by obtaining the rendered output and then using an auxiliary tool to generate spatial metadata.

[0057] Thus, when the synthesizer is used for object-based (e.g., Dolby ) production, the output must first be rendered and then spatialized in a second step. This means that any sound design decisions have been finalized, and the internal modulation signals that were used to create the output are now unavailable. This limits the available creative options and makes it difficult to consider spatialization as an integrated part of the sound design process.

[0058] In Figure 2An example of a processing chain 200 for such an object-based sound design process is illustrated in the block diagram. The processing chain 200 includes a synthesizer 205, a spatialization module 230, and an object-based rendering module (e.g., Dolby Atmos production suite) 240. The synthesizer 205 may correspond to the synthesizer 100 described above and may include a sound generation block 210 and a sound modulation block 220. The output of the synthesizer 205 is rendered into a specific (e.g., predefined) channel configuration, such as mono or stereo, or even a multi-channel configuration (e.g., 3.1.2 multi-channel output, 5.1 multi-channel output, 5.1.2 multi-channel output, 7.1.4 multi-channel output, etc.). This rendered channel-based output is then fed into the spatialization module 230. The spatialization module 230 processes the channel-based output of the synthesizer 205 to create object-based audio content therefrom. Then, the object-based audio content generated by the spatialization module 230 can be used by the object-based rendering module 240 for object-based rendering to achieve a desired (in principle, arbitrary) channel configuration (e.g., depending on the expected speaker layout).

[0059] An example of a potential problem that may occur when performing object-based rendering using the processing chain 200 is described next. According to this example, an LFO may be used to rhythmically change the filter cutoff frequency of a filter in the sound generation block 210. It may also be desirable to use the LFO to affect the spatial height of the signal so that the filter cutoff frequency is synchronized with the height. However, doing so when using the processing chain 200 can be laborious. That is, the rhythmic change in the cutoff frequency can be intended to represent or conform to the movement of a sound source (e.g., vertical movement). When generating spatial metadata for expressing this movement based on the rendered channel-based output of the synthesizer 205, the internal modulation signal that has controlled the adaptation of the cutoff frequency is no longer available. Instead, the expected spatial movement can only be indirectly inferred from the channels of the channel-based output, which can be laborious and / or inaccurate. For more complex examples (such as the case of using one or more LFOs to individually modulate the height position of synthesizer voices), generating appropriate spatial metadata will be even more laborious and even impossible because these voices will be mixed together.

[0060] The techniques according to the present disclosure relate to incorporating object spatialization within a synthesizer (or audio processor) architecture. This change allows the internal control signals (internal modulation signals) of the synthesizer that are used to shape the sound to also be used to affect spatialization. Then, the synthesizer will be able to directly generate spatial metadata and audio together without the need for an intermediate rendering step. Thus, the techniques according to the present disclosure directly output a set of one or more audio streams and associated spatial metadata. This set of signals can then be rendered or further processed and then rendered for playback.

[0061] This approach means that the internal control signals or GUI controls used to locate object outputs can also be used to adjust other aspects of sound generation (and vice versa). Further, individual oscillators can be used as object outputs, and individual synthesizer voices can also be used as object outputs.

[0062] Although this disclosure may frequently refer to synthesizers (e.g., music synthesizers), it should be understood that the techniques presented can equally apply to audio processors (e.g., audio effect processors) unless otherwise indicated. Thus, this disclosure generally can relate to apparatuses for generating or processing audio signals (sound signals).

[0063] In Figure 3 an example of an audio processing chain 300 including a synthesizer 305 conforming to the techniques of this disclosure is schematically illustrated. The processing chain 300 includes a synthesizer 305 and an object-based rendering module 340 (separate from the synthesizer), which is used to perform object-based rendering on the output of the synthesizer.

[0064] Broadly speaking, this disclosure proposes incorporating a spatial output element (e.g., a third stage described below), which generates a corresponding spatial metadata stream (or general spatial metadata) for each signal to be output, where the spatial metadata indicates to an external rendering unit how to render the signal. In particular, the metadata stream can describe the spatial characteristics of its associated audio signal over time, such as Cartesian position (e.g., in 3D space) and / or size. Thus, the spatial metadata can be metadata for object-based rendering of audio objects, where the audio track of an audio object is given by an audio signal. Dolby metadata is an example of such spatial metadata. Generally, this spatial metadata can be generated by direct literal user input. But at the same time, it can also be derived from modulation signals (control signals), which can also be used to generate and / or shape the associated audio signal. Additionally, this spatial metadata can be generated by a combination of direct literal user input and derivation from modulation signals (control signals), which can also be used to generate and / or shape the associated audio signal.

[0065] The synthesizer 305 includes a sound generation block 310, a sound modulation block 320, and a spatialization block 330. The synthesizer can also include an output stage ( Figure 3 not shown in

[0066] The sound generation block 310 may include any of the elements described above for the sound generation block 110 of the synthesizer 100 (such as oscillators, filters, and / or amplifiers), as well as other elements for shaping the sound that has already been generated by the oscillators. Thus, the sound generation of the sound generation block 310 may be performed in the same manner as the above-described sound generation block 110. This does not exclude the sound generation block 310 from including additional elements and / or having additional functions not described above. Further examples of possible elements of the sound generation block 310 will be given below.

[0067] Generally speaking, conceptually, the sound generation block 310 may be regarded as implementing a first stage for obtaining an audio signal and a second stage for modifying the audio signal (e.g., for shaping the sound represented by the audio signal).

[0068] In accordance with the above, in Figure 3 the first stage of the synthesizer 305 may include one or more oscillators (i.e., the oscillators of the sound generation block 310). The first stage may use these one or more oscillators to generate an audio signal. For example, each of the one or more oscillators may be an analog oscillator, an FM oscillator, or a wavetable oscillator. Depending on its type / implementation, these oscillators may have different operating parameters (e.g., oscillator parameters). For example, an analog oscillator may have frequency, pulse width, and / or gain / level as operating parameters. An FM oscillator may have frequency, ratio, FM depth, and / or gain / level as operating parameters. Further, a wavetable oscillator may have frequency, wave index, group index, and / or gain / level as operating parameters.

[0069] Similarly, the generation of the audio signal by one or more oscillators may be at least partially based on one or more control signals. For example, the operating parameters of the one or more oscillators (e.g., frequency, pulse width, and / or phase, and / or any of the foregoing operating parameters) may be modulated under the control of one or more control signals.

[0070] Although the above-described implementation of the first stage relates to a synthesizer (e.g., a music synthesizer), the present disclosure also relates to implementations in which the first stage receives an audio signal ( Figure 3 not shown in the figure). For example, the audio signal may be received from an external source (e.g., an audio database, a sound database). Such an implementation may relate to an audio processor, such as an effects audio processor.

[0071] Generally, a synthesizer or an audio processor (i.e., a device generally used for generating or processing audio signals) can process more than one audio signal (i.e., multiple audio signals) in parallel. For such an embodiment, a combination of (internally) generating and receiving audio signals can also be feasible, such as an embodiment where some of the audio signals are generated by an oscillator and some (other) audio signals are received from an external source.

[0072] Consistent with the above, although the synthesizer is frequently mentioned without intended limitation, embodiments of the present disclosure equally relate to an audio processor. The difference between these different embodiments lies in whether the audio signal modified and supplemented with spatial metadata subsequently is (internally) generated or received. It should be understood that any other elements / functions of the described synthesizer are equally applicable to the audio processor. That is, except for the way of obtaining the audio signal (i.e., generating or receiving), the audio processor can have the same functions as the synthesizer described throughout the present disclosure.

[0073] As described above, the second stage of the synthesizer 305 is a stage for modifying the audio signal (e.g., for shaping the sound represented by the audio signal). Modifying the audio signal is based on one or more (internal) control signals (e.g., internal modulation signals). Therefore, the control signal can also be referred to as a control signal for shaping the sound represented by the audio signal. As described above, the control signal can be time-dependent (i.e., can vary with time). Therefore, the second stage can be adapted to apply time-dependent modifications to the audio signal. The time-dependence of the modification can depend on one or more control signals.

[0074] The second stage can include filters (e.g., VCF) and / or amplifiers (e.g., VCA). Generally, the second stage can include any, some, or all of the following elements to modify the audio signal: one or more filters (e.g., VCF), one or more amplifiers (e.g., VCA), one or more LFOs, one or more audio delayers, one or more drivers, and one or more flanging effectors. For example, the second stage can also include sound shaping elements such as chorus and reverb. All of these elements can have corresponding operating parameters (e.g., synthesizer parameters or audio processor parameters), and the parameters can be modified / changed / modulated according to one or more control signals.

[0075] For example, an amplifier may have gain as an operating parameter. A filter may have a cut-off frequency and / or resonance (resonant frequency) as an operating parameter. An LFO may have rate and / or scale as an operating parameter. A delay unit (delay effect) may have time, mix, and / or feedback as an operating parameter. A driver (drive effect) may have drive and / or mix as an operating parameter. Further, a flanger (flanging effect) may have center, width, rate, regeneration, and / or mix as an operating parameter.

[0076] For each of such operating parameters that are modified, there may be a corresponding control signal that controls the modification / modulation of the operating parameter. The control signal may be generated by a corresponding modulation source. Examples of modulation sources may include an LFO and / or an envelope.

[0077] In a first non-limiting example, the second stage may include or implement a filter (e.g., a VCF), which may be applied to the audio signal output from the first stage. In this case, the characteristic frequency of the filter may be controlled by one or more control signals. In this sense, the characteristic frequency of the filter may be time-dependent. If the corresponding control signal is periodic (e.g., generated by an LFO), the characteristic frequency may change periodically. Consistent with the above, the characteristic frequency of the filter may be the cut-off frequency. Alternatively, the characteristic frequency may be the resonant frequency.

[0078] In a second non-limiting example, the second stage may include or implement an amplifier (e.g., a VCA), which may be applied to the audio signal output from the first stage. In this case, the gain of the amplifier may be controlled by one or more control signals. In this sense, the gain of the amplifier may be time-dependent. If the corresponding control signal is periodic (e.g., generated by an LFO), the gain may change periodically.

[0079] Specifically, in the second example, the second stage may use an amplifier to apply an envelope (e.g., a gain curve) to the audio signal. In this case, the corresponding control signal of the amplifier may represent the envelope.

[0080] As described above, the control signal may be generated by a modulator such as an LFO. In some embodiments, at least some of the modulators may in turn be modulated by other modulators under the control of the corresponding control signals. Additionally, there may be built-in effects that are modulated.

[0081] The (internal) control signals of the synthesizer 305 may be generated by the sound modulation block 320, which may include the same elements as the modulation block 120 of the synthesizer 100 described above. Thus, the sound modulation performed by the sound modulation 320 may be performed in the same manner as the sound modulation block 120 described above. This does not exclude the sound modulation block 320 from including additional elements and / or having additional functions not described above.

[0082] Generally, the control signals may be generated by one or more modulators that affect how the audio signal is modified. For example, one or more modulators may involve or include an LFO. As described above, the operating parameters of one or more modulators themselves may also be subject to time-dependent modulation under the control of appropriate control signals. As exemplified by the control surface (e.g., keyboard, control panel) 122 shown by Figure 1 one or more control signals may also be at least partially based on user input.

[0083] The spatialization block 330 may be regarded as involving or implementing a third stage of the synthesizer 305, which is used to generate spatial metadata related to the modified audio signal. As described above, this is done at least partially based on one or more control signals. Here and hereinafter, it should be understood that the spatial metadata may be metadata for indicating to an external device (e.g., a renderer or a rendering module) how to render the modified audio signal.

[0084] It should be understood that any control signal (e.g., the control signals described above) may be used as a basis for generating the spatial metadata. For example, these control signals may be used as a basis for determining at least one of the position of the corresponding audio object in the horizontal plane, the height of the corresponding audio object, and the size of the audio object. If the position of the audio object in a given plane (e.g., the horizontal plane) is determined based on the control signal, the parameters of the linear translation may be determined (e.g., calculated) according to the control signal (such as a control signal representing time-dependent gain or gain curve). Similarly, the parameters of the periodic motion (e.g., circular or elliptical motion) of the audio object in a given plane (e.g., the horizontal plane) may be determined (e.g., calculated) based on a periodic control signal (such as a control signal that controls the characteristic frequency of a filter).

[0085] In the first example above, an audio signal of an audio object that will be perceived as rotating around a central position may be generated by periodically changing the characteristic frequency (e.g., cut-off, resonance) of the filter applied to the audio signal. Then, the polar angle of the audio object (which may be used to derive appropriate Cartesian coordinates) may be determined (e.g., calculated) based on the control signal used to modulate the characteristic frequency. Thus, the spatial metadata is generated at least partially based on one or more control signals (specifically, the control signal that modulates the characteristic frequency of the filter).

[0086] In the second example above, an envelope can be applied to an audio signal using an amplifier to generate an audio signal that will be perceived as an audio object moving through, for example, a room. Then, a time-dependent position (which can be used to derive appropriate Cartesian coordinates), such as corresponding to a linear translation of the audio object, can be determined (e.g., calculated) based on a control signal used to modulate the gain of the amplifier. Specifically, a linear translation (or time-dependent position, or generally spatial metadata) can be generated based on the shape of the envelope (gain curve) represented by the control signal. Thus, again, spatial metadata is generated at least in part based on one or more control signals (specifically, the control signal that modulates the gain of the amplifier).

[0087] In the third example, multiple audio signals can be generated by multiple oscillators. Then, the relative spatial positions of the associated audio objects of the oscillators (i.e., relative to each other) can be modulated by an LFO, which is also used to modulate the frequencies of the oscillators. For example, when a filter applied to an audio signal is cyclically turned on and off under the control of the LFO, the objects will cyclically move closer and then farther apart from each other. This relative movement can be appropriately reflected in the metadata generated for the multiple modified audio signals. Again, spatial metadata is generated at least in part based on one or more control signals (specifically, the control signal generated by the LFO).

[0088] In the fourth example, multiple audio signals can be generated by multiple oscillators. Then, the relative spatial positions of the associated objects of the oscillators can be controlled by an envelope that also controls the cutoff of a filter applied to the audio signals. For example, the envelope can be triggered by an attached keyboard or other suitable device for receiving user input. For example, when a certain key is first pressed, the sounds of the oscillators appear to originate from the same spatial position, but when the key is held down continuously, the sound sources appear to move farther apart in space. As the distance between the sound sources increases, the filter may open more, causing the sound to become brighter. This relative movement can be equally appropriately reflected in the metadata generated for the multiple modified audio signals. Again, spatial metadata is generated at least in part based on one or more control signals (specifically, the control signal generated by the envelope).

[0089] In the fifth example, a random LFO shape can be used to control both the voice position and the wavetable of an oscillator that generates an audio signal (i.e., voice). Over time, the voice object will take on random spatial positions, with each new position corresponding to a new wavetable. The random movement of the voice object can be appropriately reflected in the metadata generated for the audio signal. Again, spatial metadata is generated at least in part based on one or more control signals (specifically, the control signal generated by the random LFO shape).

[0090] The output stage (the fourth stage) of the synthesizer is an output stage for outputting the modified audio signal together with the generated spatial metadata. Specifically, the output stage can be adapted to output one or more audio streams based on the modified audio signal together with the corresponding spatial metadata generated by the third stage for each audio stream. For example, each output audio stream can have a spatial metadata stream. When the synthesizer 305 processes more than one audio signal in parallel, each (modified) audio signal can have an output audio stream. Alternatively, the modified audio signals can be mixed into an audio stream for output. In this case, appropriate spatial metadata can be generated for the resulting audio stream. Alternatively, the spatial metadata for the output audio stream can be generated based on the spatial metadata of the individual modified audio signals mixed into the output audio stream.

[0091] Alternatively or additionally, the synthesizer 305 can process more than one audio signal serially (i.e., multiple audio signals). For example, in some such embodiments, the processing of a given audio signal can depend on an earlier processed audio signal and / or its metadata. As another example, in some embodiments, the processing can involve modifying the already existing spatial metadata of an audio signal that has already been obtained by the first stage. Consistent with the above, such modification of the already existing metadata can be based on an internal control signal (internal modulation signal).

[0092] Although the synthesizer 305 is described as outputting object-based audio content (i.e., the modified audio signal and the generated spatial metadata), it can additionally be configured to output rendered (e.g., channel-based) audio content. For example, the synthesizer 305 can be capable of rendering binaural output for preview, or for direct multi-channel input to a public address system. Thus, in addition to the output stage, the synthesizer can also include a rendering stage (or rendering unit, rendering module) for generating the rendered audio content. However, it should be understood that the rendering stage is optional, and the synthesizer can generally provide object-based audio content for external rendering.

[0093] In Figure 4 a flowchart of which schematically illustrates a corresponding method 400 for generating or processing an audio signal. The method 400 includes steps / processes S410 to S440.

[0094] In Step S410 , an audio signal is obtained. As described above, this can involve generating or receiving an audio signal.

[0095] In Step S420 , the audio signal is modified based on one or more control signals for shaping the sound represented by the audio signal.

[0096] at Step S430 ,spatial metadata associated with the modified audio signal is generated, at least in part, based on one or more control signals.

[0097] Finally, at Step S440 ,the modified audio signal is output together with the generated spatial metadata.

[0098] It should be understood that step S410 may be performed by a first stage of the above-described apparatus for generating or processing an audio signal, step S420 may be performed by a second stage, step S430 may be performed by a third stage, and step S440 may be performed by an output stage. Further, it should be understood that any statements made above regarding the respective stages equally apply to their corresponding method steps / processes, and the repetitive descriptions may be omitted for the sake of brevity.

[0099] The present disclosure similarly relates to an apparatus (e.g., a synthesizer, an audio processor) for performing the methods and techniques described throughout the present disclosure. Figure 5 An example of such an apparatus 500 is shown. The apparatus 500 includes a processor 510 and a memory 520 coupled to the processor 510. The memory 520 may store instructions for the processor 510. The processor 510 may optionally receive an input 530 from an external source such as a database. For example, the input 530 may relate to a sound signal. Further, the processor 510 may receive a user input 560, for example, via a suitable interface (e.g., a keyboard, a control panel, etc.). The user input may modify sound generation or sound shaping, as described above. The processor 510 may be adapted to perform the methods / techniques described in the present disclosure. Thus, the processor 510 may output one or more (modified) sound signals (audio streams) 540 and associated metadata (metadata streams) 550.

[0100] The present disclosure also relates to a computer program including instructions that, when executed by a computer processor, will cause the computer processor to perform the methods described throughout the present disclosure; and a computer-readable storage medium storing the computer program.

[0101] Embodiments of the present disclosure may have in common that the (same) internal control signals (internal modulation signals) for controlling the generation and / or modification / shaping of an audio signal are also used to generate spatial metadata of the audio signal. Compared with the case where control signals are not available for generating spatial metadata and where any previous shaping operations applied to the audio signal are finalized as channel-based outputs (e.g., mono, stereo, or possibly multi-channel), this allows for additional flexibility in implementing creative intentions and may help avoid laborious manual and possibly suboptimal editing of spatial metadata.

[0102] Explanation

[0103] Aspects of the systems described herein may be implemented in a suitable computer-based sound processing system for generating and / or processing sound signals. One or more of the components, blocks, processes, or other functional elements may be implemented by one or more computer programs executed by one or more processor-based computing devices of the control system. It should also be noted that various combinations of hardware, firmware, and / or data and / or instructions embodied in various machine-readable or computer-readable media may be used to describe the various functions disclosed herein in terms of behavior, register transfer, logic components, and / or other characteristics. Computer-readable media that may embody such formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0104] In particular, it should be understood that embodiments may include hardware, software, and electronic components or modules, which for discussion purposes may be illustrated and described as if most components were implemented only in hardware. However, those of ordinary skill in the art, and based on a reading of this detailed description, will recognize that in at least one embodiment, the electronic aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or an application-specific integrated circuit (“ASIC”). Accordingly, it should be noted that embodiments may be implemented using a variety of hardware- and software-based devices and a variety of different structural components. For example, the blocks or stages described herein may include one or more electronic processors, one or more computer-readable media modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.

[0105] Although one or more implementations have been described by way of example and in terms of specific embodiments, it should be understood that one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements that would be apparent to those skilled in the art. Accordingly, the scope of the appended claims should be given the broadest interpretation so as to cover all such modifications and similar arrangements.

[0106] Likewise, it should be understood that the words and terms used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and their variants is intended to cover the items listed thereafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms “mounted,” “connected,” “supported,” and “coupled” and their variants are used broadly and cover both direct and indirect mounting, connection, support, and coupling.

[0107] Enumerated exemplary embodiments

[0108] Aspects and embodiments of the present disclosure can also be understood from the following enumerated example embodiments (EEEs), which are not claims.

[0109] EEE1. A device for generating or processing an audio signal, the device comprising: a first stage for obtaining an audio signal; a second stage for modifying the audio signal based on one or more control signals for shaping the sound represented by the audio signal; a third stage for generating spatial metadata related to the modified audio signal based at least in part on the one or more control signals; and an output stage for outputting the modified audio signal together with the generated spatial metadata.

[0110] EEE2. The device according to EEE1, wherein the one or more control signals are time-dependent.

[0111] EEE3. The device according to EEE1 or EEE2, wherein the spatial metadata is metadata for indicating to an external device how to render the modified audio signal.

[0112] EEE4. The device according to any one of EEE1 to EEE3, wherein the second stage is adapted to apply time-dependent modification to the audio signal, and wherein the time-dependency of the modification depends on the one or more control signals.

[0113] EEE5. The device according to any one of EEE1 to EEE4, wherein the second stage includes at least one of the following to modify the audio signal: a filter; an amplifier; a low-frequency oscillator; an audio delay; a driver; and / or a flanger.

[0114] EEE6. The device according to any one of EEE1 to EEE5, wherein the second stage is adapted to apply a filter to the audio signal; and wherein the characteristic frequency of the filter is controlled by the one or more control signals.

[0115] EEE7. The device according to EEE6, wherein the characteristic frequency of the filter is a cut-off frequency.

[0116] EEE8. The device according to any one of EEE1 to EEE7, wherein the second stage is adapted to apply an amplifier to the audio signal; and wherein the gain of the amplifier is controlled by the one or more control signals.

[0117] EEE9. The apparatus according to EEE8, wherein the second stage is adapted to apply a capping wave to the audio signal by using the amplifier; and wherein the third stage is adapted to generate the spatial metadata at least in part based on the shape of the capping wave.

[0118] EEE10. The apparatus according to any one of EEE1 to EEE9, wherein obtaining the audio signal includes generating the audio signal by using one or more oscillators.

[0119] EEE11. The apparatus according to EEE10, wherein the audio signal is generated by the one or more oscillators at least in part based on the one or more control signals.

[0120] EEE12. The apparatus according to any one of EEE1 to EEE9, wherein obtaining the audio signal includes receiving the audio signal.

[0121] EEE13. The apparatus according to any one of EEE1 to EEE12, wherein the one or more control signals are at least in part based on a user input.

[0122] EEE14. The apparatus according to any one of EEE1 to EEE13, wherein the output stage is adapted to output one or more audio streams based on the modified audio signal together with the generated spatial metadata.

[0123] EEE15. A method for generating or processing an audio signal, the method comprising: obtaining an audio signal; modifying the audio signal based on one or more control signals for shaping the sound represented by the audio signal; generating spatial metadata associated with the modified audio signal at least in part based on the one or more control signals; and outputting the modified audio signal together with the generated spatial metadata.

[0124] EEE16. The method according to EEE15, wherein the one or more control signals are time-dependent.

[0125] EEE17. The method according to EEE15 or EEE16, wherein the spatial metadata is metadata for indicating to an external device how to render the modified audio signal.

[0126] EEE18. The method according to any one of EEE15 to EEE17, wherein modifying the audio signal includes applying a time-dependent modification to the audio signal, wherein the time-dependence of the modification depends on the one or more control signals.

[0127] EEE19. The method according to any one of EEE15 to EEE18, including modifying the audio signal by at least one of the following: a filter; an amplifier; a low-frequency oscillator; an audio delay; a driver; and / or a flanger effect.

[0128] EEE20. The method according to any one of EEE15 to EEE19, wherein modifying the audio signal includes applying a filter to the audio signal; and wherein a characteristic frequency of the filter is controlled by the one or more control signals.

[0129] EEE21. The method according to EEE20, wherein the characteristic frequency of the filter is a cut-off frequency.

[0130] EEE22. The method according to any one of EEE15 to EEE21, wherein modifying the audio signal includes applying an amplifier to the audio signal; and wherein a gain of the amplifier is controlled by the one or more control signals.

[0131] EEE23. The method according to EEE22, wherein modifying the audio signal includes applying a clip to the audio signal by using the amplifier; and wherein generating the spatial metadata is at least partially based on a shape of the clip.

[0132] EEE24. The method according to any one of EEE15 to EEE23, wherein obtaining the audio signal includes generating the audio signal by using one or more oscillators.

[0133] EEE25. The method according to EEE24, wherein the audio signal is generated by the one or more oscillators at least partially based on the one or more control signals.

[0134] EEE26. The method according to any one of EEE15 to EEE23, wherein obtaining the audio signal includes receiving the audio signal.

[0135] EEE27. The method according to any one of EEE15 to EEE26, wherein the one or more control signals are at least partially based on a user input.

[0136] EEE28. The method according to any one of EEE15 to EEE27, wherein outputting the modified audio signal includes outputting one or more audio streams based on the modified audio signal together with the generated spatial metadata.

[0137] EEE29. A computer program comprising instructions which, when executed by a computer processor, cause the computer processor to perform the method according to any one of EEE15 to EEE28.

[0138] EEE30. A computer-readable storage medium storing the computer program according to EEE29.

Claims

1. An apparatus for generating or processing an audio signal, the apparatus comprising: A first stage for obtaining an audio signal; A second stage for modifying the audio signal based on one or more control signals, the control signals being configured to shape the sound represented by the audio signal; A third stage for generating spatial metadata related to the modified audio signal based at least in part on the one or more control signals; And An output stage for outputting the modified audio signal together with the generated spatial metadata.

2. The apparatus according to claim 1, wherein, The one or more control signals are time-dependent.

3. The apparatus according to claim 1 or 2, wherein, The spatial metadata is metadata for indicating to an external device how to render the modified audio signal.

4. The apparatus according to any one of claims 1 to 2, wherein, The second stage is adapted to apply a time-dependent modification to the audio signal, wherein the time-dependence of the modification depends on the one or more control signals.

5. The apparatus according to any one of claims 1 to 2, wherein, The second stage includes at least one of the following to modify the audio signal: A filter; An amplifier; A low-frequency oscillator; An audio delay; A driver; and / or A flanger effect.

6. The apparatus according to any one of claims 1 to 2, wherein, The second stage is adapted to apply a filter to the audio signal; and wherein a characteristic frequency of the filter is controlled by the one or more control signals.

7. The apparatus according to claim 6, wherein, The characteristic frequency of the filter is a cut-off frequency.

8. The apparatus according to any one of claims 1 to 2, wherein, The second stage is adapted to apply an amplifier to the audio signal; and wherein a gain of the amplifier is controlled by the one or more control signals.

9. The apparatus according to claim 8, wherein, The second stage is adapted to apply a clipping to the audio signal by using the amplifier; and wherein the third stage is adapted to generate the spatial metadata based at least in part on a shape of the clipping.

10. The apparatus according to any one of claims 1 to 2, wherein, Obtaining the audio signal includes generating the audio signal by using one or more oscillators.

11. The apparatus according to claim 10, wherein, The audio signal is generated by the one or more oscillators based at least in part on the one or more control signals.

12. The apparatus according to any one of claims 1 to 2, wherein, Obtaining the audio signal includes receiving the audio signal.

13. The apparatus according to any one of claims 1 to 2, wherein, The one or more control signals are at least in part based on user input.

14. The apparatus according to any one of claims 1 to 2, wherein,The output stage is adapted to output one or more audio streams based on the modified audio signal together with the generated spatial metadata.

15. A method for generating or processing an audio signal, the method comprising: Obtain an audio signal; Modify the audio signal based on one or more control signals, the control signals being configured to shape the sound represented by the audio signal; Generate spatial metadata related to the modified audio signal based at least in part on the one or more control signals; And Output the modified audio signal together with the generated spatial metadata.

16. The method according to claim 15, wherein, The one or more control signals are time-dependent.

17. The method according to claim 15 or claim 16, wherein, The spatial metadata is metadata for indicating to an external device how to render the modified audio signal.

18. The method according to any one of claims 15 to 16, wherein, Modifying the audio signal includes applying a time-dependent modification to the audio signal, wherein the time-dependence of the modification depends on the one or more control signals.

19. The method according to any one of claims 15 to 16, comprising modifying the audio signal by at least one of the following: a filter; an amplifier; a low-frequency oscillator; an audio delay; a driver; and / or a flanger.

20. The method according to any one of claims 15 to 16, wherein, Modifying the audio signal includes applying a filter to the audio signal; and wherein a characteristic frequency of the filter is controlled by the one or more control signals.

21. The method according to claim 20, wherein, The characteristic frequency of the filter is a cut-off frequency.

22. The method according to any one of claims 15 to 16, wherein, Modifying the audio signal includes applying an amplifier to the audio signal; and wherein the gain of the amplifier is controlled by the one or more control signals.

23. The method according to claim 22, wherein, Modifying the audio signal includes applying a clipping wave to the audio signal by using the amplifier; and wherein the spatial metadata is generated at least in part based on the shape of the clipping wave.

24. The method according to any one of claims 15 to 16, wherein, Obtaining the audio signal includes generating the audio signal by using one or more oscillators.

25. The method according to claim 24, wherein, The audio signal is generated by the one or more oscillators at least in part based on the one or more control signals.

26. The method according to any one of claims 15 to 16, wherein, Obtaining the audio signal includes receiving the audio signal.

27. The method according to any one of claims 15 to 16, wherein, The one or more control signals are at least in part based on user input.

28. The method according to any one of claims 15 to 16, wherein, Outputting the modified audio signal includes outputting one or more audio streams based on the modified audio signal together with the generated spatial metadata.

29. A computer program comprising instructions that, when executed by a computer processor, will cause the computer processor to perform the method according to any one of claims 15 to 28.

30. A computer-readable storage medium storing the computer program according to claim 29.

Citation Information

Patent Citations

  • Apparatus for changing an audio scene and an apparatus for generating a directional function

    CN103109549A

  • Floor controller for real-time control of music signal processing, mixing, video and lighting

    US20020005111A1