Generation of audio stereo signals
The apparatus uses multiple audio beamformers to adaptively capture and position audio sources, addressing the limitations of conventional stereo microphones by enhancing spatial separation and representation, thereby improving teleconferencing performance and user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2024-04-02
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional stereo microphone systems struggle with limited spatial information, complexity, and resource-intensity, making it difficult to separate and distinguish audio sources, especially in far-field scenarios, and fail to provide an optimal user experience in teleconferencing applications.
An apparatus using multiple audio beamformers to generate a stereo audio signal by adapting beamformers to capture audio sources, determining their direction, and generating intermediate stereo signal components based on the audio source's position, allowing for improved spatial separation and representation of audio sources.
Enhances the distinction and separation of audio sources, providing a clearer spatial user experience by accurately positioning audio sources within a stereo representation, reducing sensitivity to environmental characteristics, and improving teleconferencing performance.
Smart Images

Figure 2026513759000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the generation of audio stereo signals, particularly (but not limited to) the generation of audio stereo signals for teleconference applications.
Background Art
[0002] Audio capture, particularly voice capture, has become increasingly important in the past few decades. In fact, voice capture has become increasingly important for various applications including telecommunications, teleconferencing, gaming, audio user interfaces, and the like. However, the problem in many scenarios and applications is usually that the desired voice source is not the only audio source in the environment. Rather, in a typical audio environment, there are many other audio / noise sources being captured by the microphone. One of the significant problems faced by many voice capture applications is how to optimally extract audio and voice in a noisy environment. To address this problem, various approaches for noise suppression have been proposed.
[0003] In recent years, various forms of remote voice communication applications using various communication media (particularly including the Internet) have become increasingly popular, and further expansion and improvement of the provided services, options, and user experiences are desired.
[0004] So far, remote voice communication using landline phones, mobile phones, and Internet communication has been substantially limited to the transmission of monaural voice content. However, the use of monaural signals will limit the user experience. For example, in a teleconference application based on a monaural signal, the representation of the environment where the voice is captured may be limited. In particular, the use of a monaural signal limits the ability of spatial hearing to extract individual sources such as the target speaker from a complex sound scene.
[0005] The introduction of stereo communication has been proposed for various remote audio applications, particularly internet-based remote conferencing applications. The introduction of stereo signals enables communication in spatial dimensions, potentially improving the user experience in many applications.
[0006] In general, monaural audio capture of a room tends to be considered limiting in many scenarios and not providing a sufficient user experience. Many applications, for example, would prefer to include a spatial component to facilitate the distinction between different speakers.
[0007] On the recording side, a stereo microphone can be used to capture the spatial characteristics of an environment suitable for stereo distribution. Such a stereo microphone typically consists of two microphone elements with different directivity, each capturing one channel of the stereo signal. There are various recording techniques, but basically there are two beams, one pointing to the left and the other to the right. The two outputs are the left and right channels of the stereo pair. For example, in X / Y recording, two identical cardioid microphones are placed close together, facing away from each other at a 90-degree angle. The two microphones have different radiation characteristics, and their outputs are the L and R channels of the stereo pair.
[0008] In another example, M(id) / S(ide) recording uses a Mid microphone with omnidirectional or cardioid characteristics pointing to the center of the scene, while a Side microphone with so-called figure-eight or dipole characteristics is mounted perpendicular to the Mid microphone. The sensitivity of a perfect dipole is the same for sources 180 degrees apart, but there is a 180-degree phase shift between the output signals. Using this effect, two new beams can be created by adding or subtracting a (possibly) scaled version of the dipole output to the Mid microphone output. The beams can be in any direction, including X / Y characteristics. Compared to the X / Y method, the M / S method is more flexible, allowing the stereo representation to be widened or narrowed by changing the coefficients of the mixing matrix. [Overview of the project] [Problems that the invention aims to solve]
[0009] However, while this approach may provide desirable performance in many scenarios and applications, it is not optimal in all scenarios. For example, in many scenarios, spatial information may be relatively limited, and spatially separating each sound source can be difficult, especially with far-field sound sources in stereo microphones. Also, this approach tends to be relatively complex and resource-intensive in many scenarios.
[0010] Therefore, an improved approach is advantageous, particularly one that enables reduced complexity, increased flexibility, faster implementation, lower costs, improved audio capture, improved spatial awareness / differentiation of sound sources, improved audio source isolation, improved support for remote voice applications, improved support for teleconferencing, and improved user experience and / or performance.
[0011] Therefore, the present invention preferably seeks to mitigate, reduce, or eliminate one or more of the above-mentioned drawbacks, either individually or in any combination. [Means for solving the problem]
[0012] According to one aspect of the present invention, an apparatus for generating a stereo audio signal is provided, the apparatus comprising: a plurality of audio beamformers, each audio beamformer configured to generate a monaural audio signal representing audio captured by a beam formed by the audio beamformer; an adapter configured to adapt each of the plurality of audio beamformers to capture an audio source; a direction determiner configured to determine the direction of audio captured by the beam formed by the plurality of audio beamformers; a generator configured to generate an intermediate stereo signal component for each of at least some of the monaural audio signals, wherein the first intermediate stereo signal component for a first monaural audio signal among at least some of the monaural audio signals includes a first monaural audio signal positioned at a first position in the stereo representation of the first intermediate stereo signal component, the first position depending on a first direction which is the direction of audio captured by a first beam of the first audio beamformer, which is an audio beamformer among the plurality of audio beamformers generating the first monaural audio signal; a coupler configured to generate a stereo audio signal including at least some of the intermediate stereo signal components; and a transmitter for transmitting the stereo audio signal to a remote device.
[0013] This approach can provide improvements in operation and / or performance in many embodiments. It enables an improved stereo signal with a stereo representation that can provide improved user perception and experience in many embodiments and scenarios. This approach increases the degree of freedom in generating appropriate spatial placement of audio sources within the stereo representation. This approach can typically reduce the sensitivity of the provided stereo representation to certain characteristics of the captured audio scene and certain capture options. In particular, it can increase the flexibility to provide a stereo representation that reflects the characteristics of the audio scene but is not entirely constrained by them in many scenarios and applications.
[0014] In many embodiments, this approach enables significantly improved distinction between audio sources, and in many scenarios, a listener listening to a stereo representation of a stereo signal can separate and identify individual audio sources.
[0015] In many embodiments, a more clearly defined representation of each audio source, such as a point source, can be provided.
[0016] In many embodiments, this approach facilitates or improves the capture of the spatial characteristics of an audio source, which can then be represented (including being emphasized or modified) in the stereo representation of the generated audio signal.
[0017] For example, the inventors recognized that conventional stereo information capture using stereo microphones (or especially two microphones with different directivity) tends to be insufficient for many applications. In particular, separating audio sources in the far field is often difficult, and conventional stereo signals representing audio captured by stereo microphones often fail to adequately separate / distinguish audio sources, resulting in a lack of spatial definition.
[0018] The described approach can, in many scenarios, provide stereo signals in which different audio sources are more clearly separated and distinguishable, and / or where the audio sources are spatially clearly defined.
[0019] This approach enables the generation of stereo signals that can provide an improved spatial user experience for those listening to stereo signals.
[0020] A beamformer is configured to receive multiple microphone signals from multiple microphones and can perform beamforming on the microphone signals. Multiple microphones can form a microphone array (e.g., multiple microphones arranged in a line). Beamforming can sometimes be performed by weighted addition / combination of the microphone signals. In some cases, each filter can be filtered by an adaptive filter before addition / combination. The adapter can be configured to perform beamforming by adapting coefficients / weights for weighted addition or filtering.
[0021] The adapter can be configured to adapt each of the audio beamformers of multiple audio beamformers to capture the audio source as if it were a typically different audio source.
[0022] According to an optional feature of the present invention, this device is a teleconferencing device.
[0023] This approach can, in many embodiments, provide an improved teleconferencing system. In many embodiments, the teleconferencing system can provide an improved stereo signal, an improved representation of speakers, and generally enable clearer and easier distinction between different speakers.
[0024] The monaural audio signal can be, in particular, an audio signal. The audio source can be, in particular, a speaker / participant in a telephone conference.
[0025] According to an optional function of the present invention, the generator is configured to place the first monaural audio signal in the stereo representation by scaled addition of the first monaural audio signal to the first and second channels of the first intermediate stereo signal component, and the relative scale factor of the first channel with respect to the second channel depends on a first direction.
[0026] This improves performance and / or facilitates operation and / or implementation in many embodiments. The generator can be configured to place the monaural audio signal within the stereo representation of the corresponding intermediate stereo signal component by performing amplitude panning.
[0027] The first channel can be a left channel or a right channel, and the second channel can be a complementary channel.
[0028] According to an optional function of the present invention, the generator is configured to determine the scale factor of the first channel as a function of a first direction and to determine the scale factor of the second channel such that the sum of the first scale factor and the second scale factor meets a criterion.
[0029] This improves performance and / or facilitates operation and / or implementation in many embodiments. The criterion is, for example, that the sum has a given / fixed / predefined value that does not depend on position.
[0030] According to an optional function of the present invention, the first position is a monotonic function of the first direction.
[0031] This improves performance in many scenarios and, in many cases, improves the user experience.
[0032] According to an optional feature of the present invention, the first position is a linear function of the first direction.
[0033] This improves performance in many scenarios and, in many cases, enhances the user experience.
[0034] According to an optional feature of the present invention, the first position is a nonlinear function in the first direction.
[0035] This improves performance in many scenarios and, in many cases, enhances the user experience.
[0036] According to an optional feature of the present invention, the first position is a function of the first direction, and this function provides a mapping of the intervals of the direction angles in the first direction to larger intervals of the angles of the positions in the stereo representation.
[0037] This improves performance in many scenarios and, in many cases, enhances the user experience. A larger angular interval between positions within a stereo representation can be the entire range of possible angular intervals (corresponding to the entire stereo representation).
[0038] According to an optional feature of the present invention, the first position is a function of the first direction, and this function provides a mapping of the interval of the direction angle to the smaller interval of the angles of the positions in the stereo representation.
[0039] This improves performance in many scenarios and, in many cases, enhances the user experience.
[0040] According to an optional feature of the present invention, the coupler is configured to adapt the amplitude of the first intermediate stereo signal component in the coupled stereo signal to the amplitude of the second intermediate stereo signal component representing the second monaural audio signal in the coupled stereo signal, depending on the first direction.
[0041] This improves performance in many scenarios and, in many cases, enhances the user experience.
[0042] According to an optional feature of the present invention, the coupler is configured such that the coupled stereo signal does not contain a first intermediate stereo signal component when the first direction satisfies a criterion.
[0043] This improves performance in many scenarios and, in many cases, enhances the user experience. The criterion is that the direction is within the subrange / sub-interval of the stereo representation.
[0044] According to an optional feature of the present invention, the first direction indicates the direction of arrival of the audio captured by the audio beamformer that generates the first monaural audio signal.
[0045] This improves performance in many scenarios and, in many cases, enhances the user experience.
[0046] According to one aspect of the present invention, an audio capture method for generating a stereo audio signal is provided, the method comprising the steps of: generating a monaural audio signal in which each of a plurality of audio beamformers generates an audio signal representing audio captured by the beam formed by the audio beamformers; adapting each of the plurality of audio beamformers to capture an audio source; determining the direction of the audio captured by the beam formed by the plurality of audio beamformers; generating an intermediate stereo signal component for at least some of the monaural audio signals, wherein the first intermediate stereo signal component for a first monaural audio signal among the at least some of the monaural audio signals includes a first monaural audio signal positioned at a first position in the stereo representation of the first intermediate stereo signal component, the first position depending on a first direction which is the direction of the audio captured by the first beam of the first audio beamformer, which is an audio beamformer among the plurality of audio beamformers generating the first monaural audio signal; generating a stereo audio signal to include at least some of the intermediate stereo signal components; and transmitting the stereo audio signal to a remote device.
[0047] These and other aspects, features and advantages of the present invention will become apparent from and be described with reference to the embodiments described below. [Brief explanation of the drawing]
[0048] Embodiments of the present invention will be described with reference to the drawings, merely as examples. [Figure 1] A figure showing examples of elements of a device for generating a stereo audio signal according to several embodiments of the present invention. [Figure 2] A diagram illustrating an example of an audio beamformer. [Figure 3]A diagram showing some elements of a possible configuration of a processor for implementing elements of an audio device according to some embodiments of the present invention. [Modes for carrying out the invention]
[0049] The following description focuses on embodiments of the present invention applicable to beamforming-based voice capture and communication audio systems, particularly teleconferencing equipment, but it will be understood that this approach is applicable to many other systems and scenarios for audio capture.
[0050] Figure 1 shows an example of an audio device configured to capture audio sources and generate a stereo signal representing one or more of these audio sources, where the stereo signal provides a stereo representation of the audio sources, which may have spatial characteristics. This device could specifically be a teleconferencing device, and the following description will focus on such an example, particularly focusing on the fact that the audio source is a speaker and the captured audio signal is a voice signal. However, while this approach can offer particularly favorable effects and advantages for such scenarios and applications, it will be understood that this approach is by no means limited to such applications and can be applied to the capture of other audio sources and for purposes other than teleconferencing. For example, this approach can be used to capture other types of audio sources and for other purposes and applications, such as capturing instruments in an orchestra or band and generating a stereo signal containing spatial stereo representations of different instruments.
[0051] In this example, the device includes a set of beamformers 101 coupled to a microphone array of M microphones 103. The example shows multiple beamformers 101 by a single function unit, reflecting that multiple beamformers can be implemented by a single processing / function unit that performs the operation of all beamformers (e.g., in parallel or sequential operation). Thus, a beamformer can be considered equivalent to a single beamformer that forms multiple beams (or a smaller set of beamformers, at least one of which forms multiple beams).
[0052] Each beamformer 101 is configured to form a beam and generate a monaural audio signal representing the audio captured by that beam.
[0053] The number of beamformers 101 and thus the number of monaural audio signals N can vary in different embodiments. In many embodiments, the number of beamformers N may be five or more, or even ten beamformers. In many embodiments, the number of beamformers N can exceed five, ten, or twenty beamformers.
[0054] The beamformer 101 is coupled to an adapter 105 configured to adapt the beamformer 101, specifically the adapter 105 is configured to adapt the beam being formed, which includes directing the beam towards a detected sound source and dynamically adapting the beam to follow (track) it as the sound source moves.
[0055] In many embodiments, the adapter 105 may also include functions for searching for / detecting new audio sources to which a new beam can be formed, determining when to stop tracking an audio source, and switching the beam to a new audio source. For example, the adapter 105 can control the beamformer to search for audio sources and track any detected audio source as long as the average signal level (averaged over sufficiently long time intervals to compensate for, for example, audio pauses) exceeds a threshold, otherwise stop tracking the audio source and search for a new audio source.
[0056] The beamformer 101 is therefore an adaptive beamformer whose directivity can be controlled by an adapter 105 that adapts the parameters of beamform operation.
[0057] Each beamformer 101 can specifically be a filtering-combination (and, in most embodiments, a filtering-addition) beamformer. The beamform filters can be applied to each microphone signal, and the filtered outputs can be combined. In some embodiments, the combination may include, for example, functions applied to each signal (e.g., frequency compensation to compensate for the different frequency sensitivities of each microphone, varying delay, nonlinear gain compensation, etc.). However, typically, the filter outputs are simply added together, and the filtering-combination beamformer is implemented particularly as a filtering-addition beamformer.
[0058] Figure 2 shows a simplified example of a filtering-adding beamformer based on a microphone array containing only two microphones 201. In this example, each microphone is coupled to beamform filters 203, 205, and their outputs are added in an adder 207 to produce a beamformed audio output signal. The beamform filters 203, 205 have impulse responses f1 and f2 adapted to form beams in a given direction. Typically, a microphone array contains more than two microphones, and it will be understood that the principle in Figure 2 can be easily extended to more microphones by further including beamform filters for each microphone.
[0059] Each beamformer 101 may include such filtering-adding architectures for beamforming (for example, as described in US 7 146 012 and US 7 602 926). In many embodiments, it will be understood that the microphone array 201 may include more than two microphones.
[0060] In most embodiments, each beamform filter has a time-domain impulse response that extends over time intervals typically of 2 milliseconds, 5 milliseconds, 10 milliseconds, and even 30 milliseconds or more, rather than a simple Dirac pulse (corresponding to a simple delay and therefore to gain and phase offset in the frequency domain).
[0061] The impulse response can often be implemented by a beamforming filter, which is a finite impulse response (FIR) filter with multiple coefficients. In such embodiments, adapter 105 can adapt beamforming by adapting the filter coefficients. In many embodiments, the FIR filter can have coefficients corresponding to a fixed time offset (typically a sample time offset), and adaptation is achieved by adapting these coefficient values. In other embodiments, the beamforming filter can typically have significantly fewer coefficients (e.g., only two or three), but these timings are (also) adaptable.
[0062] A particular advantage of beamform filters having an extended impulse response rather than simple variable delay (or simple frequency-domain gain / phase adjustment) is that the beamformer does not have to be adapted only to the strongest (usually direct) signal components. Rather, the beamformer can be adapted to include additional signal paths that typically correspond to reflections. Therefore, this approach enables improved performance in most real-world environments, and in particular, improved performance in reflective and / or reverberant environments and / or audio sources far from the microphone array 103.
[0063] A key element of adaptive beamformer performance is the adaptation of directivity (generally referred to as the beam, but with an extended impulse response, it is understood that this directivity has not only a spatial component but also a temporal component, i.e., the beam is formed as a temporal change such as reflection).
[0064] In the system shown in Figure 1, the adapter 105 is configured to adapt the beamform parameters of the beamformer, specifically, to adapt the coefficients of the beamform filter to provide a given (spatial and temporal) beam.
[0065] It will be understood that different adaptive algorithms may be used in different embodiments, and that various optimization parameters are known to those skilled in the art. For example, adapter 105 can adapt beamform parameters to maximize the output signal value of beamformer 101. As a concrete example, consider a beamformer in which a received microphone signal is filtered by a forward matching filter, and the filtered outputs are added together. The output signal is filtered by an inverse adaptive filter having a conjugate filter response to the forward filter (in the frequency domain corresponding to the time-inverse impulse response in the time domain). An error signal is generated as the difference between the input and output signals of the inverse adaptive filter, and the coefficients of the filter are adjusted to minimize the error signal, thereby obtaining maximum output power. This can more essentially generate a noise reference signal from the error signal. Details of such approaches can be found in US7146012 and US7602926.
[0066] It should be noted that approaches such as US7146012 and US7602926 are based on adaptations based on both the audio source signal z(n) and the noise reference signal x(n) from the beamformer, and it is understood that the same approach can be used for the beamformer in Figure 2.
[0067] In fact, the beamformer 101 can be a beamformer configuration corresponding to those disclosed in US7146012 and US7602926.
[0068] The adapter 105 can be configured to adapt beamforming to capture a desired audio source and represent it with a beamformed audio output signal. Furthermore, a noise reference signal can be generated to provide an estimate of the remaining captured audio, i.e., the noise that would be captured if the desired audio source were not present.
[0069] In an example where the beamformer 101 is a beamformer such as those disclosed in US7146012 and US7602926, the noise reference can be generated as an error signal. However, it will be understood that other approaches can be used in other embodiments. For example, in some embodiments, the noise reference can be generated as the beamformed audio output signal subtracted from the microphone signal from a (e.g., omnidirectional) microphone, or as the microphone signal itself if this noise reference microphone is far from other microphones and does not contain the desired sound. In another example, a set of beamformers 101 is configured to generate a second beam with a null in the direction of the maximum value of the beam that generates the beamformed audio output signal, and the noise reference can be generated as the audio captured by this complementary beam.
[0070] In some embodiments, post-processing, such as noise suppression, can be applied to the output of an audio capture device. This can improve the performance of voice communications and other applications. Such post-processing may involve nonlinear calculations, but for some speech recognition devices, for example, it may be more advantageous to restrict the processing to include only linear operations.
[0071] Therefore, the adapter 105 is configured to control the beamformer 101 to search for an audio source, and once an audio source is detected, the adapter 105 can dynamically adapt the beam to follow / track the audio source. In some embodiments, the adapter 105 can further adapt the beam shape, for example, by narrowing the beam after initial detection.
[0072] In a specific example, this device is intended to capture speech signals from speakers in an environment. In this example, adapter 105 can be configured to adaptively consider whether the audio source is actually a speech signal. For example, adapter 105 can extract characteristics of the captured audio (e.g., frequency distribution, activity pattern, etc.) for a given detected audio source and determine whether these characteristics match those expected of a speech signal. If they match, adapter 105 can continue tracking the audio source; otherwise, it can reject the audio source and proceed to search for another audio source representing the speaker.
[0073] Therefore, the beamformer 101 forms a beam toward the detected audio source, in particular, an audio source that may be a different speaker in the environment. Each beamformer provides a single monaural audio signal representing the audio captured by the beam formed by that beamformer. Thus, multiple monaural signals are generated by a set / bank / multiple beamformers, and typically, each monaural signal represents a different audio source / speaker in the environment.
[0074] The audio signal is supplied to a generator 107 configured to generate an intermediate stereo signal component from a monaural audio signal. In particular, the generator 107 can generate an intermediate stereo signal component for each of the monaural audio signals received from the beamformer 101.
[0075] Given a monaural audio signal, the generator 107 is configured to generate an intermediate stereo signal component such that the monaural audio signal is positioned in a predetermined location within the stereo representation of the intermediate stereo signal component. The intermediate stereo signal component includes two channels, namely a first channel and a second channel, which can correspond to the left channel and the right channel, respectively (or vice versa). Thus, rather than a monaural audio signal having a single value at any given time (specifically, one sample value per sampling time), the intermediate stereo signal component has two values, one for the left channel and one for the right channel. Consequently, the stereo representation of the two components creates a stereo representation ranging from the left position corresponding to the rendering position of the left channel to the right position corresponding to the rendering position of the right channel (and vice versa).
[0076] The device further includes a direction determiner 109 configured to determine the direction of audio captured by beams formed by multiple audio beamformers. Thus, for each beamformer / beam, the direction determiner 109 can determine the direction of audio captured by that beam. Specifically, each beamformer / beam can track one audio source, such as a speaker, and the direction determiner 109 can be configured to determine the direction to that audio source (often by considering the direction of the formed beam as the direction to the audio source).
[0077] In many embodiments, the direction of captured audio can be determined as the direction of the beam being formed. Therefore, the direction determiner 109 is configured to determine the direction of the beam formed by the beamformer, and this direction can be used to indicate the direction to the audio source captured by the beam. This provides a very accurate estimate of the direction to the audio source captured by the audio beam, as the adapter 105 can detect the audio source and then control the beam to track it. Typically, the beam is directed in the direction in which the audio reaches the microphone array 103, and therefore the beam direction reflects the (primary) direction of arrival (typically the direction of arrival on the direct path rather than reflected) of the audio from the audio source captured by the beam. The determined direction typically reflects the direction in which the audio source would be perceived as arriving by a person located at the position of the microphone array 103.
[0078] Many different approaches are known for determining the direction of the formed audio beam, and it will be understood that the direction determiner 109 can use any suitable approach. In some embodiments, the direction determiner 109 can evaluate a coefficient applied to each microphone signal or tap of a beamform filter to determine the beam direction of the beam, for example.
[0079] For example, if a beamformer forms a beam from a weighted combination of signals from multiple substantially omnidirectional microphones (usually using complex weights to represent both amplitude and phase / delay), the direction of the beam can be easily determined from the applied coefficients, as is known to those skilled in the art.
[0080] As another example, if the beamformer is based on only two microphone inputs, and the filter coefficients F1 and F2 are implemented in the frequency domains belonging to the two microphones 1 and 2 respectively, then the inverse Fourier transform of the product F1·F2* is taken, where * represents the conjugate, and from this cross-correlation the peak can be determined, from which the difference in propagation time td can be calculated. The difference in propagation time is related to the incident angle φi by the following equation. cos(φi) = td·c / d, Here, c is the speed of sound and d is the distance between the two microphones. A suitable exemplary approach is described, for example, in US6774934 or “Fast Cross-correlation for TDOA Estimation on Small Aperture Microphone Arrays” by F. Grondin et al, arXiv:2204.13622.
[0081] In some embodiments, the beam may not directly follow / track the audio source, for example, changing more slowly than the moving audio source or remaining relatively constant. Often, such discrepancies are acceptable, and the beam direction can be used directly as an estimate of the direction of arrival of the captured audio. However, in some embodiments, the direction of arrival may be determined to be not directly equal to the direction of the beam being formed. For example, in some embodiments, the direction determiner 109 may not directly consider the coefficients used by the beamformer, but may use a different algorithm to detect the direction of arrival of the audio signal captured by the microphone array 103, such an algorithm may directly consider and be based on the microphone signal, for example, independently of beamforming. The direction of arrival may be determined, for example, as the strongest audio signal detected within an angular interval corresponding to the beam angle spacing of a particular beamformer / beam. This direction of arrival can be used as the direction of the corresponding monaural signal produced by that beamformer.
[0082] The direction determiner 109 is coupled to a generator 107, which is supplied with direction estimates for each beam. The generator 107 then receives a monaural audio signal from the beamformer 101, as well as direction estimates for each beam / beamformer.
[0083] The generator 107 is configured to generate an intermediate stereo signal component from a monaural audio signal, and the monaural audio signal is positioned within a stereo representation that depends on a direction estimate.
[0084] Specifically, with respect to a given monaural audio signal from a given beamformer 101, the generator 107 can generate an intermediate stereo signal component such that the monaural audio signal has a specific position in a stereo representation, and this position is determined from the average direction of the generation of the monaural audio signal.
[0085] In many examples, an intermediate stereo signal component of a given monaural audio signal is generated to have left and right channel components containing the monaural audio signal. Typically, the phases of the monaural audio signals in the left and right channels are the same, and the position within the stereo representation is controlled by adapting the relative levels / amplitudes / signal strengths of the two channels. If the level of the left channel is non-zero and the level of the right channel is zero or near zero, the position of the monaural audio signal is perceived as being at the rendering position of the left channel within the stereo representation. If the level of the right channel is non-zero and the level of the left channel is zero or near zero, the position of the monaural audio signal is perceived as being at the rendering position of the right channel within the stereo representation. When the levels of the two channels are the same, the perceived position is in the center of the stereo representation; for other relative amplitude levels, the perceived position will be somewhere between the left and right rendering positions. Therefore, in many embodiments, the generator 107 can be configured to perform amplitude panning to position a given monaural audio signal at a desired position within the stereo representation of the intermediate stereo signal component representing that monaural audio signal.
[0086] The specific position selected for a given monaural audio signal depends on the direction estimation of the corresponding beam / beamformer. For example, the stereo representation can be expressed in angles ranging from, for example, -π / 2 to +π / 2, and the direction estimation can also be given in angles ranging from, for example, -π / 2 to +π / 2. In some embodiments, the stereo representation angle / direction can be determined as a function of the beam direction angle. In fact, in many examples, the direction of the stereo representation can be determined to be equal to the beam direction angle.
[0087] The generator 107 can be configured to generate intermediate stereo signal components such that the mono audio signal is positioned at the corresponding angle within the stereo representation by determining the relative amplitude of the mono audio signal in the left and right channels corresponding to the determined angle.
[0088] Therefore, the generator 107 typically generates multiple intermediate stereo signal components, each of which typically corresponds to one audio source, and in a particular teleconferencing application, each intermediate stereo signal component typically represents a specific speaker. Furthermore, the position in the stereo representation of each intermediate stereo signal component is determined according to the direction of the audio represented by the intermediate stereo signal component, which is typically indicated by the direction of the beam. Since speakers are usually in different positions, the direction estimations are usually different, and therefore the positions in the stereo representation are also usually different in the generated intermediate stereo signal components.
[0089] The generator 107 is coupled to a coupler 111 to which the generated intermediate stereo signal components are supplied. The coupler 111 is configured to generate an output stereo audio signal that includes at least a portion of the intermediate stereo signal components for at least some time. In many embodiments, the output stereo signal is generated to include all of the intermediate stereo signal components received from the generator 107.
[0090] In many embodiments, the coupler 111 can, for example, simply add all left channel signals / values from all intermediate stereo signal components and simply add all right channel signals / values. Thus, in many embodiments, the coupler 111 can perform channel addition of at least some of the intermediate stereo signal components. In some embodiments, the coupler 111 may not only simply add different intermediate stereo signal components but may also include, for example, adapting the relative levels of different intermediate stereo signal components, selecting a subset of intermediate stereo signal components, filtering one or more intermediate stereo signal components, and so on.
[0091] Therefore, the coupler generates an output stereo signal such that the intermediate stereo signal component retains its position within the stereo representation; that is, the position of a given monaural audio signal in the stereo representation of the output stereo signal is the same as the position of the monaural audio signal in the corresponding intermediate stereo signal component.
[0092] In particular, in some embodiments, the addition can be a weighted addition in which each intermediate stereo signal component is weighted differently (where the weights for each channel are the same for the same intermediate stereo signal component).
[0093] In a particular example, the generator 107 is coupled to a communicator 113 configured to transmit a stereo signal to a remote device. The communicator 113 can transmit the communication signal, for example, over a network, such as the Internet in particular, a wireless communication link, a dedicated link, or any link that enables the communication of audio data.
[0094] Furthermore, the communication device 113 can be configured to encode and transmit the output stereo signal in any suitable way, such as using different audio coding approaches, voice signal coding approaches, and appropriate modulation and coding techniques, as are well known to those skilled in the art.
[0095] Therefore, in a specific example of a teleconferencing application, the device can generate an output stereo signal that is sent to one or more remote devices participating in the teleconferencing. The output stereo signal contains audio corresponding to various speakers, with each speaker positioned within a stereo representation. The receiving teleconferencing device can simply render the received stereo signal to provide the user with spatial awareness and representation of the speakers. For example, the output stereo signal is simply appropriately decoded and presented using a pair of headphones / earphones.
[0096] More generally, a stereo signal can be provided that represents different audio sources, such as different instruments, at different positions in the stereo representation of the stereo signal.
[0097] Therefore, instead of pursuing the conventional approach of capturing a stereo signal for a stereo representation by capturing two channels in the environment using, for example, a stereo microphone or other stereo capture configuration, this approach can use multiple audio beams to employ a dedicated approach to select and isolate individual audio sources, such as speakers in particular, and subsequently construct a stereo representation in which the audio sources can be placed in specific locations.
[0098] Each beam formed provides a single monaural audio signal, and each audio source selected and isolated by the formed beam can be represented by a monaural audio signal that does not possess its own spatial characteristics or features. However, a stereo signal is generated by combining the individual monaural audio signals such that each signal possesses spatial characteristics in the form of a position within the stereo representation. This "reintroduction" of spatial characteristics depends on the spatial information of the captured signal, and by extension, the audio source, with the position depending on the direction of arrival (inferred particularly by the beam direction).
[0099] Therefore, this approach synergistically combines the use of beams to isolate and distinguish audio sources within a scene (and essentially extract the audio components) with the "artificial" generation of stereo representations, so that these features work together in such a way that the output spatial representation depends on the input spatial data.
[0100] This approach can offer significant advantages in many scenarios compared to conventional approaches that simply capture stereo audio representing the spatial information of a scene from the capture location. In particular, the stereo representation of the output stereo signal can be adjusted and optimized to suit specific desired performance. For example, by placing different audio sources, such as speakers, further apart, it becomes easier for listeners to distinguish them, similar to face-to-face dialogue. Also, because audio sources can be spatially clearly defined by inserting them into the stereo representation of the output stereo signal based on a single mono signal, they can be represented as point sources. This allows audio sources / speakers to be represented as point sources, even if they might be perceived as more spatially spread out when captured by, for example, a stereo microphone (due to capturing these reflections arriving at the stereo microphone from different directions). Therefore, even if the device attempts to maintain the intrinsic spatial relationship between audio sources and speakers by, for example, setting the stereo representation angle equal to the beam direction angle, it is possible to achieve a more clearly defined stereo representation with audio sources / speakers that are easier to distinguish and separate.
[0101] This approach does not simply capture the spatial characteristics of an acoustic scene, but uses the interaction between beamforming and artificial positioning to generate an output stereo representation that can be optimized for desired performance without being limited by the real-world audio characteristics of the captured environment. However, this approach allows the spatial characteristics of the environment to be reflected or preserved in the output stereo signal.
[0102] As described above, the generator 107 can be configured to position a mono audio signal in the stereo representation of the corresponding intermediate stereo signal component by performing a scaled addition of the first mono audio signal to the two channels of the intermediate stereo signal component, where the relative scale factor between the two channels (i.e., the right and left channels) depends on the position in the desired stereo representation for the mono audio signal / audio source.
[0103] In many embodiments, the generator 107 can be configured to determine separate (scalar) scale coefficients for the left and right channels of the intermediate stereo signal component. The two scale coefficients can be determined, for example, as a function of angular position indicating the arrival direction / incident angle φ, which can be specifically indicated by the beam direction / angle, or as a suitable function of the desired position / arrival direction.
[0104] As a specific example, the left (L) and right (R) channels of the intermediate stereo signal components are determined using amplitude panning as follows: Li = f(φi) * Si Ri = (1- f(φi))*Si Here, f(φi) is a function of the beam direction / incident angle φ of the beamformer i, and this function is restricted to between 0 and 1. Si is the monaural audio signal of the beam / beamformer with index i.
[0105] Assuming beamforming is based on a linear array, and defining φi as an angle in the range / interval of -π / 2 to +π / 2, then by setting f(-π / 2) to zero and f(π / 2) to 1, for example, the full range / interval of stereo representation between the left and right loudspeakers can be used during playback.
[0106] In some embodiments, the generator 107 can be configured to determine one scale factor / gain / amplitude of a channel as a function of the beam direction of a beamformer that generates a first monaural audio signal, and then determine the scale factor / gain / amplitude of a second channel such that the sum of the two scale factors / gain / amplitudes satisfies a criterion. This is equivalent to the two scale factors being determined by two functions (one for each scale factor), but the functions are designed such that the sum of the scale factors satisfies the criterion.
[0107] The criterion is that the sum of the two scale factors must equal a specific value. This value is usually a predetermined value, such as a fixed value. As a concrete example, the sum of the two functions above for the left and right channel scale factors is always 1, but in fact, the function for the left channel is given simply as 1 minus the scale factor for the right channel.
[0108] The advantage of this particular approach is that the audio source can be placed within the stereo representation with a high degree of freedom, and is not limited to a specific, precise position determined by the stereo recording.
[0109] The stereo representation of the corresponding intermediate stereo signal component and the position of a particular monaural audio signal / audio source in the output stereo signal can be determined as a function of the direction of the determined audio source (direction of arrival / beam direction). In some embodiments, the generator 107 can first determine the desired position, for example, the angular position in the stereo representation of the intermediate stereo signal component can be determined as a function of the beam direction, and then the level / scale coefficients of the left and right channels can be determined. In other embodiments, the level / scale coefficients corresponding to the desired position can be directly determined as a function of direction φ, for example, as reflected by the function f(φi) in the above example.
[0110] The relationship between the direction of the audio captured by a given beam (hereinafter referred to as the beam direction φ for brevity) and the position of the monaural audio signal in the stereo representation (hereinafter referred to as the stereo representation position) can depend on the preferences and requirements of each individual embodiment and application.
[0111] In many embodiments, the position of the stereo representation is a monotonic function of the beam direction. For example, if the beam direction φ increases monotonically from -π / 2 to +π / 2, the stereo representation positions can also be arranged to increase monotonically from -π / 2 to +π / 2. Thus, the function / relationship mapping the beam direction to the stereo representation positions is a monotonic function. When amplitude panning is used, the level / amplitude / scale coefficients are also monotonic functions of the beam direction φ, specifically, one function increases monotonically as a function of φ and the other decreases monotonically.
[0112] This allows for the provision of a desirable user experience in many situations. For example, in a teleconferencing application, detected speakers are placed within the stereo representation of the output stereo signal, where they are sufficiently separated but arranged in the same order as the captured environment. For example, speakers can be placed in fixed positions within the stereo representation, starting with the position of the left (or right) channel, and then placed at intervals of π / (N-1), but in the same order as the captured environment. Remote participants listening to the output stereo signal are provided with a spatial experience where individual speakers are sufficiently separated, but the spatial relationships are the same as those of the captured room (for example, remote participants can determine the seating order of the participants).
[0113] In some embodiments, the stereo representation position can be a linear function of the beam direction. For example, the beam direction angle can be directly mapped to the angle of the stereo representation position. This allows for a closer correspondence with the captured environment, providing remote participants with information about speaker positions, such as whether speakers are positioned close to each other.
[0114] In some cases, scaling can be included in the mapping between beam direction and stereo representation position. For example, the beam direction can be expressed again in angles within the range / interval of -π / 2 to +π / 2, and the stereo image position in angles within the range / interval of -π / 2 to +π / 2, and the function between these can be a linear function with a gradient different from 1.
[0115] For example, in many embodiments, scaling can be configured to increase / expand the stereo representation by applying a linear function with a gradient greater than 1. For instance, if the beam direction in which the desired source most likely exists is known to be in the range of -π / 3 to +π / 3, a function with a gradient of 3 / 2 can be used to map to the entire range of stereo representation from -π / 2 to +π / 2. This can be achieved, for example, by directly mapping to the scale coefficients by making one of the scale coefficients a linear function with f(-π / 3)=0, f(π / 3)=1 and the other scale factor a linear function with f(-π / 3)=1, f(π / 3)=0. It will be understood that other functions may be used to determine the scale coefficients, including nonlinear functions for nonlinear amplitude panning approaches.
[0116] Therefore, in many embodiments, this approach can expand at least a portion of the stereo representation. Thus, the mapping from beam direction to stereo representation position can be such that the range / interval of angles representing the beam direction is mapped to a wider range / angle interval of the stereo representation position (the entire range / interval of angles is the same for both the beam direction and the stereo representation position).
[0117] In some embodiments, a gradient of less than 1 can be used, and the generator 107 can be configured to compress the stereo representation of the output stereo signal compared to the captured range. For example, in many embodiments, scaling can be configured to reduce / narrow the stereo representation by applying a linear function with a gradient of less than 1. For example, if it is known that the beam direction in which the desired source exists (most likely) is in the entire range from -π / 2 to +π / 2, but it is desired to provide a remote participant with a concentrated stereo representation where all speakers are located in the range from -π / 3 to +π / 3, this can be achieved using a function with a gradient of 2 / 3. This can be achieved by directly mapping to the scaling factor.
[0118] Thus, in many embodiments, this approach can narrow at least a portion of the stereo representation. Therefore, the mapping from beam direction to stereo representation position may be such that the range / interval of angles representing the beam direction is mapped to a smaller range / interval of angles of the stereo representation position (the full range of angles is the same for both the beam direction and the stereo representation position).
[0119] In some embodiments, the stereo representation position can be a piecewise linear function of the beam direction. For example, different gradients can be used for different ranges in the beam direction. For instance, the central portion can be expanded / widened and the lateral portions compressed accordingly to ensure coverage of the entire range. Thus, the gradient may be greater than 1 for some intervals and less than 1 for others.
[0120] For example, in some embodiments, the central spacing in the beam direction from -π / 4 to +π / 4 can be extended to a spacing from -π / 3 to +π / 3, while the peripheral spacing from -π / 2 to -π / 4 is mapped to a spacing from -π / 2 to -π / 3, and the peripheral spacing from +π / 4 to +π / 2 is mapped to a spacing from +π / 3 to +π / 2.
[0121] Such an approach can, in many embodiments, provide advantageous performance and an improved user experience. For example, in many teleconferencing applications, multiple speakers are often positioned centrally relative to the teleconferencing device, with one or more speakers positioned more peripherally. Compressing the lateral stereo representation and expanding it in the center often significantly improves the experience.
[0122] In some cases, other nonlinear functions can be used; for example, a nonlinear function can provide a smoother transition between the expanded and contracted portions of the stereo representation. For instance, in many cases, a sigmoid function can be advantageously used to map the entire range of beam directions to the entire range of stereo representation positions.
[0123] In many embodiments, the coupler 111 can be configured to generate an output stereo signal so as to directly sum / include intermediate stereo signal components while maintaining the relative audio levels captured by the microphones.
[0124] However, in some embodiments, the coupler 111 can be configured to adapt the amplitude of one intermediate stereo signal component in the output stereo signal to the amplitude of another intermediate stereo signal component (from the second beam / beamformer). Thus, in some embodiments, the relative amplitudes of each beam / monaural audio signal can be adapted by the device.
[0125] For example, in some embodiments, a relative gain can be determined for each intermediate stereo signal component, and this gain is a function of the beam direction. The gain of a particular beam is determined by determining the gain for the corresponding beam direction, and the determined gain can be applied to both channels of the corresponding intermediate stereo signal component.
[0126] The gain function may be, for example, a function that has higher gain towards the center than towards the sides, or it may have a constant level except for a small range where a particular relevant or important audio source is known (or likely) to exist.
[0127] Therefore, the device can provide an improved user experience by, for example, configuring it to highlight particularly important areas or sources. For example, in a teleconferencing application, this approach allows for the highlighting of specific areas, such as, for example, a speaker facing the center, or the position of a particular speaker (e.g., the chairperson of the meeting being conducted by the teleconferencing).
[0128] In some embodiments, the coupler 111 can be specifically configured to exclude certain monaural audio signal / intermediate stereo signal components if the beam direction of the beam capturing the monaural audio signal meets a predetermined criterion. This criterion is that the beam direction falls within a predetermined exclusion interval; that is, if the beam direction of a particular monaural audio signal / intermediate stereo signal component falls within the exclusion interval, that monaural audio signal / intermediate stereo signal component is not included in the generated output stereo signal.
[0129] In fact, in some embodiments, the criteria can be dynamically adaptable, for example, in response to user input. For example, during a conference call, the chairperson can specify a given range of angles to exclude, thereby allowing the chairperson to exclude, for example, disruptive or undesirable speakers.
[0130] In some embodiments, it will be understood that the generator 107 and the coupler 111 can perform their operations sequentially. For example, the generator 107 can generate N separate stereo signals (and thus N·2 channel signals), these signals can be fed to the coupler 111, and the coupler 111 can mix them to form a single output stereo signal. However, in other embodiments, the two functions may be integrated, and the intermediate stereo signal component may be intrinsic / implicit in the combined operation. For example, in some embodiments, a single matrix multiplication can be applied to an input vector containing all the mono audio signals to generate a stereo output. For example, at each sample time point, an output stereo signal sample is generated by the matrix multiplication. TIFF2026513759000002.tif1168 Here, TIFF2026513759000003.tif1120 represents a two-component vector containing the output stereo sample, TIFF2026513759000004.tif1120 is an N-component vector containing N monaural audio signal samples. TIFF2026513759000005.tif1120 is a 2×N matrix with coefficients determined as a function of the incident direction φ, as mentioned above (the coefficients of matrix A can simply be determined scale coefficients).
[0131] Audio devices can be implemented, in particular, as one or more appropriately programmed processors. Different functional blocks can be implemented in separate processors, and / or, for example, in the same processor. Examples of suitable processors are shown below.
[0132] Figure 3 is a block diagram showing an exemplary processor 300 according to an embodiment of the disclosure. The processor 300 can be used to implement one or more processors that implement the aforementioned device or its elements (in particular, including one or more artificial neural networks). The processor 300 can be any suitable processor type, including but not limited to a microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA) / (the FPGA is programmed to form a processor), graphics processing unit (GPU), application-specific integrated circuit (ASIC) / (the ASIC is designed to form a processor), or a combination thereof.
[0133] The processor 300 may include one or more cores 302. A core 302 may include one or more arithmetic logic units (ALUs) 304. In some embodiments, a core 302 may include a floating-point logic unit (FPLU) 306 and / or a digital signal processing unit (DSPU) 308 in addition to or instead of the ALU 304.
[0134] The processor 300 may include one or more registers 312 that are communicatively coupled to the core 302. The registers 312 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 312 may be implemented using static memory. The registers can provide data, instructions, and addresses to the core 302.
[0135] In some embodiments, the processor 300 may include one or more levels of cache memory 310 that are communicatively coupled to the core 302. The cache memory 310 can provide computer-readable instructions to the core 302 for execution. The cache memory 310 can provide data for processing by the core 302. In some embodiments, computer-readable instructions may be provided to the cache memory 310 by local memory, for example, local memory attached to an external bus 316. The cache memory 310 can be implemented using any suitable cache memory type, such as static random access memory, dynamic random access memory, and / or any other suitable memory technology.
[0136] The processor 300 may include a controller 314 that can control inputs to the processor 300 from other processors and / or components included in the system, and / or outputs from the processor 300 to other processors and / or components included in the system. The controller 314 can control data paths in the ALU 304, FPLU 306, and / or DSPU 308. The controller 314 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 314 can be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0137] The registers 312 and cache memory 310 can communicate with the controller 314 and core 302 via internal connections 320A, 320B, 320C, and 320D. The internal connections can be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.
[0138] Inputs and outputs for the processor 300 may be provided via a bus 316 which may include one or more conductive wires. The bus 316 may be communicatively coupled to one or more components of the processor 300, such as a controller 314, a cache memory 310, and / or registers 312. The bus 316 may be coupled to one or more components of the system.
[0139] Bus 316 can be coupled to one or more external memories. The external memory may include read-only memory (ROM) 332. ROM 332 may be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may have random access memory 333. RAM 333 may be static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory may have electrically erasable programmable read-only memory (EEPROM) 335. The external memory may have flash memory 334. The external memory may have a magnetic storage device such as a disk 336. In some embodiments, the external memory may be included in the system.
[0140] It will be understood that in many embodiments, this approach can more favorably include an adaptive canceller configured to cancel signal components of the beamformed audio output signal that correlate with at least one noise reference signal. For example, as in the example in Figure 1, the adaptive filter takes the noise reference signal as input and its output is subtracted from the beamformed audio output signal. The adaptive filter can be configured, for example, to minimize the resulting signal level during time intervals when no sound is present.
[0141] For clarification, the above description will be understood to have illustrated embodiments of the invention with reference to different functional circuits, units, and processors. However, it will be apparent that any appropriate distribution of functions between different functional circuits, units, or processors can be used without departing from the invention. For example, functions shown to be performed by separate processors or controllers can also be performed by the same processor or controller. Thus, references to specific functional units or circuits should be considered only as references to appropriate means for providing the described functions, and not as indicating a strict logical or physical structure or organization.
[0142] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be implemented physically, functionally, and logically in any suitable manner. Indeed, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention can be implemented in a single unit or physically and functionally distributed among different units, circuits, and processors.
[0143] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term “comprising” does not preclude the existence of other elements or steps.
[0144] Furthermore, although listed individually, multiple means, elements, circuits, or method steps can be implemented, for example, by a single circuit, unit, or processor. Additionally, individual features may be included in different claims, but these can be advantageously combined in some cases, and inclusion in different claims does not mean that the combination of features is unfeasible and / or unfavorable. Also, including a feature in one category of claims does not imply limitation to that category, but rather indicates that the feature is equally applicable to other claim categories as needed. Furthermore, the order of features in a claim does not imply a specific order in which the features must operate, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Furthermore, singular references do not exclude plurals. Therefore, references to "a," "an," "first," "second," etc., do not exclude plurals. Reference numerals in a claim are provided merely as clear examples and should not be construed as limiting the scope of the claim in any way.
[0145] Generally, the examples are illustrated by the following embodiments. Embodiments: Embodiment 1. A device for generating a stereo audio signal, wherein each audio beamformer is configured to generate a monaural audio signal representing the audio captured by the beam formed by the audio beamformer. Multiple audio beamformers (101), An adapter (105) that adapts each of the multiple audio beamformers (101) to capture an audio source, A direction finder (109) configured to determine the direction of audio captured by beams formed by multiple audio beamformers (101), A generator (107) configured to generate an intermediate stereo signal component for each of at least some monaural audio signals, wherein the first intermediate stereo signal component generated for a first monaural audio signal among at least some monaural audio signals includes a first monaural audio signal located at a first position in the stereo representation of the first intermediate stereo signal component, the first position depending on a first direction which is the direction of sound captured by a first beam of a first audio beamformer, which is an audio beamformer among the plurality of audio beamformers (101) that generate the first monaural audio signal; and a combiner (111) configured to generate a stereo audio signal including at least some of the intermediate stereo signal component. Embodiment 2. The apparatus of Embodiment 1, wherein the apparatus is a teleconferencing apparatus further comprising a transmitter (113) for transmitting the stereo audio signal to a remote device. In some embodiments, the device is a teleconferencing system. In some embodiments, the device further includes a transmitter (113) for transmitting a stereo audio signal to a remote device. Embodiment 3. The apparatus of Embodiment 1 or 2, wherein the generator (107) is configured to position the first monaural audio signal within a stereo representation by scaling the first monaural audio signal into the first and second channels of the first intermediate stereo signal component, and the relative scaling factor of the first channel to the second channel is determined according to the first direction. Embodiment 4. The apparatus of Embodiment 3, wherein the generator (107) is configured to determine the scale factor of the first channel as a function of the first direction, and to determine the scale factor of the second channel such that the sum of the first and second scale factors satisfies a criterion. Embodiment 5. An apparatus of any of the previous embodiments in which the first position is a monotonic function of the first direction. Embodiment 6. An apparatus of any of the previous embodiments in which the first position is a linear function of the first direction. Embodiment 7. An apparatus according to any of Embodiments 1 to 5, wherein the first position is a nonlinear function in the first direction. Embodiment 8. Any apparatus of any previous embodiment, wherein the first position is a function of the first direction, and the function provides a mapping of the directional angle intervals of the first direction to larger angular intervals of position in a stereo representation. Embodiment 9. An apparatus of any of the previous embodiments, wherein the first position is a function of the first direction, and the function provides a mapping of the interval of the direction angle to the smaller angular interval of the position in the stereo representation. Embodiment 10. Any apparatus of any previous embodiment, wherein the coupler (111) is configured to adapt the amplitude of a first intermediate stereo signal component in a coupled stereo signal according to a first direction to the amplitude of a second intermediate stereo signal component representing a second monaural audio signal in the coupled stereo signal. Embodiment 11. An apparatus of any previous embodiment, wherein the coupler (111) is configured such that the coupled stereo signal does not contain a first intermediate stereo signal component when the first direction satisfies a criterion. Embodiment 12. An apparatus of any of the previous embodiments, wherein the first direction indicates the direction of arrival of audio captured by the first audio beamformer. Embodiment 13. An audio capture method for generating a stereo audio signal, the method comprising the steps of: each audio beamformer (101) of a plurality of audio beamformers generates a monaural audio signal representing the audio captured by the beam formed by the audio beamformer; The steps include adapting each of the multiple audio beamformers (101) to capture an audio source, The steps include determining the direction of audio captured by beams formed by multiple audio beamformers, A step of generating an intermediate stereo signal component for each of at least some monaural audio signals, wherein the first intermediate stereo signal component for a first monaural audio signal among at least some monaural audio signals includes a first monaural audio signal positioned at a first position in the stereo representation of the first intermediate stereo signal component, the first position depending on a first direction which is the direction of audio captured by a first beam of a first audio beamformer, which is an audio beamformer among a plurality of audio beamformers (101) that generate the first monaural audio signal; The process includes the step of generating a stereo audio signal that includes at least some intermediate stereo signal components. Embodiment 14. A computer program comprising computer program code, the program being adapted to be executed on a computer and to perform all the steps of Embodiment 13.
Claims
1. A device for generating stereo audio signals, A plurality of audio beamformers, each configured to generate a monaural audio signal representing the audio captured by the beam formed by the audio beamformer, An adapter configured to adapt each of the plurality of audio beamformers to capture an audio source, A direction determiner configured to determine the direction of audio captured by beams formed by the plurality of audio beamformers, A generator configured to generate an intermediate stereo signal component for each of at least some of the aforementioned monaural audio signals, wherein the first intermediate stereo signal component for a first monaural audio signal among the at least some of the monaural audio signals includes the first monaural audio signal positioned at a first position in the stereo representation of the first intermediate stereo signal component, and the first position depends on a first direction which is the direction of audio captured by a first beam of a first audio beamformer, which is an audio beamformer among the plurality of audio beamformers that generate the first monaural audio signal. A coupler configured to generate the stereo audio signal so as to include at least some of the intermediate stereo signal components, A device having a transmitter for transmitting the stereo audio signal to a remote device.
2. The apparatus according to claim 1, which is a teleconferencing device.
3. The apparatus according to claim 1 or 2, wherein the generator is configured to place the first mono audio signal within the stereo representation by scaling the addition of the first mono audio signal to the first and second channels of the first intermediate stereo signal components, and the relative scaling factor of the first channel with respect to the second channel depends on the first direction.
4. The apparatus of claim 3, wherein the generator is configured to determine a first scale coefficient of the first channel as a function of the first direction, and to determine a second scale coefficient of the second channel such that the sum of the first scale coefficient and the second scale coefficient satisfies a criterion.
5. The apparatus according to any one of claims 1 to 4, wherein the first position is a monotonic function of the first direction.
6. The apparatus according to any one of claims 1 to 5, wherein the first position is a linear function of the first direction.
7. The apparatus according to any one of claims 1 to 5, wherein the first position is a nonlinear function in the first direction.
8. The apparatus according to any one of claims 1 to 7, wherein the first position is a function of the first direction, and the function provides a mapping of the directional angle intervals of the first direction to larger angular intervals in the stereo representation.
9. The apparatus according to any one of claims 1 to 7, wherein the first position is a function of the first direction, and the function provides a mapping of the directional angle intervals of the first direction to smaller angular intervals within the stereo representation.
10. The apparatus according to any one of claims 1 to 9, wherein the coupler is configured to adapt the amplitude of the first intermediate stereo signal component in the stereo audio signal to the amplitude of the second intermediate stereo signal component representing the second monaural audio signal in the coupled stereo signal, according to the first direction.
11. The apparatus according to any one of claims 1 to 10, wherein the coupler is configured so as not to include a first intermediate stereo signal component in the stereo audio signal when the first direction satisfies a criterion.
12. The apparatus according to any one of claims 1 to 11, wherein the first direction indicates the direction of arrival of audio captured by the first audio beamformer.
13. A method for generating a stereo audio signal, The steps include: each of the multiple audio beamformers generates a monaural audio signal representing the audio captured by the beam formed by the audio beamformer; The steps include adapting each of the plurality of audio beamformers to capture an audio source, The steps include determining the direction of the audio captured by the beams formed by the plurality of audio beamformers, A step of generating an intermediate stereo signal component for each of at least some of the aforementioned mono audio signals, wherein the first intermediate stereo signal component for a first mono audio signal among the at least some of the mono audio signals includes the first mono audio signal positioned at a first position in the stereo representation of the first intermediate stereo signal component, the first position depending on a first direction which is the direction of audio captured by a first beam of a first audio beamformer, which is an audio beamformer among the plurality of audio beamformers that generate the first mono audio signal; The steps of generating the stereo audio signal so as to include at least some of the intermediate stereo signal components, A method comprising the step of transmitting the stereo audio signal to a remote device.
14. A computer program that is executed on a computer and causes the computer to perform the method described in claim 13.