Spatial audio communication

By receiving and analyzing the sound source activity information in the audio signal in the spatial audio communication and enabling corresponding spatial processing, the problem of poor sound source positioning and audio signal allocation in the prior art is solved, and clearer sound source perception and audio quality improvement are achieved.

CN120186545APending Publication Date: 2025-06-20NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411806168.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively utilize the spatial characteristics of the sound source to achieve clear sound source positioning and reasonable allocation of audio signals in spatial audio communication.

Method used

By receiving a plurality of audio signals, including at least one spatial audio signal, information related to the activity of the sound source is acquired, and spatial processing is enabled based on this information, the positioning of the sound source and the size and position of the audio signal are controlled.

Benefits of technology

It realizes clearer sound source positioning and reasonable allocation of audio signals, improving the listener's sound source perception effect and audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186545A_ABST
    Figure CN120186545A_ABST
Patent Text Reader

Abstract

The invention relates to spatial audio communication. Examples of the present disclosure relate to spatial audio communications. An apparatus receives a plurality of audio signals, where the plurality of audio signals includes at least one spatial audio signal. The apparatus obtains information related to the activity of the sound source at least for at least one spatial audio signal; and enabling spatial processing of the at least one spatial audio signal. The spatial processing is based at least in part on the obtained activity information, and the spatial processing controls localization of the sound source in accordance with the obtained activity information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Examples of the present disclosure relate to spatial audio communication. Some examples relate to spatial audio communication such as a conference call. Background Art

[0002] Spatial audio enables the presentation of the spatial characteristics of sound sources so that listeners can perceive different sounds arriving from different directions. Spatial audio can be used in communication such as a conference call. Summary of the Invention

[0003] According to various but not necessarily all examples of the present disclosure, there is provided an apparatus for spatial audio communication, the apparatus including components for the following operations:

[0004] Receiving a plurality of audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0005] Obtaining, at least for the at least one spatial audio signal, information related to the activity of the sound source; and

[0006] Enabling spatial processing of the at least one spatial audio signal, at least partially based on the obtained activity information, wherein the spatial processing controls the positioning of the sound source according to the obtained activity information.

[0007] The spatial processing can control the positioning of one or more active sources.

[0008] The activity information can be based on at least one of the following:

[0009] The position of the active source in the audio signal;

[0010] The number of active sources in the audio signal; or

[0011] The amount of activity of the active source in the audio signal.

[0012] The spatial processing can include at least one of the following: repositioning the at least one spatial audio signal; or resizing the at least one spatial audio signal.

[0013] The spatial processing can resize the at least one spatial audio signal such that an audio signal with more activity from the sound source has a larger size than an audio signal with less activity from the sound source.

[0014] The spatial processing can include: applying a weighting factor to one or more spatial audio signals; and resizing the spatial audio signal based on the weighting factor, wherein the weighting factor is at least partially based on the activity information.

[0015] Resizing the spatial audio signal can change the angular span of the audio signal.

[0016] Relocating the spatial audio signal can change the distance of the audio signal.

[0017] The spatial processing may include: relocating at least one obtained spatial audio signal such that a first audio signal is located in a first direction and a second audio signal is located in a second direction.

[0018] The spatial processing may include: relocating at least one obtained spatial audio signal such that an audio signal with more activity from a sound source is in a more prominent position than an audio signal with less activity from the sound source.

[0019] The component may be used to combine the plurality of audio signals after the spatial processing.

[0020] The plurality of audio signals may include at least one mono audio signal.

[0021] The spatial processing may include: assigning positions to the at least one mono audio signal.

[0022] The plurality of audio signals may include a first audio signal captured from a first audio scene and a second audio signal captured from a second audio scene.

[0023] The spatial audio signal may include at least one of the following:

[0024] Stereo;

[0025] Multichannel;

[0026] Ambisonics;

[0027] Parametric spatial audio.

[0028] The plurality of audio signals may be received via a plurality of channels.

[0029] According to various but not necessarily all examples of the present disclosure, a conference call system including one or more devices as described herein may be provided.

[0030] According to various but not necessarily all examples of the present disclosure, a method may be provided, including:

[0031] Receiving a plurality of audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0032] Obtaining information related to the activity of a sound source for at least the at least one spatial audio signal; and

[0033] Enable spatial processing of the at least one spatial audio signal, at least in part based on the obtained activity information, wherein the spatial processing controls the localization of the sound source according to the obtained activity information.

[0034] According to various but not necessarily all examples of the present disclosure, a computer program including instructions may be provided, which when executed by a device cause the device to at least perform:

[0035] Receive a plurality of audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0036] Obtain information related to the activity of the sound source, at least for the at least one spatial audio signal; and

[0037] Enable spatial processing of the at least one spatial audio signal, at least in part based on the obtained activity information, wherein the spatial processing controls the localization of the sound source according to the obtained activity information.

[0038] According to various but not necessarily all examples of the present disclosure, a device may be provided, including: at least one processor 902; and at least one memory 904 storing instructions, which when executed by the at least one processor 902 cause the device 900 to at least perform:

[0039] Receive a plurality of audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0040] Obtain information related to the activity of the sound source, at least for the at least one spatial audio signal; and

[0041] Enable spatial processing of the at least one spatial audio signal, at least in part based on the obtained activity information, wherein the spatial processing controls the localization of the sound source according to the obtained activity information.

[0042] Although the above examples and optional features of the present disclosure are described separately, it will be understood that they are included in the present disclosure in all possible combinations and permutations. It will be understood that various examples of the present disclosure may include any or all of the features described for other examples of the present disclosure, and vice versa. In addition, it will be understood that any one or more or all of the features (in any combination) may be implemented / included in / executed by a device, method, and / or computer program instructions as needed and appropriately. Description of the Drawings

[0043] Some examples will now be described with reference to the drawings, wherein:

[0044] Figures 1A to 1C illustrates an example system;

[0045] Figure 2 illustrates an example system;

[0046] Figure 3 illustrates an example method;

[0047] Figure 4 illustrates an example spatial audio mixer;

[0048] Figure 5 illustrates an example method;

[0049] Figure 6 illustrates an example method;

[0050] Figure 7 illustrates an example spatial audio mixer;

[0051] Figure 8 illustrates an example method; and

[0052] Figure 9 illustrates an example apparatus.

[0053] The figures are not necessarily to scale. For clarity and conciseness, particular features and views of the figures may be shown schematically or enlarged in scale. For example, the dimensions of some elements in the figures may be enlarged relative to other elements to aid in illustration. Like reference numerals are used in the figures to denote like features. For clarity, not all reference numerals may be shown in all figures. Detailed Description

[0054] Figures 1A to 1C Illustrates a system 100 that can be used to implement examples of the present disclosure. In these examples, the system 100 is a teleconference system. The teleconference system can enable voice or other similar audio content to be exchanged between different client devices 104 within the system 100. In other examples, other types of audio content can be shared between the corresponding devices.

[0055] In Figure 1A example, the system 100 includes a server 102 and a plurality of client devices 104. The server 102 can be a central server that provides communication between the corresponding client devices 104.

[0056] In Figure 1AIn the example, three client devices 104 are shown. In an implementation of the present disclosure, system 100 may include any number of client devices 104. Participants in a conference call or other communication session may use client devices 104 to listen to audio. The audio may include speech or any other suitable type of audio content or a combination of audio types.

[0057] Client device 104 includes components for capturing audio. The components for capturing audio may include one or more microphones. User device 104 also includes components for playing back audio to the participants. The components for playing back audio to the participants may include one or more speakers. In Figure 1A the example, the first client device 104A is a laptop computer, the second client device 104B is a smart phone, and the third client device 104C is a headset. In other examples, other types or combinations of other types of client devices 104 may be used.

[0058] During a conference call, the corresponding client devices 104 send data to the central server 102. The data may include audio captured by one or more microphones of client device 104. Then, server 102 combines and processes the received data and sends appropriate data to each client device 104. The data sent to client device 104 may be played back to the participants.

[0059] Figure 1B A different system 100 is shown. In this system 100, client device 104D acts as a server and provides communication between other client devices 104A-C. In this example, system 100 does not include server 102 because client device 104D performs the functions of server 102.

[0060] In this example, the client device 104D that performs the functions of server 102 is a smart phone. In other examples, other types of client devices 104 may be used to perform the functions of server 102.

[0061] Figure 1C Another different system 100 is shown, in which the corresponding client devices 104 communicate directly with each other in a peer-to-peer network. In this example, system 100 does not include server 102 because the corresponding client devices 140 communicate directly with each other.

[0062] In other examples, other arrangements for system 100 may be used.

[0063] Figure 1AExample system 100 for -C can be used to locate and merge multiple audio streams in a spatial call or spatial teleconference. This merging can occur in server 102 or client device 104, depending on the arrangement of system 100.

[0064] The transmission of audio signals in example system 100 can use data encoding, decoding, multiplexing, and demultiplexing. For example, audio signals can be encoded in various ways (such as Advanced Audio Coding (AAC) or Enhanced Voice Service (EVS)) to optimize the bit rate. Similarly, control data, spatial metadata, or any similar data can also be encoded. Immersive Voice and Audio Service (IVAS) is an example codec that can be used to encode both audio signals and corresponding spatial metadata. In addition, different encoded signals can be multiplexed into one or more combined bitstreams. Different coding systems can also be encoded in a joint manner such that the characteristics of one signal type affect the encoding of another signal type. An example in this regard is that the activity of an audio signal will affect the bit allocation for any corresponding spatial metadata encoder. When encoding and / or multiplexing has occurred at a device sending data, the receiving device then applies the corresponding decoding and demultiplexing.

[0065] Figure 2 Another example system 100 that can be used to implement the examples of the present disclosure is shown. The system 100 includes a server 102, and the server 102 is connected to a plurality of client devices 104 to enable a communication session, such as a teleconference, between the corresponding client devices 104.

[0066] Server 102 can be a spatial teleconference server. The spatial teleconference server 102 is configured to receive input audio signals from the corresponding client devices 104A - D. The server 102 processes the input audio signals to generate output spatial audio signals 204A - D. Then, the output spatial audio signals 204A - D can be sent to the corresponding client devices 104A - D.

[0067] In Figure 2 the example, the system includes four client devices 104A - D. The first client device 104A includes a laptop computer, the second client device 104B includes a smartphone, the third client device 104C includes a headset, and the fourth client device 104D includes a smartphone. In other examples, other types or combinations of client devices 104 can be used.

[0068] In Figure 2 the example system 100 shown, some client devices 104 provide input mono audio signals 200 to server 102. In Figure 2In the example, the second client device 104B and the third client device 104C send input monaural audio signals 200B, 200C to the server 102, while the first client device 104A and the fourth client device 104D send input spatial audio signals 202A, 202D to the server 102. In other examples, other arrangements for the input audio signals may be used.

[0069] The spatial audio signals 202, 204 can be any audio signals that are not monaural audio signals 200. The spatial audio signals 202, 204 can enable a participant to perceive the spatial characteristics of the audio content. The spatial characteristics can include the direction of one or more sound sources. In some examples, the spatial audio signals 202, 204 can include stereo signals, binaural signals, multi-channel signals, immersive surround sound signals, parametric spatial audio streams (e.g., metadata-assisted spatial audio (MASA) signals), or any other suitable type of signal. The MASA signal or any other suitable type of parametric audio stream can include one or more transmitted audio signals and associated spatial metadata. The metadata can be used by the client devices 104A-D to render any suitable kind of spatial audio output based on the transmitted audio signals. For example, the client devices 104A-D can use the metadata to process the transmitted audio signals to generate binaural signals or surround signals.

[0070] The server 102 can be configured to merge the received input audio signals 200, 202. The server 102 can include a spatial audio mixer that is configured to merge the received input audio signals 200, 202. In Figure 4 and 7 Examples of the spatial audio mixer are shown. Any suitable process can be used to merge the signals. The server 102 can be configured to merge the received input audio signals 200, 202 to provide a unique output spatial audio signal 204 for each client device 104. For example, the server 102 can receive the input audio signals 200, 202 and merge them such that each input audio signal 200, 202 represents a sector in the output spatial audio signal 204 that can be provided to the client device 104.

[0071] In Figure 2 the example, the server 102 performs the processing and spatial localization of the input audio signals 200, 202. In other examples, the server 102 can control other devices (e.g., the client device 104) to perform spatial localization.

[0072] Figure 3 Examples of example methods that can be used in the examples of the present disclosure are shown. A teleconference system (e.g., Figures 1A to 1C andFigure 2 The method can be implemented using the system 100 shown in [description]. Any suitable device can be used to implement the method. The device can be in the server 102 or the client device 104 or any other suitable electronic device. Figure 9 An example device that can be used in some examples is schematically shown.

[0073] At block 300, the method includes: receiving a plurality of audio signals, where the plurality of audio signals includes at least one spatial audio signal. A spatial audio signal can include any audio signal that is not a mono audio signal. A spatial audio signal can include stereo, multi-channel, panoramic surround sound, parametric spatial audio, or any other suitable type of audio signal. One or more mono audio signals can also be received along with at least one spatial audio signal.

[0074] The plurality of audio signals can include a first audio signal captured from a first audio scene and a second audio signal captured from a second audio scene. The plurality of audio signals can be received via a plurality of channels. The plurality of audio signals can be received from a plurality of client devices 104.

[0075] At block 302, the method includes: obtaining information related to the activity of a sound source for at least one spatial audio signal. Information related to the activity of a sound source can be obtained for one spatial audio signal or for a plurality of spatial audio signals.

[0076] The activity information can be related to the degree of activity of the audio signal in a communication session. This can be related to the number of sound sources within the audio signal. For example, if the sound source is a person talking, the activity information can be related to the amount of conversation within the audio signal.

[0077] In some examples, activity information can be obtained for a plurality of spatial audio signals. In some examples, activity information can be obtained for only one spatial audio signal.

[0078] The activity information can be based on the location of the active source in the audio signal, the number of active sources in the audio signal, the amount of activity of the active source in the audio signal, or any other suitable information. An active sound source can be a person talking or any other suitable type of sound source.

[0079] At block 304, the method includes: enabling spatial processing of at least one spatial audio signal. The spatial processing of at least one spatial audio signal is at least partially based on the obtained activity information and controls the localization of the sound source according to the obtained activity information.

[0080] The spatial processing can control the localization of one or more active sources. An active source can be a sound source that has been identified as active in the activity information.

[0081] Spatial processing may include at least one of the following: repositioning at least one spatial audio signal or resizing at least one spatial audio signal. In some examples, spatial processing resizes at least one spatial audio signal such that an audio signal with more activity from a sound source has a larger size than an audio signal with less activity from the sound source.

[0082] If the received plurality of audio signals includes one or more mono audio signals, spatial processing may include assigning a position to at least one mono audio signal.

[0083] In some examples, spatial processing may include: applying a weighting factor to one or more spatial audio signals; and resizing the spatial audio signals based on the weighting factor. The weighting factor may be at least partially based on activity information. For example, if the activity information indicates a higher activity level, the weighting factor may be used to resize the spatial audio signal such that the spatial audio signal is larger than other audio signals. If the activity information indicates a lower activity level, the weighting factor may be used to resize the spatial audio signal such that the spatial audio signal is smaller than other audio signals.

[0084] In some examples, resizing a spatial audio signal changes the angular span of the audio signal. For example, an angular sector may be assigned to a corresponding audio signal, and the angular span of the audio signal may be controlled by spatial processing.

[0085] In some examples, repositioning a spatial audio signal changes the distance of the audio signal. In some examples, resizing a spatial audio signal changes the depth of the audio signal.

[0086] In some examples, spatial processing may include repositioning at least one obtained spatial audio signal such that a first audio signal is positioned in a first direction and a second audio signal is positioned in a second direction.

[0087] In some examples, spatial processing may include repositioning at least one obtained spatial audio signal such that an audio signal with more activity from a sound source is in a more prominent position than an audio signal with less activity from the sound source. For example, an audio signal including the most active sound source may be repositioned to the front, while an audio signal including a less active sound source may be repositioned to the side.

[0088] In some examples, the method may include Figure 3 additional blocks not shown. For example, the method may include: combining a plurality of audio signals after spatial processing.

[0089] Accordingly, examples of the present disclosure provide a system 100 and an apparatus that can be used to control the positioning of a participant within a spatial communication session. The positioning can be controlled to improve the perceptibility and audio quality of the corresponding audio signal for a listener. For example, by providing a larger sector for an audio signal that includes more active sources, the clarity of the corresponding sources within such an audio signal can be improved.

[0090] Figure 4 An example spatial audio mixer 400 that can be used in examples of the present disclosure is shown.

[0091] The spatial audio mixer 400 receives a plurality of audio signals as inputs. The audio signals can include one or more spatial audio signals 202. In this example, the input audio signals further include one or more mono audio signals 200. One or more spatial audio signals 202 and one or more mono audio signals 200 can be received from the client device 104. In Figure 4 the example, the input audio signals can include both spatial audio signals 202 and mono audio signals 200. In other examples, the input audio signals can include only spatial audio signals 202.

[0092] The spatial audio mixer 400 is shown preparing a spatial audio output signal 204 for a single client device (e.g., the first client device 104A). The spatial audio mixer 400 can also prepare corresponding spatial audio output signals 204 for other client devices 104 in the system 100. The server 102 can also be configured to perform Figure 4 other processing not shown. For example, the server 102 can perform processing related to transmitting a video stream and / or any other suitable processing.

[0093] In Figure 4 the example, the spatial audio signals can include metadata-assisted spatial audio (MASA) signals. These signals include a mono or stereo transmitted audio signal and spatial metadata indicating spatial information (e.g., direction and direct-to-total ratio in a frequency band). The metadata can be in any suitable format. Examples of suitable formats are shown in Tables 1 and 2. In other examples, other types of spatial audio signals can be used. In other examples, other types of spatial audio can be used, such as stereo, panoramic surround sound, or surround audio.

[0094]

[0095] Table 1: MASA Format Spatial Metadata Parameters (Number Dependent on Direction)

[0096]

[0097] Table 2: MASA Format Spatial Metadata Parameters (Irrespective of the Number of Directions)

[0098] The input audio signals 200, 202 can be processed by a denoiser 402. The denoiser 402 can be configured to remove noise from the input audio signals 200, 202 and retain useful sounds, such as speech. The denoiser 402 can retain the useful sounds in their original spatial positions. The denoiser 402 can be optional. In other examples, denoising can be performed at the client device 104.

[0099] In other examples, there may not be any denoising in the signal path. For example, if one of the client devices 104 is sharing audio content other than speech. The audio content other than speech can be music or any other suitable type of content.

[0100] The operation of the denoiser 402 can depend on the type of signal. For a mono audio signal 200, the denoiser 402 can apply any suitable mono denoising process. For example, the denoising process can include transforming the mono audio signal into a time-frequency representation by using a short-time Fourier transform (STFT) or any other suitable transform. A trained machine learning model or any other suitable program can be used to determine a gain between 0 and 1 for different time-frequency regions to suppress noise in the speech. The determined gain can be applied to the signal. Then, by means of an inverse STFT or any other suitable transform, the signal can be converted back to a time-domain signal.

[0101] For a spatial audio signal 202, the denoiser 402 can apply any suitable spatial denoising process. For example, a machine learning model or any other suitable program can be used to steer an adaptive beamformer towards the source and suppress the remaining noise in the beamformed signal. The speech part of the microphone audio signal can be resynthesized by multiplying the resulting speech signal by the estimated speech steering vector. These operations can be performed in the STFT domain or any other suitable domain.

[0102] If the spatial audio signal 202 is a metadata-assisted spatial audio signal, the denoiser 402 can use the same method as that used for the mono audio signal 400 to suppress noise in the signal. In addition to this, the denoiser 402 then also modifies the spatial metadata such that the ratio parameter increases when the noise is suppressed. For example, if the ratio parameter for time index n and frequency index k is r orig (k,n), and if the denoiser 402 has determined a suppression gain g(k,n) between 0 and 1 to suppress noise in the speech, then the ratio can be modified by the following formula:

[0103]

[0104] where ∈ is a small value, such as 10 -9 . In some examples, the frequency resolution and ratio metadata of the suppression gain can be different. If the suppression gain has a high frequency resolution, it can be averaged in frequency so that it obtains the same frequency resolution as the ratio parameter, and then this average value is used in the above formula. In some examples, the metadata is not modified by the denoiser 402.

[0105] The spatial audio signal 202 and the mono audio signal 200 are provided as inputs to the spatial expansion determiner 404. The spatial expansion determiner 404 can be configured to determine the activity information of the input audio signals 200, 202. In this case, the spatial expansion determiner 404 can determine the spatial expansion of the active source. In the audio signals 200, 202, the active source can be a person who is talking or any other suitable source. Figure 5 An example method that can be implemented by the spatial expansion determiner 404 is shown.

[0106] The spatial expansion determiner 404 provides the spatial expansion information 406 as an output. In other examples, other types of activity information can be used.

[0107] The spatial expansion information 406 and the audio signals 200, 202 are provided as inputs to the relocator and combiner 408. The relocator and combiner 408 are configured to relocate the audio signals 200, 202 to control the position of the active source. This can make the distribution of the active source more evenly spread in the spatial output, thereby improving the clarity of the sound source for the listener. Figure 6 An example method that can be performed by the relocator and combiner 408 is shown.

[0108] The relocator and combiner 408 provide the spatial audio signal 204 as an output. The spatial audio signal 204 can be sent to the client device 104.

[0109] Figure 5 An example method that can be implemented by the spatial expansion determiner 404 in the examples of Figure 4 is shown.

[0110] In block 500, the method includes: obtaining parameters for the frequency bands of the input spatial audio signal 202. The obtained parameters can be related to the spatial characteristics of the spatial audio signal 202. In some examples, the obtained parameters can include a direction parameter and an energy ratio parameter or any other suitable parameters. If the spatial audio signal 202 includes a parametric spatial audio signal, the parameters can be obtained from the metadata associated with the spatial audio signal 202.

[0111] As an example, for the received spatial audio signal 202, the parameters can be represented as the azimuth azi(k,n) and the direct energy to total energy ratio r(k,n), where k is the frequency band index and n is the time index. For use cases such as teleconferences, it can be assumed that the source is mainly in the horizontal plane, and the elevation parameter (if available) can be discarded.

[0112] At block 502, the method includes: determining the signal energy E(k,n) of the input audio signal. The signal energy E(k,n) can be determined using the same resolution as that defining the resolution of the spatial metadata.

[0113] As an example, the spatial audio signal 202 can be represented as S(b,n,ch) in the short-time Fourier transform (STFT) representation, where b is the frequency bin index (of the time-frequency transform), n is the time index, in this example, the time index has the same time resolution as the metadata, and ch is the channel index.

[0114] For frequency band k, the lowest bin can be represented as b low (k), and the highest bin can be represented as b high (k). Assuming two channels, the energy can be expressed as

[0115]

[0116] At block 504, the rear direction is mirrored to the front. In this example, the rear value of azi(k,n) is mirrored to the corresponding front direction. For example, an azi(k,n) value of 150 degrees is converted to 30 degrees. This mirroring is performed because in a teleconference system, participants are typically located on the front side of the capture device. For other blocks of the method, this characteristic can be applied to the positions of other sources.

[0117] At block 506, the signal energy in the set of cumulative spatial regions is accumulated. The spatial regions can include different sectors. The sectors can have the same size. For example, there can be 18 sectors ranging from 90 degrees to -90 degrees, where the width of each sector is 10 degrees. In such an example, the sector energy is given by E sec (k,n,s idx ), where s idx is the sector index. It is expressed by the following formula:

[0118] E sec (k,n,s idx ) = E sec (k,n - 1,s idx )) * α+(1 - α)E(k,n)r(k,n)f(azi(k,n) ∈ s idx )

[0119] where f(azi(k,N)∈s idx ) is a function that is 1 if azi(k,n) lies in sector s idx and 0 otherwise, and α is a forgetting factor. The forgetting factor can take any suitable value, such as 0.99. The time-frequency domain values can be summed over frequency to obtain a broadband value:

[0120]

[0121] The summation can be performed over the entire frequency range or only over a specific frequency range. The specific frequency range can be 300 - 4000 Hz, which is the frequency range where speech energy mainly resides. In other examples, any other suitable range can be used.

[0122] In some examples, there can be frequency-dependent weighting, where higher weighting can be applied to some frequency bands. The frequency bands with higher weighting can be approximately 300 - 4000 Hz where speech energy mainly resides, or it can be any other suitable frequency range.

[0123] In block 508, the activity information of the region can be determined. In this example, the activity information can include the boundaries of the sectors containing all the sources in the spatial audio signal. This can be achieved, for example, by selecting the leftmost and rightmost peaks in the sector energy data E sec (n,s idx ) and setting these peaks as the sector boundaries. These sector boundaries provide spatial extent information, which is provided as the output from the spatial extent determiner.

[0124] In other examples, other methods can be used to determine the spatial extent information 406 or other activity information. For example, some methods can track moving sources.

[0125] The spatial extent information 406 or other activity information can be provided as an input to the relocator and combiner 408.

[0126] Figure 6 is shown that can be by Figure 4An example method performed by the relocator and combiner 408 in the example. The relocator and combiner 408 receives the spatial audio signal 202 and the mono audio signal 200 as inputs, and also receives the spatial expansion information 406 determined for the input signals. The relocator and combiner 408 is configured to generate a spatial audio signal 204 that can be output to the client device 104. The output spatial audio signal 204 can be in any suitable format. In some examples, the output spatial audio signal 204 can be in the form of spanning the front area from 90 degrees to -90 degrees and distributing different received signals within that front area.

[0127] At block 600, the method includes: determining the width of each spatial audio input. In some examples, determining the width may include: determining whether the spatial audio input is wide or narrow. This can be determined using the spatial expansion information 406. As an example, if the sector span exceeds 20 degrees (or any other threshold), the spatial audio input 202 can be considered wide. Otherwise, the spatial audio input 202 is determined to be narrow.

[0128] At block 602, determine the number of sectors value N in . The number of sectors value N in is the sum of the number of mono audio inputs 200 and spatial audio inputs 202, where the number of wide spatial audio inputs is weighted by a value. The weighting value can be any suitable value. In this example, the weighting value is 3. In other examples, other values can be used. As an example, if there are three mono audio inputs 200 and two spatial audio inputs 202 (where only one of them is considered wide), then

[0129] At block 604, the area is divided into N in sectors. In this case, the area is the front hemisphere. The area can have a uniform width. For example, if N in = 18, then the width of one sector is 10 degrees.

[0130] At block 606, sectors are assigned to the corresponding inputs. Multiple consecutive sectors can be assigned to a wide spatial input. The number of consecutive sectors that can be assigned to a wide spatial input can be associated with the weighting value. In this case, the weighting value is 3, and 3 consecutive sectors are assigned to each wide spatial input. A single sector can be assigned to any narrow input. A single sector can also be assigned to any mono audio signal 204.

[0131] Audio signals can be assigned to sectors in any suitable order. In some examples, audio signals can be randomly assigned to sectors. In other examples, some audio signals can be assigned more prominent sectors. For example, if an audio signal is known to correspond to a key speaker in a conference call, that audio signal can be given a higher priority than other audio signals and can be assigned to one or more sectors towards the center of the area.

[0132] At block 608, the audio signal can be modified to fill the corresponding assigned sector. The spatial expansion information 406 and / or any other suitable active information can be used to modify the audio signal. For example, the spatial expansion information 406 indicates how the sound sources within the audio signal are distributed. Thus, modifying the audio signal can include: modifying the spatial metadata such that the original left boundary sound and any sound to the left of the original left boundary sound are placed at the edge of the newly assigned left sector, and similarly for the right direction, and any sound between the left and right directions is mapped to the direction between the sector edges.

[0133] Modifying the audio signal can also include: modifying the energy parameter such that a narrower sector results in a higher energy ratio.

[0134] At block 610, the method includes: assigning a direction parameter to any received mono audio signal 200. The direction assigned to the mono audio signal can be the direction of the sector to which the received mono audio signal 200 has been assigned. Thus, the mono audio signal 200 is assigned to be an audio object in a specific direction.

[0135] At block 612, the method includes: combining the mono audio signal (now an object audio signal) and the modified spatial audio signal. Any suitable process for combining audio signals can be used.

[0136] At block 614, the relocator and combiner 408 provides the spatial audio signal 204 as an output. This can be the output of the spatial audio mixer 400 and can be sent to the corresponding client device 104.

[0137] Figure 7 Another example spatial audio mixer 400 that can be used in some examples of the present disclosure is shown. Corresponding reference numerals are used for features corresponding to the reference numerals in Figure 4 the reference numerals in.

[0138] In Figure 4 the example of, the spatial audio mixer 400 detects the spatial expansion of the sounds within the spatial audio signal 202 and then relocates them to new sectors. In Figure 7In the example, the spatial audio mixer 400 performs adaptive sector width so that an audio signal estimated to include a higher number of active sources is assigned a wider sector than an audio signal estimated to include a lower number of active sources.

[0139] The spatial audio mixer 400 receives multiple audio signals as input. The audio signals can include one or more spatial audio signals 202. In this example, the input audio signals also include one or more mono audio signals 200. One or more spatial audio signals 202 and one or more mono audio signals 200 can be received from the client device 104. In Figure 7 the example, the input audio signals can include both spatial audio signals 202 and mono audio signals 200. In other examples, the input audio signals can include only spatial audio signals 202. The spatial audio signals 202 can include any suitable type of audio signal.

[0140] The input audio signals 200, 202 can be processed by a denoiser 402. The denoiser 402 can be configured to remove noise from the input audio signals 200, 202 and retain useful sounds, such as speech. The denoiser 402 can retain the useful sounds in their original spatial positions. The denoiser 402 can be optional. In other examples, denoising can be performed at the client device 104.

[0141] The denoiser can use any suitable denoising process, such as the denoising process described above for Figure 4 description.

[0142] In some examples, the spatial audio mixer 400 can also be configured to perform other preprocessing steps. For example, the spatial audio mixer 400 can be configured to mirror any rear azimuth direction of the received spatial audio signal to the front, and / or any elevation data can be discarded.

[0143] The spatial audio signal 202 is provided as input to the source tracker 700. The source tracker 700 can be configured to provide an estimate of the source direction based on the input spatial audio signal 202. The source tracker 700 can be configured to determine the activity information of the input audio signals 200, 202. In Figure 7 the example, the source tracker 700 can determine the number of active sources and their corresponding directions. The source tracker 700 can use any suitable process to determine the number of active sources and their corresponding directions. For example, the source tracker 700 can examine the estimated sector energy data to determine peaks with levels above a threshold.

[0144] The source tracker 700 provides the active source location and quantity information 702 as output. In other examples, other types of activity information can be used.

[0145] The active source location and quantity information 702 may include, for 1 ≤ i s ≤ N s,j the source directions azi s (i s , j), where N s,j is the number of sources considered active for the spatial audio signal (where j is the index of the input stream). This process is performed separately for each parameterized spatial audio signal 202.

[0146] In Figure 7 the spatial audio mixer 700, the mono audio signal 200 is provided as an input to the activity determiner 704. The activity determiner 704 may be configured to determine the activity of sources within the spatial audio signal 202 and the mono audio signal 200. For example, the activity determiner 704 may be configured to determine whether a source is active within a defined period. If the source level exceeds a threshold within the defined period, the source may be determined to be active. The period may be the last 60 seconds or any other suitable period.

[0147] The activity determiner 704 provides the mono activity data 706 as an output. In other examples, other types of activity information may be used.

[0148] The mono activity data 706 may take any suitable form. In some examples, the mono activity data 706 may include the number N s,j of active sources, where for a mono input, the value is 0 or 1. The input stream index j is for all input signal types, i.e., for the mono audio signal 200 as well as the spatial audio signal 204.

[0149] The active source location and quantity information 702, the mono activity data 706, and the audio signals 200, 202 are provided as inputs to the relocator and combiner 708. The relocator and combiner 708 is configured to relocate the audio signals 200, 202 to control the location of the active sources. This may cause the distribution of the active sources to spread more evenly in the spatial output, thereby enabling improved clarity of the sound sources for the listener. Figure 8 An example method that may be performed by the relocator and combiner 708 is shown.

[0150] The relocator and combiner 708 provides the spatial audio signal 204 as an output. The spatial audio signal 204 may be sent to the client device 104.

[0151] Figure 8 An example method implemented by the relocator and combiner 708 in an example of Figure 7 is shown.

[0152] At block 800, the method includes: determining the number of active sources in the input signals 200, 202. The number of active sources can be determined by: summing the values of the number N of active sources for each received spatial audio signal 202 (this information can be included in the active source location and quantity information 702), and summing this quantity with the number of mono inputs for which the sources have been indicated as active (this information can be included in the mono active data 706). For example, s,j At block 802, the method includes: dividing the region into evenly spaced positions. In this case, the region is the front hemisphere, including the forward directions from -90 degrees to 90 degrees to spatially evenly distribute the positions, where the number of positions is the total number of active sources. For example, if 19 positions are determined, the division is positions spaced 10 degrees apart.

[0153]

[0154] At block 804, positions are assigned to the corresponding input audio signals. Each spatial audio signal 202 is assigned N

[0155] consecutive positions. For different spatial audio signals 202, the value of N s,j can be different. For one or more input spatial audio signals 202, the value of N s,j can be zero. Positions can also be assigned for each mono audio signal that has been indicated as active. s,j The audio signals can be assigned to positions in any suitable order. In some examples, the audio signals can be randomly assigned to positions. In other examples, some audio signals can be assigned more prominent positions. For example, if it is known that an audio signal corresponds to a key speaker in a conference call, that audio signal can be given a higher priority than other audio signals and can be assigned a position towards the center of the region.

[0156] At block 806, the audio signals can be modified to be located at their corresponding assigned positions. The spatial audio signals can be modified so that they match their assigned positioning. For example, the corresponding spatial audio signal 202 can be assigned N

[0157] consecutive and evenly spaced positions, while the original source position is at azi s,j (i s ,j). Thus, at block 806, the task is to modify the metadata of the spatial audio signal by a function so that the position azi s (i s ,j) is mapped to the newly assigned position. This is done so that the leftmost azi s (i s (is , j) is mapped to the leftmost of the assigned positions, the second leftmost is mapped to the second leftmost, and so on for other directions / positions. Any metadata positions in between can be mapped between the corresponding target positions. The ratio parameter can also be modified by treating the edge of the assigned position as the sector edge.

[0158] At block 808, the method includes: combining the mono audio signal (now the object audio signal) and the modified spatial audio signal. Any suitable process for combining audio signals can be used.

[0159] Any mono audio signal 200 in which no active source is identified or spatial audio signal 202 in which no sound source 202 is identified can be treated as a mono sound source located in front (or any other suitable direction) during the combining process. This can account for any estimation errors of the active talker.

[0160] At block 810, the relocator and combiner 708 provides the spatial audio signal 204 as an output. This can be the output of the spatial audio mixer 400 and can be sent to the corresponding client device 104.

[0161] Figure 6 and 8 The example method for modifying spatial metadata considers modifying a parameterized spatial audio signal. When the spatial metadata is modified, the audio signal that is normally transmitted is also modified. When the spatial audio signal is reproduced as sectors, any suitable process can be used to modify the transmitted audio signal.

[0162] During a communication session such as a conference call, some participants may be more active than others. For example, some participants will talk more than others. Thus, the width of the sector assigned to the spatial audio signal 202 can change during the communication session. At the start of the communication session, before any activity has been detected, a narrow sector can be used for the spatial audio signal 202. Subsequently, during the communication session, it can be determined that there are multiple active talkers in the spatial audio signal. In this case, a wider sector will be assigned to the spatial audio signal 202. If at another time during the communication session, the source within that position becomes inactive, the sector can be shrunk back from the wide sector to the narrow sector during the same communication session.

[0163] In some examples, when the activity of the sound source of the spatial audio signal 202 changes during a communication session, the size of the sector assigned to the spatial audio signal 202 may change, but the arrangement order of the spatial audio signal 202 does not change. This will prevent the positions of the audio signals 202 from being swapped during the communication session, which may cause interference or confusion to the listener. In other examples, the position and position size of the audio signal may be changed. This may allow the positions of the corresponding audio signals 202 to be swapped.

[0164] In the above examples, the spatial audio input signal 202 is a parametric spatial audio signal. In other examples, the spatial audio signal 202 may be in a different format, such as stereo, multi-channel, immersive surround sound, or any other suitable format. For example, if the received spatial audio signal 202 is stereo, the position of the active source within the spatial audio signal 202 can be detected by evaluating to which positions the active source has been amplitude-shifted. In some examples, this can be performed by determining the orientation of the time-frequency tiles and then determining the orientation of the active source.

[0165] Then, the stereo signal can be processed into the desired sector, for example, by first re-shifting the signal so that the active sources within the stereo reach the maximum separation, and then positioning the two channels at the left and right edges of the target sector. Any suitable method can be used to re-shift the stereo signal.

[0166] If the received spatial audio signal is immersive surround sound, the immersive surround sound signal can be converted into a parametric spatial audio signal. Any suitable method can be used to convert the immersive surround sound signal into a parametric spatial audio signal. For example, methods known from Directed Audio Coding (DirAC) can be used to estimate the metadata and generate the audio signal, which can be generated as two cardioid signals towards the left and right, or in any other form. Then, the above methods can be used for the generated parametric spatial audio signal.

[0167] In the above examples, the output of the spatial audio mixer is a parametric spatial audio signal 204. In other examples, other types of spatial audio signals 204 can be used, such as stereo signals, binaural audio signals, multi-channel audio signals, or immersive surround sound signals, or any other suitable type of signal. When positioning the received spatial audio signal 202, the corresponding type of spatial audio signal can be generated by using different shifting rules or head-related transfer function (HRTF) processing.

[0168] In the above example, the voice denoiser 402 is always active, and thus all other sounds other than the voice (or other useful sounds) are removed. However, in some cases, there may be some other useful sounds, such as music. These other useful sounds can bypass the repositioning process and can be transmitted to the client device 104 as a regular stereo signal. Such a signal can be transmitted as an additional audio stream separate from the spatial audio signal 204 containing the voice. In other examples, the voice signal and other useful sounds can be mixed into the same stereo signal via amplitude translation or binaural processing or by any other suitable process.

[0169] In some examples, other sounds may be repositioned and resized in a manner similar to that used for voice as described above. For example, the spatial extent of other useful sounds (such as music) can be determined as described above, and the remaining processing can also be as described above.

[0170] In some examples, the participants in a communication session can select between different settings for the communication session. The different settings can implement different sound source distributions based on the activity of the sound sources. For example, one setting can be such that the most active source is positioned at the center stage, while the other sources are positioned at the sides. Another setting can be such that the most active sources are positioned at the maximum separation from each other, while the other sources are positioned in between.

[0171] In some communication sessions, the number of participants can change over time. For example, due to participants joining and / or leaving a conference call or for any other reason, the number of client devices 104 sending audio signals to the server 102 in the example system can change. Due to the change in the number of participants, it may be necessary to adapt the spatial processing of the audio signals. For example, the repositioning and / or resizing may be adjusted. If new participants join the communication session, the resizing and / or repositioning methods (such as the methods shown in Figure 6 and 8 can be re-triggered to position the new participants in suitable positions. Additionally, in some cases, during a communication session, the audio signals from the participants can switch between a mono audio signal 200 and a spatial audio signal 204. In such a case, a reallocation of positions and / or sectors can be performed.

[0172] To avoid sudden changes, any panning rules and spatial metadata modification rules can be interpolated to new values over a specific period (such as over a 10-second period). For example, if the spatial sound input is repositioned to a sector with edges at specific positions, the sector edges can be slowly moved to their new positions.

[0173] In some examples, rather than using a source tracking method in server 102, server 102 can receive data (such as the number of conversants, conversant locations) and audio signals 200, 202 as metadata from client device 102. Client device 104 can use a more efficient source tracking method, such as using beamforming techniques, because the raw microphone signals are available for client device 104.

[0174] In some examples, different spatial processing can be used for different types of audio signal content. That is, the repositioning and resizing of the audio signal can depend on the type of audio content. For example, if the spatial audio signal 202 includes content such as music or other non-speech content, a wider sector can be allocated to that audio signal. In some examples, if the spatial audio signal 202 mainly includes speech, the default sector width can be used.

[0175] In some examples, server 102 can obtain information on how many participants are joining a conference session from the same acoustic space (from which the audio is captured and sent to server 102 in spatial audio signal 202). When source activity information is not yet available, the number of participants can be used as initial information to define how many adjacent sectors are to be allocated for the spatial audio signal 202 at the start of the communication session. The number of participants can be determined by server 102 or any other suitable device by detecting participants who have joined the same communication session and are in close proximity to each other.

[0176] Examples of the present disclosure can also be used in system 100 that sends video streams along with audio signals 200, 202. The video stream can be presented to the participants on a display of client device 104 or by any other suitable means. In such an example, it can be beneficial for the participants to have the virtual location of the audio signals 200, 202 from a given client device 104 in the same direction as the direction in which the video signal from that client device 104 is presented on the display. For example, if the video from another client device 104 is visible on the left side of the display, the corresponding audio signals 200, 202 will also be located to the left. In an example where the spatial audio signal 202 is received from client device 104 along with the video stream, the width of the processed audio sector can be related to the width of the embedded video visible on the display.

[0177] Figure 9 Device 900 that can be used to implement examples of the present disclosure is schematically shown. Device 900 includes at least one processor 902 and at least one memory 904. It will be understood that device 900 can include Figure 9Additional components not shown in the figure. The device 900 can be disposed within the server 102 or the client device 104 or any suitable device.

[0178] In Figure 9 example, the device 900 includes a processing device. The device 900 can be configured to process audio data. The device 900 can be configured to process audio data for a conference call or any other suitable purpose.

[0179] In Figure 9 example, the implementation of the device 900 can be as a processing circuit. In some examples, the device 900 can be implemented only in hardware, only having specific aspects in software (including firmware), or can be a combination of hardware and software (including firmware).

[0180] As Figure 9 shown, the device 900 can be implemented using instructions that implement hardware functions, for example, by using executable instructions of a computer program 906 in a general - purpose or special - purpose processor 902, and these executable instructions can be stored on a computer - readable storage medium (such as a disk, memory, etc.) to be executed by such a processor 902.

[0181] The processor 902 is configured to read from and write to the memory 904. The processor 902 may further include: an output interface through which the processor 902 outputs data and / or commands; and an input interface through which data and / or commands are input into the processor 902.

[0182] The memory 904 is configured to store a computer program 906 including computer program instructions (computer program code 908), and when these computer program instructions are loaded into the processor 902, they control the operation of the device 900. The computer program instructions of the computer program 906 provide the logic and routines that enable the device 900 to execute the methods described herein. The processor 902 can load and execute the computer program 906 by reading the memory 904.

[0183] Therefore, the device 900 includes: at least one processor 902; and at least one memory 904 storing instructions that, when executed by the at least one processor 902, cause the device 900 to at least perform:

[0184] Receiving more than 300 audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0185] Obtaining, for at least one spatial audio signal, information related to the activity of the sound source; and

[0186] Enable spatial processing of at least one spatial audio signal, at least in part based on the obtained activity information, wherein the spatial processing controls the localization of the sound source according to the obtained activity information.

[0187] As Figure 9 shown, the computer program 906 can reach the device 900 via any suitable delivery mechanism 910. The delivery mechanism 910 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a storage device, a recording medium (such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD) or a solid-state memory), an article of manufacture including or tangibly embodying the computer program 906. The delivery mechanism can be a signal configured to reliably transmit the computer program 906. The device 900 can propagate or transmit the computer program 906 as a computer data signal. In some examples, a wireless protocol (such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6LoWPan (Low-Power Personal Area Network IP v 6), ZigBee, ANT+, near field communication (NFC), radio frequency identification, wireless local area network (Wireless LAN) or any other suitable protocol) can be used to send the computer program 906 to the device 900.

[0188] The computer program 906 includes computer program instructions for causing the device 900 to at least perform the following operations or for at least performing the following operations:

[0189] Receive more than 300 audio signals, wherein the plurality of audio signals includes at least one spatial audio signal;

[0190] Obtain, for at least one spatial audio signal at least, information related to the activity of the sound source; and

[0191] Enable spatial processing of at least one spatial audio signal, at least in part based on the obtained activity information, wherein the spatial processing controls the localization of the sound source according to the obtained activity information.

[0192] The computer program instructions can be included in the computer program 906, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over multiple computer programs 906.

[0193] Although the memory 904 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable, and / or can provide permanent / semi-permanent / dynamic / cache storage.

[0194] Although the processor 902 is shown as a single component / circuit, it may be implemented as one or more separate components / circuits, some or all of which may be integrated / movable. The processor 902 may be a single-core or multi-core processor.

[0195] References to "computer-readable storage media", "computer program products", "tangibly embodied computer programs", etc. or "controllers", "computers", "processors", etc. should be understood to include not only computers having different architectures (e.g., single / multi-processor architectures and sequential (von Neumann) / parallel architectures), but also dedicated circuits such as field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc. should be understood to include software or firmware for programmable processors, such as the programmable content of a hardware device, whether instructions for a processor or configuration settings for a fixed function device, gate array, or programmable logic device, etc.

[0196] As used in this application, the term "circuit" may refer to one or more or all of the following:

[0197] (a) A pure hardware circuit implementation (e.g., an implementation in pure analog and / or digital circuits), and

[0198] (b) A combination of hardware circuits and software, e.g., (as applicable):

[0199] (i) A combination of analog and / or digital hardware circuits with software / firmware, and

[0200] (ii) A hardware processor (including a digital signal processor), any part of the software and memory, which work together to enable a device (e.g., a mobile phone or a server) to perform various functions, and

[0201] (c) A hardware circuit and / or a processor, such as a microprocessor or a part of a microprocessor, which requires software (e.g., firmware) to operate, but the software may not be present when not needed for operation.

[0202] This "circuit" definition applies to all uses of the term in this application (including in any claims). As a further example, as used in this application, the term "circuit" also encompasses a pure hardware circuit or an implementation of a processor with its accompanying software and / or firmware. The term "circuit" also encompasses (e.g., and if applicable to a particular claim element) a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or networking device.

[0203] The boxes shown in the accompanying drawings may represent steps in a method and / or segments of code in a computer program 906. The illustration of a particular order of boxes does not necessarily imply that the boxes have a required or preferred order, and the order and arrangement of the boxes may be changed. Additionally, some boxes may be omitted.

[0204] In Figure 9 the example of, apparatus 900 is shown as a single entity. In other examples, apparatus 900 may be provided as multiple distinct entities, which may be distributed within a cloud or other suitable network.

[0205] As used in this document, the term "comprising" has an inclusive meaning rather than an exclusive meaning. That is, any reference to X that comprises Y indicates that X may comprise only one Y or may comprise multiple Ys. If an exclusive meaning of "comprising" is intended, it will be explicitly indicated in the context by reference to "comprising only one" or by using "consisting of".

[0206] In this specification, the terms "connected", "coupled", and "communicate" and their derivatives refer to being connected / coupled / communicating operationally. It should be understood that any number or combination of intermediate components (including no intermediate components) may exist, i.e., so as to provide a direct or indirect connection / coupling / communication. Any such intermediate components may include hardware and / or software components.

[0207] As used herein, the term "determine" (and its grammatical variants) may include, but is not limited to: calculating, estimating, processing, deriving, measuring, investigating, identifying, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, etc. Additionally, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), obtaining, etc. Additionally, "determine" may include parsing, selecting, picking, establishing, etc.

[0208] In this specification, various examples have been referred to. The description of a feature or function with respect to an example indicates that these features or functions exist in that example. The use of the terms "example" or "for example" or "able to" or "may" in this document (whether explicitly stated or not) indicates that such a feature or function exists at least in the described example (whether described as an example or not), and they may but do not necessarily exist in some or all other examples. Thus, "example", "for example", "able to", or "may" refer to a particular instance within a class of examples. The characteristics of an instance may be the characteristics of only that instance or the characteristics of the class or a subclass of the class (which includes some but not all of the instances in the class). Thus, it is implicitly disclosed that features described with reference to one example and not with reference to another example may (if possible) be used as part of a working combination in that other example, but do not necessarily have to be used in that other example.

[0209] Although examples have been described in the preceding paragraphs with reference to various examples, it should be understood that the given examples can be modified without departing from the scope of the claims.

[0210] The features described in the foregoing description can be used in combinations other than those explicitly described above.

[0211] Although functions have been described with reference to specific features, these functions can be performed by other features, whether or not described.

[0212] Although features have been described with reference to specific examples, these features can also be present in other examples, whether or not described.

[0213] The term "a" or "the" as used in this document has an inclusive rather than an exclusive meaning. That is, any reference to X that includes a / the Y indicates that X can include only one Y or can include multiple Ys, unless the context clearly indicates the contrary. If the intention is to use "a" or "the" with an exclusive meaning, it will be made explicit in the context. In some cases, "at least one" or "one or more" may be used to emphasize the inclusive meaning, but it should not be assumed that any exclusive meaning can be inferred without these terms.

[0214] The features (or combinations of features) present in the claims refer to the feature (or combination of features) itself, and also to features (equivalent features) that achieve substantially the same technical effect. Equivalent features include, for example, features that are variants and that achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function in substantially the same way to achieve substantially the same result.

[0215] In this specification, adjectives or adjective phrases have been used to refer to various examples to describe the characteristics of the examples. This description of a characteristic with respect to an example indicates that the characteristic is present in some examples exactly as described, and is present in other examples substantially as described.

[0216] The foregoing description has described some examples of the present disclosure, but those of ordinary skill in the art will know of possible alternative structural and method features that provide functions equivalent to those of the specific examples of such structures and features described above, and that have been omitted from the foregoing description for the sake of brevity and clarity. However, the foregoing description should be understood to implicitly include a reference to such alternative structural and method features that provide equivalent functions, unless such alternative structures or method features are explicitly excluded in the foregoing description of the examples of the present disclosure.

[0217] When, in the foregoing specification, attention has been directed to those features regarded as important, it should be understood that the applicant may seek protection, by way of claim, for any patentable feature or combination of features mentioned in the foregoing and / or shown in the drawings, whether or not such features are emphasized.

Claims

1. A device for spatial audio communication, the device comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: receiving a plurality of audio signals, wherein the plurality of audio signals include at least one spatial audio signal; obtaining, at least for the at least one spatial audio signal, information related to an activity of a sound source; and Based at least in part on the obtained activity information, spatial processing of the at least one spatial audio signal is enabled, wherein the spatial processing controls the localization of the sound source in accordance with the obtained activity information.

2. The device according to claim 1, wherein: The spatial processing controls the positioning of one or more active sources.

3. The device according to claim 2, wherein: The activity information is based on at least one of the following: the location of active sources in the audio signal; the number of active sources in the audio signal; or The amount of activity of an active source in the audio signal.

4. The device according to claim 1, wherein: The spatial processing comprises at least one of: repositioning the at least one spatial audio signal; or resizing the at least one spatial audio signal.

5. The device according to claim 1, wherein: The spatial processing resizes the at least one spatial audio signal so that an audio signal having more activity from a sound source has a larger size than an audio signal having less activity from the sound source.

6. The device according to claim 1, wherein: The spatial processing includes applying weighting factors to one or more spatial audio signals and rescaling the spatial audio signals based on the weighting factors, wherein the weighting factors are based at least in part on the activity information.

7. The device according to claim 1, wherein: Resizing the spatial audio signal changes the angular span of the audio signal.

8. The device according to claim 1, wherein: The repositioning of the spatial audio signal changes the distance of the audio signal.

9. The device according to claim 1, wherein: The spatial processing comprises relocalizing the obtained at least one spatial audio signal such that a first audio signal is localized in a first direction and a second audio signal is localized in a second direction.

10. The device according to claim 1, wherein: The spatial processing includes repositioning the obtained at least one spatial audio signal so that an audio signal having more activity from a sound source is in a more prominent position than an audio signal having less activity from the sound source.

11. The device according to claim 1, wherein: The apparatus is caused to combine the plurality of audio signals after the spatial processing.

12. The device according to claim 1, wherein: The plurality of audio signals include at least one mono audio signal.

13. The device according to claim 12, wherein: The spatial processing comprises assigning a position to the at least one monophonic audio signal.

14. The device according to claim 1, wherein: The plurality of audio signals include a first audio signal captured from a first audio scene and a second audio signal captured from a second audio scene.

15. The device according to claim 1, wherein: The spatial audio signal includes at least one of the following: stereo; Multi-channel; Panoramic surround sound; Parametric spatial audio.

16. The device according to claim 1, wherein The plurality of audio signals are received via a plurality of channels.

17. A method comprising: receiving a plurality of audio signals, wherein the plurality of audio signals include at least one spatial audio signal; obtaining, at least for the at least one spatial audio signal, information related to an activity of a sound source; and Based at least in part on the obtained activity information, spatial processing of the at least one spatial audio signal is enabled, wherein the spatial processing controls the localization of the sound source in accordance with the obtained activity information.

18. The method according to claim 17, wherein: The spatial processing controls the positioning of one or more active sources.

19. The method according to claim 17, wherein: The spatial processing comprises at least one of: repositioning the at least one spatial audio signal; or resizing the at least one spatial audio signal.

20. A computer readable medium storing instructions which, when executed by an apparatus, cause the apparatus to at least perform: A plurality of audio signals are received, wherein: The plurality of audio signals include at least one spatial audio signal; obtaining, at least for the at least one spatial audio signal, information related to an activity of a sound source; as well as Based at least in part on the obtained activity information, spatial processing of the at least one spatial audio signal is enabled, wherein the spatial processing controls the localization of the sound source in accordance with the obtained activity information.