Method for generating immersive media content

The method addresses the lack of spatial properties in UGC by processing audio streams to conditionally upload data to a cloud server for advanced audio source separation and spatial localization, enhancing immersion in audio reproduction.

WO2025159984A1PCT designated stage expired Publication Date: 2025-07-31DOLBY LABORATORIES LICENSING CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012037
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-01-17
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for generating immersive media content, particularly from User Generated Content (UGC), fail to effectively capture and reproduce spatial acoustic properties, resulting in non-immersive audio experiences due to lack of spatial information in mono or stereo signals, and challenges in balancing audio components for immersive sound reproduction.

Method used

A method and system that utilizes a user device to process audio streams, determine feature information, and conditionally upload data to a cloud server for advanced audio source separation and spatial localization, generating immersive audio streams by assigning audio sources spatial locations based on estimated positions in the scene.

Benefits of technology

Enhances the perceived immersion of audio content by accurately positioning audio sources in a virtual environment, even with simple input streams, while offloading computationally intensive processing to the cloud, thus improving spatial audio reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000040_0000
    Figure 00000040_0000
  • Figure 00000041_0000
    Figure 00000041_0000
  • Figure 00000042_0000
    Figure 00000042_0000
Patent Text Reader

Abstract

An aspect relates to a method for generating immersive media content, comprising: by a user device: obtaining an audio stream of a scene; determining, based on audio feature information for the audio stream, that a cloud server upload process is to be performed; and responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server; by the cloud server: performing audio source separation to extract at least one audio source from the audio stream; and generating an immersive audio stream, wherein each respective audio source of the at least one audio source is included in an audio object and is assigned a spatial location based on an estimated location of the respective audio source in the scene, or is included in at least one channel based on an estimated location of the respective audio source in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR GENERATING IMMERSIVE MEDIA CONTENT

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003]

[0001] This application claims the benefit of priority from PCT Application No. PCT / CN2024 / 073494, filed on 22 January 2024, European Patent Application No. 24173053.0, filed on 29 April 2024 and U.S. Provisional Patent Application No. 63 / 667,532, filed on 3 July 2024 which are incorporated by reference herein in their entirety.

[0004] TECHNICAL FIELD OF THE INVENTION

[0005]

[0002] The present invention relates to a method and a system for generating immersive media content.

[0006] BACKGROUND OF THE INVENTION

[0007]

[0003] Today, media content is captured using a large variety of devices ranging from smart devices (such as smartphones, smartwatches and tablets) to professional recording devices with expensive and high-quality microphone elements. The vast majority of all audio content that is captured today is so-called User Generated Content (UGC), and most UGC is captured using microphones with limited performance and distributed (e.g. uploaded to a streaming service) with minimal or no post-processing and, because the process for capturing and distributing UGC is so simple, it has rapidly become widely adopted.

[0008]

[0004] For Professionally Generated Content (PGC) the source audio content often comprises many separate channels (e.g. one or more channels for dialogue, one or more channels for music and one or more channels for effects) and mixing engineers perform sophisticated audio processing to form well-balanced audio presentations in an immersive audio format, e.g. a surround sound format (e.g. 5.1), possibly also including height channels (e.g. 7.1.4 or 5.1.2 Dolby Atmos).

[0009]

[0005] UGC often comprises few, or only a single, captured channel, and is typically not subject to sophisticated manual post-processing by a mixing engineer. For example, in UGC, a target audio source (e.g. speech or music) may be recorded with an integrated microphone in a smartphone, held at some distance from the target audio source, wherein the microphone also captures background audio and noise that degrades the capture quality of the target audio source. Additionally, in many situations, the captured microphone signal is a mono audio signal meaning that the signal features very limited, or no, spatial properties.

[0010]

[0006] To this end, various automatic post-processing techniques for enhancing UGC have been proposed to e.g. improve intelligibility and reduce noise. Examples of such automatic post- processing includes using speech separators for extracting the speech content from an UGC signal and suppress background audio and noise to make the speech more intelligible. Further examples of automatic post-processing that can be used to enhance UGC include EQ-processing, volume adjustments and reverb processing.

[0011] GENERAL DISCLOSURE OF THE INVENTION

[0012]

[0007] A drawback with the existing solutions for processing audio content (especially UGC) is that while e.g. speech separation, noise suppression, EQ-processing, leveling, reverbprocessing etc. can enhance the perceived intelligibility and quality of the sound during playback, the resulting audio content would lack spatial acoustic properties (e.g. temporal and spectral cues for source localization) and result in a non-immersive impression when reproduced in a system capable of immersive sound reproduction. As an example, a mono audio signal recorded by a capture device carries no spatial information indicating the position of audio sources relative the capture device used to capture the audio sources. As a further example, a stereo audio signal captured by a pair of microphones of a smart device or a binaural capturing device may carry spatial information indicating the position of audio sources in a horizontal plane, however it may still be difficult to adjust the balance of different components in the recording to achieve a more balanced and immersive sound production. Additionally, a stereo audio signal captured with microphones located in a horizontal plane will still carry no information about height (elevation) of the audio sources. To this end, if attempting rendering of UGC to immersive multi-channel presentation formats (e.g. 2.0.2, 5.1.2, 5.1.4, 7.1.2, or 7.1.4 formats) the resulting presentation may lack the spatial acoustic properties that existed in the scene in which the UGC was recorded.

[0013]

[0008] It is a purpose of the present disclosure to present methods and systems for capturing and processing media content, especially UGC content, which brings back or enhances the spatial properties of the audio when reproduced.

[0014]

[0009] According to a first aspect there is provided a method for generating immersive media content.

[0015]

[0010] The method comprises, by a user device: obtaining an audio stream of a scene; processing the audio stream to determine audio feature information comprising one or more of: a signal-to-noise metric of the audio stream, a reverberation time of the audio stream, presence of a height audio source in the audio stream, an audio content type of the audio stream, an acoustic environment of the audio stream; determining, based on the audio feature information, that a cloud server upload process is to be performed; and responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server.

[0016] [Oil] The method further comprises, by the cloud server: receiving the audio stream from the user device; performing audio source separation on the audio stream to extract at least one audio source from the audio stream; and generating an immersive audio stream comprising the at least one audio source, wherein each respective audio source of the at least one audio source is included in an audio object of the immersive audio stream and is assigned a spatial location based on an estimated location of the respective audio source in the scene, or is included in at least one channel of the immersive audio stream based on an estimated location of the respective audio source in the scene.

[0017]

[0012] By the method of the first aspect, an immersive audio stream may be generated even for simple input audio streams, such as input audio streams captured by a user device such as a smart phone (e.g. mono audio signals or stereo audio signals). By using audio source separation technology one or more audio sources may be extracted from the audio stream and included in an immersive audio stream. Since each extracted audio source in the generated immersive audio stream is included as an audio object and assigned a spatial location based on an estimated location of the respective audio source in the scene, or is included in at least one channel of the immersive audio stream based on an estimated location of the respective audio source, the audio source(s) may be positioned in a virtual audio scene of the immersive audio stream, and thus enhance the perceived immersion for a listener. Various approaches for estimating a location of an audio source are set out in the following.

[0018]

[0013] An audio source may be assigned a spatial location by associating the audio object with spatial metadata (e.g. by including spatial metadata in the audio object) indicating the spatial location of the audio source in the scene (e.g. the estimated location of the respective audio source). An audio source may be included in at least one channel of the immersive audio stream based on an estimated location of the respective audio source by panning the audio source to the at least one channel based on, or in accordance with, the estimated location of the respective audio source.

[0019]

[0014] Furthermore, the method of the first aspect allows the typically computationally expensive processing involved in audio source separation and immersive audio stream generation to be offloaded from the user device (which may be constrained in terms of processing and power resources) to the cloud server. In addition to reducing the demands on the user device, it is contemplated this may allow more computationally complex audio source separations, larger models, as well as greater flexibility and easier deployment.

[0015] However, it is envisaged that not all audio streams obtained by user devices will be suitable (e.g. in terms of sound quality or type of content) for generating a convincing immersive audio stream. Therefore, the method of the first aspects implements what may be referred to as a “conditional cloud server upload process”, i.e. a conditional upload to the cloud, namely by determining, based on the audio feature information (comprising one or more of: a signal-to- noise metric, a reverberation time, presence of a height audio source, an audio content type, an acoustic environment), whether the cloud server upload process is to be performed by the user device. The audio stream may thus be uploaded to the cloud server provided (e.g., only if) the audio feature information fulfills one or more predetermined conditions.

[0020]

[0016] In some embodiments, determining that the cloud server upload process is to be performed comprises determining that the signal-to-noise metric exceeds a signal-to-noise threshold. The user device may upload the audio stream to the cloud server responsive to determining that the signal-to-noise metric exceeds the signal-to-noise threshold. Additional and alternative conditions for uploading the audio stream to the cloud server are set out herein.

[0021]

[0017] In some embodiments, the user device obtains a media stream of the scene, the media stream comprising a first video stream and the audio stream, wherein the cloud server upload process further comprises uploading the user device a second video stream to the cloud server; and the method further comprises, by the cloud server: receiving the second video stream from the user device; and determining a location of each of at least one visual object in the second video stream (e.g. at least one visual object identified in a sequence of frames of the second video stream), wherein each visual object corresponds to a respective one of the at least one audio source, and wherein the estimated location of each respective audio source is based on the location of the corresponding visual object.

[0022]

[0018] The cloud server thus generates the immersive audio stream comprising the at least one audio source, wherein each of the at least one audio source in the immersive audio stream is assigned a spatial location based on the location of the corresponding visual object. That is, each respective audio source may be included in an audio object of the immersive audio stream and assigned a spatial location based on the location of the corresponding visual object (more specifically based on the estimated location of the respective audio source which in turn is based on the location of the corresponding visual object), or may be included in at least one channel of the immersive audio stream based on the location of the corresponding visual object (more specifically based on the estimated location of the respective audio source which in turn is based on the location of the corresponding visual object).

[0023]

[0019] Audio- visual media in the form of video and audio recorded by mobile user devices (e.g. hand-held electronic devices such as smartphones or tablet computers) is an increasingly popular form of UGC, e.g. for personal moment sharing. The fast-paced development of the technical capabilities of mobile devices has enabled capturing of videos of high quality, both in terms of resolution and image quality, concurrently with capturing audio (e.g. in mono or stereo).

[0020] The cloud server may leverage this capability by estimating the location of one or more visual objects in the (visual) scene corresponding to the extracted audio sources in the (audio) scene, and using the location information when the assigning a respective spatial locations to the audio source(s) included in the immersive audio stream.

[0024]

[0021] The “conditional” cloud upload allows limiting bandwidth utilization by avoiding upload of both the second video stream and the audio stream to the cloud server if the user device estimates that the audio stream will be unsuitable for generating a convincing immersive audio stream.

[0025]

[0022] The “second video stream” may be the first video stream, wherein the video stream, as obtained by the user device may be uploaded to the cloud server. Alternatively, the “second video stream” may be a reduced bit rate version of the first video stream. A reduced bit rate version may include sufficient visual information for facilitating the estimation of the location of the respective audio sources, while requiring a smaller bandwidth to upload than the originally captured video stream.

[0026]

[0023] In some embodiments, the audio source separation is performed using a source separation model trained to extract an audio source of a predetermined height audio source type (e.g. from an input audio stream), wherein a height audio source is extracted from the audio stream, and wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

[0027]

[0024] The estimated location of the height audio source may thus be inherent to the audio source being specifically of a predetermined height audio source type, and thus be included in a height audio object. By “height audio object” is here meant an audio object assigned a spatial location of non-zero height.

[0028]

[0025] Thereby, even for simple input audio streams (e.g. mono audio signals, stereo audio signals or binaural audio signals), at least one height audio source of a height audio source type is extracted and (if included in an audio object) assigned a non-zero height, or included in at least one height channel of a (multi-channel) immersive audio stream allowing a presentation with enhanced spaciousness. UGC content in the form of a mono audio signal captured with a single microphone, or stereo typically carries no information regarding the height of captured audio sources. However, with the above method, height audio sources of a predetermined type present in the audio stream may be automatically moved (e.g. panned) to a height audio object or at least one height channel of in the immersive audio stream. This enables a distribution of audio sources between height and horizontal channels that may enhance the perceived immersion for a listener.

[0026] With the at least one height audio source being included in at least one height channel it is meant that the height audio source is included in at least one height channel of the immersive audio stream and that the height audio source optionally also is included in one or more non-height channels. In some embodiments, the at least one height audio source is included in at least one non-height channel of the immersive audio stream, in addition to the at least one height channel. Alternatively, in some embodiments the at least one height audio source is included in only one or more height channels of the immersive audio stream.

[0029]

[0027] According to a second aspect there is provided a system comprising a user device and a cloud server, each comprising one or more processors and configured to carry out the method according to the first aspect, or any embodiments thereof.

[0030]

[0028] According to a third aspect there is provided a method for generating immersive media content, the method comprising: by a user device: obtaining an audio stream of a scene; processing the audio stream to determine audio feature information of the audio stream; determining, based on the audio feature information, that a cloud server upload process is to be performed; responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server configured to generate an immersive audio stream based on an uploaded audio stream; and receiving from the cloud server, an immersive audio stream. The audio stream may be captured by the user device.

[0031]

[0029] According to a fourth aspect there is provided a user device comprising one or more processors configured to carry out the method according to the third aspect, or any embodiments thereof.

[0032]

[0030] According to a fifth aspect there is provided a method for generating immersive media content, the method comprising: obtaining a media stream comprising a video stream and an audio stream; performing audio source separation on the audio stream to extract at least one audio source from the audio stream; determining a location of each of at least one visual object in the video stream (e.g. in a sequence of frames of the video stream), wherein each visual object corresponds to a respective one of the audio sources; and generating an immersive audio stream comprising the at least one audio source, wherein each respective audio source is included in an audio object of the immersive audio stream and is assigned a spatial location based on the location of the corresponding visual object, or is included in at least one channel of the immersive audio stream based on the location of the corresponding visual object.

[0031] The method of the fifth aspect may be performed by a cloud server. The media stream may be received from a user device having obtained the media stream.

[0033]

[0032] According to a sixth aspect there is provided a cloud server comprising one or more processors configured to carry out the method according to the fifth aspect, or any embodiments thereof.

[0034]

[0033] According to a seventh aspect there is provided a computer program product comprising instructions which, when the program is executed by a processing device, causes the processing device to carry out the method according to the first, third or fifth aspect, or any embodiments thereof.

[0035]

[0034] According to an eighth aspect there is provided a computer-readable non-transitory storage medium storing the computer program product according to the seventh aspect, or any embodiments thereof.

[0036] BRIEF DESCRIPTION OF THE DRAWINGS

[0037]

[0035] The aspects of the present disclosure will be described in more detail with reference to the appended drawings, showing example implementations.

[0038]

[0036] Figure 1 is a block chart showing schematically a system for generating immersive media content.

[0039]

[0037] Figure 2 is a block chart showing schematically an example implementation of an immersive audio generation block.

[0040]

[0038] Figure 3 is a flow chart of a method for generating immersive media content.

[0041]

[0039] Figure 4 is a schematic view of a user device according to an example implementation.

[0042]

[0040] Figure 5 is an example of a frame of a video sequence of a music performance.

[0043]

[0041] Figure 6 is a block chart showing schematically a further example implementation of an immersive audio generation block.

[0044]

[0042] Figure 7 is a block chart showing a further example implementation of a cloud server.

[0045] DETAILED DESCRIPTION

[0046]

[0043] Systems and methods disclosed in the present application may be implemented as software, firmware, hardware or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; to the contrary, one physical component may have multiple functionalities, and one task may be carried out by several physical components in cooperation.

[0044] The computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, an AR / VR wearable, automotive infotainment system, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that computer hardware. Further, the present disclosure shall relate to any collection of computer hardware that individually or jointly execute instructions to perform any one or more of the concepts discussed herein.

[0047]

[0045] Certain or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions that when executed by one or more of the processors carry out at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken are included. Thus, one example is a typical processing system (e.g., computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system further may include a memory subsystem including a hard drive, SSD, RAM and / or ROM. A bus subsystem may be included for communicating between the components. The software may reside in the memory subsystem and / or within the processor during execution thereof by the computer system.

[0048]

[0046] The one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s). Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0049]

[0047] The software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to a person skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to the skilled person that communication media (transitory) typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0050]

[0048] Fig. 1 is a block chart showing schematically a system 1 in which embodiments of the present disclosure may be implemented.

[0051]

[0049] The system 1 comprises a user device 10 and a cloud server 20. The system 1 is configured to implement a method for generating immersive media content comprising, at least, an immersive audio stream. The operation of the system 1 will now be described with reference to the flow chart of Fig. 3 where steps S1-S4 are method steps performed by the user device 10 and steps S5-S7 are method steps performed by the cloud server 20.

[0052]

[0050] At step SI, the user device 10 obtains an audio stream A (in Fig. 1 schematically depicted by a waveform). As will be further described below the audio stream A may form part of a media stream M and further comprising a video stream V (in Fig. 1 schematically depicted by a sequence of video frames). The video stream V obtained by the user device 10 may in the following be referred to as a first video stream V. The audio stream A may be captured by a set of microphones of the user device 10. Fig. 4 is a schematic representation of a user device 110 which may be used to capture the audio stream A. The user device 110 is provided in the form of a portable or handheld electronic device, such as a smartphone, comprising a system of microphones 112a-b. The video stream V may be captured simultaneous to the audio stream A by a camera module 116 (interchangeably “camera”) of the user device 110. The audio stream A and / or the media stream need however not be captured by the user device 10. Alternatively, the audio stream A and / or the media stream M may be obtained from another device (e.g., a dedicated media capturing device, another smartphone) and subsequently be transferred or downloaded to the user device 10 therefrom, to generate immersive media content according to the present method.

[0053]

[0051] At step S2, the user device 10 process the audio stream A to determine audio feature information. The audio feature information may comprise one or more of: a signal-to-noise metric (SNR) of the audio stream A, a reverberation time of the audio stream, an audio content type of the audio stream A, a classification of an acoustic environment of the audio stream, and an indication of presence of a height audio source in the audio stream A. The different types of audio feature information represent features or characteristics of the audio stream being indicative of a suitability of the audio stream A to be used as an input stream for generating the immersive audio stream. An SNR metric and a reverberation time may be determined or estimated by an audio measurement block 12 of the user device 10. An audio content type, a classification of an acoustic environment, and / or presence of a height object may be determined, estimated or predicted by a metadata extraction block 14.

[0052] The metadata extraction block 14 may comprise a content classifier configured to detect or distinguish one or more types of content in the audio stream A. The content classifier may comprise a neural network trained to classify a content type of an input audio stream as a predetermined content type (e.g. music) or not (e.g. non-music).

[0054]

[0053] The metadata extraction block 14 may comprise an acoustic environment classifier (also known as a scene classifier) configured to classify an acoustic environment of the audio signal. An acoustic environment classifier may be implemented using machine-learning techniques and trained to output a prediction of an acoustic environment (e.g. a type of acoustic environment) based on an input of acoustic features extracted from the audio stream. The classifier may be implemented using Gaussian Mixture Model (GMM), a Hidden Markov Model (HMM), support vector machines (SVMs), or k-nearest neighbor (kNN) classifiers. The classifier may also be implemented using trained artificial neural networks, such as convolutional neural networks (CNN) and Recurrent Neural Networks (RNN). The metadata extraction block 14 may classify the acoustic environment as one of a predetermined set of acoustic environments. The predetermined set of acoustic environments may for instance comprise an indoor environment and an outdoor environment. However, a classifier may also be trained to provide a more granular and specific classification. For instance, the predetermined set of acoustic environments may comprise different types of indoor environments (e.g., room, hall, etc.) and / or outdoor environments (e.g., busy road or traffic junction, park, nature etc.). Training of machine-learning models and artificial neural networks to predict environments of the aforementioned types is per se known in the art and will hence not be described in further detail herein.

[0055]

[0054] The metadata extraction block 14 may comprise a height audio source detector configured to detect presence of a height audio source in the input audio stream A. The presence of one or more height audio sources in the audio stream A may for instance be indicated by a metadata in the form of a binary indicator (e.g. “1” indicating presence of a height audio source and “0” indicating absence of a height audio source). The height audio source detector may detect presence of a height object based on microphone signals of the audio stream A captured by a pair of vertically offset microphones of the user device 10.

[0056]

[0055] For example, a microphone system of a user device typically used for UGC capturing, such as the user device 110 of Fig. 4, may typically comprise at least two microphones 112a-b arranged at a top portion and a bottom portion, respectively, of the user device 110. Thus, when the user device 1 10 is held in portrait mode, oftentimes there is at least one microphone pointing upwards (e.g. 112a), and another microphone pointing downwards (e.g. 112b). With such a vertical microphone arrangement, elevation related information may be extracted from the audio signals recorded by these microphones (using e.g. beam steering), and the elevation related information can be used to detect presence of a height audio source in the audio stream A.

[0057]

[0056] For example, in frequency ranges where the upwards and downwards facing microphones 112a-b have omnidirectional microphone patterns, the microphone pair could form a fixed or adaptive beamformer with which the audio sources from non-horizontal elevations can be detected.

[0058]

[0057] As another example, in frequency ranges where a user device comprises a microphone having a polar microphone pattern, the microphone could be used to directly detect height audio sources in the audio stream A.

[0059]

[0058] At step S3, the user device 10 determines, based on the audio feature information, whether to perform a cloud server upload process (hereinafter interchangeably “upload process”). The upload process comprises uploading the audio stream.to the cloud server 20. Step S3 may be implemented by an upload decision block 16, receiving audio feature information from the audio measurement block 12, and if used, the metadata extraction block 14. As will be further described in the following, the upload process may further comprise uploading a video stream V’. Thus, Fig. 1 shows upload of a media stream M’, wherein the media stream M’ in some implementations may comprise the audio stream A, and in some implementations comprise both the audio stream A and the video stream V’.

[0060]

[0059] At step S4, responsive to a positive determination (i.e. the user device 10 determining, based on the audio feature information that the upload process is to be performed), the user device 10 performs the upload process and thus uploads the audio stream A (optionally the media stream M’ further comprising the video stream V’) to the cloud server 20.

[0061]

[0060] The audio feature information may comprise an SNR metric of the audio stream A, wherein the user device 10 may perform the upload process responsive to the SNR metric exceeding an SNR threshold. The SNR threshold may be a predetermined threshold, established based on a priori knowledge of what SNR level typically is needed to be able to generate a convincing immersive audio stream.

[0062]

[0061] Optionally, the SNR threshold may be determined in dependence on a classification of the acoustic environment of the audio stream A by the metadata extraction block 14. The upload decision block 16 may determine (e.g., set) a value of the signal-to-noise threshold in dependence on the classification of the acoustic environment of the audio stream. For instance, the acoustic environment classifier of the metadata extraction block 14 may be configured to classify the acoustic environment of the audio stream A as one of a predetermined set of acoustic environments comprising an indoor environment and an outdoor environment. The SNR threshold may be determined (e.g., by the upload decision block 16) as a first value responsive to the audio stream A being classified as an indoor environment and a second value responsive to the audio stream A being classified as an outdoor environment. The first value may be higher than the second value such that a higher tolerance for noise is provided in an outdoor environment.

[0063]

[0062] The audio feature information may comprise a reverberation time of the audio stream A, wherein the user device 10 may perform the upload process responsive to the reverberation time being less than a reverberation time threshold. The reverberation time threshold may be a predetermined threshold, established based on a priori knowledge of what reverberation time typically is needed to be able to generate a convincing immersive audio stream. Optionally, the reverberation time threshold may be determined in dependence on a classification of the acoustic environment of the audio stream A by the metadata extraction block 14. This may proceed in a manner analogous to determining the SNR threshold in dependence on the classification of the acoustic environment, as set out above. For instance, the reverberation time threshold may be determined (e.g., by the upload decision block 16) as a first value responsive to the audio stream A being classified as an indoor environment and a second value responsive to the audio stream A being classified as an outdoor environment. The first value may be higher than the second value such that a higher tolerance for reverberation is provided in an indoor environment.

[0064]

[0063] The audio feature information may comprise an audio content type of the audio stream A, wherein the user device 10 may perform the upload process responsive to the audio content type matching a predetermined audio content type. The predetermined audio content type may in some implementations be music. Another example of a predetermined audio content type is dialogue or speech content.

[0065]

[0064] The audio feature information may comprise an indication of presence of a height audio source in the audio stream A, wherein the user device 10 may perform the upload process responsive to the audio feature information indicating that a height audio source is present in the audio stream A.

[0066]

[0065] Two or more types of audio feature information may be used in combination to determine whether to perform the perform the upload process.

[0067]

[0066] For example, the SNR metric may be used in combination with the reverberation time such that both the conditions of the SNR metric exceeding an SNR threshold, and the reverberation time being less than a reverberation time threshold need to be met for the user device 10 to perform the upload process.

[0068]

[0067] In a further example, the SNR metric may be used in combination with the audio content type such that both the conditions of the SNR metric exceeding an SNR threshold, and the audio content type matching a predetermined content type (e.g. music) need to be met for the user device 10 to perform the upload process.

[0069]

[0068] In another example, the audio feature information may comprise an SNR metric and an indication of presence of a height audio source in the audio stream A, wherein the user device 10 may perform the upload process responsive to the audio feature information indicating that a height audio source is present in the audio stream A and the SNR metric exceeding an SNR threshold.

[0070]

[0069] Responsive to a negative determination (i.e. the user device 10 determining, based on the audio feature information that the audio stream A is not to perform the upload process) the method may end, i.e. be aborted without proceeding to perform the upload process. Upload of an audio stream not deemed suitable for generating an immersive audio stream may thus be avoided.

[0071]

[0070] For example, the upload decision block 16 may determine not to perform the upload process responsive to determining that the SNR of the audio stream A does not exceed the SNR threshold. In another example, the upload decision block 16 may determine to not perform the upload process responsive to determining that the reverberation time of the audio stream A exceeds the reverberation time threshold. In general, if the upload decision block 16 uses two or more types of audio feature information in combination to determine whether to perform the upload to the cloud server 20, the upload decision block 16 may determine to not perform the upload process responsive to determining that at least one of the conditions is not met, e.g., one or more of: the SNR not exceeding the SNR threshold, the reverberation time exceeding the reverberation time threshold, the audio stream A not being of the predetermined audio content type, the audio stream not including any height audio source.

[0072]

[0071] At step S5, the cloud server 20 receives the media stream M’ comprising at least the audio stream A, uploaded from the user device 10 as part of the upload process performed by the user device 10.

[0073]

[0072] At step S6, the cloud server 20 performs audio source separation on the audio stream A to extract at least one audio source from the audio stream A.

[0074]

[0073] At step S7, the cloud server 20 generates an immersive audio stream comprising the at least one audio source extracted at step S6. An audio source may be included as an audio object of the immersive audio stream and be assigned a spatial location based on an estimated location of the respective audio source in the scene. The audio object may for instance be associated with spatial metadata (e.g. by including the spatial metadata in the audio object) indicating the spatial location of the audio source in the audio scene. Alternatively, an audio source may be included in at least one channel of the immersive audio stream based on the estimated location of the respective audio source in the scene.

[0075]

[0074] According to the depicted example of Fig. 1 , the user device 10 captures a media stream M comprising, in addition to the audio stream A, the first video stream V. As mentioned above, responsive to determining to perform the upload process, the user device 10 may, as part of the upload process, upload both the audio stream A and a second video stream V’ to the cloud server 20. The second video stream V’ may be the first video stream V (i.e. V’=V), or a different video stream V’ generated based on the first video stream V. In particular, the second video stream V’ may be a reduced bit rate version of the first video stream V. A reduced bit rate version V’ of the first video stream V may be generated using various techniques, such as by down-sampling the first video stream V to generate a lower resolution version of the first video stream V, by sub-sampling the first video stream V to generate a reduced frame rate version of the first video stream V, and / or by applying a stronger compression.

[0076]

[0075] Regardless of whether the second video stream V’ is the first video stream V or a reduced bit rate version thereof, the cloud server 20 may thus as a result of the upload process receive both the audio stream A and the second video stream V’ from the user device 10.

[0077] According to this example, the cloud server 20 is thus capable of generating an immersive audio stream according to two different approaches. The cloud server 20 may generate an immersive audio stream IA1 using a “visually-aided immersive audio generation” approach (implemented by immersive audio generation block 22). Alternatively, the cloud server 20 may generate an immersive audio stream IA2 using a “height object-based immersive audio generation” approach (implemented by immersive audio generation block 24).

[0078]

[0076] The selection of the approach to be used may be based on the content type of the audio stream A uploaded to the cloud server 20. The selection may be implemented by an optional content check clock 21. The content check block 21 may determine the content type of the media stream M’, e.g. in particular the content type of the audio stream A, and determine whether the content type of the media stream M’ (e.g. the audio stream A) is of a predetermined content type. Responsive to determining that the media stream M’ is of the predetermined content type, the media stream M’ comprising the video stream V’ and the audio stream A may be forwarded to block 22 to be used in the visually-aided immersive audio generation approach. Responsive to determining that the media stream M is not of the predetermined content type, the audio stream A may be forwarded to block 24 to be used in the height object-based immersive audio generation approach.

[0079]

[0077] The determination of content type may be performed by a content classifier (e.g. comprised in the content check block 21) configured to detect or distinguish one or more types of content in the audio stream A. The content classifier may be implemented in a same manner as the content classifier of the metadata extraction block 14 of the user device 10 discussed above. The determination of content type may also or alternatively be based on metadata extracted by the metadata extraction block 14 of the user device 10. The metadata output by the metadata extraction block 14 indicating presence of an audio height source may be included in metadata of the media stream M’ uploaded to the cloud server 20.

[0080]

[0078] The predetermined content type may for instance be music. Accordingly, the visually-aided immersive audio generation approach may be used if the media stream M’ / audio stream A includes music-type content. Meanwhile, the height object-based immersive audio generation approach may be used if the media stream M / audio stream includes non-music-type content (e.g. merely dialog, environmental sound, etc.). In another example, the predetermined content type may for instance be dialogue or speech content. Accordingly, the visually-aided immersive audio generation approach may be used if the media stream M’ I audio stream A includes dialogue or speech content. Meanwhile, the height object-based immersive audio generation approach may be used if the media stream M / audio stream includes non-dialogue or speech content (e.g. merely environmental sound, etc.).

[0081]

[0079] The selection of the approach to be used may additionally or alternatively be based on whether the media stream M’ uploaded to the cloud server 20 comprises both a video stream V’ and an audio stream A, in other words whether the content type of the media stream M’ is audio-visual content. Accordingly, the content check block 21 may determine whether the received media stream M’ comprises an audio stream A and a video stream V’, and responsive to a positive determination determine that the media stream M’ may be forwarded to block 22 to be used in the visually-aided immersive audio generation approach. Responsive to determining that the media stream M’ does not include a video stream V’, (hence not being audio-visual content but only audio content), the audio stream A may be forwarded to block 24 to be used in the height object-based immersive audio generation approach. This approach may be combined with the above-mentioned determination of content type. For instance, the content check block 21 may forward the media stream M to the block 22 responsive to determining that the media stream M’ is audio-visual content and includes music, and otherwise forward the audio stream A to the block 24.

[0082]

[0080] The visually-aided immersive audio generation approach will now be described with further reference to Fig. 2 showing an example implementation of the immersive audio generation block 22.

[0083]

[0081] After receiving the media stream M’ comprising the video stream V and the audio stream A from the user device 10, the cloud server 20 (e.g. at step S6 of the flow chart of Fig. 3) performs audio source separation on the audio stream A to extract at least one audio source from the audio stream A. The audio source separation may as shown in Fig. 2 be performed by a source separation model 22a of the immersive audio generation block 22. The source separation model 22a (interchangeably “source separator”) is configured or trained to extract an audio source of at least one predetermined audio source type from the audio stream A. The source separator 22a thus performs source separation in the sense of separating individual sound sources from a mixture of multiple sounds present in the audio stream A. The source separator 22a may accordingly output a respective separated source audio signal (interchangeably “source signal”) for each of the at least one predetermined audio source type. Each source signal includes audio content of the audio stream A matching the respective predetermined audio source type but excludes audio content of the audio stream A not matching the respective predetermined audio source type.

[0084]

[0082] The source separator 22a may as per the illustrated example be trained to extract at least one of a vocal source type (i.e. vocals, a singing voice), and one or more instrument source types. Such a source separator 22a lends itself for an application wherein the media stream includes music content. For instance, the video stream V’ and the audio stream A may capture a music performance. Provided with an audio stream A comprising music including vocals, the source separator 22a may thus from the audio stream A extract a source signal comprising the vocals and excluding instrument sources in the audio stream A, and further for each instrument source type a respective source signal comprising the instrument of the respective instrument source type and excluding any further instrument source types and the vocals.

[0085]

[0083] In the depicted example the source separator 22a extracts vocals, and instrument source types of drum, bass and piano. This is however only one example and a source separator 22a for music applications may more generally be trained to extract at least one of the following instrument source types: bass, drums, keyboards, piano, guitar, violin, brass, reed.

[0086]

[0084] In another example, the source separator 22a may be trained to extract speech of one or more persons (e.g. one or more individual voices). Such a source separator 22a lends itself for an application wherein the media stream includes dialogue or speech content. For instance, the video stream V’ and the audio stream A may capture an oral presentation or an interview. Provided with an audio stream A comprising speech or dialogue, the source separator 22a may thus from the audio stream A extract a source signal comprising the speech and excluding other sources such as ambient or background sound in the audio stream A. In case of speech of one or more persons, a respective source signal may be extracted for each voice, each source signal comprising the speech of a respective participant and excluding other sources such as speech of other participants (other voices) and ambient or background sound in the audio stream A.

[0085] The source separator 22a can be implemented using machine-learning techniques, in particular neural networks NNs. Any known and conventional techniques for implementing and training NNs for source separation (e.g. for vocal and / or instrument source separation) may be used. Other approaches for realizing a source separator 22a is to provide a set of time- and / or frequency-domain filters or a filterbank, wherein each filter is tailored to present a passband matching a likely spectral signature of a respective source. While such a source separator may allow a computationally efficient implementation, it may be better suited for separation of audio sources that do not overlap, or only overlap to limited extent, in frequency domain.

[0087]

[0086] In case the audio stream A comprises more than one channel (e.g. a stereo audio stream), the source separator 22a may process each channel individually, or convert the audio stream to a mono audio stream prior to the source separation.

[0088]

[0087] After receiving the media stream comprising the video stream V’ and the audio stream A from the user device 10, the cloud server 20 (e.g. prior to, after, or in parallel to step S6 of the flow chart of Fig. 3) process the video stream V’ to determine a location of each of at least one visual object in the video stream V. The processing of the video stream V’ may as shown in Fig. 2 be performed by an image recognition model 22b. The image recognition model 22b (interchangeably “image processor”) is configured or trained to identify a visual object of at least one predetermined visual object type in the frames of the video stream V. Given a frame or a sequence of frames as input, the image processor 22b may thus identify and locate a visual object of at least one predetermined visual object type.

[0089]

[0088] The image recognition model 22b is trained to identify visual objects of one or more types, each type corresponding to a respective one of the at least one predetermined audio source type the source separator 22a is trained to extract. A correspondence between a visual object type and a predetermined audio source type (or separated audio source) may here mean that the visual object type has the visual appearance of an object which typically (e.g. with a certain level of confidence) may be the source of (e.g. emit) sound of the predetermined audio source type (i.e. the sound of the separated audio source).

[0090]

[0089] As an illustrative example, in Fig. 2 the source separator 22a is trained to extract at least one of a vocal source type (i.e. vocals, a singing voice), and one or more instrument source types, more specifically vocals, and the instrument source types of drum, bass and piano. The image processor 22b may as indicated in Fig. 2 correspondingly be trained to identify a singer or vocalist (i.e. a visual object being the likely source of vocals), and one or more visual objects corresponding to instrument source types, e.g. a bass or a bassist playing the bass (i.e. visual object being a likely source of bass sound), drums or a drummer playing the drums (i.e. a visual object being a likely source of drum sound) and a piano or a pianist playing the piano (i.e. a visual object being a likely source of piano sound). This applies correspondingly to any of the aforementioned further instrument source types the source separator 22a may be trained to separate (e.g. one or more of bass, drums, keyboards, piano, guitar, violin, brass, reed).

[0091]

[0090] The location of a visual object identified by the image processor 22b in a frame may be represented by location data, e.g. comprising one or more representative coordinates of the visual object in the frame. The location of a visual object may for instance be represented by a center of mass or centroid or the pixels depicting the visual object. According to another example the location of a visual object may be represented by coordinates of a bounding box (e.g. rectangular or polygonal) of the visual object. Either a minimum or non-minimum bounding box of the visual object may be used. According to yet another example, the location may be represented merely as coordinates of one of a set of predetermined segments of the frame (e.g. a rectangular grid) including the visual object.

[0092]

[0091] The image processor 22b can be implemented using machine-learning techniques, in particular neural networks NNs. Any known and conventional techniques for implementing and training NNs for image recognition (e.g. for identifying vocalists and / or instruments and / or performing instrumentalists) may be used.

[0093]

[0092] As the input to the image processor 22b is formed by a video stream V’ (comprising a sequence of frames) the determined location for each identified visual object may be provided on a frame-level basis, i.e. a sequence of location data for each visual object at the frame rate of the video stream V’.

[0094]

[0093] As shown in Fig. 2, the location (i.e. location data) of each visual object determined by the image processor 22b is used as an estimate of a location of the corresponding sound source in the scene (i.e. the visual scene and audio scene respectively) captured in the video stream V and the audio stream A. The location of each visual object and the corresponding audio source is provided as input to the immersive audio generator block 22c (interchangeably “audio generator block 22c”). The audio generator block 22c, based on the input, generates an immersive audio stream IA1 comprising each of the separated audio sources (e.g. vocals and one or more instruments), wherein each of the audio sources is assigned a spatial location based on the location of the corresponding visual object determined by the image processor 22b. Provided at least two audio sources are extracted from the audio stream A, the immersive audio stream IA1 may comprise a mix of the at least two audio sources.

[0095]

[0094] The assignment of a spatial location to each audio source by the audio generator 22c may comprise applying a transform mapping the respective location data of each visual object to a spatial location in a virtual audio scene. More specifically, the coordinates determined for the visual object (e.g. representative coordinates of the visual object, coordinates of a bounding box enclosing the visual object, coordinates of a predetermined segment of the frame including the visual object) may be mapped to coordinates (e.g. rectangular, polar, two- or three- dimensional) of the virtual audio scene of the immersive audio stream IA1.

[0096]

[0095] The immersive audio stream IA1 may be an object-based audio stream wherein an audio source may be included in, and thus represented by, an “audio object” and be associated with spatial metadata indicating the spatial location (which may change over time) of the audio object within the virtual audio scene.

[0097]

[0096] The immersive audio stream IA1 may also be a channel-based audio stream wherein an audio source, in the immersive audio stream IA1, may be mixed or panned into one or more channels (e.g. a front left / right channel, a center channel, rear left / right channel, a height channel) based on the location of the corresponding visual object. Combinations of object- and channel-based immersive audio streams, wherein some audio sources may be included in audio objects and some audio sources may be included in one or more channel(s) of the immersive audio stream, are also possible. The one or more channel(s) may for instance be one or more bed channel(s), i.e. audio channels that are meant to be reproduced in pre-defined, fixed locations. The immersive audio stream may for instance be in a 7.1.4 or 5.1.2 Dolby Atmos format.

[0098]

[0097] In any case, the assignment (e.g. mapping or mixing) may advantageously be such that a given audio source (or audio object) is assigned a spatial location, or is included in one or more channels, such that the audio source, when rendered on a rendering system, is perceived to originate from a position corresponding to the position of corresponding visual object in the video stream V or V’. The precision by which the mapping from the location of the visual object to the spatial location of the corresponding audio source may vary. In a simple implementation, an audio source corresponding to a visual object located e.g. on a horizontal plane (e.g. zero elevation) in a left or right portion of the (visual) scene (or frame) may be assigned a predetermined spatial location to the left or the right on the horizontal plane of the virtual audio scene. If the visual object is at a central location of the scene the audio source may be assigned a predetermined spatial location at the center of the virtual audio scene. If a visual object is located above or below the horizontal plane of the scene the audio source may be assigned a predetermined vertical spatial location above or below the horizontal plane of the audio scene. In a more elaborate implementation, a more dynamic or continuous mapping may be provided.

[0099]

[0098] Optionally, the audio generator 22c may apply a predetermined or adaptive gain for the audio sources included in the immersive audio stream. The audio generator 22c may thus achieve an altered or improved balance among the audio sources included in the immersive audio stream. The gain may for instance be based on the type of audio source. E.g. a first gain may be applied to an audio source of a vocal source type, and a second gain may be applied to an audio source of an instrument source type. It is further possible to apply different gains to audio sources of different instrument source types.

[0100]

[0099] Fig. 5 shows as an illustrative and highly schematic example a video frame of a video stream V of a scene S in the form of a music performance. The scene S includes an ensemble of a singer, a guitarist playing the guitar, a drummer playing the drums, a keyboardist playing a keyboard and a saxophonist playing the saxophone. The image processor 22b may accordingly determine the locations of the singer and instruments (and / or instrumentalists) within each frame. Meanwhile, the source separator 22a may extract the audio sources of the concurrently captured audio stream A, e.g. the song, the guitar, the drums, the keyboard, and the saxophone. The audio generator 22c may receive the determined locations of the visual objects and the separated audio sources as input and generate the immersive audio stream IA 1 including the separated audio sources assigned spatial locations corresponding to the determined locations of the visual corresponding visual objects, such that the immersive audio stream IA1 upon rendering at least approximately may reproduce the physical locations of the ensemble members during the video stream V.

[0101]

[0100] As another illustrative example, if the source separator 22a is trained to extract speech of one or more persons, the image processor 22b may be trained to identify visual objects in the form of faces with moving lips in the video stream V’ . The audio generator 22c may thus receive the determined locations of the visual object(s) and the separated audio source(s) as input and generate the immersive audio stream IA 1 including the separated audio source(s) assigned spatial locations corresponding to the determined locations of the visual corresponding visual objects (i.e. the location of a face with lip movement during non-silent periods of the respective separated audio source), such that the immersive audio stream IA1 upon rendering at least approximately may reproduce the physical locations of the talking one or more persons during the video stream V.

[0102]

[0101] After generating the immersive audio stream IA1, the cloud server 20 may optionally, at step S8, generate an immersive media stream by combining the immersive audio stream IA1 and the video stream V’. The cloud server 20 may also make the immersive audio stream IA1 available for download by the user device 10, wherein the user device 10 may download the immersive audio stream IA1 from the cloud server 20 and generate an immersive media stream by combining the immersive audio stream IA1 and the video stream V (or V’). Hence, step S8 may depending on implementation be performed by either the cloud server 20 or the user device 10.

[0103]

[0102] It may be noted that an audio stream A may further comprise other audio sources or background / ambient sound which the source separator 22a is not trained to separate from the audio stream A. The source separator 22a may output such audio content as a residual audio signal R (e.g. formed by subtracting the separated audio signals from the audio stream A). The residual signal R may optionally be included in the immersive audio stream IA1 by the audio generator 22c. The immersive audio stream IA1 may thus comprise a mix of the extracted audio sources and the residual signal R. The residual signal R may for instance be assigned one or more predetermined spatial locations or channels (e.g. corresponding to front and / or left and / or right of a stage of a music performance), or added as non-localized (i.e. background) audio. Optionally, as further described below, the residual signal R may be provided as input to the immersive audio generation block 24, to attempt extracting an audio source of at least one predetermined height audio source type.

[0104]

[0103] Determining the horizontal and vertical coordinates (e.g., the x- and y-coordinates) of a visual object in a video frame is typically straightforward. However, an accurate estimation of the depth of, or distance to, a visual object (e.g., the z-coordinate) may be more challenging. Consequently, it may be challenging to assign a proper depth coordinate to a given audio source when generating the immersive audio stream IA1. To facilitate depth assignment, the audio generator 22c may be provided with a further input in the form of focus distance information F, as shown in Fig. 2. With further reference to Fig. 4, the camera module 116 of the user device

[0105] 110 may while capturing the video stream V continually record focus distance information (e.g. on a frame-by-frame basis) as metadata of the video stream V. The metadata may further be included in the uploaded video stream V’ . The immersive audio generation block 22 of the cloud server 20 may upon receiving the video stream V’ extract the focus distance information F and forward it to the audio generator 22c. Two-dimensional coordinates of a visual object determined from the video frames may thus be supplemented with focus distance information to obtain three-dimensional location data for each visual object (e.g., x-, y- and z-coordinates wherein the z-coordinate corresponds to the focus distance information F). The audio generator 22c may thus estimate the location of each respective audio source based on the location of the corresponding visual object determined from the video stream V’ by the image recognition model 22b, and the focus distance information. The audio generator 22c may more accordingly map three- dimensional location data for each visual object to corresponding three-dimensional coordinates of the virtual audio scene.

[0106]

[0104] The height object-based immersive audio generation approach will now be described with further reference to Fig. 6 showing an example implementation of the immersive audio generation block 24.

[0107]

[0105] After receiving the media stream M’ comprising the audio stream A from the user device 10, the cloud server 20 (e.g. at step S6 of the flow chart of Fig. 3) performs audio source separation on the audio stream A to extract at least one audio source from the audio stream A. The audio source separation may as shown in Fig. 6 be performed by a source separation model 26 of the immersive audio generation block 24. The source separation model 26 (interchangeably “source separator”) is configured or trained to extract an audio source of at least one predetermined height audio source type from an input audio stream (i.e. the audio stream A). The source separator 26 thus performs height audio source separation in the sense of separating individual height sources from a mixture of multiple sounds present in the audio stream A. The source separator 26 may accordingly output a respective separated height audio source signal (interchangeably “height source signal”) for each of the at least one predetermined height audio source type. Each height source signal includes audio content of the audio stream A matching the respective predetermined height audio source type but excludes audio content of the audio stream A not matching the respective predetermined height audio source type.

[0108]

[0106] Examples of predetermined height audio source types comprises at least one of manmade sounds associated with height, sounds made by alive objects associated with height and sounds of nature associated with height. Exemplary sub-categories under manmade sounds associated with height are the sound of a blade rotating in the air (e.g. the sound of a helicopter, drone, propeller engine or jet engine), the sound caused by manmade objects flying through the air (the sound of an airplane or balloon moving through the air) and the sound of combustion (e.g. the sound of fireworks exploding). Exemplary sub-categories under alive objects associated with height are sounds associated with an animal using aerial locomotion (e.g. the vocal sounds of bats, birds or insects or the sound of bats, birds or insects moving through the air) and sounds associated sound associated arboreal animals (e.g. the vocal sounds of sloths and / or monkeys / primates and the sound of sloths and / or monkeys / primates moving in trees). Exemplary sub-categories under nature sounds associated with height are weather sounds (e.g. the sound of thunder, rain, wind and hail) and landscape feature sounds associated with height (the sound of a waterfall and the sound of rattling leaves).

[0109]

[0107] The source separator 26 may as shown comprise a respective separate source separator sub-model or module 26a-b, each trained to extract a respective predetermined height audio source type. As an example, source separator module 26a may be configured to extract a height audio source type that contains (if present in the audio stream A) manmade sounds associated with height and source separator module 26b may be configured to extract a height audio source of a more specific height audio source type, such as only the sound of a blade rotating through the air and / or the sound of insects.

[0110]

[0108] The source separator 26, or source separator blocks thereof 26a-b may like the source separator 22a be implemented using machine-learning techniques or a filter-based approach. Generally, it may be challenging to design a single source separation module 26a that reliably extracts a large variety of different height audio source types with different acoustic properties. For example, as will be described below, in a filter-based implementation, designing a single filter that covers all conceivable height audio source types (manmade and nature sounds associated as well as sounds associated with alive objects with height) without also covering one or more non-height audio object types (such as speech or traffic sounds) may be difficult. Thus, using multiple source separator modules 26a-b, wherein each source separator module 26a-b extracts a single specific type of height audio source (e.g. birdsong) or a group of specific types of height audio sources (e.g. all manmade sounds associated with height) with similar frequency characteristics, may facilitate more reliable and accurate extraction of various types of height audio objects with little or no erroneous extraction of non-height object types.

[0111]

[0109] The source separator modules 26a-b extract a respective height audio source type and each extracted type of height audio source is represented with an audio source signal comprising the sound of audio sources having the predetermined type. The audio source signals (each carrying a respective type of audio source) may optionally be provided to a height object processor 28 which may mix the audio source signals into a mixed audio height audio source signal and / or may apply a predetermined or adaptive gain for each of the audio source signals. The mixed height audio source signal may thus comprise a mix of at least two different types of height audio sources.

[0112] [HO] If only a single type of height audio source is extracted, the mixing can be skipped. Alternatively, the height object processor 28 may adjust only the gain of the single type of height audio source without performing any mixing.

[0113]

[0111] The mixed height audio source signal is provided to an immersive audio generator 30 (interchangeably “audio generator 30”) that generates an immersive audio stream IA2, wherein each height audio source signal (i.e., comprised in the mixed height audio source signal) is assigned to at least one height channel of a multi-channel immersive audio format. The mixed height audio source signal may be mixed or panned to at least one height channel of the multichannel immersive audio format. For example, if the multi-channel format is a 2.0.2 format, a 5.1.2 format or a 7.1.2 format the mixed height audio source signal (comprising a mix of two or more types of height audio objects) may be assigned to the 0.0.2 height channels.

[0114]

[0112] The height audio source signals extracted by the source separation modules 26a-b may alternatively be provided directly to the audio generator 30 which assigns the respective height audio source signals to one or more height channels of the multi-channel immersive audio format.

[0113] In an object-based immersive audio format, the audio generator 30 may instead be configured to include each height audio source signal in a height audio object (e.g. a respective height audio object), i.e. an audio object which is assigned a (respective) spatial location of nonzero height, for instance a predetermined non- zero height. For instance, an object based spatial tenderer may (e.g. during playback) render spatial audio objects to presentation channels based on the audio object’s spatial position. As an example, each audio object including a height audio source signal may be assigned a spatial location corresponding to a zenith elevation or 45° elevation so as to be perceived as coming from above the listener.

[0115]

[0114] Optionally, the mixed height audio source signal or height audio source signal(s) are provided to a cross-talk-cancellation module 29 which performs cross-talk-cancellation on the mixed height audio source signal or height audio source signal(s) to reduce or remove crosstalk between height audio sources.

[0116]

[0115] In some implementations, the source separator 26 may further be configured to extract one or more types of non-height audio objects. For example, a source separator module 26c may be configured to extract a non-height audio source of a predetermined non-height audio source type. The non-height audio source type may e.g. be speech, sounds associated with by alive objects associated with non-height (the sound of cats or the sound of dogs) or manmade sounds associated with non-height (e.g. traffic noise, background voices). The non-height audio source(s) are represented with respective non-height audio source signal(s) that optionally are provided to a non-height source processor or mixer 32 which mixes the non-height audio source signal(s) and / or applies predetermined or adaptive gains. In analogy with the discussion of multiple source separation modules 26a-b, the source separator 26 may be provided with two or more source separator modules, each configured to extract an audio source of a respective predetermined non-height audio source type.

[0117]

[0116] While separate extraction of non-height audio source types is associated with some benefits (such as allowing the relative reproduction level of non-height audio objects to be adjusted), this feature is optional. Thus, in some implementations, only one or more height audio source types are extracted and the non-height audio source type(s) are determined implicitly as any residual audio content that is present in the input audio stream A but not in any of the extracted height audio source types. The residual audio content may be provided directly to the audio generator 30 to be included in the immersive media stream IA2, e.g. in one or more audio objects (e.g. assigned predetermined spatial locations or being non-localized) and / or in one or more channels.

[0118]

[0117] As an example, if the audio stream A is recorded during an interview outdoors it may comprise a mix of the voice of the interviewer, the voice of the interviewee, voices from other people nearby, traffic noise, birdsong, the sound of an airplane passing overhead and different types of background noise, such as stationary white noise. In this case, for instance the birdsong and the sound of the airplane may be extracted as height audio sources and the other sources may define residual audio content.

[0119]

[0118] As mentioned above, the immersive audio generation block 22 implementing the visually-aided immersive audio generation approach, may optionally provide a residual stream R as input to the immersive audio generation block 24 implementing the height object-based immersive audio generation approach, as illustrated by the dashed line in Fig. 1. The residual stream R may correspond to a residual of the audio stream A remaining after performing the audio source separation (e.g., by the source separator 22a) on the audio stream A. The residual audio stream R may be determined by the block 22. The residual audio stream R may comprise a portion of the audio stream not extracted by the source separator 22a. That is, the residual audio stream R may comprise the portions of the audio stream not being extracted as an audio source by the source separator 22a. The residual stream R may for instance be determined by subtracting each extracted audio source from the audio stream A.

[0120]

[0119] The block 24 may subsequently perform audio source separation on the residual audio stream R using a source separation model 26 trained to extract an audio source of a predetermined height audio source type, such that a height audio source may be extracted from the residual audio stream R. The height audio source extraction may proceed as set out above with reference to Fig. 6.

[0121]

[0120] As further set out above, the source separator 26 may comprise a respective separate source separator sub-model or module 26a-b, each trained to extract a respective predetermined height audio source type, and optionally a source separator module 26c configured to extract a non-height audio source. The non-height audio source may be a predetermined non-height audio source type, or the non-height audio source type(s) may be determined implicitly as any residual audio content that is present in the residual audio stream R but not in any of the extracted height audio source types (i.e., the residual of the residual audio stream R).

[0122]

[0121] Accordingly, the block 24 (e.g., the source separator 24) may extract from the residual audio stream R one or more height audio sources, each of a respective height audio source type, and optionally a non-height audio source of a non-height audio source type. The block 24 may provide each extracted height audio source (separately, or as part of the mixed height audio source signal), and each non-height audio source to the immersive audio generator 22c of block 22. The block 22 may include each audio source received from the block 24 in a height audio object of the immersive audio stream, or in at least one height channel of the immersive audio stream IA1 (i.e., depending on whether the immersive audio generator 22c is configured to generate an object- or channel-based audio stream). By applying the height-object based immersive audio generation approach to a residual stream R from the immersive audio generation block 22, audio sources which the source separator 22a not is able to extract, may be processed to extract potential height audio sources therefrom, which accordingly may be included in the immersive audio stream to further contribute to an increased immersiveness. Referring again to the example of music content, the sound of an airplane, birdsong or fireworks at an outdoor scene of a music performance may be extracted from the residual stream R and included as a height object or in a height channel of the immersive audio stream IA1.

[0123]

[0122] The person skilled in the art realizes that the present invention by no means is limited to the embodiments and examples described above. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, while in the above, a system 1 comprising a cloud server 20 implementing both a visually-aided immersive audio generation block 22 and a height object-based immersive audio generation block 24, a cloud server may in other implementations implement only one of the immersive audio generation blocks 22, 24. In this case a content check block 21 may be omitted. As may be appreciated, in implementations wherein the cloud server 20 includes only a height object-based immersive audio generation block 24, upload of the video stream V’ by the user device 10 to the cloud server 20 may be optional. For instance, the user device 10 may upload a video stream V’ to the cloud server 20 in an implementation wherein the cloud server 20 is configured to combine the immersive audio stream IA2 generated by the immersive audio generation block 24 with the video stream V’ to generate an immersive media stream. In an implementation wherein the user device 10 is configured to generate an immersive media stream, the immersive audio stream IA2 may instead be downloaded from the cloud server 20 by the user device 10 wherein the user device 10 may combine the immersive audio stream IA2 with the video stream V.

[0124]

[0123] Fig. 7 shows a further example implementation of a cloud server 200 comprising a further visually-aided immersive audio generation block 222, in addition to blocks 22 and 24. The content check block 21 may in this case forward the video stream V’ and the audio stream A to either block 22 or block 222, or the audio stream A to block 24. The content check block 21 is here configured to forward the video stream V’ and the audio stream A to block 22 responsive to determining that the media stream M’ the video stream V’ and the audio stream A is of a first predetermined content type (e.g. music), and to block 222 responsive to determining that the media stream M’ the video stream V’ and the audio stream A is of a second predetermined content type (e.g. speech or dialogue content). Block 22 may accordingly comprise a source separator 22a trained to extract an audio source of at least one predetermined audio source type associated with the first predetermined content type (e.g. a vocal source type and one or more instruments source types), and an image recognition model 22b trained to identify visual objects of one or more types, each type corresponding to a respective one of the at least one predetermined audio source type the source separator 22a of block 22 is trained to extract (e.g. a singer or vocalist and visual objects corresponding to the one or more instrument source types). Block 222 may correspondingly comprise a source separator trained to extract an audio source of at least one predetermined audio source type associated with the second predetermined content type (e.g. speech or voices of one or more persons), and an image recognition model trained to identify visual objects corresponding to a respective one of the at least one predetermined audio source type the source separator of block 222 is trained to extract (e.g. faces with moving lips). This approach may be extended such that the cloud server may generate immersive audio streams for a plurality of different content types, using source separators and image recognition models specifically tailored to process a respective type of content. In analogy with the discussion of the residual audio stream R in connection with Fig. 1 , each of the blocks 22 and 222 may provide a residual audio stream R as input to block 24, to attempt extracting an audio source of at least one predetermined height audio source type therefrom.

[0125]

[0124] Furthermore, while in the above, the visually-aided immersive audio generation blocks 22, 222 has been discussed with reference to music-type content and speech- or dialoguecontent, it is also possible to apply the visually-aided immersive audio generation to other types of content. More generally, the visually-aided immersive audio generation may be used for any type audio- visual media content for which an audio source separation model and an image recognition model may be provided that are respectively capable of extracting and identifying one or more corresponding audio sources and visual objects (i.e. the visual object being a likely source of the corresponding extracted audio source).

[0126]

[0125] Example embodiments include the following enumerated example embodiments (“EEEs”):

[0127] EEE 1. A method for generating immersive media content, the method comprising: by a user device: capturing an audio stream of a scene; processing the audio stream to determine audio feature information of the audio stream; determining, based on the audio feature information, that a cloud server upload process is to be performed; and responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server; and by the cloud server: receiving the audio stream from the user device; performing audio source separation on the audio stream to extract at least one audio source from the audio stream; and generating an immersive audio stream comprising the at least one audio source, wherein each respective audio source of the at least one audio source is included in an audio object of the immersive audio stream and is assigned a spatial location based on an estimated location of the respective audio source in the scene, or is included in at least one channel of the immersive audio stream based on an estimated location of the respective audio source in the scene.

[0128] EEE 2. The method according to EEE 1, wherein the audio feature information comprises one or more of: a signal-to-noise metric of the audio stream, a reverberation time of the audio stream, presence of a height audio source in the audio stream, an audio content type of the audio stream, an acoustic environment of the audio stream

[0129] EEE 3. The method according to any one of EEEs 1-2, wherein the audio feature information comprises a signal-to-noise metric of the audio stream and wherein determining that the cloud server upload process is to be performed comprises determining that the signal-to-noise metric exceeds a signal-to-noise threshold.

[0130] EEE 4. The method according to EEE 3, wherein the audio feature information further comprises a classification of an acoustic environment of the audio stream as one of a predetermined set of acoustic environments, such as an indoor environment and an outdoor environment, wherein a value of the signal-to-noise threshold is determined in dependence on the classification of the acoustic environment of the audio stream.

[0131] EEE 5. The method according to any one of EEEs 1-4, wherein the audio feature information further comprises a reverberation time of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the reverberation time is less than a reverberation time threshold.

[0132] EEE 6. The method according to any one of the preceding EEEs, wherein the audio feature information further comprises an audio content type of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio content type matches a predetermined audio content type.

[0133] EEE 7. The method according to EEE 6, wherein the predetermined audio content type is music. EEE 8. The method according to any one of the preceding EEEs, wherein the audio feature information comprises an indication of presence of a height audio source in the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio feature information indicates that a height audio source is present in the audio stream.

[0134] EEE 9. The method according to any one of the preceding EEEs, wherein the user device captures a media stream of the scene, the media stream comprising a first video stream and the audio stream, and wherein the cloud server upload process further comprises uploading a second video stream to the cloud server, wherein the second video stream is the first video stream or a reduced bit rate version of the first video stream; and the method further comprising, by the cloud server: receiving the second video stream from the user device; determining a location of each of at least one visual object in the second video stream, wherein each visual object corresponds to a respective one of the at least one audio source, and wherein the estimated location of each respective audio source is based on the location of the corresponding visual object.

[0135] EEE 10. The method according to EEE 9, wherein the audio source separation is performed by a source separation model trained to extract an audio source of at least one predetermined audio source type, and wherein the location of each of the at least one visual object is determined by an image recognition model trained to identify a visual object of at least one predetermined visual object type.

[0136] EEE 11. The method according to EEE 10, wherein each of the at least one predetermined visual object type corresponds to a respective one of the at least one predetermined audio source type. EEE 12. The method according to any one of EEEs 10-11, wherein the at least one predetermined audio source type comprises at least one of a vocal source type, and one or more instrument source types; and wherein the at least one predetermined visual object type comprises at least one of: a singer, and one or more instrument types.

[0137] EEE 13. The method according to EEE 12, wherein the one or more instrument source types comprises at least one of: bass, drums, keyboards, piano, guitar, violin, brass, reed; and wherein the one or more instrument types comprises at least one of: bass, drums, keyboards, piano, guitar, violin, brass, reed.

[0138] EEE 14. The method according to any one of EEEs 9-13, wherein the first video stream is captured by a camera module of the user device, and the method further comprises, by the user device, obtaining focus distance information indicating a focus distance of the camera module during capturing the captured video stream; and uploading the focus distance information to the cloud server, wherein the estimated location of each respective audio source is based on the location of the corresponding visual object and on the focus distance information.

[0139] EEE 15. The method according to any one of EEEs 9-14, further comprising, by the cloud server: determining a residual audio stream comprising a portion of the audio stream not extracted by the audio source separation; performing audio source separation on the residual audio stream using a source separation model trained to extract an audio source of a predetermined height audio source type, wherein a height audio source is extracted from the residual audio stream; wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

[0140] EEE 16. The method according to any one of EEEs 1-8, wherein the audio source separation is performed using a source separation model trained to extract an audio source of a predetermined height audio source type, wherein a height audio source is extracted from the audio stream, and wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

[0141] EEE 17. The method according to any one of EEEs 15-16, wherein the height audio source type comprises at least one of manmade sounds associated with height, sounds made by alive objects associated with height and sounds of nature associated with height.

[0142] EEE 18. The method according to EEE 17, wherein the manmade sounds associated with height comprises at least one of: the sound of an airplane, the sound of a helicopter, the sound of a drone, the sound of a ceiling fan, the sound of ventilation, and the sound of fireworks, and / or wherein the sounds made by alive objects associated with height comprises at least one of: birdsong, the sound of insects, and the sound of animals known to spend time in tree canopies, and / or wherein the sounds of nature associated with height comprises thunder, the sound of wind.

[0143] EEE 19. The method according to any one of EEEs 16-18, wherein the audio source separation is performed using at least two source separation models, to extract at least two height audio sources, wherein the at least two height audio sources are extracted using a respective one of least two source separation models, each configured to extract a respective height audio source of a respective height audio source type; wherein each of the at least two height audio sources is included in a height audio object of the immersive audio stream, or is included in at least one height channel of the immersive audio stream.

[0144] EEE 20. The method according to any one of the preceding EEEs, wherein the user device captures a media stream of the scene, the media stream comprising a first video stream and the audio stream, and wherein the cloud server upload process comprises uploading the audio stream and a second video stream to the cloud server, wherein the second video stream is the first video stream or a reduced bit rate version of the first video stream; and the method further comprising, by the cloud server: receiving the second video stream from the user device; determining whether the media stream is of a predetermined content type; and responsive to determining that the media stream is of the predetermined content type: determining a location of each of at least one visual object in the video stream, wherein each visual object corresponds to a respective one of the at least one audio source, and wherein the estimated location of each respective audio source is based on the location of the corresponding visual object; or responsive to determining that the media stream is not of the predetermined content type: performing the audio source separation using a source separation model trained to extract an audio source of a predetermined height audio source type, to extract a height audio source from the audio stream, wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

[0145] EEE 21. The method according to EEE 20, wherein the predetermined content type is music content.

[0146] EEE 22. The method according to any one of the preceding EEEs, further comprising generating, by the cloud server, an immersive media stream by combining the immersive audio stream and the second video stream, or, by the user device: downloading the immersive audio stream from the cloud server and generating an immersive media stream by combining the immersive audio stream and the first video stream. EEE 23. The method according to any one of the preceding EEEs, wherein the user device comprises a camera module for capturing the video stream and a set of microphones for capturing the audio stream.

[0147] EEE 24. The method according to any one of the preceding EEEs, wherein the user device is a portable electronic device, such as a smart phone.

[0148] EEE 25. The method according to any one of the preceding EEEs, wherein at least two audio sources are extracted from the audio stream, and wherein the immersive audio stream comprises a mix of the at least two audio sources.

[0149] EEE 26. A method for generating immersive media content, the method comprising: by a user device: obtaining an audio stream of a scene; processing the audio stream to determine audio feature information of the audio stream; determining, based on the audio feature information, that a cloud server upload process is to be performed; responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server configured to generate an immersive audio stream based on an uploaded audio stream; and receiving from the cloud server, an immersive audio stream.

[0150] EEE 27. The method according to EEE 26, wherein the audio feature information comprises one or more of: a signal-to-noise metric of the audio stream, a reverberation time of the audio stream, presence of a height audio source in the audio stream, an audio content type of the audio stream, an acoustic environment of the audio stream.

[0151] EEE 28. The method according to any one of EEEs 26-27, wherein the audio feature information comprises a signal-to-noise metric of the audio stream and wherein determining that the cloud server upload process is to be performed comprises determining that the signal-to-noise metric exceeds a signal-to-noise threshold.

[0152] EEE 29. The method according to EEE 28, wherein the audio feature information further comprises a classification of an acoustic environment of the audio stream as one of a predetermined set of acoustic environments, such as an indoor environment and an outdoor environment, wherein a value of the signal-to-noise threshold is determined in dependence on the classification of the acoustic environment of the audio stream.

[0153] EEE 30. The method according to any one of EEEs 26-29, wherein the audio feature information further comprises a reverberation time of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the reverberation time is less than a reverberation time threshold.

[0154] EEE 31. The method according to any one of EEEs 26-30, wherein the audio feature information further comprises an audio content type of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio content type matches a predetermined audio content type.

[0155] EEE 32. The method according to EEE 31, wherein the predetermined audio content type is music.

[0156] EEE 33. The method according to any one of EEEs 26-32, wherein the audio feature information comprises an indication of presence of a height audio source in the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio feature information indicates that a height audio source is present in the audio stream. EEE 34. A user device comprising one or more processors configured to carry out the method according to EEE 26-33.

[0157] EEE 35. A method for generating immersive media content, the method comprising: obtaining a media stream comprising a video stream and an audio stream; performing audio source separation on the audio stream to extract at least one audio source from the audio stream; determining a location of each of at least one visual object in the video stream, wherein each visual object corresponds to a respective one of the audio sources; and generating an immersive audio stream comprising the at least one audio source, wherein each respective audio source is included in an audio object of the immersive audio stream and is assigned a spatial location based on the location of the corresponding visual object, or is included in at least one channel of the immersive audio stream based on the location of the corresponding visual object.

[0158] EEE 36. The method according to EEE 35, wherein the audio source separation is performed by a source separation model trained to extract an audio source of at least one predetermined audio source type, and wherein the location of each of the at least one visual object is determined by an image recognition model trained to identify a visual object of at least one predetermined visual object type.

[0159] EEE 37. The method according to EEE 36, wherein each of the at least one predetermined visual object type corresponds to a respective one of the at least one predetermined audio source type. EEE 38. The method according to any one of EEEs 36-37, wherein the at least one predetermined audio source type comprises at least one of a vocal source type, and one or more instrument source types; and wherein the at least one predetermined visual object type comprises at least one of: a singer, and one or more instrument types.

[0160] EEE 39. The method according to EEE 38, wherein the one or more instrument source types comprises at least one of: bass, drums, keyboards, piano, guitar, violin, brass, reed; and wherein the one or more instrument types comprises at least one of: bass, drums, keyboards, piano, guitar, violin, brass, reed.

[0161] EEE 40. The method according to any one of EEEs 35-39, obtaining focus distance information for the video stream, wherein the estimated location of each respective audio source is based on the location of the corresponding visual object and on the focus distance information.

[0162] EEE 41. The method according to any one of EEEs 35-40, further comprising, by the cloud server: determining a residual audio stream comprising a portion of the audio stream not extracted by the audio source separation; performing audio source separation on the residual audio stream using a source separation model trained to extract an audio source of a predetermined height audio source type, wherein a height audio source is extracted from the residual audio stream; wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

[0163] EEE 42. A cloud server comprising one or more processors configured to carry out the method according to EEE 35-41.

[0164] EEE 43. A computer program product comprising instructions which, when the program is executed by a processing device, causes the processing device to carry out the method according to any one of EEEs 1-33 or 35-41.

[0165] EEE 44. A computer-readable non-transitory storage medium storing the computer program product according to EEE 43.

[0166] EEE 45. A system comprising a user device and a cloud server, each comprising one or more processors and configured to carry out the method according to any of EEEs 1-25.

Claims

CLAIMS1. A method for generating immersive media content, the method comprising: by a user device: obtaining an audio stream of a scene; processing the audio stream to determine audio feature information comprising one or more of: a signal-to-noise metric of the audio stream, a reverberation time of the audio stream, presence of a height audio source in the audio stream, an audio content type of the audio stream, an acoustic environment of the audio stream; determining, based on the audio feature information, that a cloud server upload process is to be performed; and responsive to determining that the cloud server upload process is to be performed, performing the cloud server upload process, wherein the cloud server upload process comprises uploading the audio stream to a cloud server; and by the cloud server: receiving the audio stream from the user device; performing audio source separation on the audio stream to extract at least one audio source from the audio stream; and generating an immersive audio stream comprising the at least one audio source, wherein each respective audio source of the at least one audio source is included in an audio object of the immersive audio stream and is assigned a spatial location based on an estimated location of the respective audio source in the scene, or is included in at least one channel of the immersive audio stream based on an estimated location of the respective audio source in the scene.

2. The method according to claim 1 , wherein the audio feature information comprises a signal- to-noise metric of the audio stream and wherein determining that the cloud server upload process is to be performed comprises determining that the signal-to-noise metric exceeds a signal-to- noise threshold.

3. The method according to claim 2, wherein the audio feature information further comprises a classification of an acoustic environment of the audio stream as one of a predetermined set of acoustic environments, such as an indoor environment and an outdoor environment, wherein a value of the signal-to-noise threshold is determined in dependence on the classification of the acoustic environment of the audio stream.

4. The method according to any one of claims 1 -3, wherein the audio feature information further comprises a reverberation time of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the reverberation time is less than a reverberation time threshold.

5. The method according to any one of the preceding claims, wherein the audio feature information further comprises an audio content type of the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio content type matches a predetermined audio content type.

6. The method according to any one of the preceding claims, wherein the audio feature information comprises an indication of presence of a height audio source in the audio stream, and wherein determining that the cloud server upload process is to be performed comprises determining that the audio feature information indicates that a height audio source is present in the audio stream.

7. The method according to any one of the preceding claims, wherein the user device obtains a media stream of the scene, the media stream comprising a first video stream and the audio stream, and wherein the cloud server upload process further comprises uploading a second video stream to the cloud server, wherein the second video stream is the first video stream or a reduced bit rate version of the first video stream; and the method further comprising, by the cloud server: receiving the second video stream from the user device; determining a location of each of at least one visual object in the second video stream, wherein each visual object corresponds to a respective one of the at least one audio source, and wherein the estimated location of each respective audio source is based on the location of the corresponding visual object.

8. The method according to claim 7, wherein the audio source separation is performed by a source separation model trained to extract an audio source of at least one predetermined audio source type, and wherein the location of each of the at least one visual object is determined by an image recognition model trained to identify a visual object of at least one predetermined visualobject type, and optionally wherein each of the at least one predetermined visual object type corresponds to a respective one of the at least one predetermined audio source type.

9. The method according to claim 7 or claim 8, wherein the first video stream is captured by a camera module of the user device, and the method further comprises, by the user device, obtaining focus distance information indicating a focus distance of the camera module during capturing the captured video stream; and uploading the focus distance information to the cloud server, wherein the estimated location of each respective audio source is based on the location of the corresponding visual object and on the focus distance information.

10. The method according to any one of claims 7-9, further comprising, by the cloud server: determining a residual audio stream comprising a portion of the audio stream not extracted by the audio source separation; performing audio source separation on the residual audio stream using a source separation model trained to extract an audio source of a predetermined height audio source type, wherein a height audio source is extracted from the residual audio stream; wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

11. The method according to any one of claims 1-6, wherein the audio source separation is performed using a source separation model trained to extract an audio source of a predetermined height audio source type, wherein a height audio source is extracted from the audio stream, and wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

12. The method according to 10 or claim 11, wherein the height audio source type comprises at least one of manmade sounds associated with height, sounds made by alive objects associated with height and sounds of nature associated with height.

13. The method according to any of the preceding claims,wherein the audio source separation is performed using at least two source separation models, to extract at least two height audio sources, wherein the at least two height audio sources are extracted using a respective one of least two source separation models, each configured to extract a respective height audio source of a respective height audio source type; wherein each of the at least two height audio sources is included in a height audio object of the immersive audio stream, or is included in at least one height channel of the immersive audio stream.

14. The method according to any one of the preceding claims, wherein the user device obtains a media stream of the scene, the media stream comprising a first video stream and the audio stream, and wherein the cloud server upload process comprises uploading the audio stream and a second video stream to the cloud server, wherein the second video stream is the first video stream or a reduced bit rate version of the first video stream; and the method further comprising, by the cloud server: receiving the second video stream from the user device; determining whether the media stream is of a predetermined content type; and responsive to determining that the media stream is of the predetermined content type: determining a location of each of at least one visual object in the video stream, wherein each visual object corresponds to a respective one of the at least one audio source, and wherein the estimated location of each respective audio source is based on the location of the corresponding visual object; or responsive to determining that the media stream is not of the predetermined content type: performing the audio source separation using a source separation model trained to extract an audio source of a predetermined height audio source type, to extract a height audio source from the audio stream, wherein the height audio source is included in a height audio object of the immersive audio stream, or wherein the height audio source is included in at least one height channel of the immersive audio stream.

15. The method according to any one of the preceding claims, further comprising generating, by the cloud server, an immersive media stream by combining the immersive audio stream and thevideo stream, or, by the user device: downloading the immersive audio stream from the cloud server and generating an immersive media stream by combining the immersive audio stream and the video stream.

Citation Information

Patent Citations

  • Processing multiple spatial audio signals which have a spatial overlap

    EP3706432A1

  • Portable communication terminal, communication method and control program

    US20110138301A1

  • Electronic apparatus and control method thereof

    US20170201947A1

  • Network-based processing and distribution of multimedia content of a live musical performance

    US20210204003A1

  • Systems and methods for capturing and processing user consumption of information

    US20230061646A1