Processing, transmission method and device for remote conferencing

By introducing a network-based media processing server in the remote conferencing system, using audio weighting and overlay priority technology, the problems of low audio mixing efficiency and poor quality in the existing technology are solved, and efficient and high-quality audio mixing effects are achieved, improving the user experience.

CN114667727BActive Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180006331.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2021-06-22
Publication Date
2025-07-25
Estimated Expiration
2041-06-22

AI Technical Summary

Technical Problem

Existing remote conferencing systems have problems with inefficiency and poor audio quality in audio mixing and immersive media processing, especially when multimedia streams are superimposed and mixed, which are difficult to achieve high-quality audio output.

Method used

By introducing a network-based media processing server in a remote conferencing system, using audio weighting and overlay priority technologies, combining immersive and overlay media streams for mixing, high-quality mixed audio is generated, and audio mixing is performed through end user equipment or servers.

Benefits of technology

It realizes efficient and excellent audio mixing in remote meetings, reducing the processing burden of user equipment, and improving the clarity and user experience of audio output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114667727B_ABST
    Figure CN114667727B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and apparatus for remote conferencing. In some examples, the apparatus for remote conferencing includes processing circuitry. The processing circuitry of a first device receives a first media stream carrying a first audio and a second media stream carrying a second audio. The processing circuitry receives a first audio weight for weighting the first audio and a second audio weight for weighting the second audio, and generates a mixed audio by combining the weighted first audio based on the first audio weight and the weighted second audio based on the second audio weight.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 327,400, entitled "METHOD AND APPARATUS FOR TELECONFERENCE", filed on May 21, 2021, which claims the benefit of priority of U.S. Provisional Application No. 63 / 088,300, entitled "NETWORK BASED MEDIA PROCESSING FOR AUDIO AND VIDEO MIXING FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS", filed on October 6, 2020, and U.S. Provisional Application No. 63 / 124,261, entitled "AUDIO MIXING METHODS FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS", filed on December 11, 2020. The entire contents of these patent applications are incorporated herein by reference. Technical Field

[0003] Embodiments generally related to teleconferencing are described herein. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present application. To the extent that the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, has been carried out, it does not indicate that it was prior art at the time of filing of the present application, and it has never been expressly or implicitly admitted to be prior art of the present application.

[0005] Teleconference systems allow users at two or more remote locations to communicate interactively with each other via media streams such as video streams, audio streams, or both video and audio streams. Some teleconference systems also allow users to exchange digital documents such as images, text, video, applications, etc. Summary of the Invention

[0006] Aspects of the present invention provide methods and apparatuses for processing and transmitting remote conferences. In some examples, an apparatus for a remote conference includes processing circuitry. The processing circuitry of a first device (e.g., a user device or a server for network-based media processing) receives a first media stream carrying a first audio and a second media stream carrying a second audio from a second device. The processing circuitry receives a first audio weight for weighting the first audio and a second audio weight for weighting the second audio from the second device, and generates a mixed audio by combining the weighted first audio based on the first audio weight and the weighted second audio based on the second audio weight.

[0007] In some embodiments, the first device is a user device. The first device may play the mixed audio through a speaker associated with the first device.

[0008] In one example, the first device sends customization parameters to the second device for the second device to customize the first audio weight and the second audio weight based on the customization parameters.

[0009] In some examples, the second device determines the first audio weight and the second audio weight based on the sound intensities of the first audio and the second audio.

[0010] In some examples, the first audio and the second audio are superimposed audio, and the processing circuitry receives the first audio weight and the second audio weight determined by the second device based on the superimposition priorities of the first audio and the second audio.

[0011] In some examples, the second device adjusts the first audio weight and the second audio weight based on the detection of an active speaker.

[0012] In some examples, the first media stream includes immersive media content, the second media stream includes superimposed media content, and the first audio weight is different from the second audio weight.

[0013] In some embodiments, the first device is a network-based media processing device. The processing circuitry encodes the mixed audio into a third media stream and transmits the third media stream to a user device via an interface circuit of the device. In some examples, the processing circuitry sends a fourth media stream including immersive media content and the third media stream via the interface circuit. The third media stream is a superimposition of the fourth media stream.

[0014] According to some aspects of the present application, the processing circuitry of a first device (e.g., a server device for network-based media processing) receives a first media stream carrying first media content of a remote conference session and a second media stream carrying second media content of the remote conference session. The processing circuitry generates third media content by mixing the first media content and the second media content; and transmits a third media stream carrying the third media content to a second device via a sending circuit.

[0015] In some embodiments, the processing circuitry of the first device mixes a first audio in a first media content with a second audio in a second media content to generate a third audio based on a first audio weight assigned to the first audio and a second audio weight assigned to the second audio. In some examples, the first audio weight and the second audio weight are received from a host device that transmits the first media stream and the second media stream. In some examples, the first device may determine the first audio weight and the second audio weight.

[0016] In some examples, the first media stream is an immersive media stream, the second media stream is an overlay media stream, and the processing circuitry of the first device mixes the first audio and the second audio based on the first audio weight and the second audio weight having different values.

[0017] In some examples, the first media stream and the second media stream are overlay media streams, and the processing circuitry of the first device mixes the first audio and the second audio based on the first audio weight and the second audio weight having the same value.

[0018] In some examples, the first media stream and the second media stream are overlay media streams, and the processing circuitry of the first device mixes the first audio and the second audio based on the first audio weight and the second audio weight associated with the overlay priority of the first media stream and the second media stream.

[0019] According to some aspects of the present application, a first device (e.g., a host device that generates immersive media content) transmits a first media stream carrying a first audio and a second media stream carrying a second audio to a second device. The first device may determine a first audio weight for weighting the first audio and a second audio weight for weighting the second audio, and transmit the first audio weight and the second audio weight for mixing the first audio and the second audio to the second device.

[0020] In some examples, the first device receives customization parameters based on the Session Description Protocol and determines the first audio weight and the second audio weight based on the customization parameters.

[0021] In some examples, the first device determines the first audio weight and the second audio weight based on the sound intensities of the first audio and the second audio.

[0022] In some examples, the first audio and the second audio are overlay audio, and the first device determines the first audio weight and the second audio weight based on the overlay priority of the first audio and the second audio.

[0023] In some examples, the first device determines the first audio weight and the second audio weight based on the detection of an active speaker in one of the first audio and the second audio.

[0024] In some examples, the first media stream includes immersive media content and the second media stream includes overlay media content. The first device determines different values for the first audio weight and the second audio weight.

[0025] Aspects of the present invention also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for a remote conference, cause the computer to perform a method for a remote conference. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0027] Figure 1 A remote conference system according to some examples of the present application is shown.

[0028] Figure 2 Another remote conference system according to some examples of the present application is shown.

[0029] Figure 3 Another remote conference system according to some examples of the present application is shown.

[0030] Figure 4 A flowchart outlining a process according to some examples of the present application is shown.

[0031] Figure 5 A flowchart outlining a process according to some examples of the present application is shown.

[0032] Figure 6 A flowchart outlining a process according to some examples of the present application is shown.

[0033] Figure 7 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0034] Aspects of the present application provide media mixing techniques, such as audio mixing, video mixing, and similar techniques for remote conferences. In some examples, the remote conference can be an audio remote conference, and the participants in the remote conference communicate via an audio stream. In some examples, the remote conference is a video conference, and the participants in the remote conference can communicate via a media stream including video and / or audio. In some examples, media mixing is performed by a network-based media processing element (such as a server device, etc.). In some examples, media mixing is performed by an end-user device (also referred to as a user device).

[0035] According to some aspects of the present application, media mixing techniques can be performed in various remote conference systems. Figures 1 to 3 Some remote conference systems are shown.

[0036] Figure 1 FIG. Figure 1 shows a remote conferencing system (100) according to some examples of the present application. The remote conferencing system (100) includes a subsystem (110) and a plurality of user devices, such as user device (120) and user device (130). The subsystem (110) is installed at a location, such as Conference Room A. Generally, the subsystem (110) is configured to have a relatively higher bandwidth compared to the user devices (120) and (130) and can provide host services for a remote conferencing session (also referred to as a remote teleconference). The subsystem (110) can enable users or participants in Conference Room A to participate in the remote conferencing session and enable some remote users (such as User B of user device (120) and User C of user device (130)) to participate in the remote conferencing session from remote locations. In some examples, the subsystem (110), the user device (120), and the user device (130) are referred to as terminals in the remote conferencing session.

[0037] In some embodiments, the subsystem (110) includes various audio components, video components, and control components suitable for a conference room. The various audio components, video components, and control components can be integrated into one device, or the various audio components, video components, and control components can also be distributed components coupled together via appropriate communication technologies. In some examples, the subsystem (110) includes a wide-angle camera (111), such as a fisheye camera, an omnidirectional camera, and similar devices with a relatively wide field of view. For example, an omnidirectional camera can be configured to have a field of view that approximately covers the entire range, and the video captured by the omnidirectional camera can be referred to as omnidirectional video or 360-degree video.

[0038] Furthermore, in some examples, the subsystem (110) includes a microphone (112), such as an omnidirectional (also referred to as non-directional) microphone that can capture sound waves from approximately any direction. The subsystem (110) can include a display screen (114), speaker devices, etc., enabling users in Conference Room A to play multimedia corresponding to the video and audio of users located outside Conference Room A. In one example, the speaker device can be integrated with the microphone (112) or can be a separate component (not shown).

[0039] In some examples, the subsystem (110) includes a controller (113). Although Figure 1 the laptop computing device shown in FIG. Figure 1 is used as the controller (113), other suitable devices (such as a desktop computer, a tablet computer, etc.) can also be used as the controller (113). It should be noted that, in one example, the controller (113) can be integrated with other components in the subsystem (110).

[0040] The controller (113) can be configured to perform various control functions of the subsystem (110). For example, the controller (113) can be used to initiate a remote conference session and manage the communication between the subsystem (110) and the user devices (120) and (130). In one example, the controller (113) can encode the audio and / or video collected in Conference Room A (e.g., collected by the camera (111) and the microphone (112)) to generate a media stream to carry the audio and / or video, and can cause the media stream to be transmitted to the user devices (120) and (130).

[0041] Further, in some examples, the controller (113) can receive media streams carrying the audio and / or video collected on each user device in the remote conference system (100) (e.g., the user devices (120) and (130)). The controller (113) can address the received media streams and send the received media streams to other user devices in the remote conference system (100). For example, the controller (113) can receive a media stream from the user device (120), address the media stream and send the media stream to the user device (130), and can receive another media stream from the user device (130), address the other media stream and send the other media stream to the user device (120).

[0042] Further, in some examples, the controller (113) can determine appropriate remote conference parameters, such as audio and video mixing parameters, etc., and send the remote conference parameters to the user devices (120) and (130).

[0043] In some examples, the controller (113) can cause a user interface to be displayed on a screen (e.g., on the display screen (114), the screen of a laptop computing device, etc.) to facilitate user input in Conference Room A.

[0044] Each of user device (120) and user device (130) can be any suitable remote conferencing enabled device, such as a desktop computer, laptop computer, tablet computer, wearable device, handheld device, smart phone, mobile device, embedded device, game console, gaming device, personal data assistant (PDA), telecommunication device, global positioning system (GPS) device, virtual reality (VR) device, augmented reality (AR) device, implantable computing device, in-vehicle computer, network-enabled TV, Internet of Thing (IoT) device, workstation, media player, personal video recorder (PVR), set-top box, camera, integrated components (such as peripherals) included in a computing device, application or any other type of computing device.

[0045] In Figure 1 the example, user device (120) includes a wearable multimedia component to allow a user (such as user B) to participate in a remote conferencing session. For example, user device (120) includes a head-mounted display (HMD) that can be worn on the head of user B. The HMD can include display optics in front of one or both eyes of user B to play videos. In another example, user device (120) includes a headset (not shown) that can be worn by user B. The headset can include a microphone for collecting the user's voice and one or two earphones for outputting audio sounds. User device (120) also includes appropriate communication components (not shown) that can send and / or receive media streams.

[0046] In Figure 1 the example, user device (130) can be a mobile device, such as a smart phone and similar devices that integrate communication components, imaging components, audio components, etc., to allow a user (such as user C) to participate in a remote conferencing session.

[0047] In Figure 1 the example, subsystem (110), user device (120) and user device (130) include appropriate communication components (not shown) that can interact with network (101). The communication components can include one or more network interface controllers (NIC) or other types of transceiver circuits for sending and receiving communications and / or data over a network (such as network (101)).

[0048] The network (101) may include, for example, a public network (such as the Internet), a private network (such as an intranet), and / or a personal area network, or some combination of a private network and a public network. The network (101) may also include any type of wired network and / or wireless network, including but not limited to a local area network ("LAN"), a wide area network ("WAN"), a satellite network, a wired network, a Wi-Fi network, a WiMax network, a mobile communication network (such as 3G, 4G, 5G, etc.), or any combination thereof. The network (101) may utilize communication protocols, which include packet-based and / or datagram-based protocols, such as the Internet protocol ("IP"), the transmission control protocol ("TCP"), the user datagram protocol ("UDP"), or other types of protocols. Further, the network (101) may also include some devices that facilitate network communication and / or form the hardware basis of the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, backbone devices, etc. In some examples, the network (101) may also include devices that can connect to a wireless network, such as a wireless access point ("WAP").

[0049] In Figure 1 the example, the subsystem (110) may host a remote conference session using peer-to-peer technology. For example, after the user device (120) joins the remote conference session, the user device (120) may appropriately address the packets and transmit the packets to the subsystem (110) (e.g., using the IP address of the subsystem (110)), and the subsystem (110) may appropriately address the packets (e.g., using the IP address of the user device (120)) and transmit the packets to the user device (120). The packets may carry various information and data, such as media streams, acknowledgments, control parameters, etc.

[0050] In some examples, a remote conferencing system (100) can provide a remote conferencing session for an immersive remote conference. For example, during a remote conferencing session, the subsystem (110) is configured to generate immersive media, such as omnidirectional video / audio, using an omnidirectional camera and / or an omnidirectional microphone. In one example, an HMD in the user device (120) can detect the head movement of User B and determine the viewport direction of User B based on the head movement. The user device (120) can send the viewport direction of User B to the subsystem (110), and in turn, the subsystem (110) can send a viewport-related stream for playback on the user device (120), such as a video stream customized based on the viewport direction of User B (a media stream carrying a video customized based on the viewport direction of User B), an audio stream customized based on the viewport direction of User B (a media stream carrying an audio customized based on the viewport direction of User B), etc.

[0051] In another example, User C can use the user device (130) (e.g., using the touch screen of a smart phone) to enter the viewport direction of User C. The user device (130) can send the viewport direction of User C to the subsystem (110), and in turn, the subsystem (110) can send a viewport-related stream for playback on the user device (130), such as a video stream customized based on the viewport direction of User C (a media stream carrying a video customized based on the viewport direction of User C), an audio stream customized based on the viewport direction of User C (a media stream carrying an audio customized based on the viewport direction of User C), etc.

[0052] It should be noted that during a remote conferencing session, the viewport direction of User B and / or User C may change. The change in the viewport direction can be notified to the subsystem (110), and the subsystem (110) can adjust the viewport direction in the viewport-related streams sent to the user device (120) and the user device (130) respectively.

[0053] For ease of description, immersive media is used to refer to wide-angle media, such as omnidirectional video, omnidirectional audio, etc., and immersive media is used to refer to viewport-related media generated based on wide-angle media. It should be noted that in this application, 360-degree media, such as 360-degree video, 360-degree audio, etc., is used to illustrate the technology for remote conferencing, and the remote conferencing technology can be used for immersive media with less than 360 degrees.

[0054] Figure 2Another remote conferencing system (200) according to some examples of the present application is shown. The remote conferencing system (200) includes a plurality of subsystems (such as subsystems (210A) to subsystems (210Z) respectively installed in Conference Rooms A to Z), and a plurality of user devices (such as user device (220) and user device (230)). One of the subsystems (210A) to subsystems (210Z) can initiate a remote conferencing session and enable other subsystems and user devices (such as user device (220) and user device (230)) to join the remote conferencing session, so that users can participate in the remote conferencing session. For example, users in Conference Rooms A to Z, user B of user device (220), and user C of user device (230) can participate in the remote conferencing session. In some examples, the subsystems (210A) to subsystems (210Z), as well as user device (220) and user device (230) are referred to as terminals in the remote conference.

[0055] In some embodiments, the operation of each of the subsystems (210A) to subsystems (210Z) is similar to the operation of the above-mentioned subsystem (110). Further, some components used by each of the subsystems (210A) to subsystems (210Z) are the same as or equivalent to the components used in subsystem (110); descriptions of these components have been provided above, and for the sake of clarity, the descriptions of these components will be omitted here. It should be noted that the subsystems (210A) to subsystems (210Z) can be configured differently from each other.

[0056] The configuration of user device (220) and user device (230) is similar to the configuration of the above-mentioned user device (120) and user device (130), and the configuration of network (201) is similar to the configuration of network (101). Descriptions of these components have been provided above, and for the sake of clarity, the descriptions of these components will be omitted here.

[0057] In some embodiments, one of the subsystems (210A) to subsystems (210Z) can initiate a remote conferencing session, while the other subsystems of the subsystems (210A) to subsystems (210Z) and user device (220) and user device (230) can join the remote conferencing session.

[0058] According to one aspect of the present application, during a remote meeting session of an immersive remote meeting, multiple subsystems among subsystems (210A) to subsystems (210Z) can generate respective immersive media, and user device (220) and user device (230) can select one of the subsystems (210A) to subsystems (210Z) to provide immersive media. Generally, subsystems (210A) to subsystems (210Z) are configured to have a relatively high bandwidth and can each act as a host to provide immersive media.

[0059] In one example, after user device (220) joins the remote meeting session, user device (220) can select one of the subsystems (210A) to subsystems (210Z), such as subsystem (210A), as the host of the immersive media. User device (220) can address the packets and transmit the packets to subsystem (210A), and subsystem (210A) can address the packets and transmit the packets to user device (220). The packets can contain any suitable information / data, such as media streams, control parameters, etc. In some examples, subsystem (210A) can send customized media information to user device (220). It should be noted that user device (220) can change the selection of subsystems (210A) to subsystems (210Z) during the remote meeting session.

[0060] In one example, the HMD in user device (220) can detect the head movement of user B and determine the viewport direction of user B based on the head movement. User device (220) can send the viewport direction of user B to subsystem (210A), and in turn, subsystem (210A) sends a viewport-related media stream for playback on user device (220), such as a video stream customized based on the viewport direction of user B, an audio stream customized based on the viewport direction of user B, etc.

[0061] In another example, after user device (230) joins the remote meeting session, user device (230) can select one of the subsystems (210A) to subsystems (210Z), such as subsystem (210Z), as the host of the immersive media. User device (230) can address the packets and transmit the packets to subsystem (210Z), and subsystem (210Z) can address the packets and transmit the packets to user device (230). The packets can contain any suitable information / data, such as media streams, control parameters, etc. In some examples, subsystem (210Z) can send customized media information to user device (230). It should be noted that user device (230) can change the selection of subsystems (210A) to subsystems (210Z) during the remote meeting session.

[0062] In another example, User C can use the user device (230) (e.g., using the touch screen of a smart phone) to enter User C's viewport direction. The user device (230) can send User C's viewport direction to the subsystem (210Z), and in turn, the subsystem (210Z) sends a viewport-related media stream for playing on the user device (230), e.g., a video stream customized based on User C's viewport direction, an audio stream customized based on User C's viewport direction, etc.

[0063] It should be noted that during a remote conference session, the viewport directions of users (such as User B and User C) may change. For example, User B can notify the selected subsystem of the change in User B's viewport direction, and correspondingly, the selected subsystem selected by User B can adjust the viewport direction in the viewport-related stream sent to the user device (220).

[0064] For ease of description, immersive media is used to refer to wide-angle media, such as omnidirectional video, omnidirectional audio, etc., and immersive media is used to refer to viewport-related media generated based on wide-angle media. It should be noted that in this application, 360-degree media (such as 360-degree video, 360-degree audio, etc.) is used to illustrate the technology for remote conferencing, and the remote conferencing technology can be used for immersive media with less than 360 degrees.

[0065] Figure 3 Another remote conference system (300) according to some examples of the present application is shown. The remote conference system (300) includes a network-based media processing server (340), a plurality of subsystems (such as subsystems (310A) to (310Z) respectively installed in Conference Room A to Conference Room Z), and a plurality of user devices (such as user device (320) and user device (330)). The network-based media processing server (340) can set up a remote conference session and enable the subsystems (310A) to (310Z) and user devices (such as user device (320) and user device (330)) to join the remote conference session, so that users can participate in the remote conference session. For example, users in Conference Room A to Conference Room Z, User B of user device (320), and User C of user device (330) can participate in the remote conference session.

[0066] In some examples, subsystems (310A) to (310Z), as well as user equipment (320) and user equipment (330) are referred to as endpoints in a teleconference session, and a network-based media processing server (340) can bridge the endpoints in the teleconference session. In some examples, the network-based media processing server (340) is referred to as a media-aware network element. The network-based media processing server (340) can perform both media resource functions (MRF) and media control functions as a media control unit (MCU).

[0067] In some embodiments, the operation of each of the subsystems (310A) to (310Z) is similar to the operation of the above-described subsystem (110). Further, some of the components used by each of the subsystems (310A) to (310Z) are the same as or equivalent to those used in the subsystem (110); the descriptions of these components have been provided above and will be omitted here for clarity. Note that the subsystems (310A) to (310Z) can be configured differently from each other.

[0068] The configuration of the user equipment (320) and the user equipment (330) is similar to the configuration of the above-described user equipment (120) and user equipment (130), and the configuration of the network (301) is similar to the configuration of the network (101). The descriptions of these components have been provided above and will be omitted here for clarity.

[0069] In some examples, a network-based media processing server (340) can initiate a remote conference session. For example, one of subsystems (310A) to (310Z), as well as user device (320) and user device (330) can access the network-based media processing server (340) to initiate a remote conference session. Subsystems (310A) to (310Z), as well as user device (320) and user device (330) can join the remote conference session. Further, the network-based media processing server (340) is configured to provide media-related functions to bridge terminals in the remote conference session. For example, subsystems (310A) to (310Z) can each address packets carrying their respective media information such as video, audio, etc., and transmit these packets to the network-based media processing server (340). It should be noted that the media information sent to the network-based media processing server (340) is viewport-independent. For example, subsystems (310A) to (310Z) can send their respective videos (e.g., an entire 360-degree video) to the network-based media processing server (340). Further, the network-based media processing server (340) can receive viewport directions from user device (320) and user device (330), perform media processing to customize the media, and send the customized media information to the corresponding user devices.

[0070] In one example, a user device (320) joins a remote conference session. The user device (320) can be addressed by groups and transmit packets to a network-based media processing server (340), and the network-based media processing server (340) can be addressed by groups and transmit packets to the user device (320). The packets can contain any suitable information / data, such as media streams, control parameters, etc. In one example, User B can use the user device (320) to select a meeting room to view the video from a subsystem in that meeting room. For example, User B can use the user device (320) to select Meeting Room A to view the captured video from the subsystem (310A) installed in Meeting Room A. Further, the HMD of the user device (320) can detect the head movements of User B and determine the viewport direction of User B based on the head movement. The user device (320) can send the selection of Meeting Room A and the viewport direction of User B to the network-based media processing server (340), and the network-based media processing server (340) can process the media sent by the subsystem (310A) and send viewport-related streams (such as a video stream customized based on the viewport direction of User B, an audio stream customized based on the viewport direction of User B, etc.) to the user device (320) for playback on the user device (320). In some examples, when the user device (320) selects Meeting Room A, the user device (320), the subsystem (310A), and the network-based media processing server (340) can communicate with each other based on the session description protocol (SDP).

[0071] In another example, a user device (330) joins a remote conference session. The user device (330) can address packets and transmit the packets to a network-based media processing server (340), and the network-based media processing server (340) can address packets and transmit the packets to the user device (330). The packets can contain any suitable information / data, such as media streams, control parameters, etc. In some examples, the network-based media processing server (340) can send customized media information to the user device (330). For example, User C can use the user device (330) to enter a choice of a meeting room (e.g., Meeting Room Z), and the viewport direction of User C (e.g., using the touch screen of a smart phone). The user device (330) can send the selection information of Meeting Room Z and the viewport direction of User C to the network-based media processing server (340), and the network-based media processing server (340) can process the media sent by the subsystem (310Z), and send viewport-related streams (e.g., a video stream customized based on the viewport direction of User C, an audio stream customized based on the viewport direction of User C, etc.) to the user device (330) for playback on the user device (330). In some examples, when the user device (330) selects Meeting Room Z, the user device (330), the subsystem (310Z), and the network-based media processing server (340) can communicate with each other based on the Session Description Protocol (SDP).

[0072] Note that during a remote conference session, the viewport directions of users (such as User B, User C) may vary. For example, User B can notify the network-based media processing server (340) of the change in the viewport direction of User B, and correspondingly, the network-based media processing server (340) can adjust the viewport direction in the viewport-related streams sent to the user device (320).

[0073] For ease of description, immersive media is used to refer to wide-angle media, such as omnidirectional video, omnidirectional audio, etc., and immersive media is used to refer to viewport-related media generated based on wide-angle media. Note that in this application, 360-degree media (such as 360-degree video, 360-degree audio, etc.) is used to illustrate the technology for remote conferencing, and the remote conferencing technology can be used for immersive media with less than 360 degrees.

[0074] Note that the selection of the meeting room can be changed during a remote meeting session. In one example, a user device (e.g., user device (320), user device (330), etc.) can trigger a switch from one meeting room to another based on an active speaker. For example, in response to an active speaker in Meeting Room A, user device (330) can determine to switch the selection of the meeting room to Meeting Room A and send the selection of Meeting Room A to the network-based media processing server (340). Then, the network-based media processing server (340) can process the media sent by subsystem (310A) and send viewport-related streams (e.g., a video stream customized based on the viewport direction of User C, an audio stream customized based on the viewport direction of User C, etc.) to user device (330) for playback on user device (330).

[0075] In some examples, the network-based media processing server (340) can pause receiving video streams from a meeting room that has no active users. For example, if the network-based media processing server (340) determines that there are no active users in Meeting Room Z, then the network-based media processing server (340) can pause receiving video streams from subsystem (310Z).

[0076] In some examples, the network-based media processing server (340) can include distributed computing resources and can communicate with subsystems (310A) through (310Z) as well as user devices (320) and user device (330) via network (301). In some examples, the network-based media processing server (340) can be an independent system tasked with managing various aspects of one or more remote meeting sessions.

[0077] In various examples, the network-based media processing server (340) may include one or more computing devices that operate in a cluster or other grouped configuration to share resources, balance loads, improve performance, provide failover support or redundancy, or for other purposes. For example, the network-based media processing server (340) may belong to various categories of devices, such as traditional server-type devices, desktop computer-type devices, and / or mobile-type devices. Thus, although described as a single type of device (i.e., a server-type device), the network-based media processing server (340) can include a wide variety of device types and is not limited to a specific type of device. The network-based media processing server (340) can represent but is not limited to a server computer, a desktop computer, a network server computer, a personal computer, a mobile computer, a laptop, a tablet, or any other type of computing device.

[0078] According to one aspect of the present application, a network-based media processing server (340) can perform certain media functions to alleviate the processing burden at terminals (e.g., user equipment (320), user equipment (330), etc.). For example, if the media processing capabilities of user equipment (320) and / or user equipment (330) are limited or may have difficulties in encoding and presenting multiple video streams, the network-based media processing server (340) can perform media processing, such as decoding / encoding audio and video streams, etc., to alleviate the media processing burden in user equipment (320) and user equipment (330). In some examples, user equipment (320) and user equipment (330) are battery-powered devices, and when the media processing burden is transferred from user equipment (320) and user equipment (330) to the network-based media processing server (340), the battery life of user equipment (320) and user equipment (330) can be increased.

[0079] Media streams from different sources can be processed and mixed. In some examples, for instance, in International Organization for Standardization (ISO) 23090-2, an overlay can be defined as a second media presented on a first media. According to one aspect of the present application, for a remote conference session of an immersive remote conference, additional media content (e.g., video and / or audio) can be overlaid on the immersive media content. The additional media (or media content) can be referred to as the overlay media (or overlay media content) of the immersive media (or immersive media content). For example, the overlay content can be a piece of visual / audio media presented on an omnidirectional video or image item or viewport.

[0080] Use Figure 2As an example, when a participant in Conference Room A shares a presentation, in addition to being displayed on the subsystem (210A) of Conference Room A, the presentation is also broadcast as a stream (also known as an overlay stream) to other participant parties (e.g., subsystem (210Z), user device (220), user device (230), etc.). For example, user device (220) selects Conference Room A, and subsystem (210A) can transmit a first stream of immersive media (e.g., a 360-degree video captured by subsystem (210A)) and the overlay stream to user device (220). On user device (220), the presentation can be overlaid on the 360-degree video captured by subsystem (210A). In another example, user device (230) selects Conference Room Z, and subsystem (210Z) can transmit a first stream carrying immersive media (e.g., a 360-degree video captured by subsystem (210Z)) and the overlay stream to user device (220). At user device (230), the presentation can be overlaid on the 360-degree video captured by subsystem (210Z). It should be noted that in some examples, the presentation can be overlaid on a 2D video. It should be noted that in some examples, the presentation can be overlaid on a 2D video.

[0081] In another scenario, User C can be a remote speaker, and a media stream carrying the audio corresponding to User C's speech (referred to as an overlay stream) can be sent from user device (230) to, for example, subsystem (210Z) and broadcast to other participant parties, such as subsystem (210A). For example, user device (220) selects Conference Room A, and subsystem (210A) can transmit a first stream of immersive media (e.g., a 360-degree video captured by subsystem (210A)) and the overlay stream to user device (220). At user device (220), the 360-degree video captured by subsystem (210A) can be overlaid on the audio corresponding to User U's speech. In one example, the media stream carrying the audio corresponding to User C's speech can be referred to as an overlay stream, and in one example, the audio can be referred to as overlay audio.

[0082] Some aspects of the present application provide audio and video mixing techniques, as well as more specific techniques for combining audio and / or video of multiple media streams, e.g., immersive streams and one or more overlay streams. According to one aspect of the present application, audio and / or video mixing can be performed by a network-based media processing element (e.g., a network-based media processing server (340), etc.), or can be performed by an end-user device, such as user device (120), user device (130), user device (220), user device (230), user device (320), user device (330), etc.

[0083] In Figure 1In the example, the subsystem (110) is referred to as the sender, which can send multiple media streams respectively carrying media (audio and / or video), and the user devices (120) and (130) are referred to as receivers. In Figure 2 the example, the subsystems (210A) to (210Z) are referred to as the sender, which can send multiple media streams respectively carrying media (audio and / or video), and the user devices (220) and (230) are referred to as receivers. In Figure 3 the example, the network-based media processing server (340) is referred to as the sender, which can send multiple media streams respectively carrying media (audio and / or video), and the user devices (320) and (330) are referred to as receivers.

[0084] According to some aspects of the present application, in an immersive remote conference for audio mixing, a mixing level (such as audio weight) can be assigned to the overlay stream and the immersive stream. Further, in some embodiments, the audio weight can be appropriately adjusted, and the adjusted audio weight can be used for audio mixing. In some examples, audio mixing is also referred to as audio downmixing.

[0085] In some examples, such as an immersive remote conference, when overlaying overlay media on immersive media, it may be necessary to provide overlay information, such as the overlay source, overlay rendering type, overlay rendering attributes, user interaction attributes, etc. In some examples, the overlay source specifies the media, such as an image, audio, or video used as the overlay; the overlay rendering type describes whether the overlay is anchored relative to the viewport or range; and the overlay rendering attributes can include the opacity level, transparency level, etc.

[0086] In Figure 2 the example, multiple conference rooms each having an omnidirectional camera can participate in a remote conference session. A user, such as user B, can select the source of the immersive media, such as one of the multiple conference rooms each having an omnidirectional camera, via the user device (220). To add additional media (such as audio or video with the immersive media), the additional media can be separated from the immersive media and sent as an overlay stream carrying the additional media to the user device (220). The immersive media can be sent as a stream carrying the immersive media (referred to as the immersive stream). The user device (220) can receive the immersive stream and the overlay stream, and can overlay the immersive media on the additional media.

[0087] According to one aspect of the present application, a user device (e.g., user device (220), user device (230), etc.) may receive a plurality of media streams each carrying respective audio in a remote conference session. The user device may decode the media streams to retrieve the audio and mix the audio decoded from the media streams. In some examples, during a remote conference of an immersive remote conference, a subsystem in a selected conference room may send a multimedia stream and provide mixing parameters for the audio carried in the multimedia stream. For example, user B may, via user device (220), select conference room A to receive an immersive stream carrying a 360-degree immersive video captured by subsystem (210A). Subsystem (210A) may send the immersive stream with one or more overlay streams to user device (220). Subsystem (210A) may provide a mixing level for the audio carried in the immersive stream and the one or more overlay streams, for example, based on the Session Description Protocol (SDP). It should be noted that subsystem (210A) may also update the mixing level of the audio during the remote conference session and send a signal notifying the updated mixing level to user device (220) based on the SDP.

[0088] In one example, an audio mixing level is defined using audio mixing weights. For example, subsystem (210A) that sends an immersive stream and an overlay stream each carrying respective audio may determine an audio mixing weight for each audio. In one example, subsystem (210A) determines a default audio mixing weight based on sound intensity. Sound intensity may be defined as the power carried by sound waves on a unit area in a direction perpendicular to the unit area. For example, a controller of subsystem (210A) may receive an electrical signal indicating the sound intensity of each audio and determine a default audio mixing weight based on the electrical signal (e.g., based on the signal level, power, etc. of the electrical signal).

[0089] In another example, subsystem (210A) determines an audio mixing weight based on overlay priority. For example, a controller of subsystem (210A) may detect a specific media stream carrying the audio of an active speaker from the immersive stream and the overlay stream. The controller of subsystem (210A) may determine a higher overlay priority for the specific media stream and may determine a higher mixing weight for the audio carried by the specific media stream.

[0090] In another example, an end user may customize the overlay priority. For example, user B may use user device (220) to send customization parameters to subsystem (210A) based on the SDP. For example, the customization parameters may indicate, for example, a specific media stream that carries the audio that user B wishes to focus on. Then, subsystem (210A) may determine a higher overlay priority for the specific media stream and determine a higher mixing weight for the audio carried by the specific media stream.

[0091] In some embodiments, when using overlay priorities, all overlays from other senders (such as subsystem (210Z)) and their priorities in a remote conference session can be informed to a sender (such as subsystem (210A)), and the sender assigns weights accordingly. Thus, when the user equipment switches to a different subsystem, the audio mixing weights can be determined appropriately.

[0092] In some embodiments, the audio mixing weights may be customized by the end user. In one scenario, the end user may wish to listen to or focus on a specific audio carried by a media stream. In another scenario, due to reasons such as audio level changes, poor audio quality, or a poor signal-to-noise ratio (SNR) of the channel, the quality of the downmixed audio under the default audio mixing weights is unacceptable, then the audio mixing weights can be customized. In an example, user B wishes to focus on the audio from a specific media stream, then user B can use the user equipment (220) to indicate customization parameters for adjusting the audio mixing weights. For example, the customization parameter indicates an increase in the audio mixing weight for the audio in the specific media stream. During a remote conference session, the user equipment (220) can send the customization parameter to the sender of the media stream, such as subsystem (210A), based on the SDP. Based on the customization parameter, the controller of subsystem (210A) can adjust the audio mixing weights to increase the audio mixing weight for the audio in the specific media stream, and subsystem (210A) can send the adjusted audio mixing weights to the user equipment (220). Thus, the user equipment (220) can mix the audio based on the adjusted audio mixing weights.

[0093] Note that in some examples, due to user preferences, user equipment (such as user equipment (120), user equipment (130), user equipment (220), user equipment (230), user equipment (320), and user equipment (330), etc.) can rewrite the received audio mixing weights with different values.

[0094] In Figure 3In an example, multiple conference rooms each with an omnidirectional camera can participate in a remote conference session. A user (such as user B) can select a source of immersive media via a user device (320), such as one of the multiple conference rooms each with an omnidirectional camera. To add additional media (such as audio or video with immersive media), the additional media can be separated from the immersive media and sent as an overlay stream carrying the additional media to the user device (320). In some embodiments, a network-based media processing server (340) receives media streams from participant parties (such as subsystems (310A) to (310Z), user device (320), and user device (330)) in a remote conference, processes the media streams, and sends appropriate processed media streams to the participant parties. For example, the network-based media processing server (340) can send an immersive stream carrying immersive media captured at subsystem (310A) and an overlay stream carrying overlay media to user device (320). The user device (320) can receive the immersive stream and the overlay stream, and in some embodiments, the user device can overlay the overlay media with the immersive media.

[0095] According to one aspect of the present application, a user device (such as user device (320), user device (330), etc.) can receive multiple media streams carrying respective audio in a remote conference session. The user device can decode the media streams to retrieve the audio and mix the audio decoded from the media streams. In some examples, during a remote conference of an immersive remote conference, a network-based media processing server (340) can send multiple media streams to an end-user device. For example, user B via user device (320) can select conference room A to receive an immersive stream carrying a 360-degree immersive video captured by subsystem (310A). According to one aspect of the present application, audio mixing parameters (such as loudness) can be defined by the sender of the immersive media or customized by the end user. In some examples, subsystem (310A) can provide, for example, via a Session Description Protocol (SDP)-based signal, a mixing level for the audio carried in one or more overlay streams to the network-based media processing server (340). It should be noted that subsystem (310A) can also update the mixing level of the audio during the remote conference session and send a signal notifying the updated mixing level to the network-based media processing server (340) based on SDP.

[0096] In one example, an audio mixing level is defined using an audio mixing weight. For example, subsystem (310A) can determine the audio mixing weight and send it to the network-based media processing server (340) based on SDP. In one example, subsystem (310A) determines a default audio mixing weight based on sound intensity.

[0097] In another example, the subsystem (310A) determines audio mixing weights based on overlay priorities. For example, the subsystem (310A) can detect a specific media stream carrying the audio of an active speaker. The subsystem (310A) can determine a higher overlay priority for the specific media stream and can determine a higher mixing weight for the audio carried by the specific media stream.

[0098] In another example, an end user can customize the overlay priorities. For example, user B can use the user device (320) to send customization parameters to the subsystem (310A) based on SDP. For example, the customization parameters can indicate a specific media stream that carries the audio that user B wishes to focus on. Then, the subsystem (310A) can determine a higher overlay priority for the specific media stream and determine a higher mixing weight for the audio carried by the specific media stream.

[0099] In some embodiments, when using overlay priorities, all the overlays of other senders (such as the subsystem (310Z)) and the priorities of these overlays in a remote conference session are informed to the sender, and accordingly, the sender assigns weights. Thus, when the user device switches to a different subsystem, the audio mixing weights can be appropriately determined.

[0100] In some embodiments, the audio mixing weights may be customized by the end user. In one scenario, the end user may wish to listen to or focus on specific audio carried by a media stream. In another scenario, due to reasons such as audio level transformation, poor audio quality, or poor signal-to-noise ratio (SNR) of the channel, the quality of the downmixed audio under the default audio mixing weights is unacceptable, then the audio mixing weights can be customized. In an example, if user B wishes to focus on the audio from a specific media stream, user B can use the user device (320) to indicate customization parameters to adjust the audio mixing weights. For example, the customization parameters indicate an increase in the audio mixing weight for the audio in the specific media stream. During a remote conference session, the user device (320) can send the customization parameters to the sender of the media stream, such as the subsystem (310A), based on SDP. Based on the customization parameters, the subsystem (310A) can adjust the audio mixing weights to increase the audio mixing weight for the audio in the specific media stream and send the adjusted audio mixing weights to the network-based media processing server (340). In an example, the network-based media processing server (340) can send the adjusted audio mixing weights to the user device (320). Thus, the user device (320) can mix the audio based on the adjusted audio mixing weights. In another example, the network-based media processing server (340) can mix the audio according to the adjusted audio mixing weights.

[0101] In one example, a sender (e.g., one of subsystems (210A) to (210Z), one of subsystems (310A) to (310Z)) provides an immersive stream and one or more overlay streams, where N represents the number of overlay streams and N is a positive integer. Further, a0 represents the audio carried in the immersive stream; a1 - aN respectively represent the audio carried in the overlay streams; and r0 - rN respectively represent the audio mixing weights for a0 - aN. In some examples, the sum of the default audio mixing weights r0 - rN equals 1. The mixed audio (also referred to as audio output) can be generated according to Equation 1, which is:

[0102] audio output = r0 × a0 + r1 × a1 + … + rn × an Equation 1

[0103] In some embodiments, for example according to Equation 1, an end - user device (such as user device (220), user device (230), user device (320), user device (330), etc.) can perform audio mixing based on the audio mixing weights. The end - user device can decode the received media stream to retrieve the audio and the mixed audio according to Equation 1 to generate an audio output for playback.

[0104] In some embodiments, the audio mixing or a part of the audio mixing can be performed by an MRF or an MCU, for example, by a network - based media processing server (340). Refer to Figure 3, in some examples, the network-based media processing server (340) receives various media streams carrying audio. Further, the network-based media processing server (340) may perform media mixing, such as audio mixing based on audio mixing weights. Taking the subsystem (310) and the user device (330) as an example (e.g., the user device (330) selects Conference Room A), when the user device (330) is in a low-power state or has limited media processing capabilities, the audio mixing or a part of the audio mixing can be transferred to the network-based media processing server (340). In one example, the network-based media processing server (340) may receive a media stream for sending to the user device (330) and an audio mixing weight for mixing the audio in the media stream. Then, the network-based media processing server (340) may decode the media stream to retrieve the audio and mix the audio according to Equation 1 to generate the mixed audio. It should be noted that the network-based media processing server (340) may appropriately mix the video part of the media stream into the mixed video. The network-based media processing server (340) may encode the mixed audio and / or the mixed video in another stream (referred to as the mixed media stream) and send the mixed media stream to the user device (330). The user device (330) may receive the mixed media stream, decode the mixed media stream to retrieve the mixed audio and / or the mixed video, and play the mixed audio / video.

[0105] In another example, the network-based media processing server (340) receives an immersive media stream and a plurality of overlay media streams for providing media content to the user device (330), and an audio mixing weight for mixing the audio in the immersive media stream and the plurality of overlay media streams. When multiple overlay media streams need to be sent, the network-based media processing server (340) may, for example, according to Equation 2, decode the plurality of overlay media streams to retrieve the audio and mix the audio to generate the mixed overlay audio, and Equation 2 is:

[0106] mixed overlay audio = r1 × a1 + … + rn × an Equation 2

[0107] Note that the network-based media processing server (340) may appropriately mix the video portion of the overlay media stream into a mixed overlay video. The network-based media processing server (340) may encode the mixed overlay audio and / or the mixed overlay video in another stream, referred to as the mixed overlay media stream, and send the mixed overlay media stream and the immersive media stream to the user device (330). The user device (330) may receive the immersive media stream and the mixed media stream, and decode the immersive media stream and the mixed media stream to retrieve the audio (a0) of the immersive media, the mixed overlay audio, and / or the mixed overlay video. Based on the audio (a0) of the immersive media and the mixed overlay audio, the user device (330) may generate, for example, a mixed audio (also referred to as the audio output) for playback according to Equation 3, where Equation 3 is:

[0108] audio output=r0×a0 + mixed overlay audio Equation 3

[0109] In one example, when there is no background noise or interference from any audio of the overlay media stream or the immersive media stream (in some examples, the audio from the immersive media stream may be referred to as the background), or when the audio intensity levels of all media streams are the same or the variance is relatively small (e.g., less than a predefined threshold), audio mixing may be performed (e.g., using the same mixing weights of 1 respectively) by adding the audio retrieved from all streams (e.g., the overlay media stream and the immersive media stream together) to generate an aggregated audio, and the aggregated audio may be normalized (e.g., divided by the number of audios). Note that in this example, the audio mixing may be performed by the end-user device, such as the user device (120), the user device (130), the user device (220), the user device (230), the user device (320), the user device (330), and the network-based media processing server (340).

[0110] In some embodiments, audio weights may be used to select portions of the audio to be mixed. In one example, when a large amount of audio is aggregated and then normalized, it may be difficult to distinguish between two audio streams. Using the audio weights, a selected number of audios may be aggregated and then normalized. For example, when the total number of audios is 10, the weights of the 5 selected audios are 0.2, and the weights of the 5 unselected audios are 0. Note that the selection of the audio may be based on algorithm-defined mixing weights or on the overlay priority.

[0111] In some embodiments, the user equipment may choose to retrieve and mix audio by changing the corresponding audio mixing weights or even using a subset of the media streams, so as to change the audio selected from the media streams to be mixed.

[0112] In some embodiments, when the sound intensity of the audio in the media stream varies greatly, the audio mixing weights of the superimposed audio and the immersive audio can be set to the same level.

[0113] In some embodiments, the user equipment has limited resource capacity or it is difficult to distinguish the audio from different conference rooms. Therefore, the number of audio to be downmixed may be limited. If such a limitation is applied, the sender device, such as subsystems (210A) to (210Z), the network-based media processing server (340), can select the media streams to be audio downmixed based on the sound intensity or the superimposition priority. It should be noted that during a remote conference session, the user equipment can send customized parameters based on the SDP to change the selection.

[0114] In some scenarios, during a remote conference session, it is necessary to focus on the person speaking / presenting. Therefore, a larger audio mixing weight can be assigned to the media stream with the audio of the speaker, and the audio mixing weights of the other audio in the other media streams can be reduced.

[0115] In some cases, when a remote user is presenting, there is background noise in the immersive audio of the immersive media stream. The sender, such as subsystems (310A) to (310Z), the network-based media processing server (340), can reduce the audio mixing weight of the immersive audio to be less than the superimposed audio associated with the remote user. Although this can be customized by the end user already in the session by reducing the audio weight during the remote conference session, changing the default audio mixing weight provided by the sender can allow a new remote user who has just joined the conference to obtain the default audio mixing weight of the audio stream from the sender, so as to downmix the audio with good sound quality.

[0116] In one embodiment, the sender device (such as subsystems (310A) to (310Z), etc.) defines the audio mixing parameters (such as audio mixing weights). The sender device can determine the audio mixing weights to set the audio streams to the same loudness level. The audio mixing parameters (audio mixing weights) can be transmitted from the sender device to the network-based media processing server (340) via the SDP signaling.

[0117] In another embodiment, a sending device, such as subsystems (310A) to (310Z), etc., may set the audio mixing weight of the audio in the immersive media content to be higher than the audio mixing weights of other superimposed audio in the superimposed media stream. In one example, the superimposed audio may have the same audio mixing weight. The audio mixing parameter (audio mixing weight) may be transmitted from the sending device to the network-based media processing server (340) via SDP signaling.

[0118] In another embodiment, a sending device, such as subsystems (310A) to (310Z), etc., may set the audio mixing weight of the audio in the immersive media content to be higher than the audio mixing weight of the superimposed audio in the superimposed media stream. The audio mixing parameter (audio mixing weight) may be transmitted from the sending device to the network-based media processing server (340) via SDP signaling.

[0119] In some examples, such as when the end-user device may not have sufficient processing power, the network-based media processing server (340) may send the same audio stream to multiple end-user devices.

[0120] In some examples, such as when the audio mixing parameter is user-defined or user-customized, the sending device or the network-based media processing server (340) may encode a single audio stream for each user device. In one example, the audio mixing parameter may be based on the user's field of view (FoV). For example, compared to other streams, the superimposed audio stream within the field of view may be mixed with a greater loudness. The audio mixing parameter (audio mixing weight) may be negotiated by the sending device, the user device, and the network-based media processing server (340) via SDP signaling.

[0121] In one embodiment, for example, when the end device supports multimedia telephony services for the Internet Protocol Multimedia Subsystem (MTSI), but does not support MTSI immersive remote conferencing and remote presentation (ITT4RT) for remote terminals, the network-based media processing server (340) may mix audio and video to generate mixed audio and video, and provide the media stream carrying the mixed audio and video to the end-user device, thereby providing backward compatibility for MTSI terminals.

[0122] In another embodiment, for example, when the capabilities of the end device are limited, the network-based media processing server (340) may mix audio and video to generate mixed audio and video, and provide the media stream carrying the mixed audio and video to the end-user device.

[0123] In another embodiment, when the capabilities of the network-based media processing server (340) are limited and some of the end-user devices are MSTI devices with limited capabilities, the network-based media processing server (340) may mix the audio and video from the same sender device to generate mixed audio and video, and provide a media stream carrying the mixed audio and video to the end-user device, which is an MSTI device with limited capabilities.

[0124] In another embodiment, the network-based media processing server (340) may use SDP signaling to negotiate a set of common configurations for audio mixing with all or a subset of the end-user devices that are MSTI devices. This set of common configurations is used for a single video combination of immersive media and various overlay media. Then, based on this set of common configurations, the network-based media processing server (340) may perform mixed audio and / or video mixing to generate mixed audio and video, and provide a media stream carrying the mixed audio and video to all or a subset of the end devices that are MSTI devices.

[0125] Figure 4 A flowchart outlining a process (400) according to an embodiment of the present application is shown. In various embodiments, the process (400) may be executed by a processing circuit in a device, such as a processing circuit in user device (120), user device (130), user device (220), user device (230), user device (320), user device (330), network-based media processing server (340), etc. In some embodiments, the process (400) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (400). The process begins at (S401) and proceeds to (S410).

[0126] At (S410), a first media stream carrying a first audio and a second media stream carrying a second audio are received.

[0127] At (S420), a first audio weight for weighting the first audio and a second audio weight for weighting the second audio are received.

[0128] At (S430), the weighted first audio based on the first audio weight and the weighted second audio based on the second audio weight are combined to generate mixed audio.

[0129] In some examples, the device is a user device, and the processing circuitry of the user device receives a first audio weight and a second audio weight determined by a host device for immersive content (e.g., subsystem (110), subsystems (210A) to (210Z), subsystems (310A) to (310Z)), and the user device can play the mixed audio through a speaker associated with the user device. In one example, to customize the audio weights, the user device can send customization parameters to the host device so that the host device customizes the first audio weight and the second audio weight based on the customization parameters.

[0130] In some examples, the host device can determine the first audio weight and the second audio weight based on the sound intensities of the first audio and the second audio. The user device can receive the first audio weight and the second audio weight determined by the host device based on the sound intensities of the first audio and the second audio.

[0131] In some examples, the first audio and the second audio are superimposed audio, and the host device can determine the first audio weight and the second audio weight based on the superimposition priorities of the first audio and the second audio. The user device can receive the first audio weight and the second audio weight determined by the host device based on the superimposition priorities of the first audio and the second audio.

[0132] In some examples, the host device can determine the first audio weight and the second audio weight based on the detection of active speakers. The user device can receive the first audio weight and the second audio weight adjusted by the host device based on the detection of active speakers.

[0133] In some examples, the first media stream includes immersive media content, the second media stream corresponds to superimposed media content, and the host device can determine that the first audio weight is different from the second audio weight.

[0134] In some embodiments, processing (400) is performed by a network-based media processing server that performs media processing offloaded from the user device. The network-based media processing server can encode the mixed audio into a third media stream and transmit the third media stream to the user device. In some examples, processing (400) is performed by a network-based media processing server that performs superimposed media processing offloaded from the user device. The network-based media processing server can transmit the third media stream and a fourth media stream including immersive media content. The third media stream includes superimposed media content of the immersive media content.

[0135] Then, the process proceeds to (S499) and terminates.

[0136] Figure 5FIG. 0 shows a flowchart outlining a process (500) according to an embodiment of the present application. In various embodiments, the process (500) may be executed by processing circuitry in a device for network-based media processing (such as a network-based media processing server (340), etc.). In some embodiments, the process (500) is implemented as software instructions, such that when the processing circuitry executes the software instructions, the processing circuitry performs the process (500). The process starts at (S501) and proceeds to (S510).

[0137] At (S510), a first media stream carrying first media content and a second media stream carrying second media content are received.

[0138] At (S520), third media content that mixes the first media content and the second media content is generated.

[0139] In some examples, a first audio in the first media content is mixed with a second audio in the second media content to generate a third audio. The first audio is weighted based on a first audio weight assigned to the first audio, and the second audio is weighted based on a second audio weight assigned to the second audio. In one example, a host device providing immersive media content determines the first audio weight and the second audio weight, and sends the first audio weight and the second audio weight from the host device to the network-based media processing server.

[0140] In one example, the first media stream is an immersive media stream and the second media stream is an overlay media stream, and then the values of the first audio weight and the second audio weight are different; based on the first audio weight and the second audio weight having different values, the first audio is mixed with the second audio.

[0141] In one example, the first media stream and the second media stream are overlay media streams, and the values of the first audio weight and the second audio weight are the same; based on the first audio weight and the second audio weight having the same values, the first audio is mixed with the second audio.

[0142] In another example, the first media stream and the second media stream are overlay media streams, and the first audio weight and the second audio weight depend on the overlay priority of the first media stream and the second media stream; based on the first audio weight and the second audio weight, the first audio is mixed with the second audio, and the first audio weight and the second audio weight are associated with the overlay priority of the first media stream and the second media stream.

[0143] At (S530), a third media stream carrying the third media content is transmitted to a user device.

[0144] Then, the process proceeds to (S599) and terminates.

[0145] Figure 6 FIG. 600 is a flowchart outlining a process in accordance with an embodiment of the present application. In various embodiments, the process 600 may be performed by circuitry in a host device for immersive media content, such as by processing circuitry in subsystems 110, 210A - 210Z, 310A - 310Z, etc. to perform the process 600. In some embodiments, the process 600 is implemented as software instructions, such that when the processing circuitry executes the software instructions, the processing circuitry performs the process 600. The process starts at S601 and proceeds to S610.

[0146] At S610, a first media stream carrying a first audio and a second media stream carrying a second audio are transmitted.

[0147] At S620, a first audio weight for weighting the first audio and a second audio weight for weighting the second audio are determined.

[0148] In one example, the host device receives custom parameters based on the Session Description Protocol and determines the first audio weight and the second audio weight based on the custom parameters.

[0149] In some examples, the host device determines the first audio weight and the second audio weight based on the sound intensity of the first audio and the second audio.

[0150] In some examples, the first audio and the second audio are superimposed audio, and the host device may determine the first audio weight and the second audio weight based on the superimposition priority of the first audio and the second audio.

[0151] In some examples, the host device determines the first audio weight and the second audio weight based on the detection of active speakers in one of the first audio and the second audio.

[0152] In some examples, the first media stream includes immersive media content and the second media stream includes superimposed media content, and the host device determines different values for the first audio weight and the second audio weight.

[0153] At S630, the first audio weight and the second audio weight are transmitted to mix the first audio with the second audio.

[0154] Then, the process proceeds to S699 and ends.

[0155] The above techniques may be implemented as computer software that uses computer - readable instructions and is physically stored on one or more computer - readable media. For example, Figure 7 FIG. 700 shows a computer system 700 suitable for implementing some embodiments of the disclosed subject matter.

[0156] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subjected to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs) or the like or executed through decoding, microcode, etc.

[0157] Instructions can be executed on various types of computers or their components, such as including personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0158] Figure 7 The components of the illustrated computer system (700) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (700).

[0159] The computer system (700) can include certain human-machine interface input devices. Such human-machine interface input devices can respond to one or more human users through inputs such as the following: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media not necessarily directly related to human conscious inputs, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0160] The input human-machine interface devices can include one or more of the following (only one of each is shown): keyboard (701), mouse (702), touchpad (703), touch screen (710), data glove (not shown), joystick (705), microphone (706), scanner (707), camera (708).

[0161] The computer system (700) may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of one or more human users, for example, through haptic output, sound, light, and smell / taste. Such human - machine interface output devices may include haptic output devices (e.g., haptic feedback of a touch screen (710), a data glove (not shown), or a joystick (705), but may also be haptic feedback devices that are not input devices), audio output devices (e.g., speakers (709), headphones (not shown)), visual output devices (e.g., a screen (710) including a CRT screen, an LCD screen, a plasma screen, an OLED screen), each screen having or not having touch - screen input functionality, each screen having or not having haptic feedback functionality, and some of which are capable of outputting two - dimensional visual output or ultra - three - dimensional output through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).

[0162] The computer system (700) may also include human - accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (720) with media such as CD / DVD (721), thumb drives (722), removable hard disk drives or solid - state drives (723), traditional magnetic media such as tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.

[0163] Those skilled in the art should also understand that the term "computer - readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0164] The computer system (700) may also include an interface (754) to one or more communication networks (755). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicular and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial television including CANBus, and so on. Certain networks typically require an external network interface adapter (such as a USB port of the computer system (700)) connected to certain common data ports or peripheral buses (749). As described below, other network interfaces are typically integrated into the core of the computer system (700) by connecting to the system bus (such as connecting an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (700) may communicate with other entities using any of these networks. Such communication may be only one-way receiving (such as broadcast television), only one-way sending (such as CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of the networks and network interfaces.

[0165] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface may be attached to the core (740) of the computer system (700).

[0166] The core (740) may include one or more central processing units (CPUs) (741), a graphics processing unit (GPU) (742), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (743), a hardware accelerator (744) for certain tasks, a graphics adapter (750), and so on. These devices, as well as a read-only memory (ROM) (745), an internal mass storage (747) such as an internal non-user-accessible hard disk drive, SSD, etc., may be connected via a system bus (748). In some computer systems, the system bus (748) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus of the core (748) or connected to the system bus of the core via a peripheral bus (749). In some examples, the screen (710) may be connected to the graphics adapter (750). The architecture of the peripheral bus includes PCI, USB, etc.

[0167] The CPU (741), GPU (742), FPGA (743), and accelerator (744) can execute certain instructions, which can be combined to form the above computer code. The computer code can be stored in the ROM (745) or RAM (746). Temporary data can also be stored in the RAM (746), while permanent data can be stored, for example, in the internal mass storage (747). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more of the following: one or more CPUs (741), GPUs (742), mass storage (747), ROM (745), RAM (746), etc.

[0168] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be those specially designed and constructed for the purposes of the present disclosure, or the medium and the computer code can be of the type well known and available to those of ordinary skill in the computer software arts.

[0169] As a non-limiting example, it can be due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media such that there is Figure 7The architecture (700) shown, particularly the computer system of the kernel (740), provides functionality. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the kernel (740), such as the on-kernel mass memory (747) or ROM (745). Software implementing various embodiments of the present application can be stored in such devices and executed by the kernel (740). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the kernel (740), particularly the processors therein (including CPU, GPU, FPGA, etc.), to perform the specific processes or specific portions of the specific processes described herein, including defining data structures stored in the RAM (746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can be caused to provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., the accelerator (744)), which can replace the software or operate together with the software to perform the specific processes or specific portions of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to the computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both a circuit (e.g., an integrated circuit (IC)) storing software for execution and a circuit embodying logic for execution. The present application includes any suitable combination of hardware and software.

[0170] Although the present application has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present application. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and thus fall within the spirit and scope of the present application.

Claims

1. A method for a remote conference, comprising: Receiving, by a processing circuit of a first device, a first media stream carrying a first audio and a second media stream carrying a second audio from a second device; Sending customization parameters to the second device for the second device to customize a first audio weight and a second audio weight based on the customization parameters; The customization parameters are used to indicate increasing the weight of the audio carried by the first media stream or the second media stream; Receiving, from the second device, a first audio weight for weighting the first audio and a second audio weight for weighting the second audio; And Generating, by the processing circuit of the first device, a mixed audio by combining the weighted first audio based on the first audio weight and the weighted second audio based on the second audio weight.

2. The method according to claim 1, further comprising: Playing the mixed audio through a speaker associated with the first device.

3. The method according to claim 1, further comprising: Receiving the first audio weight and the second audio weight determined by the second device based on the sound intensities of the first audio and the second audio.

4. The method according to claim 1, wherein The first audio and the second audio are superimposed audios, and the method comprises: Receiving the first audio weight and the second audio weight determined by the second device based on the superimposition priorities of the first audio and the second audio.

5. The method according to claim 1, further comprising: Receiving the first audio weight and the second audio weight adjusted by the second device based on the detection of an active speaker.

6. The method according to claim 1, wherein The first media stream includes immersive media content, the second media stream includes superimposed media content of the immersive media content, and the first audio weight is different from the second audio weight.

7. The method according to claim 1, further comprising: Encoding, by the processing circuit, the mixed audio into a third media stream; And Transmitting the third media stream to a third device via an interface circuit of the first device.

8. The method according to claim 7, further comprising: Transmitting, via the interface circuit of the first device, the third media stream and a fourth media stream including immersive media content, the third media stream being a superimposed media stream of the fourth media stream.

9. A method for a remote conference, comprising: Receiving, by a processing circuit of a first device, a first media stream carrying a first media content of a remote conference session and a second media stream carrying a second media content of the remote conference session; Receiving, by the processing circuit of the first device, a first audio weight and a second audio weight from a second device that sends the first media stream and the second media stream; wherein, the first audio weight and the second audio weight are determined based on customization parameters sent by the first device to the second device, and the customization parameters are used to indicate increasing the weight of the audio carried by the first media stream or the second media stream; The processing circuit of the first device mixes the first audio in the first media content with the second audio in the second media content to generate a third audio based on the first audio weight assigned to the first audio and the second audio weight assigned to the second audio; The processing circuit of the first device generates a third media content that mixes the first media content and the second media content; and The transmission circuit of the first device transmits a third media stream carrying the third media content to a second device.

10. The method according to claim 9, wherein The first media stream includes immersive media content, and the second media stream includes superimposed media content of the immersive media content. The method further includes: The processing circuit of the first device mixes the first audio with the second audio based on the first audio weight and the second audio weight having different values.

11. The method according to claim 9, wherein The first media stream and the second media stream are superimposed media streams. The method further includes: The processing circuit of the first device mixes the first audio with the second audio based on the first audio weight and the second audio weight having the same value.

12. The method according to claim 9, wherein The first media stream and the second media stream are superimposed media streams. The method further includes: The processing circuit of the first device mixes the first audio with the second audio based on the first audio weight and the second audio weight, and the first audio weight and the second audio weight are associated with the superimposition priority of the first media stream and the second media stream.

13. A method for remote conferencing, including: A first device transmits a first media stream carrying a first audio and a second media stream carrying a second audio to a second device; The first device receives customization parameters sent by the second device, and the customization parameters are used to indicate increasing the weight of the audio carried by the first media stream or the second media stream; The first device determines a first audio weight for weighting the first audio and a second audio weight for weighting the second audio, and the first audio weight and the second audio weight are determined based on the customization parameters; And The first device transmits the first audio weight and the second audio weight for mixing the first audio and the second audio to the second device.

14. The method according to claim 13, The customization parameters are based on the Session Description Protocol.

15. The method according to claim 13, further includes: Determining the first audio weight and the second audio weight based on the sound intensity of the first audio and the second audio.

16. The method according to claim 13, wherein The first audio and the second audio are superimposed audio, and the method includes: Determining the first audio weight and the second audio weight based on the superimposition priority of the first audio and the second audio.

17. The method according to claim 13, further includes: Determining the first audio weight and the second audio weight based on the detection of the active speaker in one of the first audio and the second audio.

18. The method according to claim 13, wherein The first media stream includes immersive media content, the second media stream includes superimposed media content, and the method further includes: Determining different values for the first audio weight and the second audio weight.

19. A computer-readable medium, characterized in that, Stored with executable instructions, which when executed by a computer, implement the method according to any one of claims 1 to 18 above.

Citation Information

Patent Citations

  • Enhanced communication bridge

    CN102461139A

  • Live broadcast method, device and system

    CN105357542A