Method and apparatus for audio mixing
By downmixing multiple audio streams and optimizing audio processing based on video stream priorities, the method addresses bandwidth challenges in immersive video conferencing, enhancing the quality and efficiency of audio-visual transmission.
Patent Information
- Application Number
- JP2024111035
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2024-07-10
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2041-09-24
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing and transmitting high-quality audio and video streams, particularly in immersive video conferencing scenarios, where multiple audio streams and 360-degree video streams require significant bandwidth and processing resources.
A media-capable network element downmixes multiple audio streams into fewer channels based on audio mixing parameters and metadata, optimizing audio processing based on the field of view and priority of overlay video streams, and a processing circuit decodes these streams for immersive video conferencing.
This approach reduces bandwidth requirements and enhances the quality of immersive video conferencing by intelligently managing audio streams, ensuring efficient use of network resources and improved audio-visual experience.
Smart Images

Figure 0007765559000001 
Figure 0007765559000002 
Figure 0007765559000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims benefit of priority to U.S. Patent Application No. 17 / 478,418, entitled "METHOD AND APPARATUS FOR AUDIO MIXING," filed September 17, 2021, which in turn claims benefit of priority to U.S. Provisional Application No. 63 / 088,300, entitled "NETWORK BASED MEDIA PROCESSING FOR AUDIO AND VIDEO MIXING FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS," filed October 6, 2020, and U.S. Provisional Application No. 63 / 124,261, entitled "AUDIO MIXING METHODS FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS," filed December 11, 2020. The disclosures of the prior applications are incorporated herein by reference in their entireties.
[0002] This disclosure generally describes embodiments related to video coding. [Background technology]
[0003] The discussion of the background art provided herein is intended to provide a general context for the present disclosure. The work of the inventors named herein, to the extent that their work is described in this background art section, and aspects of the description that may not otherwise qualify as prior art at the time of filing, are not admitted, expressly or impliedly, as prior art to the present disclosure.
[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video may include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luminance samples and associated chrominance samples. The series of pictures may have a fixed or variable picture rate (also informally known as a frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920 x 1080 luminance sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.
[0005] One goal of video coding and decoding can be reducing redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, sometimes by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations of these, can be used. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed original signal. With lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that higher acceptable / tolerable distortion can result in a higher compression ratio.
[0006] Video encoders and decoders may utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.
[0007] Video codec technology may include a technique known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore may be used as the first picture in a coded video bitstream and video session or as a still image. Samples of intra-blocks may be subjected to a transform, and the transform coefficients may be quantized before entropy coding. Intra-prediction may be a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are required at a given quantization step size to represent the block after entropy coding.
[0008] For example, conventional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that rely on surrounding sample data and / or metadata obtained during the encoding and / or decoding of spatially adjacent and preceding data blocks in decoding order. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and does not use reference data from reference pictures.
[0009] Intra-prediction can take many different forms. When two or more of such techniques can be used in a given video coding technique, the technique in use can be coded in intra-prediction mode. In certain cases, a mode can have sub-modes and / or parameters, which can be coded separately or included in the mode's codeword. The codeword used for a given mode, sub-mode, and / or parameter combination can affect coding efficiency gains via intra-prediction, and therefore can also affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further refined in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark sets (BMS). Predictor blocks can be formed using neighboring sample values belonging to already available samples. The sample values of the neighboring samples are copied into the predictor block according to the direction. A reference to the direction in use may be coded in the bitstream or may itself be predicted.
[0011] Referring to FIG. 1A, depicted at the bottom right is a subset of nine known predictor directions from the 33 possible predictor directions (corresponding to the 33 angular modes of the 35 intra modes) in H.265. The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right and at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from horizontal.
[0012] 1A, a square block (104) of 4x4 samples (indicated by a thick dashed line) is depicted in the upper left. The square block (104) contains 16 samples, each labeled with "S," its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Because the block is 4x4 samples in size, S44 is located in the lower right. Reference samples, which follow a similar numbering scheme, are also shown. The reference samples are labeled R, their Y position (e.g., row index), and their X position (column index) relative to the block (104). In both H.264 and H.265, negative values need not be used because the predicted samples are adjacent to the block being reconstructed.
[0013] Intra-picture prediction may work by copying reference sample values from adjacent samples as assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction consistent with the arrow (102) for this block, i.e., the sample is predicted from one or more prediction samples located to the upper right at a 45-degree angle from the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, the values of multiple reference samples may be combined, for example by interpolation, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees.
[0015] The number of possible directions has increased as video coding technology has evolved. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and as of the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific entropy coding techniques are used to represent these likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predicted from neighboring directions used in adjacent, already decoded blocks.
[0016] Figure 1B shows a schematic diagram (105) showing 65 intra-prediction directions with JEM to illustrate the increasing number of prediction directions over time.
[0017] The mapping of intra-prediction direction bits in a coded video bitstream to represent directions may vary from one video coding technique to another, ranging, for example, from a simple direct mapping of prediction directions to intra-prediction modes to complex adaptation schemes involving codewords, most likely modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, these less likely directions are represented by more bits than more likely directions in well-performing video coding techniques.
[0018] Motion compensation may be a lossy compression technique in which blocks of sample data from a previously reconstructed picture or portion thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used to predict a newly reconstructed picture or portion thereof. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the latter may indirectly be a temporal dimension).
[0019] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, e.g., from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in decoding order. Doing so can significantly reduce the amount of data required to code the MV, resulting in the elimination of redundancy and increased compression. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because regions larger than the region to which a single MV is applicable move in similar directions and therefore, in some cases, there is a statistical likelihood that the MV can be predicted using a similar MV derived from the MVs of neighboring regions. This results in the MV found for a given region being similar or identical to the MV predicted from surrounding MVs, and as a result, after entropy coding, it can be represented with fewer bits than would be used to code the MV directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when calculating a predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, a technique hereafter referred to as "spatial merging" is described herein.
[0021] Referring to Figure 1C, the current block (111) may contain samples found by the encoder during a motion search process that are predictable from a spatially shifted previous block of the same size. Instead of coding its MV directly, the MV may be derived from metadata associated with one or more reference pictures, e.g., the most recent reference picture (in decoding order), using the MV associated with any one of five surrounding samples represented by A0, A1, and B0, B1, B2 (112 to 116, respectively). In H.265, MV prediction may use a predictor from the same reference picture as neighboring blocks. Summary of the Invention [Means for solving the problem]
[0022] An aspect of the present disclosure provides an apparatus for processing media streams. The apparatus includes a processing circuit that sends a message to a media-capable network element configured to process multiple audio streams of a conference call. The message indicates that the multiple audio streams should be downmixed by the media-capable network element. The media-capable network element is a content-capable network router that intelligently performs appropriate processing (routing, filtering, adaptation, security operations, etc.) by considering content type, content characteristics (described by metadata or extracted by protocol field analysis), network characteristics, and / or network status. Downmixing generally refers to the process of rendering more audio content channels to fewer speakers. The processing circuit receives the downmixed multiple audio streams from the media-capable network element and decodes the downmixed multiple audio streams for receiving the conference call.
[0023] In one embodiment, multiple audio streams are downmixed by a media-capable network element to one of a single stereo stream and a mono stream.
[0024] In one embodiment, multiple audio streams are downmixed based on multiple audio mixing parameters contained in a Session Description Protocol (SDP) message.
[0025] In one embodiment, a number of audio mixing parameters are set based on the field of view displayed on the user device.
[0026] In one embodiment, each of a first subset of the plurality of audio streams is associated with a respective one of the one or more 360-degree immersive video streams.
[0027] In one embodiment, each of the second subset of the plurality of audio streams is associated with a respective one of the one or more overlay video streams.
[0028] In one embodiment, audio mixing parameters for each of the second subset of the plurality of audio streams are set based on a priority of an overlay video stream associated with each one of the second subset of the plurality of audio streams.
[0029] In one embodiment, the number of the second subset of the plurality of audio streams is determined based on one or more priorities of the one or more overlay video streams associated with the second subset of the plurality of audio streams.
[0030] In one embodiment, the number of multiple audio streams to be downmixed by the media-capable network element is determined based on audio mixing parameters of the multiple audio streams.
[0031] In one embodiment, the processing circuitry performs downmixing of the downmixed audio streams and of the conference call audio streams that have not been downmixed by the media-capable network element.
[0032] Aspects of the present disclosure provide a method for processing media streams, in which a message is sent from a user device to a media-enabled network element configured to process multiple audio streams of a conference call, the message indicating that the multiple audio streams should be downmixed by the media-enabled network element, the downmixed multiple audio streams being received from the media-enabled network element and decoded for receiving the conference call.
[0033] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one or combination of methods for processing a media stream.
[0034] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0035] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 1 is a diagram of an exemplary intra-prediction direction. [Figure 1C] FIG. 1 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 2] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 3] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 5] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 6] 4 shows a block diagram of an encoder according to another embodiment; [Figure 7] 4 shows a block diagram of a decoder according to another embodiment; [Figure 8] 1 illustrates an exemplary immersive video conference call according to one embodiment of the present disclosure. [Figure 9] 1 illustrates another exemplary immersive video conference call according to an embodiment of the present disclosure. [Figure 10] 1 illustrates yet another exemplary immersive video conference call according to an embodiment of the present disclosure. [Figure 11] FIG. 1 is an exemplary flowchart according to one embodiment. [Figure 12]FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0036] I. Video Decoder and Encoder Systems Figure 2 shows a simplified block diagram of a communication system (200) according to one embodiment of the present disclosure. The communication system (200) includes multiple terminal devices that may communicate with each other, for example, via a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via the network (250). In the example of Figure 2, the first pair of terminal devices (210) and (220) perform unidirectional transmission of data. For example, the terminal device (210) may code video data (e.g., a stream of video pictures captured by the terminal device (210)) for transmission to another terminal device (220) via the network (250). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (220) may receive the coded video data from the network (250), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. One-way data transmission may be common, such as in media serving applications.
[0037] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) that perform bidirectional transmission of coded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (230) and (240) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (230) and (240) over the network (250). Each of the terminal devices (230) and (240) may also receive coded video data transmitted by the other of the terminal devices (230) and (240), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device according to the reconstructed video data.
[0038] In the example of FIG. 2 , terminal devices 210, 220, 230, and 240 may be depicted as a server, a personal computer, and a smartphone, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure are contemplated for use with laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 250 represents any number of networks that convey coded video data between terminal devices 210, 220, 230, and 240, including, for example, wired and / or wireless communication networks. Communication network 250 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 250 may not be important to the operation of the present disclosure, unless otherwise described herein.
[0039] 3 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0040] The streaming system may include a video source (301), such as a capture subsystem (313), which may include a digital camera, that creates a stream of uncompressed video pictures (302). In one example, the stream of video pictures (302) includes samples taken by the digital camera. The stream of video pictures (302), shown as a thick line to emphasize its high data volume compared to the encoded video data (304) (or coded video bitstream), may be processed by an electronic device (320) including a video encoder (303) connected to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (304) (or coded video bitstream (304)), shown as a thin line to emphasize its lower data volume compared to the stream of video pictures (302), may be stored on a streaming server (305) for future use. One or more streaming client subsystems, such as the client subsystems (306) and (308) of Figure 3, may access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) may include a video decoder (310), for example, within an electronic device (330). The video decoder (310) decodes an input copy (307) of the encoded video data and generates an output stream of video pictures (311) that can be rendered on a display (312) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.
[0041] It should be noted that the electronic devices (320) and (330) may include other components (not shown). For example, the electronic device (320) may include a video decoder (not shown), and the electronic device (330) may also include a video encoder (not shown).
[0042] 4 shows a block diagram of a video decoder (410) according to one embodiment of the present disclosure. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., a receiving circuit). The video decoder (410) may be used in place of the video decoder (310) in the example of FIG. 3.
[0043] The receiver (431) can receive one or more coded video sequences to be decoded by the video decoder (410), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (401), which can be a hardware / software link to a storage device that stores the encoded video data. The receiver (431) can receive the encoded video data along with other data, such as a coded audio data stream and / or ancillary data stream, which can be forwarded to a respective using entity (not shown). The receiver (431) can separate the coded video sequences from other data. To combat network jitter, a buffer memory (415) can be connected between the receiver (431) and the entropy decoder / parser (420) (hereinafter "parser (420)"). In certain applications, the buffer memory (415) is part of the video decoder (410). In other applications, the buffer memory may be external to the video decoder (410) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (410), e.g., to combat network jitter, plus another buffer memory (415) internal to the video decoder (410), e.g., to handle playout timing. When the receiver (431) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (415) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (415) may be required and may be relatively large, advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (410).
[0044] The video decoder (410) may include a parser (420) for reconstructing symbols (421) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (410) and, in some cases, information for controlling a rendering device (e.g., a display screen), such as the rendering device (412) that is not an integral part of the electronic device (430) but may be connected to the electronic device (430) as shown in FIG. 4. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (420) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may conform to a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (420) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (420) may also extract information such as transform coefficients, quantization parameter values, MVs, etc. from the coded video sequence.
[0045] The parser (420) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (415) to create symbols (421).
[0046] The reconstruction of the symbols (421) can involve several different units, depending on the type of video picture or portion thereof being coded (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (420). The flow of such subgroup control information between the parser (420) and the following units is not shown for clarity.
[0047] Beyond the functional blocks already described, the video decoder (410) can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0048] The first unit is a scalar / inverse transform unit (451), which receives quantized transform coefficients and control information from the parser (420) as symbols (421), including the transform to be used, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (451) can output blocks containing sample values that can be input to an aggregator (455).
[0049] In some cases, the output samples of the scaler / inverse transform (451) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (458). The current picture buffer (458), for example, buffers partially reconstructed and / or fully reconstructed current pictures. The aggregator (455) optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451).
[0050] In other cases, the output samples of the scalar / inverse transform unit (451) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (453) may access a reference picture memory (457) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (421) related to the block, these samples (in this case referred to as residual samples or residual signals) may be added to the output of the scalar / inverse transform unit (451) by an aggregator (455) to generate output sample information. The addresses in the reference picture memory (457) from which the motion-compensated prediction unit (453) fetches prediction samples may be controlled by MVs available to the motion-compensated prediction unit (453), for example, in the form of symbols (421) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (457) when sub-sample accurate MVs are used, MV prediction mechanisms, etc.
[0051] The output samples of the aggregator (455) may be subjected to various loop filtering techniques in a loop filter unit (456). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (456) as symbols (421) from the parser (420), but may also be responsive to meta-information obtained during decoding of a coded picture or previous (in decoding order) portion of the coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0052] The output of the loop filter unit (456) may be a sample stream that can be output to a rendering device (412) and stored in a reference picture memory (457) for use in future inter-picture prediction.
[0053] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (420)), the current picture buffer (458) can become part of the reference picture memory (457), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.
[0054] The video decoder (410) may perform decoding operations according to a predetermined video compression technique in a standard, such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select some tools from all tools available in the video compression technique or standard as the only tools available for use under that profile. Compliance may also be required for the complexity of the coded video sequence to be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0055] In one embodiment, the receiver (431) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0056] 5 shows a block diagram of a video encoder (503) according to one embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) may be used in place of the video encoder (303) of the example of FIG. 3.
[0057] The video encoder (503) may receive video samples from a video source (501) (not part of the electronic device (520) in the example of FIG. 5) that may capture video images to be coded by the video encoder (503). In other examples, the video source (501) is part of the electronic device (520).
[0058] The video source (501) can provide a source video sequence to be coded by the video encoder (503) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (501) can be a storage device that stores already prepared video. In a video conferencing system, the video source (501) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, each of which can contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0059] According to one embodiment, the video encoder (503) may code and compress pictures of a source video sequence into a coded video sequence (543) in real time or under any other time constraints required by the application. Enforcing the appropriate coding rate is one function of the controller (550). In some embodiments, the controller (550) controls and is operatively connected to other functional units as described below. For clarity, connections are not shown. Parameters set by the controller (550) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, and maximum MV-allowed reference region. The controller (550) may be configured with other appropriate functions associated with the video encoder (503) optimized for a particular system design.
[0060] In some embodiments, the video encoder (503) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (530) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs symbols to generate sample data in a manner similar to that generated by a (remote) decoder (when any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (534). Because decoding of the symbol stream produces bit-exact results regardless of the location (local or remote) of the decoder, the contents of the reference picture memory (534) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, for example, due to channel errors) is also used in several related technologies.
[0061] The operation of the "local" decoder (533) may be the same as the operation of a "remote" decoder, such as the video decoder (410), already described in detail above in connection with Figure 4. However, with brief reference also to Figure 4, because symbols are available and the encoding / decoding of symbols into / from a coded video sequence by the entropy coder (545) and parser (420) may be lossless, the entropy decoding portion of the video decoder (410), including the buffer memory (415) and parser (420), may not be implemented entirely in the local decoder (533).
[0062] An observation that can be made at this point is that any decoder technology other than parsing / entropy decoding present in a decoder must also necessarily be present in the corresponding encoder, in substantially the same functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology described generically. Only in certain areas is more detailed description required, and is provided below.
[0063] In operation, in some examples, the source coder (530) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this way, the coding engine (532) codes differences between pixel blocks of the input picture and pixel blocks of the reference picture(s) that may be selected as predictive reference(s) for the input picture.
[0064] The local video decoder (533) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (530). The operation of the coding engine (532) may preferably be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence may typically be a copy of the source video sequence, with some errors. The local video decoder (533) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (534). In this way, the video encoder (503) can locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures (without transmission errors) obtained by the far-end video decoder.
[0065] The predictor (535) may perform the prediction search for the coding engine (532). That is, for a new picture to be coded, the predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture MV and block shape, that may serve as suitable prediction references for the new picture. The predictor (535) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (535), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (534).
[0066] The controller (550) may manage the coding operations of the source coder (530), including, for example, setting parameters and subgroup parameters used to code the video data.
[0067] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (545), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0068] The transmitter (540) may buffer the coded video sequence(s) created by the entropy coder (545) in preparation for transmission over a communication channel (560), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (540) may merge the coded video data from the video coder (503) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0069] The controller (550) may manage the operation of the video encoder (503). During coding, the controller (550) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures are often assigned as one of the following picture types:
[0070] An intra-picture (I-picture) may be a picture that can be coded and decoded without using any other picture in a sequence as a prediction source. Some video codecs allow various types of intra-pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and functions.
[0071] A predicted picture (P picture) may be one that can be coded and decoded using intra prediction or inter prediction, which uses at most one MV and a reference index to predict the sample values of each block.
[0072] A bidirectionally predicted picture (B-picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses up to two MVs and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0073] A source picture is generally spatially subdivided into multiple sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 blocks each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0074] The video encoder (503) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (503) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0075] In one embodiment, the transmitter (540) may transmit additional data along with the coded video. The source coder (530) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0076] Video may be captured in time sequence as multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a vector vector (MV). The MV points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0077] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the current picture in the video in decoding order (but may be past and future, respectively, in display order). A block in the current picture may be coded by a first MV that points to a first reference block in the first reference picture and a second MV that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first reference block and the second reference block.
[0078] Furthermore, merge mode techniques can be used for inter-picture prediction to improve coding efficiency.
[0079] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) of 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0080] 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processed block (e.g., a predictive block) of sample values in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of the coded video sequence. In one example, the video encoder (603) is used in place of the video encoder (303) of the example of FIG. 3.
[0081] In an HEVC example, the video encoder (603) receives a matrix of sample values for a processing block, such as a predictive block, such as 8x8 samples. The video encoder (603) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, e.g., using rate-distortion optimization. If the processing block is to be coded in intra-mode, the video encoder (603) may use intra-prediction techniques to encode the processing block into a coded picture. If the processing block is to be coded in inter-mode or bi-predictive mode, the video encoder (603) may use inter-prediction or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode, and the MVs are derived from one or more MV predictors without the benefit of coded MV components outside the predictors. In certain other video coding techniques, there may be MV components applicable to the current block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.
[0082] In the example of Figure 6, the video encoder (603) includes an inter-encoder (630), an intra-encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general controller (621), and an entropy encoder (625), which are connected to each other as shown in Figure 6.
[0083] The inter-encoder (630) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-prediction information (e.g., description of redundant information by inter-encoding technique, MV, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.
[0084] The intra encoder (622) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, generate transformed and quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (622) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.
[0085] The general-purpose controller (621) is configured to determine general-purpose control data and control other components of the video encoder (603) based on the general-purpose control data. In one example, the general-purpose controller (621) determines the mode of the block and provides a control signal to the switch (626) based on the mode. For example, when the mode is intra mode, the general-purpose controller (621) controls the switch (626) to select intra-mode results for use by the residual calculator (623) and controls the entropy encoder (625) to select intra-prediction information and include the intra-prediction information in the bitstream. When the mode is inter mode, the general-purpose controller (621) controls the switch (626) to select inter-prediction results for use by the residual calculator (623) and controls the entropy encoder (625) to select inter-prediction information and include the inter-prediction information in the bitstream.
[0086] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (622) or the inter-encoder (630). The residual encoder (624) is configured to operate on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (624) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform to generate decoded residual data. The decoded residual data may be used by the intra-encoder (622) and the inter-encoder (630) as appropriate. For example, the inter-encoder (630) may generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (622) may generate decoded blocks based on the decoded residual data and intra-prediction information. In some examples, the decoded blocks are appropriately processed to generate a decoded picture, which may be buffered in a memory circuit (not shown) and used as a reference picture.
[0087] The entropy encoder (625) is configured to format a bitstream to include the encoded block. The entropy encoder (625) is configured to include various information in accordance with an appropriate standard, such as HEVC. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that, according to the disclosed subject matter, there is no residual information when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode.
[0088] 7 shows a diagram of a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (710) is used in place of the video decoder (310) of the example of FIG. 3.
[0089] In the example of Figure 7, the video decoder (710) includes an entropy decoder (771), an inter-decoder (780), a residual decoder (773), a reconstruction module (774), and an intra-decoder (772), which are connected to each other as shown in Figure 7.
[0090] The entropy decoder (771) may be configured to reconstruct, from the coding picture, certain symbols representing the syntax elements that make up the coding picture. Such symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-predictive mode, etc., the latter two being merged or other submodes), prediction information (e.g., intra-predictive information or inter-predictive information), which may identify certain samples or metadata used for prediction by the intra decoder (772) or inter decoder (780), respectively, and residual information, for example in the form of quantized transform coefficients. In one example, if the prediction mode is an inter mode or bi-predictive mode, the inter-predictive information is provided to the inter decoder (780), and if the prediction type is an intra-predictive type, the intra-predictive information is provided to the intra decoder (772). The residual information, which may undergo inverse quantization, is provided to the residual decoder (773).
[0091] The inter decoder (780) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0092] The intra decoder (772) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0093] The residual decoder (773) is configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (771) (a data path is not shown because this may be only a small amount of control information).
[0094] The reconstruction module (774) is configured to combine, in the spatial domain, the residual output by the residual decoder (773) with a prediction result (possibly output by an inter-prediction module or an intra-prediction module) to form a reconstructed block, which may be part of a reconstructed picture and, therefore, part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.
[0095] It should be noted that the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using any suitable technology. In one embodiment, the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (303), (503), and (603) and the video decoders (310), (410), and (710) may be implemented using one or more processors executing software instructions.
[0096] II. Immersive Videoconferencing and Telepresence This disclosure presents a method for audio and video mixing for immersive videoconferencing and telepresence for remote terminals (ITT4RT). According to aspects of the present disclosure, when an end client device or user device has limited capabilities and / or is not compatible with the Multimedia Telephony Service for Internet Protocol Multimedia Subsystem (MTSI), all or part of the media processing can be offloaded from the end client device to a media-capable network element, such as a multimedia resource function (MRF) and / or media control unit (MCU). The media-capable network element may be a content-capable network router that can intelligently perform appropriate processing (e.g., routing, filtering, adaptation, security actions) by considering content type, content characteristics (described by metadata or extracted by protocol field analysis), network characteristics, and / or network status. Downmixing generally refers to the process of rendering more audio content channels to fewer speakers. In some embodiments, the downmixing of audio streams for immersive and overlay streams and / or the combining of audio and video streams may be handled using network-based media processing (NBMP) in media-capable network elements.
[0097] When an omnidirectional media stream is used, only the portion of the omnidirectional media content that corresponds to the user's viewport may be rendered. For example, if a user uses a head-mounted display (HMD), the omnidirectional media content that corresponds to the user's viewpoint may provide the user with a realistic view of the media stream.
[0098] 8 illustrates an exemplary immersive video conference call according to one embodiment of the present disclosure. The exemplary immersive video conference call is organized between conference room A (801), user B (802), and user C (803). Conference room A (801) is equipped with an omnidirectional camera (804), and users B (802) and C (803) are remote participants using an HMD and a mobile device, respectively. In this embodiment, participants B (802) and C (803) can send their viewport directions to conference room A (801), which then sends viewport-dependent streams from the omnidirectional camera (804) to participants B (802) and C (803).
[0099] 9 illustrates another exemplary immersive videoconference call according to one embodiment of the present disclosure. The exemplary videoconference call is organized between multiple conference rooms (901-904), user B (905) using an HMD to receive the videoconference stream, and user C (906) using a mobile device to receive the videoconference stream. Participants B (905) and C (906) can transmit their viewport directions to one of the conference rooms (901-904), which then transmits viewport-dependent streams from its omnidirectional camera to participants B (802) and C (803).
[0100] FIG. 10 illustrates yet another exemplary immersive video conference call according to one embodiment of the present disclosure. The exemplary immersive video conference call is set up using an MRF (or MCU) (1005), which is a multimedia server that provides media-related functions for bridging terminals in a multi-party conference call. In this embodiment, multiple conference rooms (1001-1004) can send their respective videos to the MRF (or MCU). These videos are viewport-independent. That is, the entire 360-degree video of the conference room can be sent to the media server regardless of the user's viewport. The media server can receive the viewport directions of users B (1006) and C (1007) and send viewport-dependent streams to users B (1006) and C (1007) accordingly.
[0101] 9 and 10, remote users B and C can select to view one of the available 360-degree videos from conference rooms (901-904) and (1001-1004), respectively. For example, remote user B can send information about the video stream selected by user B and the viewport direction of user B to the conference room or MRF (or MCU). Remote user B can trigger switching from one room to another based on the active speaker. In addition, the media server can pause receiving video streams from conference rooms where no active users are present.
[0102] ISO 23090-2 defines an overlay as "a portion of an audiovisual medium rendered on a 360-degree video, image item, or viewport." Referring back to Figures 9-10, when a presentation is shared by a participant in a conference room, such as Conference Room A, in addition to being displayed in Conference Room A, the presentation can also be broadcast as a stream to other remote users. The stream can be overlaid on the 360-degree video of Conference Room A. In addition, overlays can also be used for 2D stream distribution.
[0103] According to aspects of the present disclosure, when a terminal device has limited capabilities or when decoding and rendering multiple video and / or audio streams is difficult, NBMP may be used to offload processing of media streams from the terminal device, thereby increasing the battery life of the terminal device.
[0104] Embodiments of the present disclosure include methods for offloading audio processing from an end user to a media-capable network element, such as an MRF (or MCU). Furthermore, embodiments of the present disclosure introduce audio mixing parameters used to downmix audio streams. In addition, embodiments of the present disclosure include methods for encoding and / or decoding audio and video streams at a media-capable network element.
[0105] When transmitting an immersive stream in a video conference, if an overlay video or image is superimposed on the 360° video, several pieces of information may be included: overlay source, overlay rendering type, rendering characteristics (e.g., opacity), user interaction characteristics, etc. The overlay source specifies the image or video to be used as an overlay. The overlay rendering type describes whether the overlay is anchored relative to the viewport or the sphere.
[0106] 9-10, each of the multiple conference rooms is equipped with an omnidirectional media capture device, such as an omnidirectional camera. A remote user can select a video stream and / or audio stream corresponding to one of the multiple conference rooms to be displayed as an immersive stream. Additional audio or video streams used with the 360-degree immersive stream may be transmitted as separate streams, referred to as overlays.
[0107] In one embodiment, upon receiving multiple audio streams, the terminal device can decode and mix the rendered multiple audio streams. A conference room sender device transmitting multiple audio streams can provide mixing levels for the different audio streams using the Session Description Protocol (SDP).
[0108] In one embodiment, the conference room sender device can use SDP to update the mixing levels of different audio streams during a video conference session.
[0109] In one embodiment, the terminal device can customize different mixing levels based on user preferences to override the mixing levels set by the sending device.
[0110] In some embodiments, upon receiving the video and audio media streams, the terminal device can decode the streams and render them on a display. When the terminal device has limited capabilities, some or all of the media processing can be offloaded to a media-capable network element, such as an MRF (or MCU), as shown in Figure 10.
[0111] In one embodiment, for audio streams, when the terminal device has limited processing capabilities, for example, due to limited battery power, or when the terminal device receives multiple media streams and has difficulty decoding and rendering the streams, the terminal device can offload media processing to the MRF (or MCU). These audio streams may include audio streams for 360-degree video and / or audio streams for overlays.
[0112] In one embodiment, when audio processing is offloaded to the MRF (or MCU), multiple audio streams can be downmixed to a single stereo or mono stream, which can be sent to the terminal device for rendering.
[0113] In one embodiment, when multiple audio streams are downmixed, audio mixing parameters such as loudness can be defined by the sending device or customized by the end user.
[0114] In one embodiment, once the audio mixing parameters are defined by the sending device, the sending device can set all audio streams to the same level, which can be sent from the sending device to the MRF (or MCU) via SDP signaling.
[0115] In one embodiment, the sending device can set the audio level of the 360-degree background video to be higher than the audio levels of the overlay audio streams. All of the overlay audio streams can have the same audio parameters. These audio levels can be sent from the sending device to the MRF (or MCU) via SDP signaling.
[0116] In one embodiment, the sending device can set the audio level of the 360-degree background video to be higher than the audio level of the overlay audio stream. The audio mixing level of the overlay stream can be set based on the priority of the overlay stream. These audio levels can be sent from the sending device to the MRF (or MCU) via SDP signaling.
[0117] In one embodiment, the MRF (or MCU) can send the same audio stream to multiple terminal devices when the multiple terminal devices may not have sufficient processing power.
[0118] In one embodiment, once audio mixing parameters are defined by the terminal device, individual audio streams can be encoded for each user by the sending device or the MRF (or MCU). The audio mixing parameters can be set based on the user's field of view (FoV). For example, audio streams for overlays within the FoV can be mixed louder than other streams. These parameters can be negotiated with the MRF (or MCU) during setup or between sessions using SDP signaling.
[0119] According to aspects of the present disclosure, when a terminal device does not support MTSI ITT4RT, media processing for video, or both audio and video, can be offloaded from the terminal device to a media-capable network element such as an MRF (or MCU), thereby providing backward compatibility for MTSI terminals.
[0120] In one embodiment, when the capabilities of the terminal device are limited, audio and video mixing can be performed in the MRF (or MCU).
[0121] In one embodiment, when the MRF (or MCU) has limited capacity, the MRF (or MCU) can generate the same audio and video streams by combining multiple audio and video streams for the same sending device and send the same audio and video streams to all MTSI devices with limited capacity.
[0122] In one embodiment, the MRF (or MCU) can use SDP signaling to negotiate with all or a subset of MTSI devices for a common set of configurations for audio mixing and single video compositing of 360-degree video, as well as various overlays.
[0123] In one embodiment, audio mixing parameters such as audio mixing weights can be defined for each audio stream. A default value for the audio mixing weights can be set based on sound intensity, which can be defined as the power carried by a sound wave per unit area in a direction perpendicular to the area. A default value for the audio mixing weights can also be set based on overlay priority. A default value for the audio mixing weights can be defined by the sending device.
[0124] In one embodiment, audio mixing weights can be customized by end users during a conference session via SDP signaling. Customizing audio mixing weights is useful in several situations, such as when a user wants to hear or focus on one particular audio stream. Another example is when the quality of the downmixed audio is unacceptable for some reason, such as audio level variations, audio quality, or poor SNR channel.
[0125] In one embodiment, the default values of audio mixing weights are r0, r1,...,rN for 360-degree video (a0) and a1, a2,...,aN for overlay video, respectively. The audio output is equal to r0*a0+r1*a1+......+rN*aN, where r0+r1+...+rN=1. The receiver or MRF (or MCU) mixes the audio sources in proportion to the corresponding weights. The default values of audio mixing weights can be changed by the sending device according to the use case.
[0126] In one embodiment, audio mixing weights can be defined based on the priority of the audio streams for the overlay. Default values for audio mixing weights can be defined by the sending device during the conference session using SDP signaling and / or customized by the end user. When overlay priority is used, the sending device can be informed of all overlays of other sending devices and the priority of these overlays during the session. The sending device can then assign audio mixing weights accordingly.
[0127] In one embodiment, the processing capabilities of the terminal device can determine the mixing of audio streams. The receiver can offload the downmixing of all streams to the MRF (or MCU) or only part of the processing. For example, the downmixing of overlay streams can be offloaded to the MRF (or MCU), and the downmixing of downmixed overlay streams and 360-degree streams can be processed on the receiver side.
[0128] In one embodiment, audio mixing can be performed by adding all audio streams from the 360-degree background video together with the overlay stream. In one example, the aggregated audio (a mix of all audio streams) can be normalized. This process can be performed in the receiver or the MRF (or MCU). This process is applicable when there is no background noise or interference from the overlay or 360-degree video audio streams, or when the intensity levels of all streams are approximately the same or the variance is below a threshold.
[0129] In one embodiment, when multiple audio streams are normalized and aggregated into a single audio stream, it may be difficult to distinguish one audio stream from another. Therefore, the sending device may choose to mix only a select number of audio streams. This selection may be based on audio mixing weights defined by an algorithm or overlay stream priorities.
[0130] In one embodiment, the receiver can choose to change the selection of audio streams to be mixed by changing the corresponding audio mixing weights or by mixing a subset of the audio streams.
[0131] In one embodiment, when the sound intensity variance of the audio streams is greater than a threshold, the default audio mixing weights for all overlay and 360-degree audio streams may be set to the same level.
[0132] In one embodiment, the number of downmixed audio streams can be limited. Such a limit can be imposed by the receiver based on its resource capacity or when it has difficulty distinguishing between audio streams from different rooms. If such a limit is imposed, the selection criteria for the downmixed audio streams can be based on sound intensity or overlay priority defined by the sending device for each audio stream. This selection can be changed by the end user during the session via SDP signaling.
[0133] In one embodiment, only one audio stream can be prioritized compared to other audio streams, including the 360-degree stream, for example, when the audio stream of a person speaking / presenting needs to be focused on, the audio mixing weight of other streams can be reduced.
[0134] In one embodiment, for example, when a remote user is presenting and the 360-degree audio stream has background noise, the audio stream of the 360-degree background video can be downmixed with an audio mixing weight that is smaller than the audio mixing weight of the overlay stream. The audio mixing weights can be customized by users already in a session by reducing the audio mixing weights during the session, but the default mixing weights may also need to be changed on the sending device. For example, when a new remote user joins a conference, the new remote user can obtain the default mixing weights of the audio streams from the sending device.
[0135] III. Flowchart 11 shows a flowchart outlining an exemplary process (1100) according to one embodiment of the present disclosure. In various embodiments, the process (1100) is performed by a processing circuit, such as a processing circuit of a terminal device (210), (220), (230), or (240), a processing circuit that performs the functions of a video encoder (303), a processing circuit that performs the functions of a video decoder (310), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of an intra-prediction module (452), a processing circuit that performs the functions of a video encoder (503), a processing circuit that performs the functions of a predictor (535), a processing circuit that performs the functions of an intra-encoder (622), or a processing circuit that performs the functions of an intra-decoder (772). In some embodiments, the process (1100) is performed in software instructions, and thus the processing circuit performs the process (1100) when the processing circuit executes the software instructions.
[0136] The process (1100) may generally begin at step (S1110), where the process (1100) sends a message to a media-capable network element configured to process multiple audio streams of the conference call. The message indicates that the multiple audio streams should be downmixed by the media-capable network element. The process (1100) then proceeds to step (S1120).
[0137] In step S1120, the process 1100 receives downmixed audio streams from the media-capable network element, after which the process 1100 proceeds to step S1130.
[0138] In step S1130, the process 1100 decodes the downmixed audio streams to receive the conference call, after which the process 1100 stops.
[0139] In one embodiment, multiple audio streams are downmixed by a media-capable network element to one of a single stereo stream and a mono stream.
[0140] In one embodiment, the multiple audio streams are downmixed based on multiple audio mixing parameters included in a Session Description Protocol (SDP) message.
[0141] In one embodiment, a number of audio mixing parameters are set based on the field of view displayed on the user device.
[0142] In one embodiment, each of a first subset of the plurality of audio streams is associated with a respective one of the one or more 360-degree immersive video streams.
[0143] In one embodiment, each of the second subset of the plurality of audio streams is associated with a respective one of the one or more overlay video streams.
[0144] In one embodiment, audio mixing parameters for each of the second subset of the plurality of audio streams are set based on a priority of an overlay video stream associated with each one of the second subset of the plurality of audio streams.
[0145] In one embodiment, the number of the second subset of the plurality of audio streams is determined based on one or more priorities of the one or more overlay video streams associated with the second subset of the plurality of audio streams.
[0146] In one embodiment, the number of multiple audio streams to be downmixed by the media-capable network element is determined based on audio mixing parameters of the multiple audio streams.
[0147] In one embodiment, the process (1100) performs downmixing of multiple downmixed audio streams and audio streams of a conference call that have not been downmixed by a media-capable network element.
[0148] IV. Computer Systems The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 12 illustrates a computer system (1200) suitable for implementing certain embodiments of the disclosed subject matter.
[0149] The computer software may be coded using any suitable machine code or computer language that can be subjected to assembly, compilation, linking, or similar mechanisms to generate code containing instructions that can be executed by one or more computer central processing units (CPUs) and graphics processing units (GPUs), etc., directly, or through interpretation and execution of microcode, etc.
[0150] The instructions may be executed by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0151] 12 for computer system 1200 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the configuration of components be interpreted as imposing any dependency or requirement on any one or combination of components shown in the exemplary embodiment of computer system 1200.
[0152] The computer system (1200) may include certain human interface input devices that may respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0153] The input human interface devices may include one or more of a keyboard (1201), a mouse (1202), a trackpad (1203), a touch screen (1210), a data glove (not shown), a joystick (1205), a microphone (1206), a scanner (1207), and a camera (1208) (only one of each is shown).
[0154] The computer system (1200) may also include some human interface output devices. Such human interface output devices may stimulate one or more human user senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., touchscreen (1210), haptic feedback via data gloves (not shown), or joystick (1205), although haptic feedback devices that do not also serve as input devices may also be present), audio output devices (e.g., speakers (1209), headphones (not shown), etc.), visual output devices (e.g., screens (1210) including CRT screens, LCD screens, plasma screens, and OLED screens, with or without touchscreen input capabilities, and with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output by means of stereoscopic video output devices, virtual reality glasses (not shown), holographic displays, smoke tanks (not shown), etc.), and printers (not shown). The above-mentioned visual output devices (such as a screen (1210)) can be connected to the system bus (1248) via a graphics adapter (1250).
[0155] The computer system (1200) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1220) with CD / DVD or similar media (1221), thumb drives (1222), removable hard drives or solid-state drives (1223), legacy magnetic media such as tape or floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0156] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0157] The computer system (1200) may also include a network interface (1254) for one or more communications networks (1255). The one or more communications networks (1255) may be, for example, a wireless network, a wired network, an optical network, etc. Furthermore, the one or more communications networks (1255) may be local, wide area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of the one or more communications networks (1255) include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable, satellite, and terrestrial television, vehicular, and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some general-purpose data port or peripheral bus 1249 (e.g., a USB port on the computer system 1200), while others are typically integrated into the computer system 1200 by attaching to the system bus, as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks, the computer system 1200 can communicate with others. Such communication can be one-way, receive-only (e.g., television broadcasts), one-way transmit-only (e.g., a CANbus to a specific CANbus device), or bidirectional, e.g., bidirectional communication with other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks, as described above, can be used with each of these networks and network interfaces.
[0158] The above-mentioned human interface devices, human-accessible storage devices and network interfaces can be attached to the central part (1240) of the computer system (1200).
[0159] The core (1240) may include one or more central processing units (CPUs) (1241), graphics processing units (GPUs) (1242), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1243), task-specific hardware accelerators (1244), and graphics adapters (1250). These devices, along with read-only memory (ROM) (1245), random access memory (1246), and internal mass storage (1247) such as internal hard drives and SSDs that are not user-accessible, may be connected via a system bus (1248). In some computer systems, the system bus (1248) may be manually operable in the form of one or more physical plugs that allow expansion by adding additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core system bus (1248) or via a peripheral bus (1249). In one example, a screen (1210) can be connected to a graphics adapter (1250). Peripheral bus architectures include PCI, USB, and the like.
[0160] The CPU (1241), GPU (1242), FPGA (1243), and accelerator (1244) can execute a number of instructions that can combine to form the above-mentioned computer code. This computer code can be stored in ROM (1245) and / or RAM (1246). Data that changes can also be stored in RAM (1246), while data that does not change can be stored, for example, in built-in mass storage (1247). Cache memory, which can be closely associated with one or more of the CPU (1241), GPU (1242), mass storage (1247), ROM (1245), RAM (1246), etc., can be used to enable fast storage and retrieval from any of the memory devices.
[0161] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0162] By way of example and not limitation, a computer system having the architecture (1200), and in particular the core (1240), may perform its functions as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media related to the user-accessible mass storage discussed above, but may also be specific storage in the core (1240) that is non-transitory, such as the core's internal mass storage (1247) or ROM (1245). Software embodying various embodiments of the present disclosure may be stored on such devices and executed by the core (1240). The computer-readable media may include one or more memory devices or chips, depending on the particular needs. The software may cause the core (1240), and in particular the processor (including a CPU, GPU, FPGA, etc.) within the core (1240), to perform particular processes or portions of particular processes described herein (including determining data structures stored in RAM (1246) and modifying such data structures according to the software-defined processes). Additionally or alternatively, the computer system may perform functions as a result of hard-wired logic or otherwise implemented in circuitry (e.g., accelerator (1244)), which may act in place of or cooperate with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to "software" may encompass logic, and vice versa. Where appropriate, references to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0163] While this disclosure has described several exemplary embodiments, there are modifications, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art can devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. Appendix A: Acronyms ALF: Adaptive Loop Filter AMVP: Advanced Motion Vector Prediction APS: Adaptation Parameter Set ASIC: Application-Specific Integrated Circuit ATMVP: Alternative / Advanced Temporal Motion Vector Prediction AV1:AOMedia Video 1 AV2:AOMedia Video 2 BMS: Benchmark Set BV: Block Vector CANBus: Controller Area Network Bus CB: Coding Block CC-ALF: Cross-Component Adaptive Loop Filter CD: Compact Disc CDEF: Constrained Directional Enhancement Filter CPR: Current Picture Referencing CPU: Central Processing Unit CRT: Cathode Ray Tube CTB: Coding Tree Block CTU: Coding Tree Unit CU: Coding Unit DPB: Decoder Picture Buffer DPCM: Differential Pulse-Code Modulation DPS: Decoding Parameter Set DVD: Digital Video Disc FPGA: Field Programmable Gate Area JCCR: Joint CbCr Residual Coding JVET: Joint Video Exploration Team GOP: Group of Pictures GPU: Graphics Processing Unit GSM: Global System for Mobile communication HDR: High Dynamic Range HEVC: High Efficiency Video Coding HRD: Hypothetical Reference Decoder IBC: Intra Block Copy IC: Integrated Circuit ISP: Intra Sub-Partitions JEM: Joint Exploration Model LAN: Local Area Network LCD: Liquid-Crystal Display LR: Loop Restoration Filter LRU: Loop Restoration Unit LTE: Long-Term Evolution MPM: Most Probable Mode MV: Motion Vector OLED: Organic Light-Emitting Diode PB: Prediction Block PCI: Peripheral Component Interconnect PDPC: Position Dependent Prediction Combination PLD: Programmable Logic Device PPS: Picture Parameter Set PU: Prediction Unit RAM: Random Access Memory ROM: Read-Only Memory SAO: Sample Adaptive Offset SCC: Screen Content Coding SDR: Standard Dynamic Range SEI: Supplementary Enhancement Information SNR: Signal Noise Ratio SPS: Sequence Parameter Set SSD: Solid-state drive TU: Transform Unit USB: Universal Serial Bus VPS: Video Parameter Set VUI: Video Usability Information VVC: Versatile Video Coding WAIP: Wide-Angle Intra Prediction [Explanation of symbols]
[0164] 101 Samples 102 Arrow 103 Arrow 104 Square Blocks 105 Schematic 111 Current Block 112~116 Ambient sample 200 Communication Systems 210 Terminal Devices 220 Terminal Devices 230 Terminal Devices 240 terminal devices 250 Communication Network 301 Video Sources 302 Video Picture Stream 303 Video Encoder 304 encoded video data 305 Streaming Server 306 Client Subsystem 307 Input copy of encoded video data 308 Client Subsystem 309 Copy of encoded video data 310 Video Decoder 311 Video Pictures 312 Display 313 Capture Subsystem 320 Electronic Devices 330 Electronic Devices 401 Channel 410 Video Decoder 412 Rendering Devices 415 Buffer Memory 420 Entropy Decoder / Parser 421 Symbol 430 Electronic Devices 431 Receiver 451 Scaler / Descaler Unit 452 Intra-picture prediction unit, intra-prediction module 453 Motion Compensation Prediction Unit 455 Aggregator 456 Loop Filter Unit 457 Reference Picture Memory 458 Current Picture Buffer 501 Video Sources 503 Video Encoder, Video Coder 520 Electronic Devices 530 Source Coder 532 Coding Engine 533 Local Video Decoder, Local Decoder 534 Reference Picture Memory, Reference Picture Cache 535 Predictor 540 Transmitter 543 Video Sequences 545 Entropy Coder 550 Controller 560 Communication Channels 603 Video Encoder 621 General-purpose controller 622 Intra Encoder 623 Residual Calculator 624 Residual Encoder 625 Entropy Encoder 626 Switch 628 Residual Decoder 630 Interencoder 710 Video Decoder 771 Entropy Decoder 772 Intra Decoder 773 Residual Decoder 774 Reconstruction Module 780 Interdecoder 801 Conference Room A 802 User B 803 User C 804 Omnidirectional Camera Conference Rooms 901-904 905 User B 906 User C Conference Rooms 1001-1004 1005 MRF (or MCU) 1006 User B 1007 User C 1100 processes 1200 Computer Systems and Architecture 1201 keyboard 1202 Mouse 1203 Trackpad 1205 Joystick 1206 Microphone 1207 Scanner 1208 Camera 1209 Speaker 1210 Touch screen, visual output device screen 1220 CD / DVD ROM / RW 1221 Medium 1222 thumb drive 1223 Solid State Drive 1240 Center 1241 Central Processing Unit (CPU) 1242 Graphics Processing Unit (GPU) 1243 Field Programmable Gate Area (FPGA) 1244 Hardware Accelerator 1245 ROM 1246 Random Access Memory 1247 Centrally-built large-capacity storage 1248 system bus 1249 Peripheral Bus 1250 graphics adapter 1254 network interface 1255 Communication Network
Claims
Claim 1: A method for processing a media stream on a user device, comprising: receiving a message from a network element configured to transmit a plurality of audio streams of a conference call, the plurality of audio streams including at least one audio stream corresponding to a 360-degree video and one or more overlay audio streams corresponding to audiovisual media rendered on the 360-degree video, the message including audio mixing parameters corresponding to each of the one or more overlay audio streams, the audio mixing parameters being generated by the network element; receiving a plurality of audio streams of the conference call from the network element; mixing the overlay audio stream according to the audio mixing parameters generated by the network element; generating an audio output for the conference call based on the mixed overlay audio stream; A method comprising:
2. The method of claim 1, further comprising a step of mixing the mixed overlay audio stream and the at least one audio stream of the 360-degree video into one of a single stereo stream and a mono stream.
3. The method described in claim 1, wherein the message is a Session Description Protocol (SDP) message.
4. The method of claim 1, wherein the audio mixing parameters are generated by the network element based on whether the audiovisual media corresponding to the overlay audio stream is within the field of view of the 360-degree video.
5. The method of claim 1, wherein each of the audio streams of the 360-degree video is associated with a respective one of one or more 360-degree immersive video streams.
6. The method of claim 1, wherein each of the overlay audio streams is associated with a respective one of one or more overlay video streams.
7. The method described in claim 6, wherein audio mixing parameters of each of the overlay audio streams are set based on the priority of the overlay video stream associated with each one of the overlay audio streams.
8. The method described in claim 6, wherein the number of overlay audio streams is determined based on one or more priorities of the one or more overlay video streams associated with the overlay audio stream.
9. The method described in claim 1, wherein the number of overlay audio streams mixed is determined based on audio mixing parameters of the overlay audio streams.
10. After receiving the message, receiving an update message, the update message including updated audio mixing parameters corresponding to each of the one or more overlaid audio streams, the updated audio mixing parameters being generated by the network element during the conference call; mixing the overlay audio stream according to the updated audio mixing parameters generated by the network element; generating an updated audio output for the conference call based on the updated audio mixing parameters; The method of claim 1 further comprising:
11. Receiving a message from a network element configured to transmit multiple audio streams of a conference call, the multiple audio streams including at least one audio stream corresponding to 360-degree video and one or more overlay audio streams corresponding to audiovisual media rendered on the 360-degree video, the message including audio mixing parameters corresponding to each of the one or more overlay audio streams, the audio mixing parameters being generated by the network element; receiving the plurality of audio streams of the conference call from the network element; mixing the overlay audio stream according to the audio mixing parameters generated by the network element; generating an audio output for the conference call based on the mixed overlay audio stream; 11. An apparatus comprising: a processing circuit configured to:
12. The device described in Claim 11, wherein the processing circuitry is further configured to mix the mixed overlay audio stream and the at least one audio stream of the 360-degree video into one of a single stereo stream and a mono stream.
13. The device of claim 11, wherein the message is a Session Description Protocol (SDP) message.
14. The device described in claim 11, wherein the audio mixing parameters are generated by the network element based on whether the audiovisual media corresponding to the overlay audio stream is within the field of view of the 360-degree video.
15. The device described in claim 11, wherein each of the audio streams of the 360-degree video is associated with a respective one of one or more 360-degree immersive video streams.
16. The device of claim 11, wherein each of the overlay audio streams is associated with a respective one of one or more overlay video streams.
17. The device described in claim 16, wherein audio mixing parameters of each of the overlay audio streams are set based on the priority of the overlay video stream associated with each one of the overlay audio streams.
18. The device described in claim 16, wherein the number of overlay audio streams is determined based on one or more priorities of the one or more overlay video streams associated with the overlay audio streams.
19. The device of claim 11, wherein the number of mixed overlay audio streams is determined based on audio mixing parameters of the overlay audio streams.
20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor of a user device, cause the at least one processor to: receiving a message from a network element configured to transmit a plurality of audio streams of a conference call, the plurality of audio streams including at least one audio stream corresponding to a 360-degree video and one or more overlay audio streams corresponding to audiovisual media rendered on the 360-degree video, the message including audio mixing parameters corresponding to each of the one or more overlay audio streams, the audio mixing parameters being generated by the network element; receiving the plurality of audio streams of the conference call from the network element; mixing the overlay audio stream according to the audio mixing parameters generated by the network element; generating an audio output for the conference call based on the mixed overlay audio stream; A non-transitory computer-readable storage medium that causes the
Citation Information
Patent Citations
Audio teleconference system
JP2004072354A
Method and system for providing media services
JP2004534457A
Device, method and program for synthesizing video image
JP2008067203A
Screen generation apparatus and screen layout sharing system
JP2009159367A
Server device, terminal, speech communication system and telephone conference system
JP2016220132A