Method, device and storage medium for signaling multiple audio mixing gains using real-time transport protocol (RTP) header extension in a conference call
By using RTP header extension to synchronously notify 360-degree background and overlay audio mixing gain in immersive teleconferencing, the problem of low audio signaling efficiency in the existing technology is solved and efficient audio mixing gain management is achieved.
Patent Information
- Application Number
- CN202280003744.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-18
- Filing Date
- 2022-03-24
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-03-24
AI Technical Summary
The existing technology has difficulty in efficiently signaling multiple audio mixing gains in immersive teleconferencing, resulting in low audio signaling efficiency.
The real-time transport protocol RTP header extension is adopted to notify the 360-degree background and superimposed audio mixing gain through a single RTP header extension, and the element identifier, length of the extension element and the value of the mixing gain of a single RTP header extension are used to realize the synchronous notification of multiple audio mixing gains.
Improved audio signaling efficiency reduces processing requirements and provides the mixed audio or video streams expected in immersive conference calls.
Smart Images

Figure CN115486058B_ABST
Abstract
Description
[0001] Cross-references
[0002] This application is based upon and claims the benefit of U.S. Provisional Patent Application No. 63 / 167,236, filed on March 29, 2021, and U.S. Patent Application No. 17 / 698,064, filed on March 18, 2022, the disclosures of which are incorporated herein by reference. Technical Field
[0003] Embodiments of the present disclosure relate to signaling audio mixing gains for Immersive Teleconferencing and Telepresence for Remote Terminals (ITT4RT), and more specifically to defining a Real-time Transport Protocol (RTP) header extension for signaling all audio mixing gains for 360-degree backgrounds and overlays through a single RTP header extension. Background Art
[0004] When using omnidirectional media streaming, only the portion of the content corresponding to the user's viewport is rendered, while a head-mounted display (HMD) is used to give the user a realistic view of the media stream.
[0005] Figure 1 A related technical scenario (Scenario 1) for an immersive teleconference call is shown, where the call is organized between Room A (101), User B (102) and User C (103). Figure 1 As shown, room A (101) represents a conference room with an omnidirectional / 360-degree camera (104), and user B (102) and user C (103) are remote participants using an HMD and a mobile device, respectively. In this case, participant user B (102) and participant user C (103) send their viewport orientation to room A (101), and room A (101) sends viewport-related streams to user B (102) and user C (103).
[0006] Figure 2AAn extended scenario (Scene 2) is shown, which includes multiple conference rooms (2a01, 2a02, 2a03, 2a04). User B (2a06) uses an HMD to view a video stream from a 360-degree camera (104), and user C (2a07) uses a mobile device to view the video stream. User B (2a06) and user C (2a07) send their viewport orientation to at least one of the conference rooms (2a01, 2a02, 2a03, 2a04), and at least one of the conference rooms (2a01, 2a02, 2a03, 2a04) sends the viewport-related stream to user B (2a06) and user C (2a07).
[0007] like Figure 2B As shown, another example scenario (Scenario 3) is when a call is established using MRF / MCU (2b05), where the Media Resource Function (MRF) and Media Control Unit (MCU) are multimedia servers that provide media-related functions for bridging terminals in a multi-party conference call. Conference rooms can send their respective videos to the MRF / MCU (2b05). These videos are viewport-independent videos, that is, the entire 360-degree video is sent to the media server (i.e., MRF / MCU), regardless of the user viewport that streams the specific video. The media server receives the viewport orientation of the users (User B (2b06) and User C (2b07)) and sends viewport-related streams to the users accordingly.
[0008] Further to scenario 3, the remote user can choose to watch one of the available 360-degree videos from the conference rooms (2a01 to 2a04, 2b01 to 2b04). In this case, the user sends information about the video to be streamed and its viewport orientation to the conference room or MRF / MCU (2b05). The user can also switch from one room to another based on active speakers. The media server can pause receiving video streams from any conference room that has no active users.
[0009] ISO 23090-2 defines an overlay as "a piece of visual media rendered over an omnidirectional video or image item or onto a viewport". When any participant in room A is sharing any presentation, in addition to being displayed in room A, the presentation is also broadcast as a stream to the other users (rooms 2a02 to 2a04, 2b02 to 2b04, user B (2b06) and / or user C (2b07). This stream can be overlaid on top of the 360 video. Additionally, overlays can be used for 2D streams as well. The default audio mixing gains for different audio streams are 360 video (a0) and overlay video (a1, a2, ..., aN ) of the audio gain (r0, r1, …, r N ), and the audio output is equal to r0*a0+r1*a1+……+rn*an, where r0+r 1+ …+r N = 1. The receiver or MRF / MCU mixes the audio source in proportion to its mixing gain. Summary of the Invention
[0010] One or more exemplary embodiments of the present disclosure provide a system and method for signaling audio mixing gains for overlay and 360-degree streams together in a single RTP header extension.
[0011] According to an embodiment, a method for signaling multiple audio mixing gains in a conference call using an RTP header extension is provided. The method may include: receiving an input audio stream from a 360-degree stream, the input audio stream including mixing gains; declaring a single Real-time Transport Protocol (RTP) header extension for the input audio stream, the single Real-time Transport Protocol (RTP) header extension including one or more extension elements; and signaling the mixing gains using the single RTP header extension. The one or more extension elements in the method include an element identifier for the single RTP header extension, a length of the extension element, and a value for the mixing gain.
[0012] According to an embodiment, a device for signaling multiple audio mixing gains using an RTP header extension in a conference call is provided. The device may include at least one memory storing instructions and at least one processor configured to read program code and operate as directed by the program code. The program code includes: receiving code configured to cause at least one processor to receive an input audio stream from a 360-degree video stream, the input audio stream including mixing gains; declaring code configured to cause at least one processor to declare a single RTP header extension for the input audio stream, the single RTP header extension including one or more extension elements, wherein the one or more extension elements include an element identifier of the single RTP header extension, a length of the extension element, and a magnitude of the mixing gains; and signaling code configured to cause at least one processor to signal the mixing gains using the single RTP header extension.
[0013] According to an embodiment, a non-volatile computer-readable medium is provided for signaling multiple audio mixing gains using an RTP header extension during a conference call. The storage medium can be connected to one or more processors and can be configured to store instructions that, when executed, cause at least one or more processors to receive an input audio stream from a 360-degree video stream, the input audio stream including mixing gains; declare a single RTP header extension for the input audio stream, the single RTP header extension including one or more extension elements; and signal the mixing gains using the single RTP header extension. The one or more extension elements of the non-volatile computer-readable storage medium include an element identifier for the single RTP header extension, a length of the extension element, and a magnitude of the mixing gains.
[0014] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other aspects, features, and aspects of embodiments of the present disclosure will become more apparent from the following description with reference to the following drawings.
[0016] Figure 1 is a schematic diagram of the ecosystem for immersive teleconferencing.
[0017] Figure 2A This is a schematic diagram of a multi-party multi-room conference call.
[0018] Figure 2B This is a diagram of a multi-party, multi-room conference call using an MRF / MCU.
[0019] Figure 3 is a simplified block diagram of a communication system in accordance with one or more embodiments.
[0020] Figure 4 is a simplified exemplary illustration of a streaming environment in accordance with one or more embodiments.
[0021] Figure 5A is a diagram illustrating signaling audio mixing gain using a one-byte RTP header extension according to one or more embodiments.
[0022] Figure 5B is a diagram illustrating signaling audio mixing gain using a two-byte RTP header extension according to one or more embodiments.
[0023] Figure 6 is a flow chart of a method for signaling multiple audio mixing gains using an RTP header extension in a conference call according to one or more embodiments.
[0024] Figure 7is a schematic diagram of a computer device according to one or more embodiments. DETAILED DESCRIPTION
[0025] The present disclosure relates to a method and apparatus for signaling audio mixing gains of overlay and 360-degree streams together in a single RTP header extension to provide users with a desired mixed audio or video stream for an immersive teleconference.
[0026] like Figure 2A and Figure 2B As shown, multiple conference rooms with omnidirectional cameras are in a conference call, and a user selects a video / audio stream from one of the conference rooms (2a01, 2a02, 2a03, 2a04) to be displayed as an immersive stream. Any additional audio stream or video stream used with the 360-degree immersive stream is sent as an overlay (i.e., as a separate stream). As soon as the terminal device receives the multiple audio streams, it decodes and mixes them for rendering to the user. The sender conference room provides the mixing gain levels of all different audio streams. The sender conference room can also update the mixing gain levels of different audio streams during the conference call session. Audio mixing gains can be defined for each audio stream. Therefore, it will be necessary to use a single header extension as detailed in the embodiments of the present disclosure to send / receive all audio gains (r0, r1, ..., r N ) and superimposed video (a1, a2, ..., a N Using this method to send multiple audio mixing gains with each header extension can reduce necessary processing and improve audio signaling efficiency.
[0027] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. However, examples of the embodiments may be implemented in various forms, and the present disclosure should not be construed as being limited to the examples described herein. On the contrary, examples of the embodiments are provided to make the technical solutions of the present disclosure more comprehensive and complete, and to fully convey the ideas of the examples of the embodiments to those skilled in the art. The accompanying drawings are merely example illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the accompanying drawings represent the same or similar components, and therefore repeated descriptions of these components are omitted.
[0028] The proposed features discussed below can be used individually or in any combination in any order. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. In addition, these embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits), or in the form of software, or in different networks and / or processor devices and / or microcontroller devices. In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.
[0029] Figure 3 1 is a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) may include at least two terminals (302, 303) interconnected via a network (305). For one-way transmission of data, a first terminal (303) may encode video data at a local location for transmission to another terminal (302) via the network (305). A second terminal (302) may receive encoded video data from the other terminal from the network (305), decode the encoded data, and display the recovered video data. One-way data transmission may be common in media service applications (such as teleconferencing).
[0030] Figure 3 A second pair of terminals (301, 304) is shown, which are provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal (301, 304) can encode video data collected at a local location for transmission to the other terminal via a network (305). Each terminal (301, 304) can also receive encoded video data transmitted by another terminal, can decode and mix the encoded data, and can display the recovered video data on a local display device.
[0031] exist Figure 3 In the embodiment of the present disclosure, the terminals (301, 302, 303, 304) can be illustrated as servers, personal computers and mobile devices, but the principles of the present disclosure are not limited thereto. The embodiments of the present disclosure are applied to laptop computers, tablet computers, HMDs, other media players and / or dedicated video conferencing equipment. The network (305) represents any number of networks that transmit encoded video data between the terminals (301, 302, 303, 304), including, for example, wired and / or wireless communication networks. The communication network (305) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks and / or the Internet. The mixing gains discussed in the embodiments of the present disclosure can be sent and received through the network (305) or the like using the network protocols explained below.
[0032] Figure 4 An example streaming environment for applications of the disclosed subject matter is shown. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, immersive teleconference calls, video teleconferencing, and telepresence.
[0033] The streaming environment may include one or more conference rooms (403), which may include a video source (401), such as a video camera, and one or more participants in a conference (402). Figure 4 The illustrated video source (401) is, for example, a 360-degree video camera that can create a video sample stream. The video sample stream can be sent to a streaming server (404) and / or stored on the streaming server (404) for future use. One or more streaming clients (405, 406) can also send their respective viewport information to the streaming server (404). Based on the viewport information, the streaming server (404) can send a viewport-related stream to the corresponding streaming client (405, 406). In another example embodiment, the streaming client (405, 406) can access the streaming server (404) to retrieve the viewport-related stream. The streaming server (404) and / or the streaming client (405, 406) can include hardware, software, or a combination thereof to enable or implement various aspects of the disclosed subject matter as described in more detail below.
[0034] In an immersive teleconference call, multiple audio streams can be sent from a sender (e.g., 403) to a streaming client (e.g., 405 and / or 406). These streams can include an audio stream of the 360-degree video and one or more superimposed audio streams. The streaming client (405, 406) can include a mixing component (407a, 407b). The mixing component can decode and mix the 360-degree video and the overlaid viewport-related streams and create an output video sample stream that can be presented on a display 408 or other reproduction device such as an HDM, a speaker, a mobile device, etc. The embodiment is not limited to this configuration, and one or more conference rooms (403) can communicate with the streaming client (405, 406) via a network (e.g., network 305).
[0035] Now, reference will be made to the examples Figure 5A and Figure 5B Describes signaling multiple audio mixing gains from a server to a client.
[0036] RTP-based solutions can be used to signal multiple audio mixing gains from the server to the client in a single RTP header extension. The packets of a 360-degree RTP audio stream can contain one or more extension elements of the RTP header extension. Each extension element in the packet indicates the mixing gain presented in the 360-degree RTP audio stream and any overlay audio. Figure 5A As shown, an RTP extension with a one-byte header format can be used. The header extension of the RTP extension with a one-byte header format can have three extension elements. These elements include ID (5a01, 5a04, 5a07, 5a12, 5a13, 5a18), length L (5a02, 5a05, 5a08, 5a09, 5a14, 5a15) and mixing gain (5a03, 5a10, 5a11, 5a06, 5a16, 5a17).
[0037] ID (5a01, 5a04, 5a07, 5a12, 5a13, 5a18) is a 4-bit ID that is a local identifier for the element. The identifier can be used to map the audio mixing gain to an overlay or 360-degree audio RTP stream. Length L (5a02, 5a05, 3a08, 5a09, 5a14, 5a15) is the 4-bit length of the data bytes of the header extension element minus 1 and follows the one-byte header. In some example embodiments, Length L may have a value of zero (0) in the number field to indicate that one byte of data follows. Additionally, a value of 15 (maximum) indicates that 16 bytes of data follow. Mixing Gain (5a03, 5a10, 5a11, 5a06, 5a16, 5a17) represents the magnitude of the mixing gain for a single byte of the header extension.
[0038] like Figure 5B As shown, an RTP extension with a two-byte header format may also be used. Figure 5B The header extension in FIG. 5 is shown as having three extension elements. These elements may include an ID (5b01, 5b03, 5b07), a length L (5b02, 5b04, 5b08), and a mixing gain (5b05, 5b09). In some example embodiments, the two-byte header format may also include a padding (5b06) byte with a value of zero (0).
[0039] ID (5b01, 5b03, 5b07) is an 8-bit ID that is the local identifier of the element. ID (5b01, 5b03, 5b07) can be used to map the audio mixing gain to an overlay or 360-degree audio RTP stream. ID (5b01, 5b03, 5b07) can also include overlay_id, which serves as the overlay identifier of the element. Length L (5b02, 5b04, 5b08) is an 8-bit length field that is the length of the extended data in bytes, excluding the ID and length fields. A value of zero (0) indicates that there is no subsequent data. Mixing Gain (5b05, 5b09) represents the magnitude of the mixing gain.
[0040] In some example embodiments, for ID values in the range 1 to 14, a one-byte header extension with the same meaning may be used. Figure 5A and Figure 5B Three extension elements are shown. However, the embodiment is not limited thereto. The RTP extension may have one or more extension elements.
[0041] In some example embodiments, the declaration and mapping of the audio mixing gain header extension is performed in the Session Description Protocol (SDP) extmap attribute. The Uniform Resource Identifier (URI) used to declare the audio mixing gain header extension in the SDP extmap attribute and map the audio mixing gain header extension to a local extension header identifier is:
[0042] urn:3gpp:rtp-hdrext:audio mixing gain
[0043] The URI identifies and describes the header extension. In some exemplary embodiments, the header extension may be present only in the first packet of the RTP audio stream and may be repeated when the mixing gain needs to be updated for optimality. Furthermore, to avoid redundancy, the header extension may be present in the first few packets of the RTP audio stream and may be repeated only when the mixing gain needs to be updated for optimality. In addition, a predetermined change in the mixing gain may be defined to determine when an update is required.
[0044] In some exemplary embodiments, the overlay audio stream and the 360-degree audio stream may be sent in a single RTP session. If the overlay audio stream and the 360-degree stream are not sent in a single RTP session, the RTP header extension for the 360-degree stream may carry the gain of the overlay audio stream as an extension element to the corresponding RTP header extension, provided that the overlay_id value is used in the ID field.
[0045] Figure 6 is a flow chart of a method 600 for signaling multiple audio mixing gains using an RTP header extension in a conference call, according to an embodiment.
[0046] like Figure 6 As shown, in operation 610, method 600 includes receiving an input audio stream, wherein the input audio stream includes a mixing gain. The input audio stream may be from a 360-degree video / audio stream in a conference call. The mixing gain included in the input audio stream may include an audio gain from the input audio stream and an audio gain from an overlay audio stream.
[0047] In operation 620, method 600 comprises the single RTP header extension that declaration is used for input audio stream.Single RTP header extension comprises one or more extension elements.Each extension element comprises the element identifier that single RTP header extends, the length of extension element and the magnitude of mixing gain.Single RTP header extension can be the form of one byte header extension or two byte header extension, and declares to identify single RTP header extension in the SDP using URI.In certain embodiments, single RTP header extension only repeats based on the variation of a mixing gain in mixing gain.
[0048] In operation 630, method 600 includes signaling the mixing gain using a single RTP header extension. When signaling the mixing gain, the single RTP header extension is present only in the first packet of the input audio stream, or only in the first of multiple consecutive packets of the input audio stream. The input audio stream can be signaled in one or more RTP sessions. If the input audio stream is signaled in more than one RTP session, the single RTP header extension carries the gain of the overlay audio stream as an extension element in the single RTP header extension, and includes the overlay identifier value in the element identifier of the extension element.
[0049] Although Figure 6 An example block diagram of the method is shown, but in some embodiments, the method may include Figure 6 Additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in the method. Additionally or alternatively, two or more blocks of the method may be executed in parallel.
[0050] The above-described techniques for signaling multiple audio mixing gains for teleconferencing and telepresence can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 7 A computer device 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0051] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or similar mechanisms to create code comprising instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0052] These instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, and the like.
[0053] Figure 7The components shown for computer device 700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. The configuration of components should not be interpreted as having any dependency or requirement on any one or combination of components illustrated in the exemplary embodiment of computer device 700.
[0054] Computer device 700 may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, tapping), visual input (such as gestures), and olfactory input. Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images, photographic images obtained from a still camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0055] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard 701 , trackpad 702 , mouse 703 , touch screen 709 , data gloves, joystick 704 , microphone 705 , camera 706 , scanner 707 .
[0056] Computer device 700 may also include certain human-computer interface output devices. Such human-computer interface output devices can stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through touch screen 709, data gloves, or joystick 704, although there may also be tactile feedback devices that are not used as input devices), audio output devices (such as: speakers 708, headphones), visual output devices (such as screen 709, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of these screens can output two-dimensional visual output or more than three-dimensional output through means such as stereo output; virtual reality glasses, holographic displays, and smoke canisters) and printers.
[0057] The computer device 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 711 including CD / DVD etc. media 710, thumb drives 712, removable hard drives or solid-state drives 713, traditional magnetic media such as tapes and floppy disks, dedicated ROM / ASIC / PLD based devices such as secure dongles, and the like.
[0058] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.
[0059] Computer device 700 may also include an interface 715 to one or more communication networks 714. Networks 714 may be, for example, wireless, wired, or optical. Networks 714 may also be local, wide-area, metropolitan, vehicular and industrial, real-time, delay-tolerant, and the like. Examples of networks 714 include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, and the like), television wired or wireless wide-area digital networks (including cable, satellite, and terrestrial broadcast television), vehicle and industrial networks (including CAN Bus), and the like. Some networks 714 typically require an external network interface adapter (e.g., a graphics adapter 725) attached to some common data port or peripheral bus 716 (such as, for example, a USB port of computer device 700); other networks are typically integrated into the core of computer device 700 by attaching to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks 714, computer device 700 can communicate with other entities. Such communication can be one-way, receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of the networks and network interfaces as those described above.
[0060] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel 717 of the computer device 700 .
[0061] The core 717 may include one or more central processing units (CPUs) 718, graphics processing units (GPUs) 719, specialized programmable processing units in the form of field programmable gate arrays (FPGAs) 720, hardware accelerators 721 for certain tasks, and the like. These devices, along with read-only memory (ROM) 723, random access memory (RAM) 724, and internal mass storage devices 722 such as internal non-user accessible hard drives, SSDs, and the like, may be connected via a system bus 726. In some computer systems, the system bus 726 may be accessible in the form of one or more physical plugs to allow expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the core's system bus 726 or through the peripheral bus 716. Peripheral bus architectures include PCI, USB, and the like.
[0062] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can execute certain instructions, the combination of which can constitute the aforementioned computer code. This computer code can be stored in ROM 723 or RAM 724. Transient data can also be stored in RAM 724, while permanent data can be stored, for example, in internal mass storage device 722. Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more of the CPU 718, GPU 719, mass storage device 722, ROM 723, RAM 724, etc.
[0063] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0064] As an example and not by way of limitation, a computer system having architecture 700 and, in particular, kernel 717, can provide functionality as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage devices as described above, as well as certain storage devices of kernel 717 having non-volatile properties (such as kernel internal mass storage devices 722 or ROM 723). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by kernel 717. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause kernel 717, and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM 724 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system may provide functionality (e.g., accelerator 721) as a result of logic hardwired or otherwise embodied in circuitry that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. References to software may encompass logic and vice versa, as appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0065] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.
Claims
1. A method for signaling multiple audio mixing gains using a Real-time Transport Protocol (RTP) header extension in a teleconference, characterized in that: The method is executed by at least one processor and includes: receiving an input audio stream from a 360-degree stream, the input audio stream including mixing gains; declaring a single real-time transport protocol (RTP) header extension of the input audio stream, the single real-time transport protocol (RTP) header extension comprising one or more of an element identifier, a length of the extension element, and a magnitude of the mixing gain; and The mixing gain is signaled using the single real-time transport protocol (RTP) header extension, wherein the single real-time transport protocol (RTP) header extension carries the gain of the overlay audio stream as an extension element to the single real-time transport protocol (RTP) header extension based on the input audio stream sent in more than one RTP session, and an overlay identifier value is included in the element identifier of the extension element.
2. The method according to claim 1, characterized in that The single real-time transport protocol RTP header extension is declared in a session description protocol SDP using a uniform resource identifier (URI) to identify the single real-time transport protocol RTP header extension.
3. The method according to claim 1, characterized in that The single Real-time Transport Protocol (RTP) header extension is present only in the first packet or a plurality of consecutive first packets of the input audio stream.
4. The method according to claim 1, wherein The single Real-time Transport Protocol (RTP) header extension is updated based on a change in one of the mixing gains.
5. The method according to claim 1, wherein The input audio stream includes a first audio gain from the input audio stream and a second audio gain from an overlay audio stream.
6. The method according to claim 5, characterized in that The input audio stream is sent in a single Real-time Transport Protocol (RTP) session.
7. A device for signaling multiple audio mixing gains using a Real-time Transport Protocol (RTP) header extension in a teleconference, characterized in that: The device comprises: at least one memory configured to store program code; and At least one processor is configured to read the program code and operate as instructed by the program code, the program code comprising: receiving code configured to cause the at least one processor to receive an input audio stream from a 360-degree video stream, the input audio stream including a mixing gain; declaration code configured to cause the at least one processor to declare a single real-time transport protocol (RTP) header extension for the input audio stream, the single real-time transport protocol (RTP) header extension comprising one or more of an element identifier, a length of the extension element, and a magnitude of the mixing gain; and signaling code configured to cause the at least one processor to signal the mixing gain using the single real-time transport protocol (RTP) header extension, wherein the single real-time transport protocol (RTP) header extension carries the gain of the overlay audio stream as an extension element to the single real-time transport protocol (RTP) header extension based on the input audio stream sent in more than one RTP session, and an overlay identifier value is included in the element identifier of the extension element.
8. The device according to claim 7, characterized in that The single real-time transport protocol RTP header extension is declared using a session description protocol SDP, and the SDP uses a uniform resource identifier URI to identify the single real-time transport protocol RTP header extension.
9. The device according to claim 7, characterized in that The single RTP header extension is present in the first packet or a plurality of consecutive first packets of the input audio stream, and The single Real-time Transport Protocol (RTP) header extension is updated based on a change in one of the mixing gains.
10. The device according to claim 7, characterized in that The input audio stream includes a first audio gain from the input audio stream and a second audio gain from an overlay audio stream.
11. The device according to claim 10, characterized in that The input audio stream is sent in a single RTP session.
12. A non-volatile computer-readable medium storing instructions, characterized in that: The instructions include: one or more instructions that, when executed by at least one processor of a device that uses a real-time transport protocol (RTP) header extension to signal multiple audio mixing gains in a conference call, cause the at least one processor to execute the method as described in any one of claims 1 to 6.
13. A computer device, characterized in that: include: processor and memory; The memory stores computer codes, and when the computer codes are executed by the processor, the processor executes the method according to any one of claims 1 to 6.