Event-Based Trigger Intervals in RTCP Viewport Signaling for Immersive Videoconferencing and Telepresence for Remote Terminals
Event-based trigger intervals for RTCP viewport feedback signaling address bandwidth and computation overhead in immersive videoconferencing by adjusting packet lengths based on viewport changes, enhancing video delivery efficiency and quality.
Patent Information
- Application Number
- JP2023213952
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-11
- Filing Date
- 2023-12-19
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2041-04-12
AI Technical Summary
Bandwidth and network limitations hinder efficient delivery of immersive video updates in real-time as the spatial orientation of head-mounted displays changes during videoconferencing, leading to increased network overhead and server computation overhead.
Implementing event-based trigger intervals for RTCP viewport feedback signaling to control the delivery of immersive video, adjusting packet lengths based on viewport changes and frequency thresholds to manage bandwidth and computation efficiently.
Reduces network overhead and server computation by optimizing video delivery based on viewport changes, ensuring high-quality immersive video experiences while adhering to bandwidth constraints.
Smart Images

Figure 0007744075000005 
Figure 0007744075000006 
Figure 0007744075000007
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 022,394, filed May 8, 2020, and U.S. Patent Application No. 17 / 095,282, filed November 11, 2020, which are expressly incorporated by reference herein in their entireties.
[0002] The present disclosure relates to the Real-Time Transport Control Protocol (RTCP), and more particularly to event-based trigger intervals in RTCP viewport feedback signaling for immersive videoconferencing and telepresence for remote terminals. [Background technology]
[0003] Immersive video conferencing provides a face-to-face high-definition video and audio experience for meetings. It supports real-time multi-connection streaming of immersive video on a head-mounted display (HMD) device / video player. Immersive video conferencing enables a lifelike communication experience with high-definition video and audio services. It aims to create an immersive experience for users participating in the meeting remotely.
[0004] VR support in Multimedia Telephony Services in IMS (MTSI) and IMS-based telepresence enables support of immersive experiences for remote terminals participating in videoconferencing and telepresence sessions. This allows for two-way audio and one-way immersive video, e.g., a remote user wearing an HMD joins a conference and receives immersive audio and video captured by an omnidirectional camera in the conference room, but transmits only audio and optionally 2D video.
[0005] Bandwidth and other technical limitations have hindered improvements in immersive video delivery with respect to updating the viewpoint margin as the spatial orientation of the HMD is updated remotely in real time.
[0006] Therefore, a technical solution to such problems involving network overhead and server computation overhead is desired. Summary of the Invention [Means for solving the problem]
[0007] To address one or more different technical problems, the present disclosure provides technical solutions for reducing network overhead and server computation overhead while delivering immersive video with respect to one or more viewport margin updates according to example embodiments. Methods and apparatuses include a memory configured to store computer program code and a processor or processors configured to access the computer program code and operate according to instructions of the computer program code, the computer program code including: control code configured to cause at least one processor to control delivery of the video conference call to a viewport; setting code configured to cause the at least one processor to set an event-based threshold for the video conference call; determining code configured to cause the at least one processor to determine whether the event-based threshold has been triggered based on an event and whether the time elapsed since another event is less than a predetermined time; and control code configured to cause the at least one processor to further control delivery of the video conference call to the viewport based on determining whether the event-based threshold has been triggered and whether the time elapsed since another event is less than a predetermined time.
[0008] According to an exemplary embodiment, the event-based threshold includes at least the degree of change in spatial orientation of the viewport.
[0009] According to an exemplary embodiment, determining whether the event-based threshold has been triggered includes determining whether the spatial orientation of the viewport has changed by more than the degree of change of the event-based threshold.
[0010] According to an example embodiment, further controlling the delivery of the video conference call to the viewport includes delivering at least an additional margin of the video conference call to the viewport when it is determined that the spatial orientation of the viewport has changed by more than an event-based threshold degree of change.
[0011] According to an example embodiment, further controlling the delivery of the video conference call to the viewport includes processing packets of different lengths depending on whether a timer is triggered or whether an event-based threshold is triggered.
[0012] According to an exemplary embodiment, the first packet of the packets of different lengths that triggers the timer is longer than the second packet that triggers the event-based threshold.
[0013] According to an exemplary embodiment, the computer program code further includes determining code configured to cause the at least one processor to determine whether a frequency at which the event triggers the event-based threshold exceeds a frequency threshold based on whether an amount of time elapsed since another event is less than a predetermined amount of time.
[0014] According to an exemplary embodiment, the computer program code further includes update code configured to cause the at least one processor to update the timer in response to determining that the frequency at which the event triggers the event-based threshold exceeds the frequency threshold.
[0015] According to an exemplary embodiment, the viewport is a display of at least one of a headset and a handheld mobile device (HMD).
[0016] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a simplified diagram according to one embodiment. [Figure 2] FIG. 1 is a simplified diagram according to one embodiment. [Figure 3] FIG. 2 is a simplified block diagram of a decoder according to one embodiment. [Figure 4] FIG. 2 is a simplified block diagram of an encoder according to one embodiment. [Figure 5] 1 is a simplified diagram of a conference call according to one embodiment. [Figure 6] FIG. 2 is a simplified diagram of a message format according to one embodiment. [Figure 7] FIG. 2 is a simplified block diagram of a picture according to one embodiment. [Figure 8] FIG. 1 is a simplified flow diagram according to one embodiment. [Figure 9] 1 is a simplified flowchart according to one embodiment. [Figure 10] FIG. 1 is a simplified graph diagram according to one embodiment. [Figure 11] FIG. 1 is a simplified graph diagram according to one embodiment. [Figure 12] FIG. 1 is a simplified graph diagram according to one embodiment. [Figure 13] 1 is a simplified graph according to one embodiment. [Figure 14] FIG. 1 is a schematic diagram according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] The proposed features described below may be used separately or combined in any order. Furthermore, the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0019] 1 shows a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The communication system 100 may include at least two terminals 102, 103 interconnected via a network 105. For one-way data transmission, a first terminal 103 may code video data at a local location for transmission to the other terminal 102 via the network 105. A second terminal 102 may receive the coded video data of the other terminal from the network 105, decode the coded data, and display the recovered video data. One-way data transmission may be common in media serving applications, for example.
[0020] 1 shows a second pair of terminals 101 and 104 provided to support two-way transmission of coded video, such as may occur during a video conference. For the two-way transmission of data, each terminal 101 and 104 may code video data captured at a local location for transmission to the other terminal over network 105. Each terminal 101 and 104 may also receive coded video data transmitted by the other terminal, may decode the coded data, and may display the recovered video data on a local display device.
[0021] In FIG. 1 , terminals 101, 102, 103, and 104 may be depicted as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 105 represents any number of networks that convey coded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communications network 105 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 105 may not be important to the operation of the present disclosure, unless otherwise described herein below.
[0022] 2 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital television, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0023] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213. The sample stream 213 may be enhanced as a higher amount of data compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may be enhanced as a lower amount of data compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to obtain copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes the incoming copy 208 of the encoded video bitstream and forms an outgoing video sample stream 210 that can be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to particular video coding / compression standards, examples of which are mentioned above and further described herein.
[0024] FIG. 3 may be a functional block diagram of a video decoder 300 according to one embodiment of the present invention.
[0025] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300. In the same or another embodiment, receiver 302 may receive one coded video sequence at a time, where decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from channel 301, which may be a hardware / software link to a storage device that stores the coded video data. Receiver 302 may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). Receiver 302 may separate the coded video sequences from other data. To combat network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter, "parser"). If receiver 302 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability or from an isochronous network, buffer 303 may not be necessary or may be small. For use in a best effort packet network such as the Internet, buffer 303 may be required and may be relatively large, and may advantageously be adaptively sized.
[0026] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder 300 and, in some cases, information for controlling a rendering device, such as a display 312, that is not an integral part of the decoder but can be coupled to it. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a subportion of a video usability information parameter set (not shown). The parser 304 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract from the coded video sequence a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to that group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0027] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode particular symbols 313. Additionally, the parser 304 may determine whether a particular symbol 313 should be provided to the motion compensated prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0028] The reconstruction of symbols 313 may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be governed by subgroup control information parsed from the coded video sequence by parser 304. The flow of such subgroup control information between parser 304 and the following units is not shown for clarity.
[0029] In addition to the functional blocks already mentioned, the decoder 300 may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:
[0030] The first unit is the scalar / inverse transform unit 305. The scalar / inverse transform unit 305 receives the quantized transform coefficients as well as control information from the parser 304 as symbol(s) 313, including the transform to use, block size, quantization coefficients, quantization scaling matrix, etc. It can output blocks containing sample values that can be input to the aggregator 310.
[0031] In some cases, the output samples of the scaler / inverse transform unit 305 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309 to generate blocks of the same size and shape as the block being reconstructed. The aggregator 310 may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 307 to the output sample information as provided by the scaler / inverse transform unit 305.
[0032] In other cases, the output samples of the scalar / inverse transform unit 305 may relate to an inter-coded and possibly motion-compensated block. In such cases, the motion-compensated prediction unit 306 may access the reference picture memory 308 to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols 313 associated with the block, these samples (in this case, referred to as residual samples or residual signals) may be added to the output of the scalar / inverse transform unit by the aggregator 310 to generate output sample information. The addresses in the reference picture memory from which the motion compensation unit fetches prediction samples may be controlled by a motion vector and made available to the motion compensation unit in the form of symbols 313, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.
[0033] The output samples of aggregator 310 may be subjected to various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to loop filter unit 311 as symbols 313 from parser 304, but may also be responsive to meta-information obtained during decoding of previous (in decoding order) portions of the coded picture or coded video sequence, and may also be responsive to previously reconstructed loop-filtered sample values.
[0034] The output of the loop filter unit 311 may be output to the rendering device 312, but may also be a sample stream that is stored in the reference picture memory 557 for use in future inter-picture prediction.
[0035] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 304), the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting reconstruction of the next coded picture.
[0036] The video decoder 300 may perform decoding operations according to a predetermined video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard as specified in the video compression technique document or standard, and specifically as specified in the profile documented therein. Compliance also requires that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and HRD buffer management metadata conveyed in the coded video sequence.
[0037] In one embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0038] FIG. 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.
[0039] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that may capture the video image(s) to be coded by the encoder 400 .
[0040] The video source 401 can provide a source video sequence to be coded by the encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source 401 can be a storage device that stores previously prepared video. In a video conferencing system, the video source 401 can also be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures that, when viewed sequentially, create motion. The pictures themselves can be organized as a spatial array of pixels, each of which can contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following discussion focuses on samples.
[0041] According to one embodiment, the encoder 400 can code and compress pictures of a source video sequence into a coded video sequence 410 in real time or under other time constraints required by the application. Achieving an appropriate coding rate is one function of the controller 402. The controller controls and is operatively coupled to other functional units, as described below. For clarity, coupling is not shown. Parameters set by the controller may include rate control-related parameters (e.g., picture skip, quantization, lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller 402 as they may be relevant to optimizing the video encoder 400 for a particular system design.
[0042] Some video encoders operate in what those skilled in the art readily recognize as a "coding loop." As an overly simplified explanation, the coding loop can consist of a coding portion, an encoder 402 (hereafter "source coder") (responsible for creating symbols based on an input picture to be coded and reference picture(s)), and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data that a (remote) decoder will also create (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream is input to a reference picture memory 405. Because decoding of the symbol stream yields bit-exact results regardless of the decoder's location (local or remote), the reference picture buffer contents are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the drift that occurs when synchronism cannot be maintained, for example due to channel errors) is well known to those skilled in the art.
[0043] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300, which has already been described in detail above in connection with Figure 3. However, briefly referring also to Figure 4, because symbols are available and the coding / decoding of the symbols into a coded video sequence by the entropy coder 408 and parser 304 may be lossless, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303 and parser 304, may not be fully implemented in the local decoder 406.
[0044] An observation that can be made at this point is that any decoder techniques other than analysis / entropy decoding that are present in a decoder must necessarily be present in the corresponding encoder in substantially identical functional form. A description of the encoder techniques can be omitted, as they are the inverse of the decoder techniques that have been comprehensively described. Only in certain areas is a more detailed description required, and this is provided below.
[0045] As part of its operation, source coder 403 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this method, coding engine 407 codes differences between pixel blocks of the input frame and pixel blocks of reference frame(s) that may be selected as predictive reference(s) for the input frame.
[0046] The local video decoder 406 may decode coded video data of frames that may be designated as reference frames based on symbols created by the source coder 403. The operation of the coding engine 407 may advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a duplicate of the source video sequence, with some errors. The local video decoder 406 may replicate the decoding process that may be performed by the video decoder on the reference frames and store the reconstructed reference frames in the reference picture cache 405. In this manner, the encoder 400 may locally store copies of reconstructed reference frames that have common content as reconstructed reference frames that would be retrieved by a far-end video decoder (barring transmission errors).
[0047] The predictor 404 may perform the prediction search of the coding engine 407. That is, for a new frame to be coded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata, such as the reference picture's motion vectors, block shape, etc., that can serve as suitable prediction references for the new picture. The predictor 404 may operate on a pixelblock-by-pixelblock basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 405.
[0048] The controller 402 may manage the coding operations of the video coder 403, including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0049] The output of all the aforementioned functional units may undergo entropy coding in entropy coder 408. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0050] The transmitter 409 may buffer the coded video sequence(s) created by the entropy coder 408 and prepare them for transmission over a communication channel 411, which may be a hardware / software link to a storage device that stores the coded video data. The transmitter 409 may also combine the coded video data from the video coder 403 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0051] A controller 402 may manage the operation of the encoder 400. During coding, the controller 405 may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as any of the following frame types:
[0052] An intra picture (I-picture) may be one that can be coded and decoded without using other frames in a sequence as a source of prediction. Some video codecs allow for various types of intra pictures, such as Independent Decoder Refresh Pictures. Those skilled in the art are aware of variations of I-pictures and their respective uses and functions.
[0053] A predictive picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0054] A bidirectionally predicted picture (B picture) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0055] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0056] Video coder 400 may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In its operation, video coder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0057] In one embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source coder 403 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0058] FIG. 5 illustrates a call 500, such as a 360-degree conference call, according to an exemplary embodiment. Referring to FIG. 5, the conference call is taking place in a room 501. The room 501 consists of people physically present in the room 501, an omnidirectional camera 505, and a viewing screen 502. Two other people 503 and 504 participate in the conference; according to an exemplary embodiment, person 503 may be using a VR and / or AR headset, and person 504 may be using a smartphone or tablet. Persons 503 and 504 receive a 360-degree view of the conference room via the omnidirectional camera 505, and the received view may be for each person 503 and 504, regardless of whether they are looking at different portions of the 360-degree view relative to the direction of a particular viewing screen, for example. Remote participants, e.g., persons 503 and 504, also have the option to focus on each other's camera-sourced images. In Figure 5, people 503 and 504 send viewport information 507 and 508, respectively, to room 501 or other network devices associated with omnidirectional camera 505 and view screen 502, which in turn send viewport-dependent videos 509 and 510, respectively.
[0059] A remote user participating in the conference, for example, person 503 wearing a head-mounted display (HMD), remotely receives stereo or immersive voice / audio and immersive video from the conference room captured by an omnidirectional camera. Person 504 may also wear an HMD or use a handheld mobile device such as a smartphone or tablet.
[0060] According to an exemplary embodiment, the Omnidirectional Media Format (OMAF) defines two types of media profiles: (i) viewport-independent and (ii) viewport-dependent. When viewport-independent streaming (VIS) is used, the entire video is transmitted in high quality regardless of the user's viewport. When VIS is used, there is no latency during HMD movement. However, bandwidth requirements may be relatively high.
[0061] Streaming the entire high-resolution immersive video at a desired quality may be inefficient due to network bandwidth limitations, decoding complexity, and computational constraints on end devices, as the user's field of view (FoV) may be limited. Therefore, according to an exemplary embodiment, viewport-dependent streaming (VDS) is defined in the Omnidirectional Media Format (OMAF). When VDS is used, only the user's current viewport is streamed at high quality, and the rest is streamed at a relatively lower quality. This helps save a significant amount of bandwidth.
[0062] While using the VDS, remote users, e.g., persons 503 and 504, can send viewport direction information via RTCP reports. These reports can be sent at fixed intervals, event-based triggers, or using a hybrid method that includes both fixed intervals and event-based triggers.
[0063] According to an exemplary embodiment, event-based feedback is triggered whenever the viewport changes and immediate feedback is sent. The frequency of RTCP reports depends on the speed of the HMD and increases as the HMD speed increases.
[0064] Here, the HMD speed is large and the feedback trigger angle is relatively large. small In this case, a large number of event-based RTCP reports are generated and sent to the server. This can cause the required bandwidth to exceed the RTCP 5% bandwidth limit. For example, referring to Figure 10, for diagram 1000 of a 20 Mbps point-to-point scenario, if the feedback trigger is 0.1 degrees and the HMD velocity exceeds 125 degrees per second, the bandwidth required for the RTCP reports to be sent exceeds the RTCP bandwidth limit of 1 Mbps. Figure 10 shows the event-based RTCP feedback generated for triggers of 0.1 degrees, 0.5 degrees, 1 degree, and 2 degrees.
[0065] In addition to the usual multimedia telephony services for IMS (MTSI) call signaling, during call setup, the user's initial viewport direction, decoding / rendering metadata, and captured field of view are signaled in the Session Description Protocol (SDP) between conference room 801 and remote user 802, such as at S803 in Figure 8. After call establishment, the remote party sends viewport direction information via an RTCP report.
[0066] RTCP feedback may follow the 5% bandwidth rule according to exemplary embodiments. Therefore, the frequency of RTCP depends on the group size or the number of remote participants. As the group size increases, feedback can be sent less frequently to comply with bandwidth usage restrictions. When the number of remote participants is small, immediate feedback can be used. As the number of participants increases, earlier RTCP feedback can be used. However, as the group size increases, regular RTCP feedback should be sent. In accordance with Internet Engineering Task Force (IETF) Request for Comments (RFC) 3550, the minimum RTCP transmission interval may be 5 seconds. As the group size increases, the RTCP feedback interval also increases, potentially introducing additional delay. According to exemplary embodiments herein, for use in immersive video, for example, RTCP reports may be sent either on a fixed interval basis or an event basis, which may be triggered by a change in viewport orientation for each remote person, such as person 503 and / or person 504 in FIG. 5.
[0067] According to an exemplary embodiment, the RTCP feedback packet may be a composite packet consisting of a status report and a feedback (FB) message. Additionally, sender report (SR) / receiver report (RR) packets contain status reports that are sent at regular intervals as part of a composite RTCP packet that contains a source description in addition to other messages.
[0068] According to an exemplary embodiment, the order of the RTCP packets within a compound RTCP packet containing an FB message is: - an optional encrypted prefix, - Required SR or RR, - Mandatory SDES, - One or more FB messages is.
[0069] In a compound packet, the FB message may be placed after the RR and Source Description RTCP packets (SDES).
[0070] The two compound RTCP packets carrying feedback packets can be described as a minimal compound RTCP feedback packet and a full compound RTCP feedback packet:
[0071] The RTCP feedback message is specified in IETF 4585. It can be identified by PT (Payload Type) = PSFB (206), where PSFB refers to a payload-specific feedback message. According to an example embodiment, the feedback message can include both regular interval and event-based signaling of viewport information.
[0072] When any remote participant, such as one of person 503 and person 504 in FIG. 1, changes their respective viewport (e.g., by changing the spatial orientation of their respective display device), RTCP viewport feedback should be delivered in a timely manner; otherwise, it will cause delays and affect the user's high-quality VR experience. As the number of remote participants increases, the RTCP feedback interval increases. If a regular RTCP feedback interval is sent alone, such as every 5 seconds, it may become delayed as the number of remote participants increases. Therefore, according to an exemplary embodiment, the RTCP interval may be a combination of a regular feedback interval and an event-based interval to improve such technical deficiencies.
[0073] According to an exemplary embodiment, the normal RTCP feedback interval should be transmitted as a compound RTCP packet that conforms to the RTP rule that the minimum RTCP interval (Tmin) between successive transmissions should be 5 seconds. This RTCP interval can be derived from the RTCP packet size and the available RTCP bandwidth. The full compound packet includes any additional RTCP packets, such as additional receiver reports, additional SDES items, etc.
[0074] A change in the viewport triggers event-based feedback. In this case, a minimal compound RTCP packet can be sent. The minimal compound RTCP feedback packet contains only essential information, such as the required encryption prefix, exactly one SR or RR, exactly one SDES (only the CNAME item is present), and an FB message. This helps minimize the RTCP packets sent for feedback, and therefore has minimal impact on bandwidth. Event-based feedback is not affected by group size, unlike the regular RTCP feedback interval.
[0075] As shown in Figure 7, when a user changes the viewport 701, event-based feedback is triggered, and the regular feedback interval must start after the minimum interval (Tmin). Now, if the user changes the viewport 701 before the minimum interval, event-based feedback is triggered again. This may affect the 5% bandwidth rule if these events occur consecutively. However, imposing a minimum interval constraint on event-based feedback would degrade the user experience. Therefore, there should be no minimum interval defined for event-based feedback to be sent. To take into account the bandwidth usage of RTCP feedback, the interval for regular feedback can be increased, and should therefore depend on the frequency of event-based feedback and their intervals.
[0076] In consideration of the exemplary embodiments described herein, a user may request an additional, higher-quality margin, such as 702 in illustration 700 of FIG. 7 , around viewport 701 to minimize delays, such as M2HQ delay, and improve the user experience. Viewport 701 may be the viewport of either of the devices of remote persons 503 and 504 of FIG. 5 . This is very useful when one of remote persons 503 and 504 is performing a small head movement perturbation. However, during a call, such as S805 of FIG. 8 , if a user moves their head only a small (negligible) amount that is outside the viewport margin 702, event-based feedback should not be triggered because the out-of-margin viewport area is relatively negligible, and therefore regular feedback should wait to be sent. Therefore, a certain tolerance 703 may be defined for yaw, pitch, and roll before event-based feedback can be triggered. For example, referring to S804 of FIG. 8 , tolerance information may be sent from remote user 802 to conference room 801. The degree of this tolerance 703, which is tolerance information including the event-based feedback tolerance margin 705, can be defined as one or more of the rotation angles yaw (tyaw), pitch (tpitch), and roll (troll) from the user's viewport. Such information can be negotiated during the initial SDP session S804 as in Figure 8 or during the session according to an embodiment such as S806.
[0077] FIG. 6 illustrates a format 600 of such a feedback message described herein, according to an example embodiment.
[0078] Figure 6 shows an RTCP feedback message format 600. In Figure 6, FMT indicates the feedback message type and PT indicates the payload type. For an RTCP feedback message, FMT may be set to the value "9" and PT is set to 206. The FCI (Feedback Message Control Information) contains viewport information and is composed of the following parameters according to an exemplary embodiment: Viewport azimuth, Viewport elevation, Viewport tilt, Viewport azimuth range, Viewport elevation range, Viewport stereoscopic.
[0079] Figure 9 shows a flowchart 900. S901 includes call setup 901, which includes initialization such as in S803 of Figure 8, and in one embodiment, the user's viewport orientation, decoding / rendering metadata, and captured field of view are signaled in the Session Description Protocol (SDP) during call setup in addition to the normal multimedia telephony services for IMS (MTSI) call signaling. In this S901, the degree of tolerance can be configured, as described above.
[0080] After call establishment, in S902, the remote parties send their viewport orientation information via RTCP reports upon the start of the call. Next, in S903, it can be determined whether event-based feedback is triggered and whether a regular feedback interval is triggered. The regular feedback interval can be triggered by determining whether a certain period of time, such as 5 seconds, has passed as time-based feedback. Event-based feedback can be determined by determining whether the remote user's viewport has changed in spatial orientation, and if so, whether the change is within or beyond a preset tolerance range, such as with respect to tolerance margin 705 in FIG. 7.
[0081] If time-based feedback is determined in S905, a compound RTCP packet can be transmitted. According to an exemplary embodiment, the normal RTCP feedback interval should be transmitted as a compound RTCP packet conforming to the RTP rule that the minimum RTCP interval (Tmin) between successive transmissions should be 5 seconds. This RTCP interval can be derived from the RTCP packet size and the available RTCP bandwidth. The full compound packet includes any additional RTCP packets, such as additional receiver reports, additional SDES items, etc. Then, in S904, it can be determined whether there is any user input or other input of updates to the tolerance information; if not, processing can loop or otherwise proceed with the call in S902 according to any communication received regarding the compound RTCP packet in S905. This S905 can also reset a timer to count another period, such as 5 seconds.
[0082] If event-based feedback is determined in S906, a minimum RTCP packet may be sent. For example, a viewport change triggers event-based feedback. In this case, a minimum compound RTCP packet may be sent. The minimum compound RTCP feedback packet contains only essential information, such as the required encryption prefix, exactly one SR or RR, exactly one SDES (only the CNAME item is present), and an FB message. This helps minimize the RTCP packets sent for feedback, thus minimizing the impact on bandwidth. Unlike the normal RTCP feedback interval, event-based feedback is not affected by group size. Then, in S904, it may be determined whether there is any user input or other input for any updates to the tolerance information. If not, processing may loop or otherwise proceed with the call in S902 according to any communication received regarding the minimum RTCP packet in S906. This S906 may also reset a timer to count another period, such as 5 seconds. Additionally, if it is also determined in S906 that the frequency of the event-based trigger exceeds a threshold, such as contributing to exceeding the RTCP 5% bandwidth rule according to embodiments further described with respect to Figures 10, 11, 12, and 13, then it may also be determined in S904 whether to update the timer to count the increased elapsed time from S906.
[0083] Exemplary embodiments introduce a parameter to define the minimum interval between two consecutive event-based RTCP feedbacks, e.g., S906 to S906 without intermediate S905, so that the bandwidth requirement does not exceed the RTCP bandwidth limit and can be updated or otherwise taken into account in S904.
[0084] If a relatively short feedback trigger is used for large HMD movements in S906, the RTCP bandwidth requirements may exceed the RTCP bandwidth limit. Therefore, the bandwidth requirements for event-based RTCP reporting may depend on the HMD speed and the degree of feedback triggering according to an exemplary embodiment.
[0085] The event-based feedback interval is the time interval between two consecutive triggers. As the HMD speed increases, the event-based feedback interval decreases and the bandwidth requirement increases. Event-based feedback can be defined as follows:
[0086]
number
[0087] Therefore, to limit the bandwidth requirement to comply with the 5% RTCP rule, a threshold parameter is defined, which may depend on the event-based feedback interval.
[0088] According to an exemplary embodiment, the following assumptions can be made: Bandwidth (bps) = B (Equation 2) RTCP allocated bandwidth (bps) = R B (Formula 3) RTCP packet size (bytes) = P (Equation 4) RTCP minimum interval=I min (Formula 5) The RTCP bandwidth should not exceed 5% bandwidth according to the RTCP rules. Therefore, R B =0.05B (Equation 6) On the other hand, I min can be expressed as follows:
[0089]
number
[0090] Assuming a total bandwidth of 20 Mbps, when the feedback degree is 0.1 degrees and the HMD velocity exceeds 125 degrees / second, the bandwidth value exceeds the RTCP bandwidth limit by 5%. However, as can be seen from chart 1000 in Figure 10, this value is well within the limit when the trigger is increased to 0.5, 1, and 2.
[0091] The number of event-based RTCP feedbacks sent per second for 0.1 degree, 0.5 degree, 1 degree, and 2 degree triggers is shown in chart 1100 of Figure 11. Therefore, considering the RTCP bandwidth limitation, the exemplary embodiment increases the RTCP event-based feedback interval, which can be defined as the minimum interval between two consecutive RTCP feedbacks:
number
[0092] As the HMD speed increases, the number of triggers per second also increases, resulting in a decrease in the trigger interval and an increase in the RTCP bit rate. min , it should not be decreased any further and thus reaches the maximum number of triggers per second, as shown by the dotted curve in diagram 1200 of Figure 12. Figure 12 illustrates a graph of I, according to one embodiment. min Figure 1 shows the event-based RTCP feedback generated for 0.1 degree, 0.5 degree, 1 degree, and 2 degree triggers after the parameter is introduced. B This minimum is reached when the bandwidth is close to, but not exceeding, 5% of the bandwidth. Therefore, I min The parameter depends on the allowed RTCP bandwidth. The bit rates before and after the introduction of the minimum parameter for the 0.1 degree trigger are shown in diagram 1300 of Figure 13. Therefore, at the minimum point I min After reaching , the graph flattens out.
[0093] Further referring to Figure 13, I minis calculated for a constant head speed, and I min It refers to the bit rate of the 0.1 degree trigger before and after the parameter is introduced. However, because the head movement time is relatively short, the difference between the average head velocity and the constant head velocity can be ignored.
[0094] According to an exemplary embodiment, when such a hybrid reporting scheme consisting of regular interval and event-based triggers is used: (i) The normal interval must be greater than or equal to (usually a multiple of) the RTCP minimum interval. (ii) The trigger threshold angle and RTCP minimum interval should be selected as follows:
[0095]
number
[0096] One or more of such calculations described above with respect to FIGS. 10, 11, 12, and 13 may be performed at S904 in FIG.
[0097] Therefore, according to the exemplary embodiments described herein, the technical problems noted above may be advantageously improved by one or more of these technical solutions. For example, a parameter I, defined as the minimum interval between two consecutive triggers, may be used to improve the accuracy of the parameter I. min should be introduced for event-based RTCP feedback. This helps to limit the bandwidth requirements by limiting the number of event-based triggers sent per second. Thus, as head movement increases and the RTCP interval reaches its minimum (Imin) value, the bit rate saturates and does not increase further.
[0098] The techniques described above can be implemented using computer-readable instructions, as computer software physically stored on one or more computer-readable media, or by one or more specially configured hardware processors. For example, Figure 14 illustrates a computer system 1400 suitable for implementing certain embodiments of the disclosed subject matter.
[0099] Computer software can be coded using any suitable machine code or computer language and can be subjected to assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly, or through interpretation, microcode execution, etc.
[0100] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0101] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 1400.
[0102] The computer system 1400 may include certain human interface input devices that can respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices can also be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0103] The input human interface devices may include one or more of a keyboard 1401, a mouse 1402, a trackpad 1403, a touchscreen 1410, a joystick 1405, a microphone 1406, a scanner 1408, and a camera 1407 (only one of each is shown).
[0104] The computer system 1400 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., touchscreen 1410 or haptic feedback via joystick 1405, although there may be haptic feedback devices that do not function as input devices), audio output devices (speakers 1409, headphones (not shown), etc.), visual output devices (screens 1410, including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, and each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0105] The computer system 1400 may also include human-accessible storage devices and their associated media, such as CD / DVD 1411 or CD / DVD ROM / RW 1420 with similar media, thumb drives 1422, removable hard drives or solid state drives 1423, legacy magnetic media such as tape and floppy disks (not shown), optical media including specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0106] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitory signals.
[0107] The computer system 1400 may also include an interface 1499 to one or more communications networks 1498. The network 1498 may be, for example, wireless, wired, optical, or the like. Furthermore, the network 1498 may be local, wide area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks 1498 include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable, satellite, and terrestrial broadcast, vehicular, and industrial networks including CANBus, etc. Particular networks 1498 generally require external network interface adapters coupled to particular general-purpose data ports or peripheral buses (1450 and 1451) (e.g., USB ports on computer system 1400), while others are generally integrated into the core of computer system 1400 by coupling to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 1498, computer system 1400 can communicate with other entities. Such communications may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., from a CANbus to a particular CANbus device), or bidirectional, e.g., to other computer systems using local-area or wide-area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.
[0108] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the core 1440 of the computer system 1400.
[0109] The core 1440 may include one or more central processing units (CPUs) 1441, graphics processing units (GPUs) 1442, graphics adapters 1417, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 1443, hardware accelerators for specific tasks 1444, etc. These devices may be connected via a system bus 1448, along with read-only memory (ROM) 1445, random access memory 1446, and internal mass storage devices 1447, such as an internal hard drive or SSD, that are not user accessible. In some computer systems, the system bus 1448 may also be accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, etc. Peripheral devices may be coupled to the core's system bus 1448 directly or via a peripheral bus 1451. Peripheral bus architectures include PCI, USB, etc.
[0110] The CPU 1441, GPU 1442, FPGA 1443, and accelerator 1444 can combine to execute specific instructions that may constitute the aforementioned computer code, which may be stored in ROM 1445 or RAM 1446. Persistent data may be stored, for example, in internal mass storage device 1447, while transient data may also be stored in RAM 1446. Cache memory, which may be closely associated with one or more of the CPU 1441, GPU 1442, mass storage device 1447, ROM 1445, RAM 1446, etc., may be used to enable fast storage and retrieval to any of the memory devices.
[0111] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0112] By way of example, and not limitation, computer system 1400 having the architecture, and particularly a computer system having core 1440, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices, as introduced above, as well as media associated with specific storage devices of the core 1440 that are non-transitory in nature, such as core internal mass storage device 1447 or ROM 1445. Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core 1440. The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software can cause the core 1440, and particularly the processors therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM 1446 and modifying such data structures through software-defined operations. Additionally, or alternatively, a computer system may provide functionality as a result of logic circuitry hardwired or otherwise embodied in circuitry (e.g., accelerator 1444) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software may include logic circuitry, and vice versa, as appropriate. References to computer-readable media may encompass, where appropriate, circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software.
[0113] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure, i.e., are within its spirit and scope. [Explanation of symbols]
[0114] 100 Communication Systems 101, 102, 103, 104 terminals 105 Network 201 Camera / Video Source 202 Encoder 203 Capture Subsystem 204 Video Bitstream 205 Streaming Server 206 Copying Video Bitstreams 207 Streaming Client 208 Copying Video Bitstreams 209 Display 210 outgoing video sample streams 211 Video Decoder 212 Streaming Client 213 Sample Stream 300 decoder 301 Channel 302 Receiver 303 Buffer 304 Parser 305 Scaler / Descaler Unit 306 Motion Compensation Prediction Unit 307 Intra Prediction Unit 308 Reference Picture Buffer 309 Pictures 310 Aggregator 311 Loop Filter Unit 312 Display 313 Symbol 400 Encoder 401 Video Source 402 Controller 403 Source Coder 404 Predictor 405 Reference Picture Memory 406 decoder 407 Coding Engine 408 Entropy Coder 409 Transmitter 410 Video Sequence 411 Communication Channel 500 calls 501 rooms 502 view screens 503,504 people 505 Omnidirectional Camera 507,508 viewport information 509,510 viewport-dependent videos 557 Reference Picture Memory 600 Feedback Message Formats 701 viewport 702 viewport margins 703 Tolerance 705 Event-Based Feedback Tolerance Margin Conference Room 801 802 users 900 Flowchart 1400 Computer Systems 1401 keyboard 1402 Mouse 1403 Trackpad 1405 Joystick 1406 Mike 1407 Camera 1408 Scanner 1409 Speaker 1410 Touchscreen 1411 CDs / DVDs 1420 CD / DVD ROM / RW 1417 Graphics Adapter 1422 thumb drive 1423 Solid State Drive 1440 cores 1441 Central Processing Unit (CPU) 1442 Graphics Processing Unit (GPU) 1443 Field Programmable Gate Array (FPGA) 1444 Accelerator 1445 Read-Only Memory (ROM) 1446 Random Access Memory (RAM) 1447 Mass Storage 1448 System Bus 1450,1451 surrounding buses 1498 Network 1499 Interface
Claims
1. 1. A video signaling method for a video call, executed by at least one processor, comprising: determining whether a viewport-based event-based trigger has occurred and whether the time elapsed since another event indicating a normal feedback interval is less than a predetermined time; controlling delivery of the video call to the viewport based on a determination of whether the trigger has occurred; A video signaling method comprising:
2. 10. The video signaling method of claim 1, wherein the trigger is based at least on a degree of change in spatial orientation of the viewport.
3. The step of determining whether the trigger has occurred comprises: determining whether the change in the spatial orientation of the viewport is greater than a predetermined tolerance range; 3. The video signaling method of claim 2, comprising:
4. The step of controlling the delivery of the video call to the viewport comprises: delivering at least an additional margin of the video call to the viewport if the change in the spatial orientation of the viewport is determined to be greater than the predetermined tolerance range.
4. The video signaling method of claim 3, comprising:
5. The step of controlling the delivery of the video call to the viewport comprises: processing RTCP feedback packets of different lengths depending on whether the time elapsed since the other event has exceeded the predetermined time or whether the trigger has occurred.
2. The video signaling method of claim 1, comprising:
6. 6. The video signaling method of claim 5, wherein, among the RTCP feedback packets of different lengths, a first RTCP feedback packet when the time elapsed since the other event exceeds the predetermined time is longer than a second RTCP feedback packet when the trigger occurs.
7. determining whether the frequency with which the trigger occurs exceeds a frequency threshold; 10. The video signaling method of claim 1, further comprising:
8. limiting the number of event-based triggers per second in response to determining that the frequency at which the triggers occur exceeds the frequency threshold.
8. The video signaling method of claim 7, further comprising:
9. 10. The video signaling method of claim 1, wherein the viewport is a display of at least one of a headset and a handheld mobile device.
10. The video signaling method of claim 1 , wherein the video call includes 360-degree video data from an omnidirectional camera.
11. A video signaling device, comprising: at least one memory configured to store computer program code; at least one processor configured to access said computer program code and to act according to the instructions of said computer program code; wherein the computer program code causes the at least one processor to perform the video signaling method of any one of claims 1 to 10. Video signaling equipment.
12. A computer program product for causing a computer to carry out the video signaling method according to any one of claims 1 to 10.
Citation Information
Patent Citations
System and method for end-to-end call quality indication
US20130250786A1
Interactive video conferencing
US20180013980A1
Rectilinear viewport extraction from a region of a wide field of view using messaging in video transmission
US20180192001A1
New radio (NR) for spectrum sharing
US20190335337A1
Method for transmitting 360-degree video, method for receiving 360-degree video, apparatus for transmitting 360-degree video and apparatus for receiving 360-degree video
US20190364261A1