Viewport stabilization

The described method and apparatus for viewport stabilization in 360-degree video streaming address the challenge of unwanted motion by processing key points, motion vectors, and rotation angles, resulting in a smoother and more stable video streaming experience.

WO2025131487A1PCT designated stage expired Publication Date: 2025-06-26NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082771
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing video streaming technologies face challenges in stabilizing 360-degree video viewports, particularly in addressing unwanted 'jitter' caused by jerky movements during recording, which traditional video stabilization techniques struggle to handle effectively.

Method used

The proposed solution involves an apparatus and method for viewport stabilization in 360-degree video streaming, which includes processing steps such as extracting key points, estimating motion vectors, calculating global 2D motion data, updating rotation angles, rotating the 360 video frame, and cropping it to a viewport image for viewport-dependent delivery (VDD). This process is designed to compensate for motion and provide a smoother viewing experience.

Benefits of technology

The technical solution effectively reduces unwanted motion in 360-degree video viewports, enhancing the stability and smoothness of the video streaming experience by adapting to viewer movements and optimizing video delivery based on the viewer's field of view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000018_0001
    Figure IMGF000018_0001
  • Figure 00000033_0000
    Figure 00000033_0000
  • Figure 00000033_0001
    Figure 00000033_0001
Patent Text Reader

Abstract

Various embodiments describe methods, apparatuses, and computer program products. An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two- dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.
Need to check novelty before this filing date? Find Prior Art

Description

VIEWPORT STABILIZATIONTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to video streaming and, more particularly, to viewport stabilization.BACKGROUND

[0002] It is known to provide video streaming.SUMMARY

[0003] Example 1: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

[0004] Example 2: The apparatus of example 1, wherein the apparatus is further caused to perform: encoding and signaling viewport dependent delivery data.

[0005] Example 3 : The apparatus of any of examples 1 or 2, wherein the viewport dependent delivery data comprises the global 2D motion data.

[0006] Example 4: The apparatus of any of examples 1 to 3, wherein the apparatus is further caused to perform: rotating an original 360 video frame to a three-dimensional (3D) sphere from the viewport direction to obtain a rotated 360 video frame; mapping the rotated 360 video frame to a new equirectangular projection (ERP) frame in 2D for which a viewport is located in a center of the field of view; and selecting a pixel from the new ERP frame.

[0007] Example 5: The apparatus of any of the examples 1 to 3, wherein the apparatus further comprises: selecting a pixel from a three-dimensional space mapping from the ERP to the 2D image.

[0008] Example 6: The apparatus of example 1, wherein the apparatus is further caused to perform: compensating rotation 2D angels from a direction of a viewer with the global 2D motiondata to a 3D Euler rotation delta to original viewport angles; and signaling the global 2D motion data when a size of the cropped 360 video frame or a cropped equirectangular projection (ERP) is larger than a size of the viewport.

[0009] Example 7 : The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: negotiating motion compensation as a session negotiation parameter.

[0010] Example 8: The apparatus of any of the previous examples, wherein the apparatus is further caused to perform: signaling viewport dependent delivery data via a header extension. In an example, the viewport delivery data includes motion vector and / or updated viewport data.

[0011] Example 9: The apparatus of example 8, wherein the header extension comprises: a local identifier of the header extension; and a length field for identifying length of the header extension.

[0012] Example 10: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

[0013] Example 11: The apparatus of example 10, wherein the apparatus is further caused to perform: signaling viewport data via a header extension.

[0014] Example 12: The apparatus of example 11, wherein the header extension comprises zooming ratio header extension for identifying zooming ratio.

[0015] Example 13: A method comprising: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

[0016] Example 14: The method of example 13 further comprising encoding and signaling viewport dependent delivery data.

[0017] Example 15: The method of any of examples 13 or 14, wherein the viewport dependentdelivery data comprises the global 2D motion data.

[0018] Example 16: The method of any of the examples 13 to 15 further comprising: rotating an original 360 video frame to a three-dimensional (3D) sphere from the viewport direction to obtain a rotated 360 video frame; mapping the rotated 360 video frame to a new equirectangular projection (ERP) frame in 2D for which a viewport is located in a center of the field of view; and selecting a pixel from the new ERP frame.

[0019] Example 17: The method of any of the examples 13 to 15 further comprising: selecting a pixel from a three-dimensional space mapping from the ERP to the 2D image.

[0020] Example 18: The method of example 13 further comprising: compensating rotation 2D angels from a direction of a viewer with the global 2D motion data to a 3D Euler rotation delta to original viewport angles; and signaling the global 2D motion data when a size of the cropped 360 video frame or a cropped equirectangular projection (ERP) is larger than a size of the viewport.

[0021] Example 19: he method of any of the examples 13 to 18 further comprising: negotiating motion compensation as a session negotiation parameter.

[0022] Example 20: The method of any of the examples 13 to 19 further comprising: signaling viewport dependent delivery data via a header extension. In an example, the viewport delivery data includes motion vector and / or updated viewport data.

[0023] Example 21: The method of example 20, wherein the header extension comprises: a local identifier of the header extension; and a length field for identifying length of the header extension.

[0024] Example 22: A method comprising: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

[0025] Example 23: The method of example 22 further comprising: signaling viewport data via a header extension.

[0026] Example 24: The method of example 23, wherein the header extension comprises zooming ratio header extension for identifying zooming ratio.

[0027] Example 25: An apparatus comprising: means for starting a 360 video viewport delivery processing for a viewport; means for extracting key points in the viewport; means forestimating motion vectors of the key points on the viewport surface; means for estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; means for calculating updated rotation angles from the global motion data in the pixel; means for rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and means for cropping the 360 video frame to a viewport image for the viewportdependent delivery (VDD) process.

[0028] Example 26: The apparatus of example 25, wherein the apparatus further comprises means for performing the methods as described in any of the examples 14 to 21.

[0029] Example 27: An apparatus comprising: means for starting a 360 video viewport delivery processing for a viewport; means for extracting key points in the viewport; means for estimating zoom motion of the key points in the viewport surface, based on zooming ration; and means for updating a field of view by scaling the field of view based on the zoom motion.

[0030] Example 28: The apparatus of example 27 further comprising means for performing the methods as described in any of the examples 23 or 24.

[0031] Example 29: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

[0032] Example 30: The computer readable medium of example 29, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0033] Example 31: The computer readable medium of any of examples 29 or 30 comprises program instructions which, when executed, cause the apparatus to perform the methods as described in any of the examples 14 to 21.

[0034] Example 32: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

[0035] Example 33: The computer readable medium of example 32, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0036] Example 34: The computer readable medium of any of examples 32 or 33 comprises program instructions which, when executed, cause the apparatus to perform the methods as described in any of the examples 23 or 24.BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0038] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.

[0039] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.

[0040] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.

[0041] FIG. 4 shows schematically a block chart of an encoder used for data compression on a general level.

[0042] FIG. 5 illustrates inconsistent motion in 360 equirectangular projection (ERP) frame and viewport video frames;

[0043] FIG. 6 illustrates a 360 frame re-rotation in ERP mode based on a viewing direction and a field of view (FoV);

[0044] FIG. 7 illustrates projecting a viewport region in an 360 ERP frame to a 360 unit sphere based on a viewing direction and a FoV;

[0045] FIG. 8 illustrates local motion vs global motion with simplified motion pattern;

[0046] FIG. 9 illustrates motion estimation techniques;

[0047] FIG. 10 illustrates stabilized motion;

[0048] FIG. 11 illustrates an example of 3D viewport-dependent stabilization process;

[0049] FIG. 12 illustrates an example of 2 viewports being projected on a unit sphere;

[0050] FIG. 13 illustrates an example of 2D viewport-dependent stabilization process;

[0051] FIG. 14 illustrates an example of key point extraction;

[0052] FIG. 15 illustrates an example of reversed projection with motion compensation from2D offsets to 3D angular degrees;

[0053] FIG. 16 illustrates updated FoV when zoom motion is estimated;

[0054] FIG. 17. illustrates an example of a one-byte header extension;

[0055] FIG. 18 is an example apparatus configured to implement the examples described herein.

[0056] FIG. 19 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0057] FIG. 20 is an example method, based on the examples described herein.

[0058] FIG. 21 is another example method, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0059] Described herein is a method, apparatus and a computer program for viewport stabilization, for example, in 360 degree viewport-dependent video delivery (VDD).

[0060] The following describes in detail a suitable apparatus and possible mechanisms for viewport stabilization, for example, in 360 degree viewport-dependent video delivery (VDD). In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.

[0061] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.

[0062] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 further may comprise a display 32 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0063] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analog signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analog audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0064] The apparatus 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0065] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.

[0066] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0067] The apparatus 50 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.

[0068] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.

[0069] The system 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the examples described herein.

[0070] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0071] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, or a head-mounted apparatus, which head-mounted apparatus may be a head-mounted display (HMD), or glasses having a camera or other device used for processing images and / or video. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.

[0072] The embodiments may also be implemented in a set-top box; e.g. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operatingsystems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0073] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.

[0074] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0075] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0076] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).

[0077] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presentsan encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 may comprise similar elements for encoding incoming pictures. The encoder sections 500, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404. FIG. 4 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406 (Pinter), an intrapredictor 308, 408 (Pintra), a mode selector 310, 410, a filter 316, 416 (F), and a reference frame memory 318, 418 (RFM). The pixel predictor 302 of the first encoder section 500 receives 300 base layer images (Io,n) of a video stream to be encoded at both the inter-predictor 306 (which determines the difference between the image and a motion compensated reference frame 318) and the intrapredictor 308 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 310. The intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer picture 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives 400 enhancement layer images (Ii,n) of a video stream to be encoded at both the inter-predictor 406 (which determines the difference between the image and a motion compensated reference frame 418) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra- predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intraprediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer picture 400.

[0078] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 306, 406 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 310, 410. The output of the mode selector is passed to a first summing device 321, 421. The first summing device may subtract the output of the pixel predictor 302, 402 from the base layer picture 300 / enhancement layer picture 400 to produce a first prediction error signal 320, 420 (Dn) which is input to the prediction error encoder 303, 403.

[0079] The pixel predictor 302, 402 further receives from a preliminary reconstructor 339, 439 the combination of the prediction representation of the image block 312, 412 (P’n) and the output 338, 438 (D’n) of the prediction error decoder 304, 404. The preliminary reconstructed image 314,414 (I’n) may be passed to the intra-predictor 308, 408 and to the filter 316, 416. The filter 316, 416 receiving the preliminary representation may fdter the preliminary representation and output a final reconstructed image 340, 440 (R’n) which may be saved in a reference frame memory 318, 418. The reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer picture 300 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be the source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 318 may also be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations.

[0080] Filtering parameters from the filter 316 of the first encoder section 500 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be the source for predicting the filtering parameters of the enhancement layer according to some embodiments.

[0081] The prediction error encoder 303, 403 comprises a transform unit 342, 442 (T) and a quantizer 344, 444 (Q). The transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain. The transform is, for example, the DCT transform. The quantizer 344, 444 quantizes the transform domain signal, e.g. the DCT coefficients, to form quantized coefficients.

[0082] The prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414. The prediction error decoder 304, 404 may be considered to comprise a dequantizer 346, 446 (Q ’), which dequantizes the quantized coefficient values, e.g. DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448 (T ’), which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.

[0083] The entropy encoder 330, 430 receives the output of the prediction error encoder 303, 403 and may perform a suitable entropy encoding / variable length encoding on the signal to provideerror detection and correction capability. The outputs of the entropy encoders 330, 430 may be inserted into a bitstream e.g. by a multiplexer 508 (M).

[0084] In some embodiments, the terms “picture”, “image”, and “frame” may be used interchangeably.

[0085] Various devices / apparatuses as described in any of FIGs. 1 to 4 may be used to implement video stabilization / viewport stabilization.

[0086] Video stabilization is a technique used to reduce or eliminate unwanted shaky or jerky motion in videos. It aims to provide a smoother and more stable viewing experience for the audience. Shaky footage can occur due to various factors such as hand movements while recording, vibrations, or other external disturbances. Various embodiment described herein, focus on software-based video stabilization techniques. It provides flexibility and may be applied to videos captured with any camera. Typically, it analyzes the video frames and applies corrective measures to reduce the apparent motion.

[0087] Viewport-dependent delivery (VDD) in omnidirectional 360-degree video refers to a streaming technique that optimizes the delivery of video content based on the viewer's field of view (FOV) in a 360-degree video. It is specifically designed to address the challenges of streaming high- resolution 360° (360) videos efficiently over limited bandwidth networks. A viewport refers to the portion of the 360 video representing a 3D spherical space or scene. Typically, a viewport is presented by the viewing orientation in the 3D space or coordinate system. The size of the viewport is typically determined by the field of view (FoV) of the viewing device. In the process of the viewport generation from a 360 video frame, the unwanted “jitter” associated with jerky and inadvertent movements of the original 360 video requires viewport-dependent stabilization to compensate for the unwanted movement. Typically, 360 video is projected and represented in 2D surface in a format called equirectangular projection (ERP). However, traditional video stabilization techniques or optical flow motion estimation may not work directly due to the fact that the image represents the whole 360 direction. For example, a motion in one viewing direction may be reversed in its opposite direction (see FIG. 5). For example, a horizontal motion in one viewport 550 (the right picture in FIG. 5) may be seen as two or more different local motions 552, 554 in its original 360 video frame 556 in equirectangular projection (ERP) mode. As a result of this motion inconsistence, in VDD streaming, the stabilization must be viewport dependent and in an example, may be done at the server side since only the viewport of the viewer is being streamed and there may be no visual information outside the viewport boundaries at the viewer side for motion stabilization. Also as illustrated in FIG. 5 motion as indicated on the right image 550, from side to side as shown by the 2-way arrow; results in an up and down motion on the left image 556 as indicated by the up arrow and down arrow.

[0088] The stabilization is done usually by cropping and warping the pixels in the new video frames from the old frames. In the VDD processing, the warping may be skipped as every viewport may be re-generated from the original 360 video frames with 2D-to-Sphere re-projection and cropping. Various embodiments describe the idea of motion-compensated viewport cropping and signaling with the viewport stabilization feature.

[0089] Viewport-dependent delivery (VDD) takes advantage of the fact that viewers generally focus their attention on a smaller portion of the 360 video, known as the viewport or FOV. One of VDD techniques is to produce the viewport region based on the constant FOV tracking and send the viewport content to the viewer based only on the viewer's current viewing direction (translated relatively in the 360-video coordinate system). During the viewport generation, the various embodiments work as following:

[0090] 2D motion estimation in viewport frames: The movement is estimated from the motion vectors of the current viewport of interests against the previous number of viewport frames (1 or n). A global movement (direction and magnitude) of the video (camera) movement from calculated individual motion vectors per smaller blocks are detected. It is important to avoid false-positive motion estimation from a natural and desired movement of the video device (e.g., a camera). Based on how to key points in the viewport are selected, there are 2 example approaches:- A pixel is selected from the ERP 2D plane, after the original 360 video frame 602 is rotated first to the 3D sphere from the viewport direction, and mapped back to a new ERP frame 604 based on the field of view where the viewport area (606) is located in the center (as illustrated in FIG. 6). Each viewport frame is generated from the FoV information. The original 360 frame is rotated first based on the viewing orientation (yaw / pitch / roll) to the center before cropping. The process may be performed in a graphics processing unit (GPU) to have fast possible speed with minimal delay; or- The pixel is selected from the 3D sphere 704 space mapping from the ERP 2D plane 701 (as illustrated in FIG. 7). Unlike the 2D viewport plane re-generation, the motion calculation may be done directly on the viewport surface (705) projected from the 2D viewport (702) on the 360 frame 701. In the example, it is straightforward to back-project a degree in a unit sphere (e.g., x / y degrees) to the x / y pixel value in the original 360 ERP plane, which may be referred to as the 2D motion compensation'a / y,

[0091] 2D to 3D viewport rotation adjustment: A direct viewport cropping on the 360 ERP space may not ensure the best result. The estimated motion vector needs to be translated into the 3D rotation angles (e.g., new yaw / pitch / roll) for the new ERP-to-sphere reprojection and viewport cropping 608.- The rotation 2-dimensional angles (x / y) from the viewer’s direction may then be compensated with the global motion estimated on ERP coordinate to a 3D Euler rotation delta to the original viewport angles.- Optionally, the cropped viewport size may also be larger than original viewport size with the motion magnitude considered. In this case, the motion data needs to be signaled to the receiver for rendering the real viewport content (client / rendering stabilization).

[0092] Viewport coding and packetization: The content which size comprises the viewport size plus extra margin is then encoded and packetized into a transport format such as RTP packets (FIG. 17) and sent to the receiver for VDD-based playback. Optionally, the motion data needs to be signaled per frame as RTP Header extension to the receiver when the viewport size is larger than the FOV size signaled from the viewer. The RTP extension or similar auxiliary signaling of the carriage of the motion data can be defined in the video streaming session description protocol such as SDP, or as, for example, attributes in an HTML document.

[0093] 360 viewport processing

[0094] Video stabilization is a technique to reduce or eliminate the shakiness or unwanted jittery motion in videos. A 360-degree video (360 video) is typically known as a 360 omnidirectional video that captures the full spherical view of the surroundings from typically one or more camera lenses. The 360 video produced by a 360 cameras is post-processed (or stitched) into a single 2D video output format using some sphere to 2D plane projection techniques. A commonly used mapping format is known as equirectangular projection (ERP) that covers the 360 sphere space.

[0095] In viewport-dependent 360 video streaming, a smaller resolution video content is delivered based on the viewer’s current viewport or screen size, available bandwidth; and / or device capabilities. Aligning the viewport content along with any motion patterns pattern preset the 360 video, aims to improve the overall viewing experience by making the video appear smoother and more professional. Shaky footage may result from various factors, such as handheld camera movements, vibrations, or other disturbances during recording. Video stabilization algorithms workto correct these issues and create a more stable final video.

[0096] United states patent US 11,711,505, hereinafter incorporated as a reference, describes viewport processing modes, for example, sphere-locked (SU) and viewport-locked (VU). In these examples the mode and its viewing angles are signaled from the sender to the receivers. In both modes, the viewport-only output requires ERP transformation from 2D ERP surface to a 3D sphere (a unit sphere) and apply the rotation based on the viewing angles (e.g., namely Viewport_Azimuth, Viewport_Elevation, Viewport_Tilt, plus other optional ones). The steps or rotation result is illustrated in FIG. 6 (azimuth, elevation, and tilt may also be represented as yaw, pitch, and roll). The final viewport may be cropped (606) based on the Field of View (FoV) size. It is worth noting that the cropped viewport segment is still mapped in the 360 ERP coordinate, not the surface in the unit sphere. Usually the surface in the sphere is referred to as a “rectified picture” using Gnomonic projection available from [https: / / matliworld.wolfram.com / GnomoiiicProjectioii.hbnl (last accessed on November 14, 2023) and [https: / / en.wikipedia.org / wiki / Gnomonic_projection (last accessed on November 14, 2023)].

[0097] Viewport motion estimation and motion-compensated viewport-dependent (VDD) streaming

[0098] Traditional vision-based video stabilization is based on the whole video frames. It is not appliable to the viewport-dependent solution as the motion is different per individual viewport. The stabilization has to be applied independently to every viewport generation process. So does the motion estimation.

[0099] FIG. 8 illustrates the way to distinguish any local motions, e.g., 802-1, 802-2, 802-3, and 802-04 caused by any moving objects versus a true global motion, e.g., 804-1, 804-2, 804-3, 804-04 noise caused by the environment. As there are many different motion patterns, e.g., straight, circular, and the like. In the case of video streaming, the frame per second is for example a good measurement. Given a 30 frames per second, each frame takes about 33 milliseconds. The proper timing window can be selected by the number of frames to identify the movement of objects or features between consecutive frames in a video sequence. A motion may be defined as the motion vector with the displacement (e.g., Euclidean norm) in x / y (2D coordinate system) from various motion estimation techniques.

[0100] The motion estimation process is illustrated in the diagram (FIG. 9. Following describes an example motion estimation algorithm based on the optical flow technique:- Generate a target viewport frame with given yaw / pitch / roll and FoV values) from a 360 ERP frame;- Use any method such as OpenCV’s available from[https: / / docs.opencv.Org / 3.4 / d4 / dee / tutorial optical flow.html (last accessed on November 20, 2023) to calculate the motion vectors of the viewport frame;- Do a global motion estimation by calculating the histogram of the x and y distances and their standard deviations (SD) of the motion histogram and do the filtering using a configurable global motion threshold; andCalculate global motion magnitude: Once a motion estimation is confirmed, the motion offsets magnitude in pixel can be estimated from the moving average of the motion vectors (902). Alternatively, a Seasonal Decomposition algorithm such as available from [https: / / ote.xts.com / fpp2 / decomposition.html (last accessed on November 20, 2023) may also be used to estimate the motion residuals too (904). FIG. 10 illustrates a motion estimation in one axis (e.g., y)

[0101] The estimation is done based on a timing window (the window size is >=1, the unit can be the consecutive frame numbers, or timing in millisecond, for example), which size is the time or the number of the viewport frames for the motion estimation.

[0102] FIG. 11 illustrates an example of 3D viewport-dependent stabilization process. FIG. 11 includes an example of key point extraction in motion calculation. FIG. 12 illustrates an example of 2 viewports being projected on a unit sphere. As shown in FIG. 12, the key points in FIG. 11 may be one-to-one mapped to the spherical surface on the unit sphere so that each point is more precise representing the motion situation in the real world. FIG. 11 illustrates an example of 2D viewportdependent stabilization process. In FIG. 11 key points are calculated directly from the cropped 2D viewport images, where some distortion may be introduced from the rotation processing step when the latitude of the viewport is more towards north or south poles in the ERP frame.

[0103] As illustrated in FIG. 11, at 1102, the method 1100 includes starting a 360 video viewport delivery processing with viewport parameters, for example, rotation (e.g., yaw / pitch / roll), and field of view (e.g., width / height). The whole ERP image is first projected (mapped) to a 3D sphere (for example, a unit sphere that is a sphere of unit radius in Euclidean space that are a distance of 1 from a given center point). At 1104, the method 1100 includes starting a sphere-based motion estimation analysis with a fixed or dynamic time window (e.g., from 1 to n frames) on viewport surface only (the surface size is determined by the FoV size). At 1106, the method 1100 includes extracting key points, for example, each key point is calculated on the 3D viewport surface directly, without the need of the 3D rotation to a fixed direction (for example, a front view from a viewer’sperspective). At 1108, the method 1100 includes estimating a global 2D motion data (x / y) in pixels of the viewport from all motion vectors of the key points on the viewport surface. At 1110, the method 1100 includes calculating new rotation angles (e.g., yaw2 / pitch2 / roll2) from the global motion data (e.g., x / y) in a pixels, considering the compensation of the spherical surface. At 1112, the method 1100 includes a rotating, based on the updated angles, to position the viewport region to the center of the 2D ERP image for the viewport creation. The rotation can be represented as a rotation matrix from the 3 angles as a 3x3 rotation matrix, which can be used to represent a rotation of a pixel in spherical coordinate in 3D space. In the case of a sphere, the combined rotation matrix (Rcombmed) can be created by multiplying the individual rotation by Euler angles in the rotation order of XYZ (Pitch, Yaw, and Roll), for example. At 1114, the method 1100 includes, cropping the newly created equirectangular projection (ERP) image (1112) to a 2D viewport image for viewportdependent delivery (VDD) process, for example, encoding and packaging processes for final delivery. Following are example of roll, pitch, yaw, and combined rotation matrices:

[0104] Turning to FIG. 13, at 1302, the method 1300 includes starting a 360 video viewport delivery processing with viewport parameters, for ex ample, rotation (e.g., yaw / pitch / roll), and FoV (e.g., width / height). At 1304, the method 1300 includes rotating an equirectangular projection (ERP) video frame to make the viewport to the center of the ERP surface (e.g., converting 2D-to-Sphere-to-2D). At 1306, the method 1300 includes extracting key points, e.g., each key point may be calculated directly from the viewport frame cropped from the rotated ERP video frame. At 1308, the method 1300 includes estimating a global 2D motion data (e.g., x / y) in pixels of the viewport from motion vectors of the key points in the viewport region. At 1310, the method 1300 includes, calculating new rotation angles (yaw2 / pitch2 / roll2) from the global motion data (x / y) in pixels. At 1312, the method 1300 includes, rotating based on the updated angles and FoV. At 1314, the method 1300 includes cropping the ERP to a 2D image for viewport-dependent delivery (VDD)process.

[0105] The 2D ERP approach illustrated in FIG. 13 may require an extra step to improve the stabilization accuracy to calculate an updated rotation angle (the vector3). The updated rotation data is done by a reversed projection as described in FIG. 13.

[0106] The key difference between key point extractions in sphere and plane is that the motion vectors generated are on the spere or plane surface. FIG. 14 is an illustration of the 1306, where the viewports from frame i and i+1 are cropped directly from the 360 frame i and frame i+1. In sphere surface, as described in FIG. 11, the motion estimated may be used directly to compensate the 3D viewport rotation vector; but the motion estimated in FIG. 13 requires a reversed calculation from the 2D ERP coordinate to the 3D viewport rotation of FIG. 15.

[0107] Special motion pattern

[0108] There is one possible special motion pattern, which motion may be difficult to be presented globally. In a certain viewing angle, a motion in a viewport video may result in the “zoom motion” pattern. There are, again, many zoom motion estimation techniques. The proposed embodiments can split the viewport region into 3 sub-regions in the global motion estimation. The process for this “zoom motion” is different from general straight motion transformation. When a zoom motion is detected, unlike using the updated viewing rotation, the field of view (FoV) gets updated by scaling proportionally to the zooming ratio as described in FIG. 16.

[0109] Motion signaling for effective receiver-side stabilization

[0110] It is important to allow effective viewport-dependent stabilization on the receiver side. Stabilization process is expensive alone in terms of computational resources (e.g., memory and CPU). This motion compensation may save the client-side stabilization process without computing the key points (and visual features) but simply apply the motion compensation from the sender (e.g., a server).[oni] As described in the Unites States patent US 11,711,505, the sphere-locked mode allows the delivery of some extra margin space, which is greater than the real viewport region size. The extra margin space tolerates some small degree of viewing motion in any direction without requesting a new viewport upon the small motion for a much better user experience. The amount of extra margin can enhance the stabilization efforts on the receiver side by compensating for motion as necessary.

[0112] In addition to the viewport margin signaled to the receiver in the US 11,711,505, the signaled compensation data for stabilization is needed. The change of the viewport with motioncompensation needs to be signaled to the receiver when a 360 VDD is streaming over the network or playing back, when any receiver-side adaptation process may be necessary. The signaling may be out of band using any transport protocols. In practice, an in-band approach is usually preferred as it can be perfectly synchronized with the motion data and the updated viewport contents. For example, a motion supplemental enhancement information (SEI) message in video codecs such as ITU-T AVC / HEVC / VVC or in a similar way) that tells where in the motion data applied to the particular frame corresponds to may be included in the video stream. Alternative, the sender may signal them as RTP header extension () when the coded viewport stream is packed into the RTP format.

[0113] RTP header extensions are discussed in IETF RFC 8285 available from [https: / / datatracker.ietf.org / doc / html / rfc8285 (last accessed on November 20, 2023)1: “A General Mechanism for RTP Header Extensions”, D. Singer, H. Desineni, which is hereinafter incorporated by reference in its entirety. RTP payload formats may specify an RTP payload header, which may include metadata of the payload comprising compressed video data. The RTP payload header may have an extensible format, thus allowing for RTP payload header extensions. For example, the RTP payload format for High-Efficiency Video Coding, specified in IETF RFC 7798, specifies a payload header extension structure (PHES). In an example embodiment, an RTP payload header extension may carry sphere rotation metadata and / or delivered viewport FOV metadata, as described in United States patent US 11,711,505.

[0114] An RTP header extension for motion compensation (MC) based on RFC 5285 may comprise the updated viewport. An example urn to signal MC in SDP may be: a=extmap:5 urn: motion-compensation where the value “5” in the example may be any value in the range 1-14. The one-byte RTP extension header format is used in this example too.

[0115] In an example embodiment, the motion compensation may be negotiated as a session negotiation parameter, for example, as a session description protocol (SDP). An SDP is used during the session negotiation phase to inform the receiver side that any RTP header extension data are included in the RTP packets.

[0116] Both motion vector and updated viewport data may be signaled via the SEI or RTP header extension.

[0117] FIG. 17 illustrates an example of a One-Byte Header extension 1700. The 4-bit “ID” is the local identifier of this element in the range 1-14 inclusive. The 4-bit “len” (length) is thenumber, minus one, of data bytes of this header extension element following the one-byte header. In the follow example header extension, which is extended from the header extension described in United States patent US 11,711505, where the viewport data is followed by the motion compensation x / y bytes (total 2 x 16 bits).

[0118] The Motion X and Motion Y represents the 2D motion compensation values that are relative offset values in pixels (ERP plane coordinate relative to the viewport frame), or in 3 Euler angles relative to the rotation values (azimuth, tilt, and elevation).

[0119] When a special motion pattern like zooming compensation is detected as described in special motion pattern, the RTP header extension may include zooming ratio header extension as well. It may affect the field of view (FoV) handling on the receiver how to render the viewport frame to the sphere (projection or rectification rendering).

[0120] FIG. 18 is an example apparatus 1800, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 1800 comprises at least one processor 1802 (e.g., an FPGA and / or CPU), one or more memories 1804 including computer program code 1805, the computer program code 1805 having instructions to carry out the methods described herein, wherein the at least one memory 1804 and the computer program code 1805 are configured to, with the at least one processor 1802, cause the apparatus 1800 to implement circuitry, a process, component, module, or function (implemented with control module 1806) to implement the examples described herein, including viewport stabilization, for example, a server-driven viewport stabilization in viewport-dependent omnidirectional video streaming. Optionally included encoder 1830 of the control module 1806 performs encoding, and optionally included decoder 1840 implements decoding. The memory 1804 may be a non-transitory memory, a transitory memory, a volatile memory (e.g., RAM), or a non-volatile memory (e.g., ROM).

[0121] The apparatus 1800 includes a display and / or I / O interface 1808, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 1800 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 1810. The communication I / F(s) 1810 may be wired and / or wireless and communicate over the Intemet / other network(s) via any communication technique including via one or more links 1824. The communication I / F(s) 1810 may comprise one or more transmitters or one or more receivers.

[0122] The transceiver 1816 comprises one or more transmitters 1818 and one or morereceivers 1820. The transceiver 1816 and / or communication I / F(s) 1810 may comprise standard well- known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 1814 used for communication over wireless link 1822.

[0123] The control module 1806 of the apparatus 1800 comprises one of or both parts 1806-1 and / or 1806-2, which may be implemented in a number of ways. The control module 1806 may be implemented in hardware as control module 1806-1, such as being implemented as part of the one or more processors 1802. The control module 1806-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 1806 may be implemented as control module 1806-2, which is implemented as computer program code (having corresponding instructions) 1805 and is executed by the one or more processors 1802. For instance, the one or more memories 1804 store instructions that, when executed by the one or more processors 1802, cause the apparatus 1800 to perform one or more of the operations as described herein. Furthermore, the one or more processors 1802, one or more memories 1804, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.

[0124] The apparatus 1800 to implement the functionality of control 1806 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 1800 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 1800 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.

[0125] The apparatus 1800 may also be distributed throughout the network (e.g. internet 28) including within and between apparatus 1800 and any network element (such as a base station 24 and / or apparatus 50).

[0126] Interface 1812 enables data communication and signaling between the various items of apparatus 1800, as shown in FIG. 18. For example, the interface 1812 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 1805, including control 1806 may comprise object-oriented software configured to pass data or messages between objects within computer program code 1805. The apparatus 1800 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 1800 may at least partially reside in a common housing 1828, or a subset of the various components of apparatus 1800may at least partially be located in different housings, which different housings may include housing 1828.

[0127] FIG. 19 shows a schematic representation of non-volatile memory media 1900a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 1900b (e.g. universal serial bus (USB) memory stick) and 1900c (e.g., cloud storage for downloading instructions and / or parameters 1902 or receiving emailed instructions and / or parameters 1902) storing instructions and / or parameters 1902 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein.

[0128] FIG. 20 is an example method 2000, based on examples described herein. At 2002, the method 2000 includes starting a 360 video viewport delivery processing for a viewport. At 2004, the method 2000 includes extracting key points in the viewport. At 2006, the method 2000 includes estimating motion vectors of the key points on the viewport surface. At 2008, the method 2000 includes estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface. At 2010, the method 2000 includes calculating updated rotation angles from the global motion data in the pixel. At 2012, the method 2000 includes rotating the 360 video frame based on the updated rotation angles and a field of view (FoV). At 2014, the method 2000 includes cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

[0129] The method 2000 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 4, FIG. 18, or any other apparatus described herein.

[0130] FIG. 21 is an example method 2100, based on examples described herein. At 2102, the method 2100 includes starting a 360 video viewport delivery processing for a viewport. At 2104, the method 2100 includes extracting key points in the viewport. At 2106, the method 2100 includes estimating zoom motion of the key points in the viewport surface, based on zooming ration. At 2108, the method 2100 includes updating a field of view by scaling the field of view based on the zoom motion.

[0131] The method 2100 may be performed with an apparatus described herein, for example, any apparatus of FIG. 1 to FIG. 4, FIG. 18, or any other apparatus described herein.

[0132] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0133] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0134] In the above, some example embodiments have been described with reference to an SEI message or an SEI NAL unit. It needs to be understood, however, that embodiments may similarly be realized with any similar structures or data units, such as metadata OBUs. Where example embodiments have been described with SEI messages included in a structure, any independently parsable structures could likewise be used in embodiments. Specific SEI NAL unit and SEI message syntax structures have been presented in example embodiments, but it needs to be understood that embodiments generally apply to any syntax structures with a similar intent as SEI NAL units and / or SEI messages.

[0135] In the above, some example embodiments have been described with reference to postfilters. It is to be understood that embodiments may similarly be realized with in-loop filters. Furthermore, it is to be understood that some embodiments have been described with reference to syntax structures, such as SEI messages, that are suitable for post-filters. It is to be understood that embodiments for in-loop filters may be similarly realized with other syntax structures, such as parameter sets (e.g., video, sequence, picture, and / or adaptation parameter sets) and / or headers of different structural level(s) (e.g., sequence, group of pictures, picture, and / or slice header).

[0136] In the above, some embodiments have been described with reference to neural-network filters. It is to be understood that embodiments may similarly be realized with any filters.

[0137] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, etc.

[0138] As used herein, the term ‘circuitry’, ‘circuit’ and variants may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable):(i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and one or more memories that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device. Circuitry or circuit may also be used to mean a function or a process used to execute a method.

[0139] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.

[0140] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):2D two-dimensional3D three-dimensional3GPP 3rd generation partnership project4G fourth generation of broadband cellular network technology5G fifth generation cellular network technology802.x family of IEEE standards dealing with local area networks and metropolitan area networksASIC application specific integrated circuitAVC advanced video codingBD Bjontegaard delta (e.g. BD-rate)BT.2020 set of specifications covering various aspects of video broadcastingBT.709 standard developed by ITU-R for image encoding and signal characteristics of high-definition televisionCDMA code-division multiple accessCLVS coded layer video sequenceCPU central processing unitCTU coding tree unitDCT discrete cosine transformDPB decoded picture bufferDSNN decoder-side neural networkDSP digital signal processorFDMA frequency division multiple accessFPGA field programmable gate arrayGSM global system for mobile communicationsH.222.0 MPEG-2 systems, standard for the generic coding of moving pictures and associated audio informationH.2xx family of video coding standards in the domain of the ITU-T (e.g. H.263,H.264, H.266)HEVC high efficiency video codingHMD head-mounted displayIBC intra block copyIEC International Electrotechnical CommissionIEEE Institute of Electrical and Electronics EngineersI / F interfaceIMD integrated messaging deviceIMS instant messaging serviceI / O input / output loT internet of thingsIP internet protocolISO International Organization for StandardizationISOBMFF ISO base media file formatITU International Telecommunication UnionITU-R ITU Radiocommunication SectorITU-T ITU Telecommunication Standardization SectorJVET Joint Video Experts TeamJVET-O1150 scalable coding proposalLTE long-term evolutionMAE mean absolute error mAP mean average precisionMMS multimedia messaging serviceMOTA multi-object tracking accuracyMPEG-2 moving picture experts group, H.222 / H.262 as defined by the ITUMSE mean squared errorNAL network abstraction layerNN neural networkNNPF neural-network post-processing filterNNPFA neural-network post-filter activationNNPFC neural-network post-filter characteristicsN / W networkPC personal computerPDA personal digital assistantPID packet identifierPLC power line communicationPSNR peak signal-to-noise ratioQP quantization parameterRAM random access memoryRFID radio frequency identificationRFM reference frame memoryROI region of interestSEI supplemental enhancement informationSMS short messaging serviceSNR signal to noise ratioSON self-organizing / optimizing networkSSIM structural similarity index measureTCP-IP transmission control protocol-internet protocolTDMA time divisional multiple accessTS transport streamTV televisionUHDTV ultra-high-definition televisionUI user interfaceUICC universal integrated circuit cardUMTS universal mobile telecommunications systemURI uniform resource identifierURL uniform resource locatorUSB universal serial busVCM video coding for machinesVPS video parameter setVSEI versatile supplemental enhancement information vvc versatile video coding

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

2. The apparatus of claim 1, wherein the apparatus is further caused to perform: encoding and signaling viewport dependent delivery data.

3. The apparatus of any of claims 1 or 2, wherein the viewport dependent delivery data comprises the global 2D motion data.

4. The apparatus of any of claims 1 to 3, wherein the apparatus is further caused to perform: rotating an original 360 video frame to a three-dimensional (3D) sphere from the viewport direction to obtain a rotated 360 video frame; mapping the rotated 360 video frame to a new equirectangular projection (ERP) frame in 2D for which a viewport is located in a center of the field of view; and selecting a pixel from the new ERP frame.

5. The apparatus of any of the claims 1 to 3, wherein the apparatus further comprises: selecting a pixel from a three-dimensional space mapping from the ERP to the 2D image.

6. The apparatus of claim 1, wherein the apparatus is further caused to perform: compensating rotation 2D angels from a direction of a viewer with the global 2D motion data to a 3D Euler rotation delta to original viewport angles; andsignaling the global 2D motion data when a size of the cropped 360 video frame or a cropped equirectangular projection (ERP) is larger than a size of the viewport.

7. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: negotiating motion compensation as a session negotiation parameter.

8. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: signaling viewport dependent delivery data via a header extension.

9. The apparatus of claim 8, wherein the header extension comprises: a local identifier of the header extension; and a length field for identifying length of the header extension.

10. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

11. The apparatus of claim 10, wherein the apparatus is further caused to perform: signaling viewport data via a header extension.

12. The apparatus of claim 11, wherein the header extension comprises zooming ratio header extension for identifying zooming ratio.

13. A method comprising: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel;rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

14. The method of claim 13 further comprising encoding and signaling viewport dependent delivery data.

15. The method of any of claims 13 or 14, wherein the viewport dependent delivery data comprises the global 2D motion data.

16. The method of any of the claims 13 to 15 further comprising: rotating an original 360 video frame to a three-dimensional (3D) sphere from the viewport direction to obtain a rotated 360 video frame; mapping the rotated 360 video frame to a new equirectangular projection (ERP) frame in 2D for which a viewport is located in a center of the field of view; and selecting a pixel from the new ERP frame.

17. The method of any of the claims 13 to 15 further comprising: selecting a pixel from a three- dimensional space mapping from the ERP to the 2D image.

18. The method of claim 13 further comprising: compensating rotation 2D angels from a direction of a viewer with the global 2D motion data to a 3D Euler rotation delta to original viewport angles; and signaling the global 2D motion data when a size of the cropped 360 video frame or a cropped equirectangular projection (ERP) is larger than a size of the viewport.

19. The method of any of the claims 13 to 18 further comprising: negotiating motion compensation as a session negotiation parameter.

20. The method of any of the claims 13 to 19 further comprising: signaling viewport dependent delivery data via a header extension.

21. The method of claim 20, wherein the header extension comprises: a local identifier of the header extension; and a length field for identifying length of the header extension.

22. A method comprising: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

23. The method of claim 22 further comprising: signaling viewport data via a header extension.

24. The method of claim 23, wherein the header extension comprises zooming ratio header extension for identifying zooming ratio.

25. An apparatus comprising: means for starting a 360 video viewport delivery processing for a viewport; means for extracting key points in the viewport; means for estimating motion vectors of the key points on the viewport surface; means for estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; means for calculating updated rotation angles from the global motion data in the pixel; means for rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and means for cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

26. The apparatus of claim 25, wherein the apparatus further comprises means for performing the methods as claimed in any of the claims 14 to 21.

27. An apparatus comprising: means for starting a 360 video viewport delivery processing for a viewport; means for extracting key points in the viewport; means for estimating zoom motion of the key points in the viewport surface, based on zooming ration; and means for updating a field of view by scaling the field of view based on the zoom motion.

28. The apparatus of claim 27 further comprising means for performing the methods as claimedin any of the claims 23 or 24.

29. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating motion vectors of the key points on the viewport surface; estimating a global two-dimensional (2D) motion data in pixel of the viewport from the motion vectors of the key points on the viewport surface; calculating updated rotation angles from the global motion data in the pixel; rotating the 360 video frame based on the updated rotation angles and a field of view (FoV); and cropping the 360 video frame to a viewport image for the viewport-dependent delivery (VDD) process.

30. The computer readable medium of claim 29, wherein the computer readable medium comprises a non-transitory computer readable medium.

31. The computer readable medium of any of claims 29 or 30 comprises program instructions which, when executed, cause the apparatus to perform the methods as claimed in any of the claims 14 to 21.

32. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform: starting a 360 video viewport delivery processing for a viewport; extracting key points in the viewport; estimating zoom motion of the key points in the viewport surface, based on zooming ration; and updating a field of view by scaling the field of view based on the zoom motion.

33. The computer readable medium of claim 32, wherein the computer readable medium comprises a non-transitory computer readable medium.

34. The computer readable medium of any of claims 32 or 33 comprises program instructions which, when executed, cause the apparatus to perform the methods as claimed in any of the claims 23

Citation Information

Patent Citations

  • Viewport dependent delivery methods for omnidirectional conversational video

    US11711505B2

  • Method and apparatus for stabilising 360 degree video

    GB2562529A

  • Viewport Dependent Delivery Methods For Omnidirectional Conversational Video

    US20220021864A1

  • Method, an apparatus and a computer program product for video conferencing

    US20230033063A1