360-degree video coding using face continuities
By converting 360-degree videos to cubemap format using geometric padding and chroma subsampling, the inefficiencies in encoding and decoding are addressed, resulting in improved coding efficiency and reduced bandwidth for 360-degree video delivery.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- INTERDIGITAL VC HOLDINGS INC
- Filing Date
- 2018-04-10
- Publication Date
- 2026-07-29
AI Technical Summary
Existing video codecs struggle to efficiently encode and decode 360-degree videos due to the uneven spherical sampling density and complex motion fields in equirectangular projection, leading to inefficient compression and increased bandwidth requirements, particularly in areas less focused by viewers.
The use of cubemap projection and geometric padding techniques to convert 360-degree videos into a cubemap format, combined with chroma subsampling and adaptive interpolation, enhances encoding efficiency by maintaining spherical geometry and reducing bandwidth needs.
This approach improves coding efficiency and reduces bandwidth requirements by aligning with viewer behavior, providing high-quality 360-degree video experiences with reduced computational resources.
Smart Images

Figure 112023085296098-PAT00076_ABST
Abstract
Description
Technology Field
[0001] Cross-reference regarding related applications
[0002] This application claims the benefits of U.S. provisional application serial number 62 / 484,218 filed April 11, 2017 and U.S. provisional application serial number 62 / 525,880 filed June 28, 2017, the contents of which are incorporated herein by reference. Background Technology
[0003] Virtual reality (VR) has begun to enter our daily lives. For example, VR has many applications in fields including, but not limited to, healthcare, education, social networking, industrial design / training, gaming, movies, shopping, and / or entertainment. VR can enhance the viewer's experience, for example, by creating a virtual environment that surrounds the viewer and by generating a genuine sense of "being there" for the viewer. For example, the user experience may rely on providing a full real feeling within the VR environment. For instance, VR systems may support interaction through posture, gestures, eye gaze, and / or voice. Systems may also provide haptic feedback to users to allow them to interact with objects in a natural way within the VR world. VR systems may also use 360-degree video to provide users with the ability to view scenes from a 360-degree horizontal angle and / or a 180-degree vertical angle.
[0004] A coding device (e.g., a device that may or may not include an encoder and / or decoder) may receive a frame-packed picture of a 360-degree video. The frame-packed picture may include a number of faces and a current block. The coding device may identify a face in the frame-packed picture to which the current block belongs. The coding device may determine that the current block is located at the exiting boundary of the face to which the current block belongs. For example, the coding device may determine that the current block is located at the exiting boundary of the face to which the current block belongs according to the coding order of the frame-packed picture. The exiting boundary of the face to which the current block belongs may be located in the same direction as the coding order with respect to the current block.
[0005] Frame-packed pictures may be coded in coding order. In the example, the coding order may have a left-to-right orientation relative to the current block associated with the frame-packed picture. In the example, the coding order may have a top-to-bottom orientation relative to the current block. In the example, the coding order may have a left-to-right and top-to-bottom orientation relative to the current block. For example, if the coding order has a left-to-right orientation relative to the current block, the exit boundary of the face may be located on the right (e.g., the rightmost side of the face to which the current block belongs). For example, if the coding order has a top-to-bottom orientation relative to the current block, the exit boundary of the face may be located on the bottom side (e.g., the bottommost side of the face to which the current block belongs). For example, if the coding order is oriented from left to right and from top to bottom relative to the current block, the exit boundary of the face may be located on the right and bottom sides (e.g., the far right and far bottom sides of the face to which the current block belongs).
[0006] When it is determined that the current block is located at the exit boundary of the face to which the current block belongs, the coding device may use cross-face boundary neighboring blocks located on a face that shares a boundary with the exit boundary of the face to which the current block belongs in order to code the current block. For example, the coding device may identify multiple spherical neighboring blocks of the current block. For example, the coding device may identify multiple spherical neighboring blocks of the current block based on the spherical characteristics of a 360-degree video. The coding device may also identify cross-face boundary neighboring blocks associated with the current block. For example, the coding device may identify cross-face boundary neighboring blocks among the identified multiple spherical neighboring blocks of the current block. In an example, the cross-face boundary neighboring blocks may be located on a face that shares a boundary with the exit boundary of the face to which the current block belongs. In an example, the cross-face boundary neighboring blocks may be located on the opposite side of the face boundary to which the current block belongs, or may be located in the same direction of the coding order relative to the current block.
[0007] The coding device may determine whether to use a plane boundary cross-neighbor block to code the current block. For example, a block within a frame-packed picture corresponding to a plane boundary cross-neighbor block may be identified. A block within a frame-packed picture corresponding to a plane boundary cross-neighbor block may be identified based on frame-packing information of a 360-degree video. The coding device may determine whether to use an identified block within a frame-packed picture corresponding to a plane boundary cross-neighbor block to code the current block based on the availability of the identified block within the frame-packed picture. For example, the availability of the identified block within the frame-packed picture may be determined based on whether the identified block has been coded. The coding device may code the current block based on the decision to use the identified block within the frame-packed picture, or it may code the current block using an identified available block corresponding to a plane boundary cross-neighbor block.
[0008] When used herein, 360-degree video may include or be spherical video, omnidirectional video, virtual reality (VR) video, panoramic video, immersive video (e.g., light field video which may include 6 degrees of freedom), point cloud video, and / or the like. Brief explanation of the drawing
[0009] FIG. 1a depicts exemplary sphere sampling of longitude (φ) and latitude (θ). Figure 1b depicts an exemplary sphere projected onto a 2D plane using equirectangular projection (ERP). Figure 1c depicts an exemplary picture generated using ERP. Figure 2a depicts an exemplary 3D geometry structure in cubemap projection (CMP). FIG. 2b depicts an exemplary 2D plane having a 4×3 frame packing and six faces. Figure 2c depicts an exemplary picture generated using CMP. Figure 3 depicts an exemplary workflow for a 360-degree video system. Figure 4a depicts an exemplary picture generated by repetitive padding boundaries using ERP. Figure 4b depicts an exemplary picture generated by iterative padding boundaries using CMP. Figure 5a depicts exemplary geometry padding for an ERP representing padding geometry. Figure 5b depicts exemplary geometry padding for an ERP showing a padded ERP picture. Figure 6a depicts an exemplary geometry padding process for a CMP representing padding geometry. Figure 6b depicts an exemplary geometry padding process for a CMP representing a padded CMP face. Figure 7 is an exemplary diagram of a block-based video encoder. Figure 8 is an exemplary diagram of a block-based video decoder. Figure 9 depicts an exemplary reference sample used for high efficiency video coding (HEVC) intra prediction. Figure 10 depicts an exemplary representation of the intra-predicted direction in HEVC. Figure 11a depicts an exemplary projection of a left reference sample to extend the top reference row to the left. Figure 11b depicts an exemplary projection of an upper reference sample to extend the left reference column upward. FIGS. 12a through 12d depict exemplary boundary prediction filtering for (A) intra mode 2; (B) intra mode 34; (C) intra modes 3-6; and (D) intra modes 30-33. Figure 13 illustrates exemplary spatial neighborhoods used for the most possible modes in an HEVC intra-angular process. Figure 14 depicts exemplary locations of samples used to derive α and β in cross-component linear model prediction. Figure 15 depicts an exemplary inter prediction with a single motion vector. Figure 16 depicts exemplary padding for a reference sample outside the picture boundary. Figure 17 depicts exemplary spatial neighbors used for merge candidates in the HEVC merge process. Fig. 18a depicts an exemplary 3D representation of a CMP. FIG. 18b depicts an exemplary 3×2 frame packing configuration of a CMP. Figure 19 depicts an exemplary reconstructed sample used to predict the current block in intra and inter coding. FIGS. 20a through 20c depict exemplary spatial neighbors at (A) the right face boundary; (B) the bottom face boundary; and (C) the bottom right face boundary. FIGS. 21a through 21c depict exemplary availability of a reconstructed sample at (A) a right-side boundary; (B) a bottom-side boundary; and (C) a bottom-right-side boundary. FIGS. 22a and FIGS. 22b depict exemplary additional intra-prediction modes at (A) the right face boundary; and (B) the bottom face boundary. FIGS. 23a through 23d depict exemplary additional intra-prediction modes at (AB) right-side boundary; and (CD) bottom-side boundary. FIGS. 24a through 24d depict exemplary bidirectional intra predictions at (AB) right face boundary; and (CD) bottom face boundary. FIGS. 25a through 25h depict exemplary boundary prediction filtering at the right-side boundary for: (A) Intra Mode 2; (B) Intra Modes 3-6; (C) Intra Modes 7-9; (D) Intra Mode 10; (E) Intra Modes 11-17; (F) Intra Mode 18; (G) Intra Modes 19-21; and (H) Intra Mode 22. FIGS. 26a through 26h depict exemplary boundary prediction filtering at the bottom face boundary for: (A) Intra Mode 14; (B) Intra Modes 15-17; (C) Intra Mode 18; (D) Intra Modes 19-25; (E) Intra Mode 26; (F) Intra Modes 27-29; (G) Intra Modes 30-33; and (H) Intra Mode 34. FIGS. 27a through 27c depict exemplary locations of samples used for component cross-linear model prediction at (A) the right-side boundary; (B) the bottom-side boundary; and (C) the bottom-right-side boundary. FIG. 28 illustrates exemplary mathematical formulas for calculating linear model parameters (e.g., mathematical formulas (38), (41), (43) and (44)). FIGS. 29a and 29b depict examples of block processing sequences for a CMP 3×2 packing configuration: (A) raster scan sequence; and (B) face scan sequence. FIG. 30 illustrates an exemplary coding tree unit (CTU) and block partitioning. FIG. 31 depicts an exemplary 3×2 packing configuration. Dotted lines may indicate CTU boundaries, and arrows may indicate shared boundaries between two faces. FIG. 32 depicts an exemplary 3×2 packing configuration as described herein. Dotted lines may indicate CTU boundaries, and arrows may indicate shared boundaries between two faces. FIG. 33a is a system drawing illustrating an exemplary communication system in which one or more disclosed embodiments may be implemented. FIG. 33b is a system drawing illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 33a according to an embodiment. FIG. 33c is a system diagram illustrating an exemplary radio access network (RAN) and an exemplary core network (CN) that may be used within the communication system illustrated in FIG. 33a according to an embodiment. FIG. 33d is a system diagram illustrating additional exemplary RAN and additional exemplary CN that may be used within the communication system illustrated in FIG. 33a according to an embodiment. Specific details for implementing the invention
[0010] Now, a detailed description of exemplary embodiments will be given with reference to the various drawings. Although this description provides detailed examples of possible embodiments, it should be noted that the details are intended to be illustrative and are not intended to limit the scope of this application in any way.
[0011] VR systems and / or 360-degree video may be intended for media consumption that surpasses, for example, Ultra High Definition (UHD) services. Improving the quality of 360-degree video in VR and / or standardizing the processing chain for client interoperability may have been focused by one or more groups. For example, an ad-hoc group may have been established within ISO / IEC / MPEG to work on requirements and / or technologies for omnidirectional media application formats. For example, an ad-hoc group may have conducted exploratory experiments on 360-degree 3D video applications. An ad-hoc group may have tested systems based on 360-degree video (e.g., omnidirectional video) and / or multi-view based systems. The Joint Video Exploration Team (JVET) from MPEG and ITU-T, which is researching technologies for next-generation video coding standards, has announced requirements for test sequences including VR. An ad-hoc group (AHG8) was established, and the mandate of the AHG8 group is to create general test conditions, test sequence formats, and evaluation criteria for 360-degree video coding. AHG8 may also study conversion software as well as the impact on compression when different projection methods are applied. One or more companies willingly provided several 360-degree videos as test sequences to develop 360-degree video coding techniques. To conduct experiments following a set of general test conditions and evaluation procedures, reference software 360Lib was established by JVET to perform projection format conversions and measure objective 360-degree video quality metrics.In light of interest in 360-degree video coding, JVET agreed to include 360-degree video in a preliminary joint request for evidence of video compression with performance superior to HEVC.
[0012] The quality of one or more aspects of the VR processing chain, including capture, processing, display, and / or applications, and / or the user experience may be enhanced. For example, on the capture side, the VR system may use one or more cameras to capture scenes from one or more different views (e.g., 6 to 12 views). The different views may be stitched together to form a high-resolution (e.g., 4K or 8K) 360-degree video. For example, on the client and / or user side, the VR system may include a computing platform, a head-mounted display (HMD), and / or head tracking sensors. The computing platform may receive and / or decode the 360-degree video and generate a viewport for display. Two pictures, one for each eye, may be rendered for the viewport. The two pictures may also be displayed on the HMD for stereo viewing. Lenses may be used to magnify the image displayed on the HMD for a better field of view. The head tracking sensor may continuously track the viewer's head orientation (e.g., continuously) and may supply orientation information to the system to display a viewport picture for that orientation. The VR system may provide a touch device (e.g., a specialized touch device) for the viewer to interact with objects in the virtual world, for example. In an example, the VR system may be powered by a workstation with GPU support. In an example, the VR system may use a smartphone as a computing platform, an HMD display, and / or a head tracking sensor. The spatial HMD resolution may be, for example, 2160×1200. The refresh rate may be 90 Hz, and the field of view (FOV) may be 110 degrees.The sampling rate for head tracking sensors may be 1000 Hz, which can capture fast (e.g., very fast) movements. Examples of VR systems may use a smartphone as a computing platform and may include lenses and / or cardboard. There may also be 360-degree video streaming services.
[0013] The quality of the experience, such as interaction and / or haptic feedback, can be enhanced in a VR system. For example, the HMD may be too large and / or uncomfortable to wear. The resolution provided by the HMD (e.g., 2160×1200 for stereoscopic views) may not be sufficient and may cause dizziness and / or discomfort to the user. The resolution may be increased. The sensations from vision in a VR environment may be combined with real-world feedback (e.g., force feedback) to enhance the VR experience. A VR roller coaster may be an example of such a combined application.
[0014] 360-degree video delivery may represent 360-degree information using, for example, a spherical geometry structure. For example, one or more synchronized views captured by one or more cameras may be stitched together as a single structure on the sphere. The spherical information may be projected onto a 2D planar surface through a geometry transformation process. For example, equirectangular projection (ERP) and / or cubemap projection (CMP) may be used to exemplify the projection format.
[0015] The ERP may also map the latitude and / or longitude coordinates of a spherical object to the horizontal and / or vertical coordinates of a grid (e.g., directly). Fig. 1a depicts an example of sphere sampling at longitude (φ) and latitude (θ). Fig. 1b depicts an example of a sphere projected onto a 2D plane using the ERP, for example. Fig. 1c depicts an example of a projected picture through the ERP. Longitude (φ) in the range [-π,π] may be yaw, and latitude (θ) in the range [-π / 2,π / 2] may be pitch in aeronautics. π may be the ratio of the circumference of a circle to its diameter. In Figs. 1a and 1b, (x, y, z) may represent the coordinates of a point in 3D space, and (ue, ve) may represent the coordinates of a point in a 2D plane. ERP may also be expressed mathematically as shown in mathematical formula (1) and / or mathematical formula (2):
[0016]
[0017] Here, W and H may be the width and height of the 2D planar picture. As illustrated in FIG. 1a, point P, which is the intersection point between longitude L4 and latitude A1 on the sphere, may be mapped to an intrinsic point q on the 2D plane (e.g., FIG. 1b) using Equation (1) and / or (2). Point q on the 2D plane may be back-projected to point P on the sphere through inverse projection. The field of view (FOV) in FIG. 1b may illustrate an example where the FOV on the sphere may be mapped to the 2D plane with a field of view along the X-axis of approximately 110 degrees.
[0018] 360-degree video may be mapped to 2D video using ERP. For example, 360-degree video may be encoded using video codecs such as H.264 and / or HEVC. The encoded 360-degree video may be delivered to the client. On the client side, equirectangular video may be decoded. Equirectangular video may be rendered based on the user's viewport, for example, by projecting and / or displaying the portion of the FOV within the equirectangular picture onto the HMD. Spherical video may be converted into a 2D planar picture for encoding using ERP. The characteristics of the equirectangular 2D picture may differ from those of a 2D picture (e.g., linear video).
[0019] FIG. 1c depicts an exemplary picture generated using ERP. As illustrated in FIG. 1c, the top and / or bottom portions of the ERP picture (e.g., the North and / or South Poles, respectively) may be stretched relative to the middle portion of the picture (e.g., the equator). The stretching of the top and / or bottom portions of the ERP picture may indicate that the spherical sampling density may be uneven for the ERP format. The motion field, which may explain the temporal correlation between neighboring ERP pictures, may be more complex than in 2D video.
[0020] Video codecs (e.g., MPEG-2, H.264, or HEVC) may use transformation models to describe the motion field. Video codecs may not be able to represent shape-changing motion in an equirectangular projected 2D planar picture (e.g., may not be able to represent it efficiently). As illustrated in Fig. 1c, areas closer to the poles (e.g., North and / or South Pole) in the ERP may be less interesting to viewers and / or content providers. For example, viewers may not focus on the top and / or bottom areas for long durations. Based on the warping effect, the stretched areas may become a large portion of the 2D plane after the ERP, and compressing these areas may require many bits. Equirectangular picture coding may be improved by applying preprocessing, such as smoothing, to the polar areas to reduce the bandwidth required to code the polar areas. To map a 360-degree video onto multiple faces, one or more geometric projections may be used. For example, one or more geometric projections may include, but are not limited to, cubemaps, equal-area, cylinders, pyramids, and / or octahedrons.
[0021] Cube Map Projection (CMP) may be a compression-friendly format. A CMP includes six faces. For example, a CMP may include six square faces. The faces may be planar squares. Fig. 2a depicts an exemplary 3D geometric structure in a CMP. If the radius of the tangent sphere is 1 (e.g., Fig. 2a), the lateral length of one or more faces of the CMP (e.g., square faces) may be 2. Fig. 2b depicts an exemplary 2D packing method for arranging the six faces into a rectangular picture that may be used for encoding and / or transmission. Fig. 2c depicts an exemplary picture generated using a CMP. The shaded area shown in Fig. 2c may be a padded area to fill the rectangular picture. In the case of the faces, the picture may look identical to a 2D picture. The boundaries of the faces may not be continuous. For example, a straight line crossing two adjacent faces may be curved and / or may be multiple line segments (e.g., two line segments) at the boundary between the two faces. Motion at the face boundary may be discontinuous.
[0022] One or more objective quality metrics have been proposed for the coding efficiency of one or more different geometry projection methods. For example, peak signal-to-noise ratio (PSNR) measurements may include spherical PSNR (S-PSNR) and viewport PSNR. In S-PSNR, distortion may be measured using the mean square error (MSE) calculated across a predefined set of samples (e.g., these may be uniformly distributed on a sphere). Latitude-based PSNR (L-PSNR) may also be used. L-PSNR may take into account the viewer's viewing behavior by weighting one or more samples based on the latitude of the samples. The weights may be derived by tracking the viewer's field of view when the viewer observes the training sequence. If it is viewed frequently, the weights may be greater. From statistics, the weights around the equator may be greater. For example, the weight around the equator may be greater than the weight near the pole(s) because interesting content may be located around the equator. For viewport PSNR, the viewport may be rendered, and the PSNR may be calculated on the rendered viewport. A portion of the sphere may be considered, for example, for distortion measurement. The average viewport PSNR may be calculated for multiple viewports covering different parts of the sphere. S-PSNR may consider multiple samples. For example, S-PSNR may consider samples that may be uniformly distributed on the sphere. Weighted to spherically uniform PSNR (WS-PSNR) may also be used. WS-PSNR may calculate the PSNR using one or more (e.g., all) samples available on the 2D projection plane.For one or more locations on a 2D projection plane, distortion may be weighted by the spherical area covered by that sample location. WS-PSNR may be calculated, for example, directly on the projection plane. Different weights may be derived for different projection formats. Craster Parabolic Projection (CPP) may be used to project a 360-degree image and / or PSNR may be calculated on the projected image. This approach may be CPP-PSNR.
[0023] The equirectangular format may be supported via a 360-degree camera and / or stitching procedure. Encoding 360-degree video in cubemap geometry may use a conversion of the equirectangular format to the cubemap format. The equirectangular may also be related to the cubemap. In FIG. 2a, there are six faces (e.g., PX, NX, PY, NY, PZ, and NZ) and three axes (e.g., X, Y, and Z) extending from the center of the sphere (e.g., O) to the center of the faces. "P" may represent positive, and "N" may represent negative. PX may be a direction along the positive x-axis from the center of the sphere, and NX may be the inverse of PX. Similar concepts may be used for PY, NY, PZ, and NZ. The six faces (e.g., PX, NX, PY, NY, PZ, and NZ) may correspond to the front face, back face, top face, bottom face, left face, and right face, respectively. The faces may be indexed from 0 to 5 as follows (e.g., PX (0), NX (1), PY (2), NY (3), PZ (4), and NZ (5)). Ps(X_s, Y_s, Z_s) may be a point on a sphere with radius 1. Ps may also be expressed in terms of yaw (φ) and pitch (θ) as follows:
[0024]
[0025]
[0026] Pf may be a point on the cube where the line extends from the center of the sphere to Ps, or Pf may lie on face NZ. The coordinates of Pf, (X_f, Y_f, Z_f), can be calculated as follows:
[0027]
[0028] Here, |x| may be the absolute value of the variable x. The coordinates of Pf (uc, vc) in the 2D plane of plane NZ can also be calculated as follows:
[0029]
[0030] By using one or more of the equations (3) through (10), there may be a relationship between coordinates (uc, vc) on a cubemap on a specific plane and coordinates (φ, θ) on a sphere. The relationship between a point (φ, θ) on a sphere and an equirectangular point (ue, ve) may be known from equations (1) and / or (2). There may be a relationship between the equirectangular geometry and the cubemap geometry. The geometry mapping from the cubemap to the equirectangular may be expressed. For example, a point (uc, vc) may be given on a plane on the cubemap. An output (ue, ve) on the equirectangular plane may be calculated. For example, the coordinates of a 3D point P_f on a plane may be calculated using (uc, vc) based on equations (9) and (10). The coordinates of a 3D point P_s on a sphere may be calculated using P_f based on equations (6), (7), and (8). (φ, θ) on the sphere may also be calculated using P_s based on mathematical formulas (3), (4), and (5). The coordinates of the point (ue, ve) on the equirectangular picture may also be calculated from (φ, θ) based on mathematical formulas (1) and (2).
[0031] 360-degree video may be represented in a 2D picture. For example, 360-degree video may be presented in a 2D picture using a cubemap. The six faces of the cubemap may be packed into a rectangular area. This may be frame packing. A frame packing picture may be treated as a 2D picture (e.g., coded). Different frame packing configuration(s) may be used (e.g., 3×2 and / or 4×3 packing configurations). In a 3×2 configuration, the six cubemap faces may be packed into two rows, with three faces in one row. In a 4×3 configuration, four faces (e.g., PX, NZ, NX, and PZ) may be packed into one row (e.g., the center row), and faces PY and NY may be packed into two different rows (e.g., the top and bottom rows) (e.g., packed individually). FIG. 2c depicts an example of a 4×3 frame packing corresponding to the equirectangular picture in FIG. 1c.
[0032] A 360-degree video in an equirectangular format may be input or converted to a cubemap format. For each (e.g., each) sample position (uc, vc) in the cubemap format, the corresponding coordinates (ue, ve) in the equirectangular format may be calculated. If the calculated coordinates (ue, ve) in the equirectangular format are not at an integer sample position, an interpolation filter may be used. For example, the interpolation filter may be used to obtain a sample value (fractional position) at a fractional position using samples from neighboring integer positions.
[0033] FIG. 3 illustrates an exemplary workflow for a 360-degree video system. The workflow may include 360-degree video capture using one or more cameras (e.g., covering the entire spherical space). The videos may be stitched together in a geometry structure (e.g., using an equirectangular geometry structure). The equirectangular geometry structure may be converted into another geometry structure (e.g., a cubemap or other projection format) for encoding (e.g., encoding using a video codec). The encoded video may be delivered to a client, for example, via dynamic streaming and / or broadcasting. The video may be decoded. For example, the video may be decoded at a receiver. The decompressed frames may be unpacked to display geometry (e.g., equirectangular). The geometry may be used for rendering (e.g., via viewport projection based on the user's field of view).
[0034] The chroma component may be subsampled to a smaller resolution, for example. For example, the chroma component may be subsampled to a resolution smaller than that of the luma component. Chroma subsampling may reduce the amount of video data used for encoding, save bandwidth and / or computing power, and do so without affecting video quality (e.g., without significantly affecting it). In a 4:2:0 chroma format, both chroma components may be subsampled to 1 / 4 of the luma resolution (e.g., 1 / 2 horizontally and 1 / 2 vertically). After chroma subsampling, the chroma sampling grid may differ from the luma sampling grid. In FIG. 3, throughout the processing flow, the 360-degree video being processed at each stage may be in a chroma format in which the chroma component may be subsampled.
[0035] Video codec(s) may be designed with 2D video captured on a plane in mind. If motion compensation prediction uses one or more samples outside the boundaries of the reference picture, padding may be performed by copying sample values from the picture boundaries. For example, iterative padding may be performed by copying sample values from the picture boundaries. FIGS. 4a and 4b depict examples of extended pictures generated by iterative padding for ERP (e.g., FIG. 4a) and CMP (e.g., FIG. 4b). In FIGS. 4a and 4b, the original picture may be within the dotted box, and the extended boundary may be outside the dotted box. The 360-degree video may contain video information over the entire sphere and may have circular properties. Considering the circular properties of a 360-degree video, the reference picture of the 360-degree video may not have boundaries because the information contained in the picture of the 360-degree video may be wrapped around a sphere. The circular properties may be maintained for one or more different projection formats, or frame packing is used to represent the 360-degree video on a 2D plane. Geometry padding may also be used for 360-degree video coding. For example, geometry padding may be used for 360-degree video coding by padding samples and / or by considering the 3D geometry structures represented in the 360-degree video.
[0036] Geometric padding for the ERP may also be defined on a sphere having longitude and / or latitude. For example, given a point (u, v) to be padded (e.g., outside the ERP picture), the point (u', v') used to derive the padding sample may be calculated as follows:
[0037]
[0038] Here, W and H may be the width and height of the ERP picture. Fig. 5a depicts an exemplary geometry padding process for the ERP. To pad the samples outside the left boundary of the picture, i.e., samples A, B, and C in Fig. 5a, the samples may be padded with corresponding samples located inside the right boundary of the picture, at A', B', and C'. To pad the samples outside the right boundary of the picture (e.g., samples D, E, and F in Fig. 5a), the samples may be padded with corresponding samples located inside the left boundary of the picture, at D', E', and F'. For samples located outside the top boundary, i.e., samples G, H, I, and J in Fig. 5a, the samples are padded with corresponding samples located inside the top boundary of the picture, at G', H', I', and J', with an offset of half the width. For samples located outside the bottom boundary of the picture (e.g., samples at K, L, M, and N in FIG. 5a), the samples may be padded with corresponding samples at K', L', M', and N' located inside the bottom boundary of the picture, with an offset of half the width. FIG. 5b depicts an exemplary extended ERP picture using geometry padding. The geometry padding shown in FIG. 5a and FIG. 5b may provide meaningful samples and / or improve the continuity of neighboring samples for areas outside the ERP picture boundary.
[0039] When the coded picture is in CMP format, the face of the CMP may be extended by using geometry padding by projecting samples from neighboring faces onto the extended area of the current face. Fig. 6a depicts an example of how geometry padding may be performed for a given CMP face in 3D geometry. In Fig. 6a, point P may be on face F1 or outside the boundary of face F1. Point P may be padded. Point O may be the center of the sphere. R may be the left boundary point closest to P or inside face F1. Point Q may be the projection point of point P on face F2 from the center point O. Geometry padding may use sample values from point Q to fill sample values at point P (for example, instead of using sample values from point R to fill sample values at point P). Fig. 6b depicts an exemplary extended face with geometry padding for a CMP 3×2 picture. The geometry padding shown in Figs. 6a and 6b may also provide meaningful samples for the area outside the CMP face boundary.
[0040] FIG. 7 illustrates an exemplary block-based hybrid video encoding system (600). An input video signal (602) may be processed block by block. To compress a high-resolution (e.g., 1080p and / or higher) video signal, an extended block size (e.g., a coding unit or CU) may be used (e.g., in HEVC). A CU may have up to 64×64 pixels (e.g., in HEVC). A CU may also be divided into prediction units or PUs, for which a distinct prediction method may be applied. Spatial prediction (660) and / or temporal prediction (662) may be performed on an input video block (e.g., a macroblock; MB or CU). Spatial prediction (e.g., "intra-prediction") may use pixels from already coded neighboring blocks in the same video picture and / or slice to predict the current video block. Spatial prediction may reduce inherent spatial redundancy in the video signal. Temporal prediction (e.g., "inter prediction" or "motion compensation prediction") may use pixels from already coded video pictures to predict the current video block. Temporal prediction may reduce inherent temporal redundancy in the video signal. A temporal prediction signal for a given video block may be signaled by one or more motion vectors indicating the amount and / or direction of motion between the current block and its reference block. If multiple reference pictures are supported (e.g., in H.264 / AVC or HEVC), the reference picture index of the video block may be signaled to the decoder. The reference picture index may be used to identify which reference picture in the reference picture store (664) the temporal prediction signal may originate from.
[0041] After spatial and / or temporal prediction, the mode determination (680) within the encoder may select a prediction mode based, for example, on rate-distortion optimization. The prediction block may be subtracted from the current video block at 616. The prediction residual may be decorrelated using a transform module (604) and a quantization module (606) to achieve a target bit rate. The quantized residual coefficients may be inversely quantized at 610 and inversely transformed at 612 to form a reconstructed residual. The reconstructed residual may be added back to the prediction block at 626 to form a reconstructed video block. Before the reconstructed video block is placed in the reference picture storage (664), at 666, an in-loop filter, such as a deblocking filter and / or an adaptive loop filter, may be applied to the reconstructed video block. A reference picture in the reference picture storage (664) may be used to code future video blocks. An output video bitstream (620) may be formed. Coding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients may be transmitted to an entropy coding unit (608), compressed and packed to form the bitstream (620).
[0042] FIG. 8 illustrates an exemplary block-based hybrid video decoder. The decoder in FIG. 8 may correspond to the encoder in FIG. 7. The video bitstream (202) may be received, unpacked, and / or entropy decoded at the entropy decoding unit (208). The coding mode and / or prediction information may be transmitted to the spatial prediction unit (260) (e.g., in the case of intra-coding) and / or to the temporal prediction unit (262) (e.g., in the case of inter-coding). A prediction block may be formed at the spatial prediction unit (260) and / or the temporal prediction unit (262). Residual transformation coefficients may be transmitted to the inverse quantization unit (210) and the inverse transformation unit (212) to reconstruct the residual block. Then, the prediction block and the residual block may be added at 226. The reconstructed block may pass through filtering (266) within the loop or be stored in the reference picture storage (264). The reconstructed video in the reference picture storage (264) may be used to drive a display device and / or to predict future video blocks.
[0043] Video codec(s) such as H.264 and / or HEVC may be used to encode 2D planar linear video(s). Video coding may utilize spatial and / or temporal correlation(s), for example, to eliminate information redundancy. Various prediction techniques, such as intra-prediction and / or inter-prediction, may be applied during video coding. Intra-prediction may predict sample values using neighboring reconstructed samples. FIG. 9 depicts an exemplary reference sample that may be used to intra-prediction a current transform unit (TU). The current TU described herein may be a current block, and the two terms may be used interchangeably. As described herein, the reference sample may include a reconstructed sample located above and / or to the left of the current TU.
[0044] One or more intra prediction modes may be selected. For example, HEVC may specify 35 intra prediction modes, including planar (0), DC (1), and / or angular prediction (2 to 34), as illustrated in FIG. 10. Planar prediction may generate a first-order approximation for the current block, for example, using top and / or left reconstructed samples. Top right and bottom left sample values may be copied along the right column and bottom row, respectively (for example, due to raster scan order). A vertical predictor may be formed for one or more locations within the block, for example, using the weighted average of the corresponding top and bottom samples. A horizontal predictor may be formed using the corresponding left and right samples. A final predictor may be formed, for example, by averaging the vertical and horizontal predictors. The bottom right sample value may be extrapolated by the average of the top right and bottom left sample values. The right column (e.g., bottom row) may also be extrapolated using the top right and bottom right samples (e.g., bottom left and bottom right samples).
[0045] Directional prediction may be designed to predict a directional texture. For example, in HEVC, the intra-directional prediction process may be performed by extrapolating sample values from a reconstructed reference sample utilizing a given direction. One or more (e.g., all) sample locations within the current block may be projected onto a reference row or column (e.g., depending on the angular mode). If a projected pixel location has a negative index, the reference row may be extended to the left by projecting the left reference column for vertical prediction, while the reference column may be extended upward by projecting the top reference row for horizontal prediction. Figures 11a and 11b depict exemplary projections for the left reference sample (e.g., Figure 11a) and the top reference sample (e.g., Figure 11b). In Figures 11a and 11b, the thick arrows may indicate the prediction direction, and the thin arrows may indicate the reference sample projection. Figure 11a depicts an exemplary process for extending the top reference row using samples from the left reference column. Predicted samples may also be filtered at block boundaries to reduce blocking artifacts (e.g., intra-predicted blocks are generated). For vertical intra mode (e.g., in HEVC), the leftmost column s of the predicted samples i,j is the left-based column R i,j It can also be adjusted as follows using:
[0046]
[0047] For horizontal intra modes, the top-most row of predicted samples may be adjusted using a similar process. Figures 12a through 12d depict examples of boundary prediction filtering for different directional modes. For example, Figure 12a depicts exemplary boundary prediction filtering for intra mode 2. For example, Figure 12b depicts exemplary boundary prediction filtering for intra mode 34. For example, Figure 12c depicts exemplary boundary prediction filtering for intra modes 3-6. For example, Figure 12d depicts exemplary boundary prediction filtering for intra modes 30-33. An appropriate intra prediction mode may be selected at the encoder side by minimizing the distortion between the predictions generated by one or more intra prediction modes and the original samples. The most probable mode (MPM) may be used for intra coding (for example, to efficiently encode the intra prediction mode). The MPM may also reuse the intra-directional mode of a spatial neighboring PU. For example, the MPM may reuse the intra-directional mode of a spatial neighboring PU so that it may not code the intra-directional mode for the current PU.
[0048] Figure 13 depicts exemplary spatial neighbors (e.g., bottom left (BL), left (L), above right (AR), top (A), above left (AL)) used for MPM candidate derivation. Selected MPM candidate indices may be encoded. The MPM candidate list may be configured on the decoder side in the same manner as on the encoder side. An entry with a signaled MPM candidate index may be used as the intra-directional mode of the current PU. For example, RGB to YUV color conversion may be performed to reduce correlation between different channels. A correlation may exist between the lumina channel and the chroma channel. The component cross-linear model prediction utilizes this correlation to predict the chroma channel from the lumina channel using a liner model, as follows (e.g., assuming an N×N sample chroma block and following the same notation as in Figure 9), the downsampled reconstructed lumina sample value From chroma sample value p i,j It is also possible to predict:
[0049]
[0050] Downsampled luma samples can also be calculated as follows:
[0051]
[0052] The parameters of the linear model may be derived by minimizing the regression error between neighboring reconstructed samples at the top and left. The parameters of the linear model may also be calculated as follows:
[0053]
[0054] FIG. 14 depicts exemplary positions of neighboring reconstructed samples at the top and left used for the derivation of α and β. FIG. 15 depicts an exemplary inter-prediction with a single motion vector (MV). Blocks B0' and B1' in the reference picture may be reference blocks of blocks B0 and B1, respectively. Reference block B0' may be partially outside the picture boundary. A padding process (e.g., in HEVC / H.264) may be configured to fill unknown samples outside the picture boundary. FIG. 16 depicts exemplary padding for a reference sample (e.g., block B0') outside the picture boundary in HEVC / H.264. Block B0' may have four parts, e.g., P0, P1, P2, and P3. Parts P0, P1, and P2 may be outside the picture boundary and may be filled, e.g., through a padding process. Part P0 may be filled with the top-left sample of the picture. Part P1 may be filled with vertical padding using the top row of the picture. Part P2 may be filled with horizontal padding using the leftmost column of the picture. Motion vector prediction and merging modes may be used for intercoding to encode motion vector information. Motion vector prediction may use motion vectors from neighboring PUs or temporally collocated PUs as predictors for the current MV. The encoder and / or decoder may form a list of motion vector predictor candidates, for example, in the same way. The index of a selected MV predictor from the candidate list may be coded and signaled to the decoder. The decoder may construct a list of MV predictors, and an entry with the signaled index may be used as a predictor for the current PU's MV.The merge mode may reuse MV information from spatial and temporal neighbor PUs, and as a result, the merge mode may not code motion vectors for the current PU. The encoder and / or decoder may, for example, form a list of motion vector merge candidates. FIG. 17 depicts exemplary spatial neighbors (e.g., bottom left, left, top right, top, top left) used for deriving merge candidates. The selected merge candidate indices may be coded. The list of merge candidates may be constructed on the decoder side, for example, in the same way as the encoder. An entry with a signaled merge candidate index may be used as the MV of the current PU.
[0055] Frame-packed 360-degree video may have characteristics different from 2D video. 360-degree video may contain 360-degree information about the environment surrounding the viewer. This 360-degree information may indicate that the 360-degree video possesses intrinsic circular symmetry. 2D video does not possess this symmetry characteristic. Video codec(s) may be designed for 2D video and may not take into account the symmetry features of 360-degree video. For example, codec(s) may process the video in a coding order (e.g., coding). For example, codec(s) may process the video signal in blocks using a coding order such as raster scan order, which codes blocks from top to bottom and / or left to right. Information for coding the current block may be inferred from the blocks located above and / or to the left of the current block.
[0056] In the case of 360-degree video, neighboring blocks in a frame-packed picture may not, for example, be related to coding the current block. Blocks that are neighbors of the current block in a frame-packed picture may be frame-packed neighbors. Blocks that are neighbors of the current block in 3D geometry may be face neighbors or spherical neighboring blocks. Figures 18a and 18b depict examples of CMP. Figure 18a depicts an exemplary 3D representation of CMP. Figure 18b depicts an exemplary 3×2 frame-packing configuration of CMP. In Figures 18a and 18b, Block A may be a frame-packed neighbor located above Block C. Considering 3D geometry, Block D may be an exact face neighbor (e.g., or spherical neighboring block) located above Block A. Additional face neighbor blocks may also be used in 360-degree video coding. For example, if the current block is located at the right and / or bottom face boundary associated with the face of the frame-packed picture, the right and / or bottom face neighbor block may be used. The right and / or bottom face neighbor block may be located on a face on the other side of the boundary (e.g., on the opposite side or across the face). For example, the right and / or bottom face neighbor block may share a boundary located at the right and / or bottom face boundary of the face to which the current block belongs. The face arrangement and / or scan processing order (e.g., raster scan order) may be used to determine which block may be used to code the current block as described herein. In the example depicted in FIG. 18b, block B may be a right face neighbor in relation to block A. If block B is a right face neighbor in relation to block A, the right face neighbor may be matched with a right frame-packed neighbor.When blocks are scanned from left to right (e.g., a raster scan with a scan order moving from left to right), when coding block A, block B may not be coded, and right-side neighbors may not be available (e.g., may not yet be coded). When encoding block E, its right-side neighbors (e.g., block F using the inherent spherical characteristics of 360-degree video) may be coded and may be used to code (e.g., predict) block E. When encoding block G, its bottom-side neighbors (e.g., block H using the inherent spherical characteristics of 360-degree video) may be coded and may be used to code (e.g., predict) block G. In FIG. 18a and FIG. 18b, a block located at the face boundary of one of the hatched areas (e.g., FIG. 18b) may use its right and / or bottom face neighbor block(s) in the current block's coding (e.g., prediction) process, for example, because its right and / or bottom neighbors have already been coded (e.g., because they are available), taking into account the spherical properties of the 360-degree video. The right and / or bottom face neighbor blocks may also be used as reference blocks.
[0057] For example, left (L), above (A), above right (AR), above left (AL), and / or below left (BL) neighbors may be used to infer information in 2D video coding (e.g., as illustrated in FIGS. 13 and 17) (e.g., due to scan order, such as in raster scan processing). In 360-degree video, if the current block is at the right face boundary, the right (R) and below right (BR) face neighbor blocks may be used to infer attributes to derive motion vector candidates in motion vector prediction and / or merge modes (e.g., to derive the MPM list in intra prediction). If the current block is at the bottom face boundary, the below (B) and below right (BR) face neighbor blocks may be used to infer attribute(s). One or more additional spatial candidates may be used to infer attribute(s) from the neighbor blocks.
[0058] For example, a reconstructed sample located above and / or to the upper left of the current block may be used in 2D video coding to predict the current block (e.g., as illustrated in FIG. 9) (e.g., due to raster scan processing). In 360-degree video, if a neighboring reconstructed sample is located outside the face to which the current block belongs, the sample may be extrapolated, for example, using geometry padding. For example, if the block is on the right-side boundary, sample R N+1,0 , ..., R 2N,0 It may also be obtained using geometry padding. A reconstructed sample located on the upper right side of the block (e.g., R shown in FIG. 19) N+1,0 , ..., R N+1,2N) may also be used. FIG. 19 depicts an exemplary reconstructed sample used to predict the current block in intra and inter coding. If the block is at the bottom face boundary, the sample located on the bottom side of the block (e.g., R 0,N+1 , ..., R 2N,N+1 ) may also be used. As described herein, reconstructed samples (e.g., additional and / or more meaningful reconstructed samples) may also be used in different prediction methods (e.g., intra-prediction, component cross-linear model prediction, boundary prediction filtering, and / or in-loop filtering, DC, planar, and / or directional modes).
[0059] 360-degree video coding may use spatial neighbors and reconstructed samples for intra-prediction and / or inter-prediction. The blocks described herein may include the current block or sub-blocks and may be used interchangeably. If a block is located on the right (e.g., or bottom) face boundary, its right and right-bottom (e.g., or bottom and right-bottom) face neighbor blocks may be considered as candidate(s) for different procedure(s) for inferring attributes from spatial neighbors (e.g., for deriving MPM in intra-directional processes, for deriving merge modes in inter-prediction, for motion vector prediction, and / or etc.). Blocks outside the current face may be obtained from neighboring face(s). The location of spatial candidates at the face boundary may be described herein.
[0060] For intra and / or inter prediction, if a block (e.g., the current block) is on a right (e.g., or bottom) face boundary, reconstructed samples from its right and right-bottom (e.g., bottom or right-bottom) face neighbor blocks may be used to code (e.g., predict) the current block. The reconstructed samples may be outside the current face and may be obtained, for example, using geometry padding. The location of the reconstructed samples at the face boundary may be identified as described herein.
[0061] In the case of intra-prediction, if a block lies on a right (e.g., or bottom) face boundary, reconstructed samples from its right and right-bottom (e.g., or bottom and right-bottom) face neighbor blocks may be used to derive a reference sample. If a block lies on a right (e.g., or bottom) face boundary, one or more additional horizontal (e.g., or vertical) directional prediction modes may be defined. The reference sample derivation process and one or more additional directional modes at the face boundary may also be described herein.
[0062] In the case of intra-directional prediction, if the block lies on the right (e.g., or bottom) face boundary, boundary filtering may be applied at the block's right (e.g., or bottom) boundary. For example, boundary filtering may be applied at the block's right (e.g., bottom) boundary to reduce discontinuities that may appear at the intersection between the interpolated sample and the reconstructed sample. Boundary prediction filtering at face boundaries may also be described herein.
[0063] In 2D video coding, the top, right, bottom, and / or left picture boundaries may not be filtered, for example, during the in-loop filtering process. In the case of deblocking, sample(s) outside the boundaries (e.g., top, right, bottom, and / or left) may not exist. In the case of 360-degree video coding, the top, right, bottom, and / or left boundaries of a face may be connected to other face boundaries. For example, in the case of 360-degree video coding, due to the inherent circular nature of 360-degree video, the top, right, bottom, and / or left boundaries of a face may be connected to other face boundaries. In-loop filtering may be applied across one or more (e.g., all) face boundaries. The in-loop filtering process at face boundaries may be described herein.
[0064] For component cross-linear model prediction, if a block is located on a right (e.g., or bottom) face boundary, reconstructed samples from its right (e.g., or bottom) face neighbor blocks may be used to estimate the parameters of the linear model. The reconstructed samples may be located outside the current face, for example, and may be obtained using geometry padding. The location of the reconstructed samples on the face boundary, the downsampling of the reconstructed luminance samples, and / or the derivation of the linear model parameters may be described herein.
[0065] One or more faces may be processed (e.g., sequentially) by using a scan order (e.g., raster scan order) within the face. In a face scan order, the availability of face neighbor blocks may be increased. A face scan order is described herein.
[0066] For CMP and / or related cube-based geometry, one or more faces may be packed using the configurations described herein. For example, one or more faces may be packed, for example, with a 3×2 packing configuration. The 3×2 packing configuration described herein may maximize the availability of face neighbor blocks.
[0067] A coding device (e.g., a device that may be an encoder and / or decoder, or may include these) may use one or more additional neighbor blocks based, for example, on the position of the current block within a geometry plane. For example, the coding device may increase the number of candidates for inferring information from neighbor block(s) by using one or more additional neighbor blocks based on the position of the current block within a geometry plane. The coding device may infer information from neighbor block(s) using MPM in intra-prediction, motion estimation in inter-prediction, and / or a merge mode in inter-prediction.
[0068] In the example, a coding device (e.g., a device that may be an encoder and / or decoder, or may include these) may receive a coded frame-packed picture in coding order. The current block may be located at the exit boundary of a face in the frame-packed picture to which the current block belongs. For example, the coding device may determine that the current block is located at the exit boundary of the face to which the current block belongs according to the coding order of the frame-packed picture. The exit boundary of the face to which the current block belongs may be located in the same direction as the coding order with respect to the current block.
[0069] In the example, the coding block may have a left-to-right orientation relative to the current block. If the coding block has a left-to-right orientation relative to the current block, the exit boundary of the face in the frame-packed picture to which the current block belongs may be located at the right face boundary (e.g., the rightmost face boundary) to which the current block belongs (e.g., which may be in the same orientation as the coding order). In the example, the coding block may have a top-to-bottom orientation relative to the current block. If the coding block has a top-to-bottom orientation relative to the current block, the exit boundary of the face in the frame-packed picture to which the current block belongs may be located at the bottom face boundary (e.g., the bottommost face boundary) to which the current block belongs (e.g., which may be in the same orientation as the coding order). In the example, the coding block may have a left-to-right and top-to-bottom orientation relative to the current block. If the coding block has a left-to-right and top-to-bottom orientation relative to the current block, the exit boundary of the face in the frame-packed picture to which the current block belongs may be located at the right and bottom face boundaries (e.g., the far right and bottommost face boundaries) to which the current block belongs (e.g., which may be in the same orientation as the coding order).
[0070] If the coding device determines that the current block is located at the exit boundary of the face to which the current block belongs (e.g., the rightmost and / or bottommost face boundary to which the current block belongs), the coding device may identify one or more (e.g., multiple) spherical neighbor blocks of the current block. For example, the coding device may identify the spherical neighbor blocks of the current block based on the spherical characteristics of the 360-degree video.
[0071] The coding device may identify a face boundary cross-neighbor block located on a face (e.g., another face) among the identified spherical neighbor blocks. The face to which the face boundary cross-neighbor block belongs may share the boundary of the face to which the current block belongs (e.g., right and / or bottom face boundary). For example, the face boundary cross-neighbor block may be located outside the current block, across the face to which the current block belongs, and / or on the opposite side of the face boundary from the current block. The face boundary cross-neighbor block may be located in the same direction of the coding order relative to the current block. For example, the face boundary cross-neighbor block may be the right (R) block, bottom (B) block, and / or right bottom (BR) block of the current block.
[0072] In the example, if the current block is located at the right boundary of the face to which the current block belongs (e.g., the far right boundary), the coding device may determine whether the identified block corresponding to the cross-face neighboring block (e.g., the right (R) block and / or the right bottom (BR) block) can be used as a candidate(s) (e.g., additional candidate(s)) as depicted in FIG. 20a.
[0073] FIGS. 20a through 20c depict exemplary spatial neighbors at the right-side boundary (e.g., FIG. 20a), bottom-side boundary (e.g., FIG. 20b), and bottom-right-side boundary (e.g., FIG. 20c) of the face to which the current block belongs. The block(s) depicted using the hatched pattern in FIGS. 20a through 20c may be located outside the current face. If the current block is at the bottom-side boundary, the bottom (B) and / or right-bottom (BR) (e.g., already coded neighboring blocks) may be used as candidates (e.g., additional candidates) as depicted on FIG. 20b, for example, to predict the current block. The current block located at the bottom-side boundary may follow an approach similar to that described herein for the right-side boundary. If the current block is located at the bottom right face boundary, the right, bottom, and / or bottom right (e.g., already coded neighboring blocks) may be used as candidates (e.g., additional candidates) to predict the current block, for example, as depicted in FIG. 20c. The current block located at the bottom right face boundary may follow an approach similar to that described herein for the right face boundary. If the neighboring block location is outside the current face, the corresponding block may be obtained from the corresponding neighboring face (e.g., by mapping sample locations to derive spatial neighboring blocks as described herein).
[0074] When identifying a face boundary cross-neighbor block, the coding device may identify a block within a frame-packed picture corresponding to the face cross-neighbor block. For example, the coding device may identify a block within a frame-packed picture corresponding to the face cross-neighbor block based on frame-packing information of a 360-degree video.
[0075] The coding device may determine whether to use an identified block within a frame-packed picture corresponding to a face-crossing neighbor block to code the current block. For example, the coding device may determine whether to use an identified block corresponding to a face-crossing neighbor block to code the current block based on, for example, the availability of an identified block within a frame-packed picture. If the identified block is coded, the identified block within the frame-packed picture may be considered available. The coding device may also determine whether an identified block corresponding to a face-crossing neighbor block (e.g., right and / or right-bottom block(s)) is coded and / or available for coding the current block. If the coding device determines that an identified block corresponding to a face-crossing neighbor block (e.g., right and / or right-bottom block(s)) is available (e.g., already coded), the coding device may use the identified block corresponding to the face-crossing neighbor block.
[0076] The coding device may determine not to use the identified block corresponding to the face-crossing neighbor block. In an example, the coding device may determine that the identified block corresponding to the face-crossing neighbor block is not coded and / or is not available to code the current block. In an example, the coding device may determine that the current block is located within the face to which the current block belongs. In an example, the coding device may determine that the current block is located at the entering boundary of the face to which the current block belongs. The entering boundary of the face to which the current block belongs may be located according to the coding order for the frame-packed picture. If the coding device determines not to use the identified block corresponding to the face-crossing neighbor block (e.g., if the face-crossing neighbor block is not available and / or not coded, if the current block is located within the face, or if the current block is located at the entering boundary of the face to which the current block belongs), the coding device may use one or more spherical neighbor block(s) that are coded to code the current block. The spherical neighboring block(s) described herein may include a frame-packed or neighboring block of 3D geometry (e.g., already coded). For example, a coding device may use at least one of a left (L), top (A), and / or top-left block as one or more spherical neighboring block(s) to code (e.g., predict) the current block. In the example, the coding device may use geometry padding to code (e.g., predict) the current block.
[0077] The coding device may use one or more additional blocks (e.g., associated with an identified block corresponding to a face-crossing neighbor block described herein) to increase the number of available samples associated with one or more additional blocks for predicting the current block (e.g., intra-prediction). For example, the coding device may use one or more additional reconstructed samples based, for example, on the location of the current block within a geometry face. For example, the coding device may use one or more reconstructed samples (e.g., additional reconstructed samples) associated with an identified block corresponding to a face-crossing neighbor block described herein and / or illustrated in FIG. 19.
[0078] In the example, the coding device may not use one or more (e.g., all) reconstructed samples (e.g., additional reconstructed samples as described herein). For example, if the coding device determines that the current block is within a face and / or if the current block is located at a face boundary and the current block is located in the opposite direction of the coding order relative to the current block (e.g., at the top face boundary and / or at the left face boundary), the coding device may use one or more reconstructed samples as depicted in FIG. 9. For example, the coding device may use one or more available (e.g., coded) reconstructed samples. One or more available reconstructed samples may be located to the left and / or above neighboring block(s) of the current block (e.g., spatial neighboring block(s)). One or more samples located to the right and / or below neighboring blocks of the current block may not be available (e.g., may not be coded and / or reconstructed). If one or more reconstructed samples are outside the current block (e.g., the current block associated with the current face), the coding device may acquire one or more reconstructed samples using geometry padding, for example.
[0079] As described herein, the coding device may determine to use one or more samples from identified blocks corresponding to face-crossing neighbor blocks for coding the current block. In an example, if the coding device determines that the current block is located on the right-side boundary of the face associated with the frame-packed picture, the coding device may (e.g., R, each) one or more reconstructed samples associated with the block(s) located on the left and / or top of the current block, for example, as depicted in FIG. 21a. 0,0 , ..., R 0,2N and / or R 0,0, ..., R N,0 In addition to ) one or more reconstructed samples located to the upper right of the current block (e.g., R N+1,0 , ..., R N+1,2N ( ) may also be used. FIGS. 21a through 21c depict examples of one or more (e.g., additional) available reconstructed samples at the right face boundary (e.g., FIG. 21a), the bottom face boundary (e.g., FIG. 21b), and the bottom right face boundary (e.g., FIG. 21c). The reconstructed sample(s) depicted using the hatched pattern in FIGS. 21a through 21c may be located outside the current block (e.g., or the current face). The coding device may apply preprocessing to one or more of the reconstructed samples. For example, preprocessing may include, but is not limited to, filtering, interpolation, and / or resampling. If the coding device applies preprocessing (e.g., filtering, interpolation, and / or resampling) to one or more reconstructed samples located on the left side of the current block, the coding device may apply similar (e.g., the same) preprocessing to one or more reconstructed samples located on the right side of the current block.
[0080] In the example, if the coding device determines that the current block is located at the bottom face boundary of the face associated with the frame-packed picture, the coding block is (as depicted in FIG. 21b, for example, one or more reconstructed samples associated with the block(s) located to the left and / or above the current block (e.g., each, R 0,0 , ..., R 0,N and / or R 0,0 , ..., R 2N,0 In addition to ) one or more reconstructed samples located below the current block (e.g., R 0,N+1 , ..., R 2N,N+1) may also be used. If the coding device applies preprocessing (e.g., filtering, interpolation, and / or resampling) to the reconstructed sample(s) located above the current block, the coding device may also apply similar (e.g., the same) preprocessing to the reconstructed sample(s) located below the current block.
[0081] In the example, if the coding device determines that the current block is located at the bottom right side boundary of the face associated with the frame-packed picture, the coding block is (as depicted in FIG. 21c, for example, one or more reconstructed samples associated with the block(s) located to the left and / or above the current block (e.g., each, R 0,0 , ..., R 0,N and / or R 0,0 , ..., R N,0 In addition to ) one or more reconstructed samples located to the right and below the current block (e.g., respectively, R N+1,0 , ..., R N+1,N+1 and / or R 0,N+1 , ..., R N+1,N+1 ) may also be used. If the coding device applies preprocessing (e.g., filtering, interpolation, and / or resampling) to the reconstructed sample(s) located to the left and / or above the current block, the coding device may also apply similar (e.g., the same) preprocessing to the reconstructed sample(s) located to the right and / or below the current block.
[0082] The coding device may use one or more reference sample lines in one or more (e.g., all) cases described herein. The coding device may also apply one or more cases described herein to rectangular block(s).
[0083] If the current block associated with the face is located on the right face boundary of the face, the sample(s) located on the face boundary neighbor blocks (e.g., located to the right of the current block) (e.g., R N+1,0 , ..., R N+1,2N ) may also be used as depicted in FIGS. 21a to 21c. Face boundary intersecting neighbor blocks (e.g., R N+1,0 , ..., R N+1,2N The predicted reference sample(s) derived from ) are the current block (e.g., R N+1,0 , ..., R 2N,0 It may be located closer to the sample situated on the upper right (AR) side of ). Reference sample(s) (e.g., R N+1,0 , ..., R N+1,2N ) may also be filtered. For example, one or more reference samples (e.g., R N+1,0 , ..., R N+1,2N ) may be filtered before performing a prediction (similar to the intra prediction process in HEVC, for example).
[0084] (For example, as described herein) face boundary cross-neighbor blocks may be used to predict the current block. Reference rows or reference columns may be used (for example, depending on the directionality of the selected prediction mode). The upper reference row may be extended to the right, for example, by projecting the right reference column, as depicted in FIG. 22a. FIG. 22a and FIG. 22b depict examples of intra prediction at the right face boundary (e.g., FIG. 22a) and the bottom face boundary (e.g., FIG. 22b). The thick arrows shown in FIG. 22a and FIG. 22b may indicate the prediction direction, and the thin arrows may indicate the reference sample projection. The reference samples depicted using dashed lines in FIG. 22a and FIG. 22b may be located outside the current face. Considering the intra-directional prediction direction defined in FIG. 10, sample R N+1,N , ..., R N+1,2N may not be used to extend the top reference row to the right. Sample R N+1,N , ..., R N+1,2N It can also be used to filter the right-hand reference column.
[0085] Sample located at the bottom right of the current block (e.g., R N+1,N+1 , ..., R N+1,2N ) when considering the intra-directional prediction direction, for example, as depicted in FIG. 10 It may also be used to cover a range wider than the range. The right reference column may be used as depicted in FIG. 23a, or it may be extended upward by projecting the upper reference row as depicted in FIG. 23b. FIGS. 23a through 23d depict examples of additional intra-prediction modes at the right face boundary (e.g., FIGS. 23a-b) and the bottom face boundary (e.g., FIGS. 23c through 23d). The thick arrows shown in FIGS. 23a through 23d may indicate the prediction direction, and the thin arrows may indicate the projection of the reference sample. The reference sample, depicted using dashed lines in FIGS. 23a through 23d, may be located outside the current face.
[0086] For the horizontal angular direction, blending between the left and right reference columns may be performed to predict the current block sample, as depicted in FIGS. 24a and 24b. For example, linear weighting or a similar process may be performed by considering the distance between the sample to be predicted and the reference sample. In the example, the projected pixel location may have a negative index. If the projected pixel location has a negative index, the left reference column may be extended upward by projecting the upper reference row, for example, as depicted in FIG. 24a, and / or the right reference column may be extended upward by projecting the upper reference row, for example, as depicted in FIG. 24b. FIGS. 24a through 24d depict exemplary bidirectional intra predictions at the right face boundary (e.g., FIGS. 24a and 24b) and the bottom face boundary (e.g., FIGS. 24c and 24d). In FIGS. 24a through 24d, the thick arrows may indicate the prediction direction, and the thin arrows may indicate the reference sample projection. The reference sample, depicted using dashed lines in FIGS. 24a through 24d, may be located outside the current face.
[0087] In the case of DC mode, if the current block is at the right-side boundary, the sample located to the right of the current block may be used to calculate the DC predictor:
[0088]
[0089] In planar mode, when the current block is at the right face boundary, for example, sample R obtained using geometry padding N+1,1 , ..., R N+1,N It can also be used for horizontal predictors. For vertical predictors, sample R i,N+1The values of (i = 1, ..., N) are, for example, R using linear weighting or a similar process that considers the distance to these two samples. 0,N+1 and R N+1,N+1 It can also be interpolated from. R N+1,N+1 The value of may also be obtained from a corresponding available reconstructed sample.
[0090] If the current block is at the bottom face boundary, the sample located below the current block (e.g., R 0,N+1 , ..., R 2N,N+1 ) may also be used as depicted in FIG. 21b. In this way, a reference sample derived from a block associated with a face boundary intersecting neighbor block as described herein (e.g., derived from a reconstructed sample) may be closer to the current block sample to be predicted. Reference sample R 0,N+1 , ..., R 2N,N+1 It may be filtered, for example, before performing a prediction. For example, reference sample R 0,N+1 , ..., R 2N,N+1 It may be filtered before performing a prediction similar to the intra prediction process (e.g., in HEVC).
[0091] For example, the left reference column may be extended downward (e.g., by projecting the bottom reference row, as depicted in Fig. 22b). Considering the intra-directional prediction direction in Fig. 10, sample R N,N+1 , ..., R 2N,N+1 may not be used to extend the left reference column downward. Sample R N,N+1 , ..., R 2N,N+1 It can also be used to filter the bottom reference row.
[0092] Sample located at the bottom right of the current block (e.g., R N+1,N+1 , ..., R 2N,N+1) when considering the intra-directional prediction direction, for example, as depicted in FIG. 10 It may also be used to cover a range wider than the range. In this case, the bottom reference row may be used, as depicted in FIG. 23c. The bottom reference row may be extended to the left by projecting the left reference column, for example, as depicted in FIG. 23d.
[0093] For the vertical angular direction, as depicted in FIGS. 24c and 24d, a blending of the upper and lower reference rows may be performed to predict the current block sample. In this case, linear weighting or a similar process may be performed, for example, by considering the distance between the sample to be predicted and the reference sample. In some cases, the projected pixel location may have a negative index. If the projected pixel location has a negative index, the upper reference row may be extended to the left by projecting the left reference column, for example, as depicted in FIG. 24c. The lower reference row may be extended to the left by projecting the left reference column, for example, as depicted in FIG. 24d.
[0094] In the case of DC mode, if the current block is at the bottom face boundary, samples located below the current block may also be used to calculate the DC predictor:
[0095]
[0096] In planar mode, when the current block is at the bottom face boundary, sample R obtained using geometry padding 1,N+1 , ..., R N,N+1 It can also be used for vertical predictors. For horizontal predictors, sample R N+1,jThe values of (j = 1, ..., N) are, for example, R using linear weighting or a similar process that considers the distance to these two samples. N+1,0 and R N+1,N+1 It can also be interpolated from. R N+1,N+1 The value of may also be obtained from a corresponding available reconstructed sample.
[0097] If the current block is at the bottom right boundary, the sample located to the upper right of the current block (e.g., R N+1,0 , ..., R N+1,N+1 ) may also be used. Samples located below the current block (e.g., R 0,N+1 , ..., R N+1,N+1 ) may also be used as depicted in FIG. 21c. As described herein, the reference sample derived from the face boundary cross-neighbor block may be closer to the current block sample to be predicted.
[0098] In the case of DC mode, if the current block is at the bottom right boundary, samples located to the upper right and / or lower right of the current block may also be used to calculate the DC predictor:
[0099]
[0100] In planar mode, when the current block is at the bottom-right face boundary, the sample R obtained using geometry padding N+1,1 , ..., R N+1,N It can also be used for horizontal predictors, and sample R obtained using geometry padding 1,N+1 , ..., R N,N+1 It can also be used for vertical predictors.
[0101] One or more reference sample lines may be used in one or more (e.g., all) cases described herein, and rectangular blocks may be configured to use reconstructed samples as described herein.
[0102] If the current block is located at the right, bottom, and / or bottom-right side boundary of a face in a frame-packed picture associated with, for example, a 360-degree video, additional boundary prediction filtering(s) may be applied (e.g., after intra prediction). For example, additional boundary prediction filtering(s) after intra prediction may be applied to reduce discontinuities at the face boundary. The filtering described herein may also be applied to the top of the boundary prediction filtering. In an example, the filtering described herein may be applied to the top row(s) of the block (e.g., top row(s)) and / or left column(s) (e.g., leftmost column(s)).
[0103] For example, in the case of horizontal intra-mode(s) that may be nearly horizontal, if the current block is at the right-face boundary, the predicted sample(s) of the block located in the right column (e.g., the far right column) i,j is, for example, as follows, right reference column R i,j It can also be adjusted using:
[0104]
[0105] FIGS. 25a through 25h depict exemplary boundary prediction filtering at the right-side boundary for intra-mode 2 (e.g., FIG. 25a), intra-modes 3-6 (e.g., FIG. 25b), intra-modes 7-9 (e.g., FIG. 25c), intra-mode 10 (e.g., FIG. 25d), intra-modes 11-17 (e.g., FIG. 25e), intra-mode 18 (e.g., FIG. 25f), intra-modes 19-21 (e.g., FIG. 25g), and intra-mode 22 (e.g., FIG. 25h). The reference sample, depicted using dashed lines in FIGS. 25a through 25h, may be located outside the current face.
[0106] For other intra-directional mode(s), if the current block is at the right-face boundary, the predicted sample(s) of the block located in the right column(s) (e.g., the rightmost column(s)). i,j is, for example, as follows, right reference column R i,j It can also be filtered using:
[0107] In the case of Mode 2 (e.g., FIG. 25a).
[0108]
[0109] For modes 3-6 (e.g., FIG. 25b) and / or modes 7-9 (e.g., FIG. 25c)
[0110]
[0111] In the case of mode 10 (e.g., FIG. 25d).
[0112]
[0113] For modes 11-17 (e.g., FIG. 25e)
[0114]
[0115] In the case of mode 18 (e.g., FIG. 25f).
[0116]
[0117] For modes 19-21 (e.g., Fig. 25g)
[0118]
[0119] In the case of mode 22 (e.g., Fig. 25h).
[0120]
[0121] Here, D may be a parameter controlling the number of the rightmost column and may be filtered. Weights 'a', 'b', and 'c' may be selected, for example, based on the intra mode and / or distance to the reference sample. For example, a lookup table (LUT) may be used to obtain the parameter(s) as a function of the intra mode. In the LUT, depending on the intra mode, higher weights may be given to reference sample(s) that may be closer to the current position. For example, for modes 2 and 18, the weights defined in Table 1 may be used to filter the rightmost column of the block (e.g., Equations (23) and (27) respectively).
[0122]
[0123] The weights defined in Table 1 may be used for horizontal modes (e.g., mode 10). For modes 19 through 22, the weights defined in Table 2 may be used.
[0124]
[0125] For other mode(s), the weights may be determined such that the values of 'b' and / or 'c' may be mapped to positions in the right reference sample column. The predicted sample may be projected considering the opposite angular direction, and this value may be weighted (e.g., equally weighted) against the predicted sample value. For example, for modes 3 through 9, the weights are It could be determined as, here and Each may be a horizontal and vertical component in the angle direction.
[0126] For example, in the case of vertical intra mode(s) that may be near vertical, if the current block is at the bottom face boundary of a face in a frame-packed picture related to, for example, 360-degree video, the predicted samples s of the block located in the bottom row i,jis, for example, as follows, bottom reference row R i,j It can also be adjusted using:
[0127]
[0128] Filtering (e.g., boundary prediction filtering) may be applied to the current block using face boundary cross-neighbor blocks as described herein, for example. FIGS. 26a through 26h depict exemplary boundary prediction filtering at the bottom face boundary for intra mode 14 (e.g., FIG. 26a), intra modes 15-17 (e.g., FIG. 26b), intra mode 18 (e.g., FIG. 26c), intra modes 19-25 (e.g., FIG. 26d), intra mode 26 (e.g., FIG. 26e), intra modes 27-29 (e.g., FIG. 26f), intra modes 30-33 (e.g., FIG. 26g), and intra mode 34 (e.g., FIG. 26h). Reference sample(s), depicted using dashed lines in FIGS. 26a through 26h, may be located outside the current face.
[0129] For other intra-directional mode(s), if the current block is at the bottom face boundary, the predicted sample(s) of the block located in the bottom row(s) (e.g., bottommost row(s)). i,j is, for example, as follows, bottom reference row R i,j It can also be filtered using:
[0130] In the case of mode 14 (e.g., FIG. 26a).
[0131]
[0132] For modes 15-17 (e.g., FIG. 26b)
[0133]
[0134] In the case of mode 18 (e.g., FIG. 26c).
[0135]
[0136] For modes 19-25 (e.g., FIG. 26d)
[0137]
[0138] In the case of mode 26 (e.g., FIG. 26e).
[0139]
[0140] For modes 27-29 (e.g., FIG. 26f) and modes 30-33 (e.g., FIG. 26g)
[0141]
[0142] In the case of mode 34 (e.g., FIG. 26h).
[0143]
[0144] Here, D may be a parameter controlling the number of bottom rows or may be filtered. Weights 'a', 'b', and 'c' may be selected, for example, based on the intra mode and / or distance to the reference sample.
[0145] In the case of intra-directional mode(s), if the current block is at the bottom right side boundary of a face in a frame-packed picture associated with, for example, a 360-degree video, the predicted samples of the block located in the right column(s) (e.g., the far right column(s)) may be filtered using, for example, the right reference column, and the blocks located in the bottom row(s) (e.g., the bottom row(s)) may be filtered using, for example, the bottom reference row as described herein.
[0146] In the case of DC mode, if the current block is at the right side boundary of a face in a frame-packed picture associated with a 360-degree video, for example, the predicted sample(s) of the block located in the right column(s) (e.g., the far right column(s)) may be filtered using the right reference column, for example (e.g., according to Equation (25). If the current block is at the bottom side boundary of a face in a frame-packed picture associated with a 360-degree video, for example, the predicted sample(s) of the block located in the bottom row(s) (e.g., the bottom row(s)) may be filtered using the bottom reference row, for example (e.g., according to Equation (35)). If the current block is, for example, at the bottom right side boundary of a frame-packed picture associated with a 360-degree video, the predicted sample(s) of the block located in the right column(s) (e.g., the far right column(s)) may be filtered using, for example, the right reference column, and the block located in the bottom row(s) (e.g., the bottom row(s)) may be filtered using, for example, the bottom reference row (e.g., according to mathematical formulas (25) and (35).
[0147] The filtering process may be implemented, for example, using fixed-point precision and / or bit shift operations. For example, when considering finer intra-angular granularity and / or rectangular blocks, similar filtering operations may be applied to blocks located in the right column(s) (e.g., the rightmost column(s)) and / or blocks located in the bottom row(s) (e.g., the bottommost row(s)).
[0148] In the case of filtering within a loop, filtering may be applied across one or more (e.g., all) face boundaries. For example, filtering may be applied across face boundaries including right and / or bottom face boundaries. If the current block is, for example, on the left (e.g., or top) face boundary of a face in a frame-packed picture associated with a 360-degree video, the reconstructed sample(s) on the left (e.g., or top) may be used to filter blocks located in the leftmost column(s) (e.g., or top row(s)), even though the block may be on the frame-packed picture boundary. If the block is, for example, on the top-left face boundary of a face in a frame-packed picture associated with a 360-degree video, the reconstructed sample(s) on the left and top may be used to filter blocks located in the leftmost column(s) and top row(s), respectively, even though the block may be on the frame-packed picture boundary. If a block is located on the right (e.g., or bottom) face boundary of a face in a frame-packed picture associated with a 360-degree video, for example, the right (e.g., or bottom) reconstructed sample(s) may be used to filter blocks located in the rightmost column(s) (e.g., or bottom row(s)), even though the block may be located on the frame-packed picture boundary. If a block is located on the bottom right face boundary of a face in a frame-packed picture associated with a 360-degree video, for example, the right and bottom reconstructed sample(s) may be used to filter blocks located in the rightmost column(s) and bottom row(s), respectively, even though the block may be located on the frame-packed picture boundary. The reconstructed sample(s) may be located outside the current face, and the reconstructed sample(s) may be obtained using geometry padding, for example.
[0149] For component cross-linear model prediction, reconstructed sample(s) (e.g., additional resource sample(s)) may be used. For example, the reconstructed sample(s) (e.g., additional resource sample(s)) may be based on the location of the current block within the geometry plane.
[0150] FIGS. 27a through 27c depict exemplary locations of samples used for component cross-linear model prediction at the right-side boundary (e.g., FIG. 27a), the bottom-side boundary (e.g., FIG. 27b), and the bottom-right-side boundary (e.g., FIG. 27c). Reconstructed samples, depicted using dashed lines in FIGS. 27a through 27c, may be located outside the current face. If the current block is located at the right-side boundary of a face in a frame-packed picture, for example, a 360-degree video, reconstructed sample(s) located on the right side of the current block may be used to predict the parameter(s) of the linear model, for example, in addition to the reconstructed sample(s) located on the left side of the current block and / or the reconstructed sample(s) located above the current block, as depicted in FIG. 27a. In this case, the linear model parameter(s) may be calculated as follows (e.g., Equations (38) through (40)):
[0151] Mathematical formula (38) is illustrated in FIG. 28 and
[0152]
[0153] Here It may be, for example, a downsampled reconstructed luma sample located to the upper right of the current block. It may also be calculated as follows, for example, considering the availability of reconstructed luminance samples and / or chroma locations:
[0154]
[0155] One or more downsampling filters may be applied, for example, using face boundary cross-neighbor block(s). If preprocessing (e.g., filtering) is applied to a reconstructed sample located to the left of the current block, similar (e.g., identical) preprocessing may be applied to a reconstructed sample located to the right of the current block.
[0156] If the current block is located at the bottom face boundary of a face in a frame-packed picture associated with a 360-degree video, for example, as depicted in FIG. 27b, to predict the parameters of the linear model, for example, a reconstructed sample located below the current block may be used in addition to a reconstructed sample located to the left and / or above the current block. In this case, the linear model parameters may be calculated as follows (e.g., Equations (41) and (42)):
[0157] Mathematical formula (41) is illustrated in FIG. 28 and
[0158]
[0159] Reconstructed luma samples located below the current block may be downsampled (e.g., according to Equation (16)). One or more downsampling filters may be applied, for example, using face boundary cross-neighbor block(s). If preprocessing (e.g., filtering) is applied to the reconstructed samples located above the current block, similar (e.g., identical) preprocessing may be applied to the reconstructed samples located below the current block.
[0160] If the current block is located at the bottom right side boundary of a face in a frame-packed picture associated with a 360-degree video, for example, as depicted in FIG. 27c, reconstructed samples located to the right and below the current block may be used to predict the parameters of the linear model (for example, in addition to reconstructed samples located to the left and above the current block and reconstructed samples located above the current block). The linear model parameters may also be calculated as follows (for example, Equations (43) and (44)):
[0161] Mathematical formulas (43) and (44) are illustrated in FIG. 28.
[0162] For rectangular block(s), neighboring samples of the longer boundary may be subsampled, for example, using face boundary cross neighbor block(s). For example, neighboring samples of the longer boundary of a rectangular block may be subsampled to have the same number of samples on the shorter boundary. The component cross-linear model predictions described herein may be used to predict between two chroma components (e.g., in the sample domain or in the residual domain). Multiple cross-component linear models may be used at face boundaries, wherein one or more component cross-linear model predictions may be defined over a range of sample values and applied as described herein. If the reconstructed sample(s) are outside the current face, the reconstructed sample may be obtained, for example, using geometry padding.
[0163] One or more available blocks and / or samples located on the other side of the right and / or bottom face boundary (e.g., on the side of the face boundary opposite to the current block location and in the same direction of the coding order with respect to the current block or face boundary intersecting neighbor blocks) may be used for prediction. The availability of face neighbor blocks and / or samples may depend on the coding order in which the blocks of the frame-packed picture are processed. For example, FIG. 29a depicts an exemplary raster scan order (e.g., from top to bottom and / or left to right) for a CMP 3×2 packing configuration. Blocks may be processed face-by-face using a raster scan order within one or more faces, for example (as exemplified in FIG. 29b). Using the face scan order shown in FIG. 29b may increase the availability of face neighbor blocks and / or samples. To achieve similar results, one or more different frame-packing configurations may be used. For example, in the situation depicted in FIG. 29a, if a 6×1 packing configuration is used (e.g., instead of or in addition to a 3×2 packing configuration), the raster scan order may process one or more faces one by one. The block coding order (e.g., processing order) may vary depending on the face arrangement used.
[0164] Constraint(s) may be applied, for example, during block partitioning. For example, constraint(s) may be applied during block partitioning to reduce blocks that overlap across two or more faces. If one or more Coding Tree Units (CTUs) are used, the CTUs may be configured such that one or more (e.g., all) coded blocks within the CTU belong to the same face. If the face size is not a multiple of the CTU size, an overlapping CTU may be used where blocks within the face to which the CTU belongs may be coded. FIG. 30 depicts an exemplary CTU and block partitioning for a face whose face size is not a multiple of the CTU size. Solid lines shown in FIG. 30 may represent face boundaries. Dotted lines may represent CTU boundaries, and dotted lines may represent block boundaries. Blocks depicted using hatched patterns may be located outside the current face. One or more different block scan orders may be used for intra- and inter-coded frames. For example, intra-coded frames may use a face scan order (e.g., FIG. 29b). Inter-coded frames may use a raster scan order (e.g., FIG. 29a). For example, different scan order(s) may be used between different faces and / or within different faces based on the coding mode of the faces (e.g., prediction mode).
[0165] Loop filter operation may be enabled or disabled (for example, as described herein, using neighbor samples from a face already coded in the loop filtering operation to the right and / or bottom boundary). For example, loop filter operation may be enabled or disabled based on the face scan order and / or frame packing configuration. If appropriate neighbor samples in the 3D geometry are not used in the deblocking filter or other in-loop filters, unpleasant visual artifacts in the form of face seams may be visible in the reconstructed video. For example, if the reconstructed video is used to render a viewport and is displayed to a user, for example, via a head-mount device (HMD) or a 2D screen, face seams may be visible in the reconstructed video. For example, FIG. 18b illustrates a 3×2 CMP example. The three faces in the top half shown in FIG. 18b may be horizontally continuous in the 3D geometry. The three faces in the bottom half may be horizontally continuous in the 3D geometry. The top half and the bottom half may be discontinuous in the 3D geometry. A 3×2 CMP picture may be coded using two tiles (e.g., a tile for the top half of the 3×2 CMP picture and a tile for the bottom half), and loop filtering may be disabled across the tile boundaries. For example, loop filtering across the tile boundaries may be disabled by setting the value of the picture parameter set (PPS) syntax element loop_filter_across_tiles_enabled_flag to 0.Loop filtering may be disabled to prevent deblocking and / or other in-loop filters from being applied across discontinuous edges (e.g., a horizontal edge separating the top and bottom halves).
[0166] The face scan order may be used to encode and / or decode blocks in a frame-packed picture. The six faces shown in FIG. 18b may be processed using the order shown in FIG. 29b. The face scan order may be achieved by aligning the six faces with six tiles. In this case, setting an indicator (e.g., setting loop_filter_across_tiles_enabled_flag to 0) may cause the deblocking and in-loop filter to be disabled across horizontal edges between tiles (e.g., these may be discontinuous and disabled) and across vertical edges (e.g., these may be continuous and not disabled). Whether an edge loop filter is applied or not may be specified. The type of loop filter may be considered. For example, a loop filter, such as a deblocking and / or adaptive loop filter (ALF), may be an N-tap filter that uses neighboring samples in the filtering process. One or more of the neighboring samples used for the filter may cross discontinuous edges. A loop filter, such as a Sample Adaptive Offset (SAO), may modify the sample value decoded at the current position by adding an offset. In the example, the loop filter (e.g., SAO) may not use neighboring sample values in the filtering operation. Deblocking and / or ALF may be disabled across some tile and / or face boundaries. SAO may be enabled. For example, deblocking and / or ALF may be disabled across some tile and / or face boundaries, while SAO may be enabled.
[0167] Extensions for an indicator (e.g., loop_filter_across_tiles_enabled_flag) may be indicated to a coding device (e.g., in a bitstream). The syntax element of the indicator (e.g., loop_filter_across_tiles_enabled_flag) may be separated. For example, the loop_filter_across_tiles_enabled_flag syntax element may be separated into two or more syntax elements (e.g., two syntax elements). In the example, whether to apply the loop filter to horizontal edges may be indicated, for example, via a syntax element. In the example, whether to apply the loop filter to vertical edges may be indicated, for example, via a syntax element.
[0168] The syntax element of an indicator (e.g., loop_filter_across_tiles_enabled_flag) may be separated into two or more syntax elements and may be indicated to a coding device (e.g., in a bitstream). Whether to enable or disable a loop filter across a given edge may be indicated, for example, through two or more separate syntax elements. For example, a frame-packed projection format containing M×N faces (e.g., Fig. 18b where M = 3 and N = 2) may be considered. In a frame-packed projection format containing M×N faces, there may be (M-1)×N vertical edges between faces in the picture and M×(N-1) horizontal edges between faces in the picture. In this case, (M-1)×N + M×(N-1) indications (e.g., flags) may specify whether to enable or disable one or more of the horizontal and / or vertical edges. The semantics of the indicator (e.g., loop_filter_across_tiles_enabled_flag) syntax element may be adapted to disable or enable the loop filter across edges between continuous faces. In this case, for example, to prevent the occurrence of seams, the loop filter may be disabled across edges between discontinuous faces. Signaling may be configured to specify which edges are between continuous faces and / or which edges are between discontinuous faces.
[0169] The syntax element of an indicator (e.g., loop_filter_across_tiles_enabled_flag) may be separated into two or more syntax elements. Two or more syntax elements may be used to control different types of loop filters. For example, an indicator (e.g., a flag) may be used to enable or disable deblocking. An indicator (e.g., a flag) may be used to enable or disable ALF, for example. An indicator (e.g., a flag) may be used to enable or disable SAO, for example. If more loop filters are used in a video encoder and / or decoder (e.g., illustrated in FIGS. 7 and 8), more indicators (e.g., flags) may be used. The indicator (e.g., loop_filter_across_tiles_enabled_flag) syntax element may be separated (e.g., into two syntax elements). For example, an indication (e.g., flag) may control a loop filter that uses neighboring samples. For example, a different indication (e.g., flag) may control a loop filter that does not use neighboring samples. The indicator (e.g., loop_filter_across_tiles_enabled_flag) syntax element may indicate whether to disable or enable a loop filter that uses neighboring samples. In this case, a loop filter that does not use neighboring samples may be enabled.
[0170] The extensions described herein may be combined. For example, when multiple markings (e.g., flags) are used, one or more markings (e.g., flags) may be used to control the type of loop filter, and the semantics of one or more markings (e.g., flags) may be adapted accordingly. For example, the semantics of one or more markings may be adapted to control whether the filter is applied to edges across continuous surfaces, and the filter may be disabled for edges across discontinuous surfaces (e.g., to prevent the occurrence of seams).
[0171] Extensions to the indicators described herein (e.g., loop_filter_across_tiles_enabled_flag) may involve the use of tiles in an exemplary manner. Those skilled in the art will recognize that they may also be applied to other face levels (e.g., slices). For example, pps_loop_filter_across_slices_enabled_flag and / or slice_loop_filter_across_slices_enabled_flag, which control whether to apply a loop filter across a slide, may be used as indicators (e.g., different indicators). Because coded blocks within a CTU belong to the same tile and / or slice, the tile and / or slice sizes may be multiples of the CTU size. The face sizes may not be multiples of the CTU size. The tiles and / or slices described herein may be replaced with faces (e.g., to prevent the use of tiles and / or slices and to change the semantics described herein). For example, the syntax of the indicator described herein (e.g., loop_filter_across_tiles_enabled_flag) and one or more extensions thereof may be replaced by a different indicator (e.g., loop_filter_across_faces_enabled_flag) (e.g., or may be used in conjunction with it). When tiles and / or slices are enabled, face-based loop filter control indicators (e.g., flags) may be activated. In this case, such face-based loop filter control indicators (e.g., flags) may be applicable to edges on face boundaries, and tile and / or slice-based loop filter control indicators (e.g., flags) may be applicable to edges on tile and / or slice boundaries.
[0172] One or more (e.g., all) coded blocks within a CTU belonging to the same face, and a CTU that overlaps so that blocks within the face to which the CTU belongs may also be coded. In this way, tiles and / or slices may be used (e.g., even if the face size is not a multiple of the CTU size).
[0173] For CMP and / or related cube-based geometry, a 3×2 packing configuration may be used. For example, the 3×2 packing configuration may be used for a representation (e.g., a compact representation) or to form a rectangular frame-packed picture. The 3×2 packing configuration may omit filling empty areas with default (e.g., void) samples to form the rectangular frame-packed picture (e.g., a 4×3 packing configuration) shown in FIG. 2b. One or more faces of the 3×2 packing configuration may be arranged so that discontinuity between two adjacent faces in the frame picture may be reduced. The 3×2 packing configuration may be defined such that, as depicted in FIG. 31, the right face (front face), front face, and left face may be arranged in the top row (e.g., in this particular order), and the bottom face, rear face, and top face may be arranged in the bottom row (e.g., in this particular order). FIG. 31 illustrates an exemplary 3×2 packing configuration (e.g., having a right face, front face, and left face placed in the top row, and a bottom face, rear face, and top face placed in the bottom row). The dashed lines in FIG. 31 may indicate CTU boundaries. Arrows may indicate shared boundaries between two faces. Within a face row (e.g., each face row), faces may be rotated. For example, a face row may be rotated to minimize discontinuity between two adjacent faces. The face size may not be a multiple of the CTU size. Part of the second face row may be coded, for example, before the first face row is fully coded. When coding one or more blocks of the bottom face, rear face, and top face, adjacent blocks of the right face, front face, and / or left face may not be available.For example, in the case of the 3×2 packing configuration depicted in FIG. 31, when encoding a block of the first (e.g., partial) CTU row of the bottom face (e.g., shaded area in FIG. 31), information may not be inferred from neighboring blocks on the front face and left face because the corresponding neighboring block may not yet be encoded. Similarly, in the case of the rear face, for the first block within the width of size h = mod(face size, CTU size) in the first (e.g., partial) CTU row (e.g., shaded area in FIG. 31) (where mod(x, y) may be a modulo operator), information may not be inferred from neighboring blocks on the left face because the corresponding neighboring block may not yet be encoded.
[0174] A 3×2 packing configuration may be used and / or signaled in a bitstream, for example, to / from a coding device. For example, when the 3×2 packing configuration described herein is used, information from one or more adjacent faces of a first row of faces may be inferred for a face located in a second row of faces. The 3×2 packing configuration may be defined such that, as depicted in FIG. 32, a right face, a front face, and a left face are arranged in the first row of faces (e.g., in this particular order), and a top face, a rear face, and a bottom face are arranged in the second row of faces (e.g., in this particular order). FIG. 32 illustrates an exemplary 3×2 packing configuration (e.g., having a right face, a front face, and a left face in the first row of faces, and a top face, a rear face, and a bottom face in the second row of faces). Within a row of faces (e.g., each row of faces), the faces may be rotated, for example, to minimize discontinuity between two adjacent faces. In the case of the 3×2 packing configuration illustrated in FIG. 32, one or more faces of the second row of faces may be rotated 180 degrees (e.g., compared to the 3×2 packing configuration illustrated in FIG. 31). In the configuration illustrated in FIG. 32, when encoding the top face, rear face, and bottom face, adjacent blocks of the right face, front face, and left face may be available.
[0175] The definitions of front face, back face, left face, right face, top face, and / or bottom face described herein may be relative. Rotation may be applied to the cube to obtain an arrangement similar to that described herein.
[0176] FIG. 33a is a drawing illustrating an exemplary communication system (100) in which one or more disclosed embodiments may be implemented. The communication system (100) may be a multiple access system that provides content such as voice, data, video, messaging, broadcast, etc. to multiple wireless users. The communication system (100) may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system (100) may utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), and others.
[0177] As illustrated in FIG. 33a, the communication system (100) may include wireless transceiver units (WTRUs) (102a, 102b, 102c, 102d), a RAN (104 / 113), a CN (106 / 115), a public switched telephone network (PSTN) (108), the Internet (110), and other networks (112), but it will be recognized that the disclosed embodiments consider any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs (102a, 102b, 102c, 102d) may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU (102a, 102b, 102c, 102d)—any of which may be referred to as a “station” and / or “STA”—may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile station, fixed or mobile subscriber unit, subscription-based unit, pager, cellular phone, personal digital assistant (PDA), smartphone, laptop, netbook, personal computer, wireless sensor, hotspot or Mi-Fi device, Internet of Things (IoT) device, watch or other wearable, head-mounted display (HMD), vehicle, drone, medical device and application (e.g., remote surgery), industrial device and application (e.g., robot and / or other wireless device operating in the context of an industrial and / or automated processing chain), consumer electronic device, device operating on a commercial and / or industrial wireless network, and the like.Any of WTRU(102a, 102b, 102c and 102d) may be interchangeably referred to as UE.
[0178] The communication system (100) may also include a base station (114a) and / or a base station (114b). Each of the base stations (114a, 114b) may be any type of device configured to wirelessly interface with at least one of the WTRUs (102a, 102b, 102c, 102d) to facilitate access to one or more communication networks, such as a CN (106 / 115), the Internet (110), and / or other networks (112). For example, the base stations (114a, 114b) may be a base transceiver station (BTS), Node-B, eNode B, Home Node B, Home eNode B, gNB, NR Node B, site controller, access point (AP), wireless router, and the like. Although each base station (114a, 114b) is described as a single element, it will be recognized that the base station (114a, 114b) may include any number of interconnected base station and / or network elements.
[0179] A base station (114a) may be part of a RAN (104 / 113) which may also include other base station and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. Base station (114a) and / or base station (114b) may be configured to transmit and / or receive radio signals on one or more carrier frequencies and may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed spectrum and unlicensed spectrum. A cell may provide coverage for a radio service in a specific geographical area, which may be relatively fixed or may vary over time. A cell may be further divided into cell sectors. For example, a cell associated with a base station (114a) may be divided into three sectors. Accordingly, in one embodiment, the base station (114a) may include three transceivers, that is, one transceiver for each sector of the cell. In one embodiment, the base station (114a) may utilize multiple-input multiple output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0180] Base stations (114a, 114b) may communicate with one or more of WTRUs (102a, 102b, 102c, 102d) via an air interface (116) which may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface (116) may be established using any suitable radio access technology (RAT).
[0181] More specifically, as mentioned above, the communication system (100) may be a multiple access system and may utilize one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and so on. For example, base stations (114a) and WTRUs (102a, 102b, 102c) within a RAN (104 / 113) may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish a radio interface (115 / 116 / 117) using wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0182] In one embodiment, the base station (114a) and WTRU (102a, 102b, 102c) may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA) which may establish a radio interface (116) using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0183] In one embodiment, the base station (114a) and WTRU (102a, 102b, 102c) may implement a wireless technology such as NR wireless access, which may establish a wireless interface (116) using New Radio (NR).
[0184] In one embodiment, the base station (114a) and the WTRU (102a, 102b, 102c) may implement multiple wireless access technologies. For example, the base station (114a) and the WTRU (102a, 102b, 102c) may implement LTE wireless access and NR wireless access together, for example, using the dual connectivity (DC) principle. Accordingly, the wireless interface utilized by the WTRU (102a, 102b, 102c) may be characterized by transmissions and / or multiple types of wireless access technologies transmitted to / from multiple types of base stations (e.g., eNB and gNB).
[0185] In another embodiment, the base station (114a) and WTRU (102a, 102b, 102c) may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, IS-2000 (Interim Standard 2000), IS-95 (Interim Standard 95), IS-856 (Interim Standard 856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and others.
[0186] The base station (114b) of FIG. 33a may be, for example, a wireless router, home node B, home eNode B, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in localized areas such as workplaces, homes, vehicles, campuses, industrial facilities, air corridors (for use by drones, for example), roads, and so on. In one embodiment, the base station (114b) and WTRU (102c, 102d) may implement wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, the base station (114b) and WTRU (102c, 102d) may implement wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). Still in other embodiments, the base station (114b) and WTRU (102c, 102d) may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As illustrated in FIG. 33a, the base station (114b) may have a direct connection to the Internet (110). Thus, the base station (114b) may not need to access the Internet (110) via the CN (106 / 115).
[0187] The RAN (104 / 113) may communicate with a CN (106 / 115), which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services, including one or more of the WTRUs (102a, 102b, 102c, 102d). The data may have various quality of service (QoS) requirements, such as throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and so on. The CN (106 / 115) may provide call control, billing services, mobile location-based services, prepaid calling, internet connectivity, video distribution, and so on, and / or perform high-level security functions such as user authentication. Although not illustrated in FIG. 33a, it will be recognized that the RAN (104 / 113) and / or CN (106 / 115) may communicate directly or indirectly with other RANs utilizing the same RAT or different RAT as the RAN (104 / 113). For example, in addition to being connected to the RAN (104 / 113) which may utilize NR wireless technology, the CN (106 / 115) may also communicate with other RANs (not illustrated) utilizing GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi wireless technology.
[0188] CN (106 / 115) may also serve as a gateway for WTRUs (102a, 102b, 102c, 102d) to access the PSTN (108), the Internet (110), and / or other networks (112). The PSTN (108) may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet (110) may include a global system of interconnected computer networks and devices using common communication protocols such as the transmission control protocol (TCP), user datagram protocol (UDP), and / or internet protocol (IP) in the TCP / IP Internet protocol suite. The network (112) may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network (112) may include other CNs connected to one or more RANs that may utilize the same RAT or different RAT as the RAN (104 / 113).
[0189] Some or all of the WTRUs (102a, 102b, 102c, 102d) in the communication system (100) may include multimode capabilities (for example, the WTRUs (102a, 102b, 102c, 102d) may include multiple transceivers for communicating with different wireless networks through different wireless links). For example, the WTRU (102c) shown in FIG. 33a may be configured to communicate with a base station (114a) that may utilize cellular-based wireless technology and a base station (114b) that may utilize IEEE 802 wireless technology.
[0190] FIG. 33b is a system diagram illustrating an exemplary WTRU (102). As illustrated in FIG. 1b, the WTRU (102) may include, among other things, a processor (118), a transceiver (120), a transmit / receive element (122), a speaker / microphone (124), a keypad (126), a display / touchpad (128), non-removable memory (130), removable memory (132), a power supply (134), a global positioning system (GPS) chipset (136), and / or other peripherals (138). It will be recognized that the WTRU (102) may include any noncombination of the aforementioned elements while still conforming to one embodiment.
[0191] The processor (118) may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, and the like. The processor (118) may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU (102) to operate in a wireless environment. The processor (118) may be coupled to a transceiver (120) which may be coupled to a transmit / receive element (122). Although FIG. 33b depicts the processor (118) and transceiver (120) as separate components, it will be recognized that the processor (118) and transceiver (120) may be integrated together in an electronic package or chip.
[0192] The transmitting / receiving element (122) may be configured to transmit a signal to a base station (e.g., base station (114a)) or to receive a signal from the base station via the wireless interface (116). For example, in one embodiment, the transmitting / receiving element (122) may be an antenna configured to transmit and / or receive an RF signal. In another embodiment, the transmitting / receiving element (122) may be an emitter / detector configured to transmit and / or receive an IR, UV, or visible light signal, for example. Still in another embodiment, the transmitting / receiving element (122) may be configured to transmit and / or receive both RF and optical signals. It will be recognized that the transmitting / receiving element (122) may be configured to transmit and / or receive any combination of wireless signals.
[0193] Although the transmit / receive element (122) is depicted as a single element in FIG. 33b, the WTRU (102) may include any number of transmit / receive elements (122). More specifically, the WTRU (102) may utilize MIMO technology. Thus, in one embodiment, the WTRU (102) may include two or more transmit / receive elements (122) (e.g., multiple antennas) for transmitting and receiving wireless signals through the wireless interface (116).
[0194] The transceiver (120) may be configured to modulate a signal to be transmitted by the transmitting / receiving element (122) and to demodulate a signal received by the transmitting / receiving element (122). As mentioned above, the WTRU (102) may have multimode capabilities. Accordingly, the transceiver (120) may include multiple transceivers to enable the WTRU (102) to communicate through multiple RATs, such as NR and IEEE 802.11, for example.
[0195] The processor (118) of the WTRU (102) may be coupled to a speaker / microphone (124), a keypad (126), and / or a display / touchpad (128) (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input data from these. The processor (118) may also output user data to the speaker / microphone (124), the keypad (126), and / or the display / touchpad (128). Additionally, the processor (118) may access information in any type of suitable memory, such as non-removable memory (130) and / or removable memory (132), and may store data in any type of suitable memory. The non-removable memory (130) may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory (132) may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In another embodiment, the processor (118) may access information in memory that is not physically located on the WTRU (102), such as memory on a server or home computer (not shown), and may store data in that memory.
[0196] The processor (118) may receive power from the power source (134) and may be configured to distribute power to other components of the WTRU (102) and / or control the power. The power source (134) may be any suitable device for supplying power to the WTRU (102). For example, the power source (134) may include one or more dry cell batteries (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li ion), etc.), solar cells, fuel cells, and the like.
[0197] The processor (118) may also be coupled to a GPS chipset (136) which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU (102). Additionally, in addition to, or instead of, information from the GPS chipset (136), the WTRU (102) may receive location information from a base station (e.g., base station (114a, 114b)) via the wireless interface (116) and / or determine its location based on the timing of signals being received from two or more nearby base stations. It will be recognized that the WTRU (102) may obtain location information through any suitable location determination method while remaining consistent with one embodiment.
[0198] The processor (118) may also be coupled to other peripherals (138) which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals (138) may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photography and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. The peripheral device (138) may include one or more sensors, and the sensors may be one or more of a gyroscope, accelerometer, hall effect sensor, magnetometer, orientation sensor, proximity sensor, temperature sensor, time sensor; geolocation sensor; altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0199] The WTRU (102) may include a full-duplex radio in which the transmission and reception of some or all of the signal (e.g., associated with a specific subframe for both the UL (e.g., for transmission) and the downlink (e.g., for reception) may be simultaneous and / or concurrent. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference through signal processing via a processor (e.g., a separate processor (not shown) or processor (118)) or through a hard (e.g., a choke). In one embodiment, the WRTU (102) may include a half-duplex radio, in which case the transmission and reception of some or all of the signal (e.g., associated with a specific subframe for either the UL (e.g., for transmission) or the downlink (e.g., for reception)
[0200] FIG. 33c is a system diagram illustrating a RAN (104) and a CN (106) according to one embodiment. As mentioned above, the RAN (104) may utilize E-UTRA radio technology to communicate with a WTRU (102a, 102b, 102c) via a wireless interface (116). The RAN (104) may also communicate with a CN (106).
[0201] It will be recognized that the RAN (104) may include eNode-Bs (160a, 160b, 160c), but the RAN (104) may include any number of eNode-Bs while remaining consistent with one embodiment. Each of the eNode-Bs (160a, 160b, 160c) may include one or more transceivers for communicating with the WTRU (102a, 102b, 102c) via the wireless interface (116). In one embodiment, the eNode-Bs (160a, 160b, 160c) may implement MIMO technology. Thus, the eNode-B (160a) may use multiple antennas, for example, to transmit a wireless signal to the WTRU (102a) and / or to receive a wireless signal from the WTRU (102a).
[0202] Each of the eNode-B (160a, 160b, 160c) may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, and so on. As shown in FIG. 33d, the eNode-B (160a, 160b, 160c) may communicate with each other via an X2 interface.
[0203] The CN (106) illustrated in FIG. 33c may include a mobility management entity (MME) (162), a serving gateway (SGW) (164), and a packet data network (PDN) gateway (or PGW) (166). Although each of the aforementioned elements is depicted as part of the CN (106), it will be recognized that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0204] The MME (162) may be connected to each of the eNode-B (162a, 162b, 162c) within the RAN (104) via the S1 interface and may act as a control node. For example, the MME (162) may be responsible for authenticating users of the WTRU (102a, 102b, 102c), enabling / disabling bearers, selecting a specific serving gateway during the initial connection of the WTRU (102a, 102b, 102c), and so on. The MME (162) may also provide a control plane function for switching between the RAN (104) and another RAN (not shown) utilizing other wireless technologies such as GSM and / or WCDMA.
[0205] The SGW (164) may be connected to each of the eNode B (160a, 160b, 160c) within the RAN (104) via the S1 interface. The SGW (164) may generally route and forward user data packets to and from the WTRU (102a, 102b, 102c). The SGW (164) may also perform other functions, such as anchoring the user plane during inter-eNode B handover, triggering paging when DL data is available for the WTRU (102a, 102b, 102c), managing and storing the context of the WTRU (102a, 102b, 102c), and so on.
[0206] The SGW (164) may be connected to a PGW (166) which may provide the WTRU (102a, 102b, 102c) with access to a packet-switched network such as the Internet (110) to facilitate communication between the WTRU (102a, 102b, 102c) and an IP-enabled device.
[0207] CN (106) may facilitate communication with other networks. For example, CN (106) may provide WTRU (102a, 102b, 102c) with access to a circuit-switched network, such as PSTN (108), to facilitate communication between WTRU (102a, 102b, 102c) and traditional terrestrial circuit communication devices. For example, CN (106) may include an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between CN (106) and PSTN (108), or may communicate with it. Additionally, CN (106) may provide WTRU (102a, 102b, 102c) with access to another network (112), which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0208] Although the WTRU is described as a wireless terminal in FIGS. 33a to 33d, it is considered that in certain representative embodiments, such a terminal may use a wired communication interface with a communication network (e.g., temporarily or permanently).
[0209] In a typical embodiment, the other network (112) may be a WLAN.
[0210] A WLAN in Infrastructure Basic Service Set (BSS) mode may include an Access Point (AP) for the BSS and one or more Stations (STAs) associated with the AP. The AP may access or interface with a distribution system (DS) or other types of wired / wireless networks that carry traffic into and / or out of the BSS. Traffic originating from outside the BSS to a STA may reach or be forwarded to the STA through the AP. Traffic originating from a STA to a destination outside the BSS may be transmitted to the AP to be forwarded to the respective destination. Traffic between STAs within the BSS may be transmitted, for example, through the AP, in which case the source STA may transmit the traffic to the AP and the AP may forward the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be transmitted between a source STA and a destination STA (e.g., directly between them) via a direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode may sometimes be referred to herein as an "ad-hoc" communication mode.
[0211] When using an 802.11ac infrastructure operating mode or a similar operating mode, the AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., 20 MHz broadband) or a width dynamically set through signaling. The primary channel may be the operating channel of the BSS or may be used by the STA to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. In the case of CSMA / CA, STAs including the AP (e.g., all STAs) may sense the primary channel. If the primary channel is sensed / detected by a specific STA and / or determined to be in use, that specific STA may be backed off. A single STA (e.g., only one station) may transmit at any given time in a given BSS.
[0212] A High Throughput (HT) STA may form a 40 MHz wide channel by using a 40 MHz wide channel for communication through a combination of a 20 MHz main channel and an adjacent or non-adjacent 20 MHz channel, for example.
[0213] Very High Throughput (VHT) STAs may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels may be formed, for example, by combining consecutive 20 MHz channels. A 160 MHz channel may be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, the combination of two non-consecutive 80 MHz channels may be referred to as an 80 + 80 configuration. In the case of an 80 + 80 configuration, data may pass through a segment parser that may split the data into two streams after channel encoding. Inverse fast Fourier transform (IFFT) processing and time domain processing may be performed individually on each stream. The stream may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operation described above for the 80 + 80 configuration may be reversed, and the combined data may be transmitted to the Medium Access Control (MAC).
[0214] Sub-1 GHz operating modes may be supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah may support Meter Type Control / Machine-Type Communication, such as MTC devices, in a macro coverage area. The MTC device may have limited performance, including, for example, support for a certain and / or limited bandwidth (for example, support only for a certain and / or limited bandwidth). The MTC device may also include a battery having a battery life exceeding a threshold (for example, to maintain a very long battery life).
[0215] A WLAN system that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that may be designated as the main channel. The main channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs within the BSS. The bandwidth of the main channel may be set and / or limited by one STA among all STAs operating in the BSS that support the smallest bandwidth operating mode. In the example of 802.11ah, even if the AP and other STAs within the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes, the main channel may be 1 MHz wide for a STA (e.g., an MTC type device) that supports the 1 MHz mode (e.g., only the 1 MHz mode). Carrier detection and / or Network Allocation Vector (NAV) settings may depend on the state of the main channel. For example, if the main channel is in use due to a STA transmitting to the AP (which supports only 1 MHz operating mode), the entire available frequency band may be considered in use, even if most of the frequency band remains idle and available.
[0216] In the United States, the available frequency band that may be used by 802.11ah is from 902 MHz to 928 MHz. In Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total available bandwidth for IEEE 802.11ah is 6 MHz to 26 MHz, depending on the country code.
[0217] FIG. 33d is a system diagram illustrating a RAN (113) and a CN (115) according to one embodiment. As mentioned above, the RAN (113) may utilize NR radio technology to communicate with WTRUs (102a, 102b, 102c) via a wireless interface (116). The RAN (113) may also communicate with the CN (115).
[0218] It will be recognized that the RAN (113) may include gNBs (180a, 180b, 180c), but the RAN (113) may include any number of gNBs while remaining consistent with one embodiment. Each gNB (180a, 180b, 180c) may include one or more transceivers for communicating with the WTRU (102a, 102b, 102c) via the wireless interface (116). In one embodiment, the gNBs (180a, 180b, 180c) may implement MIMO technology. For example, the gNB (180a, 108b) may utilize beamforming to transmit a signal to the gNB (180a, 180b, 180c) and / or receive a signal from it. Accordingly, the gNB (180a) may, for example, transmit a radio signal to the WTRU (102a) using multiple antennas and / or receive a radio signal from it. In one embodiment, the gNB (180a, 180b, 180c) may implement carrier aggregation technology. For example, the gNB (180a) may transmit multiple component carriers to the WTRU (102a) (not shown). A subset of these component carriers may be on unauthorized spectrum, while the remaining component carriers may be on authorized spectrum. In one embodiment, the gNB (180a, 180b, 180c) may implement Coordinated Multi-Point (CoMP) technology. For example, WTRU (102a) may receive cooperative transmissions from gNB (180a) and gNB (180b) (and / or gNB (180c)).
[0219] The WTRU (102a, 102b, 102c) may also communicate with the gNB (180a, 180b, 180c) using transmissions associated with scalable numerology. For example, the OFDM symbol interval and / or OFDM subcarrier interval may vary for different transmissions, different cells, and / or different parts of the radio transmission spectrum. The WTRU (102a, 102b, 102c) may also communicate with the gNB (180a, 180b, 180c) using subframes or transmission time intervals (TTIs) of varying or scalable lengths (e.g., containing a varying number of OFDM symbols and / or sustaining an absolute time of varying lengths).
[0220] The gNB (180a, 180b, 180c) may be configured to communicate with the WTRU (102a, 102b, 102c) in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRU (102a, 102b, 102c) may also communicate with the gNB (180a, 180b, 180c) without accessing other RANs (e.g., eNode-B (160a, 160b, 160c)). In a standalone configuration, the WTRU (102a, 102b, 102c) may utilize one or more of the gNB (180a, 180b, 180c) as a mobility anchor point. In a standalone configuration, the WTRU (102a, 102b, 102c) may communicate with the gNB (180a, 180b, 180c) using unauthorized band signals. In a non-standalone configuration, the WTRU (102a, 102b, 102c) may also communicate with the gNB (180a, 180b, 180c) while communicating with / connecting to other RANs such as the eNode-B (160a, 160b, 160c). For example, the WTRU (102a, 102b, 102c) may implement the DC principle to communicate substantially simultaneously with one or more gNBs (180a, 180b, 180c) and one or more eNode-Bs (160a, 160b, 160c). In a non-standalone configuration, eNode-B(160a, 160b, 160c) may serve as mobility anchors for WTRU(102a, 102b, 102c) and gNB(180a, 180b, 180c) may provide additional coverage and / or throughput to service WTRU(102a, 102b, 102c).
[0221] Each of the gNBs (180a, 180b, 180c) may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, duplex connectivity, interoperability between NR and E-UTRA, routing of user plane data toward User Plane Functions (UPF) (184a, 184b), routing of control plane information toward Access and Mobility Management Functions (AMF) (182a, 182b), and the like. As shown in FIG. 33d, the gNBs (180a, 180b, 180c) may communicate with each other via an Xn interface.
[0222] The CN (115) illustrated in FIG. 33d may include at least one AMF (182a, 182b), at least one UPF (184a, 184b), at least one Session Management Function (SMF) (183a, 183b), and possibly a Data Network (DN) (185a, 185b). Although each of the aforementioned elements is depicted as part of the CN (115), it will be recognized that any of these elements may be owned and / or operated by an entity other than the CN operator.
[0223] The AMF (182a, 182b) may be connected to one or more of the gNBs (180a, 180b, 180c) within the RAN (113) via the N2 interface and may act as a control node. For example, the AMF (182a, 182b) may be responsible for authenticating users of the WTRU (102a, 102b, 102c), supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting a specific SMF (183a, 183b), managing the registration area, terminating NAS signaling, managing mobility, and so on. Network slicing may be used by the AMF (182a, 182b) to customize CN support for the WTRU (102a, 102b, 102c) based on the type of service being utilized in the WTRU (102a, 102b, 102c). For example, different network slices may be established for different use cases such as services relying on Ultra-Reliable Low Latency (URLLC) access, services relying on Enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, and / or etc. The AMF (162) may also provide control plane functions for switching between different RANs (not shown) and RAN (113) utilizing other wireless technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0224] SMF (183a, 183b) may be connected to AMF (182a, 182b) of CN (115) via the N11 interface. SMF (183a, 183b) may also be connected to UPF (184a, 184b) of CN (115) via the N4 interface. SMF (183a, 183b) may select and control UPF (184a, 184b) and configure the routing of traffic through UPF (184a, 184b). SMF (183a, 183b) may also perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and so on. The PDU session type may be IP-based, non-IP-based, Ethernet-based, and so on.
[0225] The UPF (184a, 184b) may be connected to one or more of the gNBs (180a, 180b, 180c) in the RAN (113) via an N3 interface that may provide the WTRU (102a, 102b, 102c) with access to a packet-switched network such as the Internet (110) to facilitate communication between the WTRU (102a, 102b, 102c) and the IP-compatible device. The UPF (184, 184b) may also perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, and so on.
[0226] CN (115) may facilitate communication with other networks. For example, CN (115) may include an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that acts as an interface between CN (115) and PSTN (108), or may communicate with it. Additionally, CN (115) may provide WTRU (102a, 102b, 102c) with access to other networks (112), which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRU (102a, 102b, 102c) may be connected to the local data network (DN) (185a, 185b) through the N3 interface to the UPF (184a, 184b) and through the N6 interface between the UPF (184a, 184b) and the data network (DN) (185a, 185b).
[0227] With reference to FIGS. 33a through 33d and the corresponding descriptions of FIGS. 33a through 33d, one or more of the functions described herein in relation to one or more of the WTRU (102a-d), base station (114a-114), eNode-B (160a-c), MME (162), SGW (164), PGW (166), gNB (180a-c), AMF (182a-b), UPF (184a-b), SMF (183a-b), DN (185a-b), and / or any other device(s) described herein may be performed by one or more emulation devices (not shown). The emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, the emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.
[0228] An emulation device may be designed to implement one or more tests of another device in a laboratory environment and / or an operator network environment. For example, one or more emulation devices may perform one or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test another device within the communication network. One or more emulation devices may perform one or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device may be directly coupled to another device for testing purposes and / or may perform testing using over-the-air wireless communication.
[0229] One or more emulation devices may perform one or more functions, including all functions, while not being temporarily implemented or deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in testing scenarios of non-deployed (e.g., testing) wired and / or wireless communication networks and / or testing laboratories to implement testing of one or more components. One or more emulation devices may be test devices. Direct RF coupling and / or wireless communication through RF circuitry (e.g., which may include one or more antennas) may be used by the emulation devices to transmit and / or receive data.
Claims
Claim 1 A method for decoding 360-degree video comprises: acquiring a frame-packed picture from video data; acquiring a plurality of boundary location indications—the plurality of boundary location indications representing corresponding locations of a plurality of boundaries in the frame-packed picture—; acquiring a cross boundary loop filter indicator—the cross boundary loop filter indicator indicating that loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture and that loop filtering is enabled across one or more continuous boundaries in the frame-packed picture—the location of a boundary in the frame-packed picture where the loop filtering is to be disabled—based on the cross boundary loop filter indicator indicating that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture and based on the acquired plurality of boundary location indications—where the loop filtering A method for decoding 360-degree video, comprising: determining that the location of the boundary in the frame-packed picture to be disabled includes discontinuous boundaries; and disabling loop filtering across the determined location of the boundary in the frame-packed picture. Claim 2 A method for decoding a 360-degree video according to claim 1, wherein the boundary in the frame-packed picture comprises one or more of a face boundary, a tile boundary, or a slice boundary. Claim 3 A method for decoding 360-degree video according to claim 1, wherein the boundary cross loop filter indicator is a cross boundary loop filter flag configured to indicate that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture. Claim 4 A method for decoding 360-degree video according to claim 1, wherein the boundary cross-loop filter indicator is obtained from a picture parameter set (PPS). Claim 5 A method for decoding a 360-degree video according to claim 1, comprising: identifying a plurality of boundaries in a frame-packed picture based on the plurality of boundary location indicators; and determining the location of the boundary where loop filtering is disabled from the plurality of identified boundaries in the frame-packed picture. Claim 6 A method for decoding 360-degree video according to claim 1, wherein the loop filtering is associated with one or more of a sample adaptive offset (SAO) filter, a deblocking filter, or an adaptive loop filter (ALF). Claim 7 A method for decoding a 360-degree video according to claim 1, comprising: identifying the plurality of boundaries in the frame-packed picture based on the plurality of boundary location indicators; determining the location of the boundary where the loop filtering is enabled—the location of the boundary where the loop filtering is enabled includes the continuous boundary—based on the boundary cross-loop filter indicator indicating that the loop filtering is enabled across the one or more continuous boundaries in the frame-packed picture and based on the plurality of boundaries in the frame-packed picture; and applying the loop filtering across the determined location of the boundary where the loop filtering is enabled. Claim 8 A method for decoding a 360-degree video according to claim 1, wherein the frame-packed picture includes a horizontal boundary and a vertical boundary associated with the frame-packed picture. Claim 9 delete Claim 10 An apparatus for decoding 360-degree video comprises a processor, wherein the processor comprises: acquiring a frame-packed picture in video data; acquiring a plurality of boundary location indications—the plurality of boundary location indications representing corresponding locations of a plurality of boundaries in the frame-packed picture—; acquiring a cross boundary loop filter indicator—the cross boundary loop filter indicator representing whether loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture and whether loop filtering is enabled across one or more continuous boundaries in the frame-packed picture—the location of a boundary in the frame-packed picture where loop filtering is to be disabled—based on the cross boundary loop filter indicator indicating that loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture and based on the acquired plurality of boundary location indications—the loop An apparatus for decoding 360-degree video, configured to determine that the position of the boundary in the frame-packed picture to be disabled includes discontinuous boundaries; and to disable loop filtering across the determined position of the boundary in the frame-packed picture. Claim 11 A device for decoding 360-degree video according to claim 10, wherein the boundary in the frame-packed picture comprises one or more of a face boundary, a tile boundary, or a slice boundary. Claim 12 An apparatus for decoding 360-degree video, wherein, in paragraph 10, the boundary cross loop filter indicator is a cross boundary loop filter flag configured to indicate whether the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture. Claim 13 In claim 10, the above boundary cross-loop filter indicator is obtained from a picture parameter set (PPS), a device for decoding 360-degree video. Claim 14 An apparatus for decoding 360-degree video, wherein, in paragraph 10, the processor is configured to: identify the plurality of boundaries in the frame-packed picture based on the plurality of boundary location indicators; and determine the location of the boundary where the loop filtering is disabled from the plurality of identified boundaries in the frame-packed picture. Claim 15 An apparatus for decoding 360-degree video, wherein, in paragraph 10, the loop filtering is associated with one or more of a sample adaptive offset (SAO) filter, a deblocking filter, or an adaptive loop filter (ALF). Claim 16 An apparatus for decoding 360-degree video, wherein, in paragraph 10, the processor is configured to: identify the plurality of boundaries in the frame-packed picture based on the plurality of boundary location indicators; determine the location of the boundary where the loop filtering is enabled—the location of the boundary where the loop filtering is enabled includes the continuous boundary—based on the boundary cross-loop filter indicator indicating that the loop filtering is enabled across the one or more continuous boundaries in the frame-packed picture and based on the plurality of boundaries in the frame-packed picture; and apply the loop filtering across the determined location of the boundary where the loop filtering is enabled. Claim 17 A device for decoding 360-degree video, wherein, in paragraph 10, the frame-packed picture comprises a horizontal boundary and a vertical boundary associated with the frame-packed picture. Claim 18 delete Claim 19 A computer-readable storage medium comprising instructions for decoding 360-degree video, wherein the instructions cause a processor to perform the method of any one of claims 1 through 8. Claim 20 A method for video encoding a 360-degree video comprises: a step of acquiring a frame-packed picture from video data; a step of acquiring a plurality of boundary location indications—the plurality of boundary location indications representing corresponding locations of a plurality of boundaries in the frame-packed picture—and a step of determining, based on the acquired plurality of boundary location indications, a location of a boundary in the frame-packed picture to which loop filtering is disabled or enabled—the location of the boundary in the frame-packed picture to which loop filtering is disabled includes discontinuous boundaries, and the location of the boundary in the frame-packed picture to which loop filtering is enabled includes continuous boundaries. A method for video encoding a 360-degree video, comprising the step of including a boundary cross-loop filter indicator in the video data based on the determined position of the boundary in the frame-packed picture—the boundary cross-loop filter indicator indicates that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture or that the loop filtering is enabled across one or more continuous boundaries in the frame-packed picture. Claim 21 A method for video encoding a 360-degree video, wherein, in paragraph 20, the boundary cross-loop filter indicator is obtained from a picture parameter set (PPS), and the loop filtering is associated with one or more of a sample adaptive offset (SAO) filter, a deblocking filter, or an adaptive loop filter (ALF). Claim 22 A method for video encoding a 360-degree video, wherein, in paragraph 20, the boundary in the frame-packed picture comprises one or more of a face boundary, a tile boundary, or a slice boundary, and the boundary cross-loop filter indicator is a boundary cross-loop filter flag configured to indicate that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture or that the loop filtering is enabled across one or more continuous boundaries in the frame-packed picture. Claim 23 delete Claim 24 An apparatus for video encoding 360-degree video comprises a processor, wherein the processor comprises: acquiring a frame-packed picture from video data; acquiring a plurality of boundary position indications—the plurality of boundary position indications representing corresponding positions of a plurality of boundaries in the frame-packed picture—and, based on the acquired plurality of boundary position indications, determining a boundary position in the frame-packed picture to which loop filtering is disabled or enabled—wherein the boundary position in the frame-packed picture to which loop filtering is disabled includes discontinuous boundaries, and wherein the boundary position in the frame-packed picture to which loop filtering is enabled includes continuous boundaries; An apparatus for video encoding a 360-degree video, configured to include, based on the determined position of the boundary in the frame-packed picture, a boundary cross-loop filter indicator in the video data—the boundary cross-loop filter indicator indicates that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture or that the loop filtering is enabled across one or more continuous boundaries in the frame-packed picture. Claim 25 An apparatus for video encoding 360-degree video, wherein, in paragraph 24, the boundary cross-loop filter indicator is obtained from a picture parameter set (PPS), and the loop filtering is associated with one or more of a sample adaptive offset (SAO) filter, a deblocking filter, or an adaptive loop filter (ALF). Claim 26 An apparatus for video encoding a 360-degree video, wherein, in paragraph 24, the boundary in the frame-packed picture comprises one or more of a face boundary, a tile boundary, or a slice boundary, and the boundary cross-loop filter indicator is a boundary cross-loop filter flag configured to indicate that the loop filtering is disabled across one or more discontinuous boundaries in the frame-packed picture or that the loop filtering is enabled across one or more continuous boundaries in the frame-packed picture. Claim 27 delete Claim 28 A computer-readable storage medium comprising instructions for encoding a 360-degree video, wherein the instructions cause a processor to perform the method of any one of claims 20 to 22.