Capture and playback of 360-degree videos
The system addresses the challenges of capturing and playing back 360-degree videos by merging and encoding them into multiple formats, enabling efficient rendering and user-controlled viewing angles with improved compression efficiency.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
- Filing Date
- 2017-09-29
- Publication Date
- 2026-04-30
AI Technical Summary
Existing systems face challenges in effectively capturing, processing, and playing back 360-degree videos, particularly in merging and rendering these videos efficiently across different projection formats.
A system that includes a video capture device configured to capture 360-degree videos, a stitching device to merge them into multiple projection formats, an encoding device to encode the stitched video with signaling for different formats, and a rendering device to render the video using user-selected or suggested viewing angles, along with pre-calculated coordinate transformations and unrestricted motion compensation to minimize spatial discontinuities.
Enhances the merging and rendering of 360-degree videos, improving compression efficiency and user control over viewing angles, while reducing rendering complexity and memory bandwidth requirements.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present disclosure relates to the capture and playback of videos, or the capture, processing and playback of 360-degree videos, in particular a system with a video capture device.
[0002] 360-degree videos, also known as surround videos, full-sphere videos, and / or panoramic videos, are video recordings of a panorama from the real world, where the view in every direction is recorded simultaneously using an omnidirectional camera or a collection of cameras. During playback, the viewer controls the angles of the field of view (FOV) and the viewing direction (a form of virtual reality).
[0003] Document US 2004 / 0247173 A1 describes a device for capturing videos over 360 degrees.
[0004] Document US 2010 / 0157016 A1 discloses a multi-view camera system for video conferencing that uses scalable video encoding.
[0005] The publication EP 2 490 179 A1 describes a system for video transmission in which a panoramic image is encoded with low quality, and certain sections thereof with higher quality.
[0006] The invention aims to provide a novel system with a video capture device that can capture 360-degree videos, in particular a novel system with a video capture device in which the merging of the captured 360-degree video can be improved.
[0007] The invention achieves this or further objectives through the subject matter of the independent claim.
[0008] Advantageous embodiments of the invention are specified in particular in the dependent claims.
[0009] The assembly device is expediently configured for the following: Calculating a normalized projection plane size using field-of-view angles; Calculating a coordinate in a normalized rendering coordinate system from a coordinate system of an output rendering image using the normalized projection plane size; Mapping the coordinate onto a viewing coordinate system using the normalized projection plane size; Converting the coordinate from the observation coordinate system to a recording coordinate system using a coordinate transformation matrix; Converting the coordinate from the acquisition coordinate system to the intermediate coordinate system using the coordinate transformation matrix; Converting the coordinate from the intermediate coordinate system into a normalized projection system; and Mapping the coordinate from the normalized projection system to the coordinate system for input images.
[0010] It is advantageous to pre-calculate the coordinate transformation matrix using viewing direction angles and global rotation angles.
[0011] Conveniently, the global rotation angles are signaled in a message contained in the 360-degree video bitstream.
[0012] The coding device is also expediently configured for the following: Encoding the stitched 360-degree video into a multitude of view sequences, with each of the multitude of view sequences corresponding to a different view region of the 360-degree video. It is advisable that at least two of the numerous view sequences are encoded in a different format for the projection layout.
[0013] The system also conveniently features the following: a rendering device configured to receive the multitude of view sequences as input and to render each of the multitude of view sequences using a rendering control input.
[0014] Advantageously, the rendering device is further configured to select at least one of the multitude of view sequences for rendering and to exclude at least one of the multitude of view sequences from rendering.
[0015] Advantageously, the encoding device is further configured to include signaling for unrestricted motion compensation in the 360-degree video bitstream to indicate one or more pixels of a view that lie outside an image boundary of the view.
[0016] It is advantageous to place the signaling for unrestricted motion compensation in a sequence header or in an image header of the 360-degree video bitstream.
[0017] It is advantageous to define a relationship between the intermediate coordinate system and the coordinate system for input images by means of counterclockwise rotation angles about one or more axes.
[0018] The assembly device is also conveniently configured for the following: Converting an input projection format of the 360-degree video into an output projection format that differs from the input projection format.
[0019] According to one form of appearance, a system comprises the following: a video capture device configured to capture 360-degree video; a device for merging, configured to merge the captured 360-degree video into at least two different projection formats using a projection format decision; and a coding device configured for the following: Encoding the stitched 360-degree video into a 360-degree video bitstream, wherein the 360-degree video bitstream includes signaling that indicates the at least two different projection formats; and Preparing the 360-degree video bitstream to be played back for transmission.
[0020] The projection format decision is expediently made on the basis of raw data statistics associated with the 360-degree video and encoding statistics from the encoding device.
[0021] The system also conveniently features the following: a decoding device configured to receive the 360-degree video bitstream as input and to decode the projection format information from the signaling in the 360-degree video bitstream, wherein the projection format information specifies the at least two different projection formats.
[0022] The system also conveniently features the following: a rendering device configured to receive the 360-degree video bitstream from the decoding device as input and to render view sequences from the 360-degree video bitstream using at least two different projection formats.
[0023] According to one form of appearance, a system comprises the following: a video capture device configured to capture 360-degree videos; a device for piecing together the captured 360-degree video with a variety of viewing angles; an encoding device configured to encode the stitched 360-degree video into a 360-degree video bitstream; a decoding device configured to decode the 360-degree video bitstream associated with the multitude of viewing angles; and a rendering device configured to render the decoded bitstream using one or more suggested viewing angles from the multitude of possible viewing angles.
[0024] The decoding device is also conveniently configured for the following: Extracting one or more suggested viewing angles from the 360-degree video bitstream.
[0025] The rendering device is also conveniently configured for the following: Receiving one or more user-selected viewing angles as input; and Making a selection between one or more suggested viewing angles and one or more user-selected viewing angles to render the decoded bitstream.
[0026] The rendering device is also conveniently configured for the following: Rendering one or more view sequences of the decoded bitstream with one or more user-selected viewing angles, wherein one or more view sequences are switched back from one or more user-selected viewing angles to one or more suggested viewing angles after a predetermined period without user activity. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Certain features of the claimed technology are set forth in the accompanying claims. However, for illustrative purposes, one or more implementations of the claimed technology are shown in the following figures. Fig. Figure 1 illustrates an exemplary network environment in which the capture and playback of a 360-degree video can be implemented according to one or more implementations. Fig. Figure 2 schematically illustrates an example of an equirectangular projection format. Fig.Figure 3 schematically illustrates an example of an equirectangular projection using a map of the Earth. Fig. Figure 4 schematically illustrates an example of a 360-degree video with equirectangular projection. Fig. Figure 5 schematically illustrates examples of 360-degree images with equirectangular projection layout. Fig. Figure 6 schematically illustrates an exemplary definition of a six-sided cube. Fig. Figure 7 schematically illustrates an example of a cube projection format. Fig. Figure 8 schematically illustrates examples of 360-degree images in the cube projection layout. Fig. Figure 9 schematically illustrates an example of a normalized projection plane size determined using field of view angles. Fig. Figure 10 schematically illustrates an example of viewing direction angles. Fig.Figure 11 illustrates a schematic diagram of a coordinate mapping between an output rendering image and a 360-degree input video image. Fig. Figure 12 schematically illustrates an example of mapping a point in the normalized rendering coordinate system to the normalized projection coordinate system using the equirectangular projection format. Fig. Figure 13 schematically illustrates an example of mapping a point in the normalized rendering coordinate system to the normalized projection coordinate system using the cube projection format. Fig. Figure 14 schematically illustrates an example of a two-dimensional layout 1400 from samples of a 360-degree input video image, which was projected to render a 360-degree video. Fig.Figure 15 schematically illustrates an example of global rotation angles between the acquisition coordinate system and the coordinate system for 360-degree video projection. Fig. Figure 16 schematically illustrates an example of an alternative 360-degree view projection in the equirectangular format. Fig. Figure 17 illustrates a schematic diagram of an example of a coordinate mapping process modified with the coordinate system for 360-degree video projection. Fig. Figure 18 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos using a layout format consisting of six views. Fig. Figure 19 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos using a layout format consisting of two views. Fig.Figure 20 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos using a two-view layout format with a view sequence for rendering. Fig. Figure 21 schematically illustrates examples of multiple layouts for the cube projection format. Fig. Figure 22 schematically illustrates an example of unrestricted movement compensation. Fig. Figure 23 schematically illustrates examples of several layouts for 360-degree video projection formats. Fig. Figure 24 schematically illustrates an example of extended unrestricted motion compensation. Fig. Figure 25 illustrates a schematic diagram of another example of a system for capturing and playing back 360-degree videos. Fig. Figure 26 schematically illustrates an example of a cube projection format. Fig. Figure 27 schematically illustrates an example of a fisheye projection format. Fig. Figure 28 schematically illustrates an example of an icosahedral projection format. Fig. Figure 29 illustrates a schematic diagram of an example of a 360-degree video in several projection formats. Fig. Figure 30 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos with an adaptive projection format. Fig. Figure 31 illustrates a schematic diagram of another example of a system for capturing and playing back 360-degree videos with an adaptive projection format. Fig. Figure 32 illustrates a schematic diagram of an example of a projection format decision. Fig. Figure 33 schematically illustrates an exemplary project format transition without interprediction beyond the boundaries of the projection format transition. Fig. Figure 34 schematically illustrates an exemplary project format transition with interprediction beyond the boundaries of the projection format transition. Fig. Figure 35 illustrates an exemplary network environment 3500 in which proposed views within a 360-degree video can be implemented according to one or more implementations. Fig. Figure 36 schematically illustrates examples of equirectangular and cube projections. Fig. Figure 37 schematically illustrates an example of a 360-degree video rendering. Fig. Figure 38 illustrates a schematic diagram for extraction and rendering with suggested viewing angles. Fig. Figure 39 schematically illustrates an electronic system with which one or more implementations of the claimed technology can be implemented.
[0028] The accompanying annex, which was included to provide a better understanding of the claimed technology and is incorporated into this patent application and forms a part thereof, illustrates manifestations of the claimed technology and, together with the description, serves to explain the principles of the claimed technology. DETAILED DESCRIPTION
[0029] The description set forth below is intended to represent various configurations of the claimed technology and is not meant to represent the only configurations in which the claimed technology can be implemented in practice. The accompanying drawings are incorporated into this document and form part of the detailed description. The detailed description contains specific details intended to facilitate a better understanding of the claimed technology. However, it is clear and obvious to those skilled in the art that the claimed technology is not limited to the specific details set forth in this document and that it can be implemented in practice using one or more of these implementations.In one or more cases, generally known structures and components are shown in the form of block diagrams to prevent the concepts of the claimed technology from becoming incomprehensible.
[0030] In a 360-degree video capture and playback system, 360-degree videos can be captured, stitched, encoded, stored or transmitted, decoded, rendered, and played back. In one or more implementations, a stitching device can be configured to stitch the 360-degree video using an intermediate coordinate system between an input image coordinate system and a capture coordinate system. In one or more implementations, the stitching device can be configured to stitch the 360-degree video into at least two different projection formats using a projection format decision, and an encoding device can be configured to encode the stitched 360-degree video with a signal indicating the at least two different projection formats.In one or more implementations, the merging device can be configured to merge the 360-degree video with multiple viewing angles, and a rendering device can be configured to render the decoded bitstream using one or more suggested viewing angles.
[0031] In the claimed system, a system message containing the presets for the (recommended) viewing angles, field of view angles, and / or rendering image size can be displayed along with the content of the 360-degree video. A 360-degree video playback device (not shown) can leave the rendering image size unchanged but intentionally reduce the active rendering area to decrease rendering complexity and memory bandwidth requirements. The 360-degree video playback device can save the 360-degree video rendering settings (for example, field of view angles, viewing angles, rendering image size, etc.) immediately before playback stops or a switch to another program channel, so that the saved rendering settings can be used when playback resumes on the same channel.The 360-degree video playback device can include a preview mode in which the viewing angles can be automatically changed every N frames to make it easier for viewers to select their desired viewing direction. The 360-degree video capture and playback device can calculate the projection image during processing (for example, block by block) to save memory bandwidth. In this case, the projection image cannot be loaded from the chip's external memory. With the claimed system, different views can be assigned different information to ensure fidelity.
[0032] Fig.Figure 1 illustrates an exemplary network environment 100 in which the capture and playback of 360-degree videos can be implemented according to one or more implementations. Not all components shown may be used; however, one or more implementations may include additional components not shown in the figure. Variations in the arrangement and type of components are possible without derogating from the essence or scope of protection of the claims set forth in this document. Additional components, different components, or fewer components may be provided.
[0033] The exemplary network environment 100 comprises a 360-degree video capture device 102, a 360-degree video stitching device 104, a video encoding device 106, a transmission link or storage media, a video decoding device 108, and a 360-degree video rendering device 110. In one or more implementations, one or more of the devices 102, 104, 106, 108, and 110 may be combined in the same physical device. For example, the 360-degree video capture device 102, the 360-degree video stitching device 104, and the video encoding device 106 may be combined in a single device, and the video decoding device 108 and the 360-degree video rendering device 110 may be combined in a single device.In some manifestations, the network environment 100 may include a storage device 114 in which the encoded 360-degree video (such as on DVDs, Blu-ray, in digital video recordings in the cloud or in a gateway / set-top box, etc.) is stored and then played back on a display device (for example, 112).
[0034] The network environment 100 may further include a device for converting the projection format for 360-degree videos (not shown), which can perform a conversion of the projection format for 360-degree videos before video encoding by the video encoding device 106 and / or after video decoding by the video decoding device 108. The network environment 100 may also include a device for converting the projection format for 360-degree videos (not shown) that is inserted between the video decoding device 108 and the 360-degree video rendering device 110. In one or more implementations, the video encoding device 106 may be communicatively coupled to the video decoding device 108 via a transmission link, such as a network.
[0035] In the claimed system, the 360-degree video stitching device 104 can use an additional coordinate system that allows more flexibility on the 360-degree video capture side when the captured 360-degree video is projected onto a coordinate system for 2D input images for storage or transmission. The 360-degree video stitching device 104 can also support multiple projection formats for storing, compressing, transmitting, decoding, rendering, etc., 360-degree videos. For example, the video stitching device 104 can remove overlapping areas captured by a camera mount and output, for example, six view sequences, each covering a 90° × 90° viewport.The device for converting the projection format for 360-degree videos (not shown) can convert an input projection format for a 360-degree video (for example, the cube projection format) into an output projection format for a 360-degree video (for example, the equirectangular format).
[0036] The Video Coding Device 106 can minimize spatial discontinuities (for example, the number of side boundaries) in the composite image to achieve better spatial prediction and thus improved compression efficiency during video compression. For cube projection, for example, a preferred layout should have a minimized number of side boundaries, such as four, within a composite 360-degree video image. To achieve better compression efficiency, the Video Coding Device 106 can implement Unrestricted Motion Compensation (UMC).
[0037] In the claimed system, the 360-degree video rendering device 110 can derive a chroma projection image from a luma prediction image. The 360-degree video rendering device 110 can also select the rendering image size to optimize the video's display quality. Furthermore, the 360-degree video rendering device 110 can select the horizontal field of view α along with the vertical field of view β to minimize rendering distortion. The 360-degree video rendering device 110 can also control the field of view to achieve real-time rendering of the 360-degree video within the limits of the available memory bandwidth budget.
[0038] In Fig.1. The 360-degree video is captured using a camera mount and assembled into the equirectangular format. The video is then compressed into any suitable video compression format (for example, MPEG / ITU-T AVC / H.264, HEVC / H.265, VP9, etc.) and transmitted via the transmission link (for example, cable, satellite, terrestrial transmission, internet streaming, etc.). On the receiving end, the video is decoded (for example, 108) and stored in the equirectangular format, then rendered (for example, 110) and displayed (for example, 112) according to the viewing angles and the field of view. In the claimed system, the end users have control over the field of view and the viewing angles to view the video from the desired angles. Coordinate systems
[0039] There are several coordinate systems that are applied to the claimed technology, including but not limited to the following: • (x, y, z) - 3D coordinate system for 360-degree video capture (camera coordinate system) • (x', y', z') - 3D coordinate system for viewing 360-degree videos • (x p , y p ) -Normalized 2D projection coordinate system with x p ∈ [0,0: 1,0] and y p ∈ [0,0: 1,0]. • (X p , Y p ) - Coordinate system for 2D input images with X p ∈ [0: inputPicWidth - 1] and Y p ∈ [0: inputPicHeight - 1], where inputPicWidth x inputPicHeight is the size of the input images of a color component (for example, Y, U, or V). • (x c , y c ) - Normalized 2D rendering coordinate system with x c ∈ [0,0: 1,0] and y c ∈ [0,0: 1,0]. • (X e , Yc ) - Coordinate system for 2D output rendering images with X c ∈ [0: renderingPicWidth - 1] and Y c ∈ [0: renderingPicHeight - 1], where picWidth x picHeight is the output rendering image size of a color component (for example, Y, U, or V). • (x r , y r , z r ) - 3D coordinate system for 360-degree video projection
[0040] Fig. Figure 2 schematically illustrates an example of an equirectangular projection format 200. The equirectangular projection format 200 is a standard method for texture mapping a sphere in computer graphics. It is also known as equidistant cylindrical projection, geographic projection, plate map, or plate carrée. As shown in Fig. As shown in Figure 2, to project a point p(x, y, z) on the surface of a sphere (for example, 202) onto a sampling point p'(x) p , y p) in the normalized projection coordinate system (for example 204) both the geographic longitude ω and the geographic latitude φ for p(x, y, z) are calculated according to equation 1. {ω=arc tant2(x,z)φ=arcsin(yx2+y2+z2) , where ω ∈ [-π: π] and φ∈[−π2:π2] This applies. π is the ratio of a circle's circumference to its diameter, usually approximated as 3.1415926.
[0041] The equirectangular projection format 200 can be defined as in equation 2: {xp=ω2π+0,5yp=−φπ+0.5 where x p ∈ [0,0: 1,0] and y p ∈ [0,0: 1,0] holds. (x p , y p ) is the coordinate in the normalized projection coordinate system.
[0042] Fig.Figure 3 schematically illustrates an example of an equirectangular projection layout 300 using a map of the Earth. In the equirectangular projection layout 300, the image only has a 1:1 mapping along the equator; everywhere else it is stretched. The greatest distortion occurs at the north and south poles of a sphere (for example, 302), where a single point is mapped onto a line of sampling points on the equirectangular projection image (for example, 304), resulting in a large amount of redundant data in the composite 360-degree video using the equirectangular projection layout 300.
[0043] Fig.Figure 4 schematically illustrates an example of a 360-degree video with an equirectangular projection layout (400). To utilize the existing video delivery infrastructure, which employs a single-layer video codec, the 360-degree video footage (for example, 402) captured by multiple cameras at different angles is normally stitched together and assembled into a single video sequence, which is stored in the equirectangular projection layout. As shown in Fig.As shown in Figure 4, in the equirectangular projection layout 400, the left, front, and right video footage of the 360-degree video is projected into the center of the image; the rear video footage is divided equally and placed on the right and left sides of the image; the top and bottom video footage are placed in the upper and lower parts of the image, respectively (for example, Figure 404). All video footage is stretched, with the top and bottom footage being stretched the most. Fig. Figure 5 schematically illustrates examples of 360-degree video images with equirectangular projection layout. Cube projection
[0044] Fig. Figure 6 schematically illustrates an exemplary definition of a six-sided cube 600. Another common projection format for storing the 360-degree view is to project the video footage onto the sides of a cube. As in Fig. Figure 6 shows the six sides of a cube labeled Front, Back, Left, Right, Top and Bottom.
[0045] Fig. Figure 7 schematically illustrates an example of a 700 cube projection format. Fig. 7 includes the cube projection format 700, which maps a surface point p(x, y, z) of a sphere onto one of six cube faces (for example, 702), where both the ID of the cube face and the coordinate (x) p , y p ) in the normalized coordinate system for cube projection (for example, 704).
[0046] Fig. Figure 8 schematically illustrates examples of 360-degree video images with cube projection layout 800. The projection rule for cube projection is described in Table 1, which provides pseudo-code for mapping a surface point p(x, y, z) of a sphere onto a face of a cube. Field of view and viewing direction angle
[0047] To display a 360-degree video, a portion of each 360-degree video frame must be projected and rendered. The field-of-view angles define how large the displayed portion of a 360-degree video frame is, while the viewing direction angles define which part of the 360-degree video frame is shown.
[0048] To display a 360-degree video, imagine that the video is projected onto the surface of a unit sphere. A viewer sitting at the center of the sphere can view a rectangular screen, and the screen has four corners located on the surface of the sphere. Here, (x', y', z') is called the 360-degree view viewing coordinate system, and (x c , y c ) is the normalized rendering coordinate system.
[0049] Fig.Figure 9 schematically illustrates an example of a normalized projection plane size of 900, determined using field-of-view angles. As in Fig. As shown in Figure 9, in the viewing coordinate system (x',y',z'), the center of the projection plane (that is, the rectangular screen) lies on the z'-axis and is parallel to the x'y'-plane. Therefore, the size of the projection plane w × h and its distance to the center of the sphere d can be calculated as follows: {w=2tata2+tb2+1h=2tbta2+tb2+1d=1ta2+tb2+1 , where ta=tan(α2) and tb=tan(β2) and α ∈ (0: π) specify the horizontal field of view angle and β ∈ (0: π] the vertical field of view angle.
[0050] Fig.Figure 10 schematically illustrates an example for a viewing direction angle of 1000°. The viewing direction is defined by the rotation angle of the 3D viewing coordinate system (x', y', z') relative to the 3D acquisition coordinate system (x, y, z). As shown in Fig. As shown in Figure 10, the viewing direction is determined by the clockwise rotation angle θ around the y-axis (for example, 1002, yaw angle), the counterclockwise rotation angle γ around the x-axis (for example, 1004, pitch angle), and the counterclockwise rotation angle ε around the z-axis (for example, 1006, roll angle).
[0051] The coordinate mapping between the coordinate systems (x,y,z) and (x',y',z') is defined as follows: [xyz]=[cos εsin ε0−sin εcos ε0001][cos θ0sin θ010−sin θ0cos θ][1000cos γsin γ0−sin γcos γ][x'y'z']
[0052] Therefore: [xyz]=[cos ε cos θ−cos ε sin θ sin γ+sin ε cos γcos ε sin θ cos γ+sin ε sin γ−sin ε cos θsin ε sin θ sin γ+cos ε cos γ−sin ε sin θ cos γ+cos ε sin γ−sin θ−cos θ sin γcos θ cos γ][x'y'z']
[0053] Fig. Figure 11 illustrates a schematic diagram of a coordinate mapping 1100 between an output rendering image and an input image. Using the field of view and viewing direction angles defined above, the coordinate mapping between the coordinate system for output rendering images (X) can be determined. c , Y c ) (that is, the rendering image for display) and the coordinate system for input images (X p , Y p ) (that is, the input 360-degree video image). At a given sampling point (X) c , Y c ) in the rendering image, the coordinate of the corresponding sampling point (X) can be displayed. p , Y p ) in the input image, as in Fig.11 shown, can be derived by means of the following steps: • Calculating the normalized projection plane size and the distance to the center of the sphere based on the field-of-view angles (α, β) (i.e., Equation 3); calculating the coordinate transformation matrix between the viewing and acquisition coordinate systems based on the viewing direction angles (ε, θ, γ) (i.e., Equation 4). • Normalizing (X c , Y c ) based on the rendering image size and the normalized projection plane size. • Mapping the coordinate (X) c , y c ) in the normalized rendering coordinate system to the 3D viewing coordinate system (x', y', z'). • Convert the coordinate into the 3D acquisition coordinate system (x,y,z). • Differentiating the coordinate (x p , y p ) in the normalized projection coordinate system. • Converting the derived coordinate into an integer position in the input image based on the size of the input image and the format for the projection layout.
[0054] Fig. Figure 12 schematically illustrates an example of mapping a point in the normalized rendering coordinate system (for example, p(x)). c , y c )) to the normalized projection coordinate system (for example, p'(x) p , y p )) using the equirectangular projection format 1200.
[0055] In one or more implementations, the projection is performed starting from the equirectangular input format. For example, if the input image is in the equirectangular projection format, the following steps can be applied to determine a sampling point (X). c , Y c ) in the rendering image at a sampling point (X) p , Y p) to be displayed in the input image. • Calculating the projection plane size of a normalized display based on the field of view angles: {w=2tata2+tb2+1h=2tbta2+tb2+1d=1ta2+tb2+1, where the following applies: ta=tan(α2) and tb=tan(β2) • Mapping of (X e , Y c ) to the normalized rendering coordinate system: {xc=XcwrenderinaPicWidthyc=YchrenderingPichHeight • Calculating the coordinate of p(x) c , y c ) in the coordinate system (x',y',z') {x'=xc−w2y'=−yc+h2 z'=d • Converting the coordinate (x',y',z') into the acquisition coordinate system (x, y, z) based on the viewing direction angles: [xyz]=[cos ε cos θ−cos ε sin θ sin γ+sin ε cos γcos ε sin θ cos γ+sin ε sin γ−sin ε cos θsin ε sin θ sin γ+cos ε cos γ−sin ε sin θ cos γ+cos ε sin γ−sin θ−cos θ sin γcos θ cos γ][x'y'z'] • Projecting p(x,y,z) onto the normalized projection coordinate system p'(x) p ,y p ): {xp=arc tant2(x,z)2π+0.5yp=−arcsin(yx2+y2+z2)+0.5π • Mapping p'(x p ,y p ) onto the (equirectangular) input image coordinate system (X p , Y p ) {Xp=(int)(xp∗inputPicWidth)Yp=(int)(yp∗inputPicHeight) , where the following applies: • α, β are field of view angles and ε, θ, γ are viewing direction angles. • renderingPicWidth x renderingPicHeight is the rendering image size. • inputPicWidth x inputPicHeight is the size of the input image (in the equirectangular projection format).
[0056] Fig. Figure 13 schematically illustrates an example of mapping a point in the normalized rendering coordinate system (for example, p(x)). c , y c)) to the normalized projection coordinate system (for example, p'(x) p , y p )) using the cube projection format 1300.
[0057] In one or more implementations, the projection is performed starting from the input format for cube projection. For example, if the input image is in cube projection format, the following similar steps can be applied to determine a sampling point (X). c , Y c ) in the rendering image at a sampling point (X) p , Y p ) to be displayed in the input image. • Calculating the projection plane size of a normalized display based on the field of view angles: {w=2tata2+tb2+1h=2tbta2+tb2+1d=1ta2+tb2+1 where the following applies: ta=tan(α2) and tb=tan(β2) • Mapping of (X c , Y c ) to the normalized rendering coordinate system: {xc=XcwrederingPicWidthyc=YchrendetringPicHeight • Calculating the coordinate of p(x) c , y c ) in the coordinate system (x', y', z'): {x'=xc−w2y'=−yc+h2 z'=d • Converting the coordinate (x',y',z') into the acquisition coordinate system (x, y, z) based on the viewing direction angles: [xyz]=[cos ε cos θ−cos ε sin θ sin γ+sin ε cos γcos ε sin θ cos γ+sin ε sin γ−sin ε cos θsin ε sin θ sin γ+cos ε cos γ−sin ε sin θ cos γ+cos ε sin γ− sin θ−cos θ sin γcos θ cos γ][x'y'z'] • Projecting p(x,y,z) onto the normalized cube coordinate system p'(x) p , y p ) based on the pseudo-code defined in Table 1. • Mapping p'(x p ,y p ) on the input cube coordinate system (X p ,Y p ) (assuming that all sides of the cube have an identical resolution) {Xp=(int)(xp∗inputPicWidth3)+Xoffset[faceID]Yp=(int)(yp∗inputPicHeight2)+Yoffset[faceID] , where the following applies: • α, β are field of view angles and ε, θ, γ are viewing direction angles. • renderingPicWidth x renderingPicHeight is the rendering image size. • inputPicWidth x inputPicHeight is the input image size (in the cube projection format). • f(Xoffset[faceID],Yoffset[faceID]) | faceID (Side ID) = Front, Back, Left, Right, Top and Down} gives the coordinate offsets of a cube face in the input cube projection coordinate system. {Xoffset[6]={0,inputPicWidth3,2inputPicWidth3, 0, inputPicWidth3,2inputPicWidth3}Yoffset[6]={0, 0, 0, inputPicHeight2,inputPicHeight2,inputPicHeight2}
[0058] For the in Fig.In the 13 illustrated cube projection layout, the faceID uses the following sequence: Front, Back, Left, Right, Top and then Down to access the array of coordinate offsets. Rendering the sampling points for the display
[0059] When projecting 360-degree videos for display, multiple sampling points on an input image for a 360-degree video (for example, in the equirectangular format or in the cube projection format) can be placed at the same integer position (X). e , Y c ) are projected into the rendering image. To achieve smooth rendering, not only the integer pixel positions but also their subpixel positions are projected into the rendering image in order to find corresponding sampling points in the input image.
[0060] Fig.Figure 14 schematically illustrates an example of a two-dimensional layout 1400 from samples of a 360-degree input video image, which was projected to render a 360-degree video. If, as in Fig. Figure 14 shows the accuracy of the projection. 1n Subpixels in the horizontal direction and 1m If the subpixel in the vertical direction is, the sample value of the rendering image at position (X) can be c , Y c ) are rendered as follows: renderingImg[Xc,Yc]=∑i=0m−1∑j=0n−1inputImg[mapping_func(XC+j+0.5n,YC+i+0.5m)]+mn2mn , where the following applies: • (X p , Y p ) = mapping_func(X c ,Y c ) is the coordinate mapping function from the rendering image to the input image of the 360-degree video, as defined in the sections above (for example, using equirectangular projection or in the cube projection format). • inputlmg[X p, Y p ] is the sample value at position (X p , Y p ) in the input image. • renderinglmg [X c , Y c ] is the sample value at position (X c , Y c ) in the output rendering image.
[0061] Instead of the coordinate mapping between the coordinate system for output rendering images (X c , Y c ) and the coordinate system for input images (X p , Y p To avoid having to calculate the coordinate mapping during operation, the coordinate mapping can also be pre-calculated and saved as a projection for the entire rendered image. Since the viewing direction and field of view angles may not change from image to image, the pre-calculated projection can be used for rendering multiple images together.
[0062] It is projectMap[n * X c + j,m * Y c+ i] the pre-calculated projection mapping with X c = 0.1, renderingPicWidth - 1, Y c = 0,1, ..., renderingPicHeight - 1,j = 0,1, ..., n - 1 and i = 0,1, ..., m - 1. Each entry of the projection mapping stores the pre-calculated coordinate value (X p , Y p ) of the coordinate system for input images for a subpixel position (Xc+j+0,5n,Yc+i+0,5m) in the rendering image. The rendering can be expressed as follows: renderingImg[Xc,Yc]=∑i=0m−1∑j=0n−1inputImg[projectMap[n∗Xc+j,m∗YC+i]]+mn2mn
[0063] An image can contain multiple color components, such as YUV, YCbCr, and RGB. The rendering process described above can be applied to individual color components independently.
[0064] Fig.Figure 15 schematically illustrates an example of a global rotation angle of 1500 between the acquisition coordinate system and the coordinate system for 360-degree video projection. Instead of directly projecting the 360-degree video from the 3D coordinate system for 360-degree video acquisition (x, y, z) onto the coordinate system for input images (X p , Y p To project ) we introduce an additional coordinate system in the process for capturing and stitching 360-degree videos, which serves as the 3D coordinate system for 360-degree video projection (x r , y r , z r ) is denoted. The relationship between (x r , y r , z r ) and (x, y, z) is determined by counterclockwise rotation angles (Kr,φr,ωr) specified around the z-, y- and x-axes, as in Fig. Figure 15 shows the coordinate transformation between the two systems as follows: [xryrzr]=[cos Krsin Kr0−sin Krcos Kr0001][cos φr0−sin φr010sin φr0cos φr][1000cos ωrsin ωr0−sin ωrcos ωr][xyz]
[0065] In one or more implementations, equation 8 can be transformed into equation 9: [xryrzr]=[cos φr cos Krcos ωr sin Kr+sin ωr sin φr cos Krsin ωr sin Kr−cos ωr sin φr cos Kr−cos φr sin Krcos ωr cos Kr−sin ωr sin φr sin Krsin ωr cos Kr+cos ωr sin φr sin Krsin φr− sin ωr cos φrcos ωr cos φr][xyz]
[0066] Fig. Figure 16 schematically illustrates an example of an alternative 360-degree view projection 1600 in the equirectangular projection format. The additional coordinate system (x r , y r , z rThis allows more flexibility on the 360-degree video capture side when the captured 360-degree video is projected onto the coordinate system for 2D input images for storage or transfer. Taking the equirectangular projection format as an example, it may sometimes be desirable to project the front or back view onto the South or North Pole, as in Fig. 16 shown, instead of the Top or Bottom view, as in Fig. 4 shown. Since the equirectangular projection better preserves data around the south and north poles of a sphere, allowing alternative layouts, such as the one in Fig. 16, potentially offer advantages under certain circumstances (such as better compression efficiency).
[0067] Since the 360-degree video data is derived from the 3D coordinate system (x r, y r, z r) instead of (x, y, z) can be projected onto the coordinate system for input images, before projecting the 360-degree video data onto the coordinate system for input images, both in Fig. 2 as well as in Fig. 7 an additional step to convert the coordinate (x, y, z) to (x r , y r , z r ) based on Equation 9 is needed. Accordingly, in Equation 1, Equation 5 and Table 1 (x, y, z) must be replaced by (x r , y r , z r ) will be replaced.
[0068] Fig. Figure 17 illustrates a schematic diagram of an example of a coordinate mapping process modified with the coordinate system for 360-degree video projection. The coordinate mapping from the coordinate system for output rendering images (X) c , Y c ) to the coordinate system for input images (X p , Y p ) can therefore, as in Fig.Figure 17 illustrates how the coordinate transformation matrix can be modified. The calculation of the coordinate transformation matrix takes into account both the viewing direction angles (ε, θ, γ) and the global rotation angles. (Kr,φr,ωr) modified. By cascading equation 9 and equation 4, the coordinate can be directly converted from the viewing coordinate system (x',y',z') to the 3D coordinate system for 360-degree video projection (x) using the following equation. r , y r , z r ) be converted: [xryrzr]=[cos φr cos Krcos ωr sin Kr+sin ωr sin φr cos Kr sin ωr sin Kr−cos ωr sin φr cos Kr−cos φr sin Kr cos ωr cos Kr−sin ωr sin φr sin Krsin ωr cos Kr+cos ωr sin φr sin Krsin φr−sin ωr cos φrcos ωr cos φr] [cos ε cos θ−cos ε sin θ sin γ+sin ε cos γcos ε sin θ cos γ+sin ε sin γ−sin ε cos θ sin ε sin θ sin γ+cos ε cos γ−sin ε sin θ cos γ+cos ε sin γ−sin θ−cos θ sin γcos θ cos γ][x'y'z']
[0069] In one or more implementations, the 3x3 coordinate transformation matrix can be pre-calculated using the equation mentioned above.
[0070] Fig. Figure 17 illustrates a schematic diagram of an example of a coordinate mapping process modified with the coordinate system for 360-degree video projection (for example, (x r , y r , z r )). In Fig. 17 can determine the global rotation angles (Kr,φr,ωr) The system can signal this through any suitable means. For example, by defining a SEI (Supplemental Enhancement Information) message that is carried in the elementary video bitstream, such as AVC-SEI messages or HEVC-SEI messages. A standard view can also be specified and carried.
[0071] Fig.Figure 18 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos using a layout format consisting of six views. Instead of assembling a 360-degree video into a single view sequence, as in Fig. Using the equirectangular format, a system for capturing and playing back 360-degree videos can support a format with multiple view layouts for storing, compressing, transmitting, decoding, rendering, etc., 360-degree videos. As in Fig.As shown in Figure 18, the 360-degree video stitching block (for example, 104) can simply remove the overlapping area captured by the camera mount and output, for example, six view sequences (that is, Top, Bottom, Front, Back, Left, Right), each covering a 90° × 90° viewport. The sequences may or may not have the same resolution, but they are each compressed (for example, 1802) and decompressed (for example, 1804) separately. The 360-degree video rendering engine (for example, 110) can accept multiple sequences as input and render sequences for display using the rendering control.
[0072] Fig. Figure 19 illustrates a schematic diagram of an example system for capturing and playing back 360-degree videos using a two-view layout format. Fig.For example, the 360-degree video merging block (104) can generate two view sequences. One sequence covers, for example, 360° × 30° of the data from the Top and Bottom views, and the other covers 360° × 120° of the data from the Front, Back, Left, and Right views. The two sequences can have different resolutions and different projection layout formats. For example, the Front+Back+Left+Right view can use the equirectangular projection format, while the Top+Bottom view can use a different projection format. In this particular example, two encoders (1902) and two decoders (1904) are used. The 360-degree video rendering block (110) can render sequences for display by accepting two sequences as input.
[0073] Fig.Figure 20 illustrates a schematic diagram of an example system for capturing and playing back 360-degree videos in 2000, using a two-view layout format with a view sequence for rendering. The multi-view layout format can also provide scalable rendering functionality when rendering 360-degree videos. As shown in Fig. As shown in 20, it can be specified for the 360-degree video that the upper or lower 30° area of the video (for example, 2002, 2004) will not be rendered due to the limited capabilities of the rendering engine (for example, 110) or if the bitstream packets of the Top+Bottom view are lost during transmission.
[0074] Even if the 360-degree video stitching block (for example, 104) generates multiple view sequences, a single video encoder and decoder can still be used in the 360-degree video capture and playback system 2000. For example, if the output of the 360-degree video stitching block (for example, 104) is six 720p@30 view sequences (that is, sequences of 720 pixels each at 30 frames per second), the output can be stitched together into a single 720p@180 sequence (that is, a sequence of 720 pixels at 180 frames per second) and compressed / decompressed using a single codec.Alternatively, the six independent sequences, for example, can be compressed / decompressed using a single video encoder and decoder instance without having to combine them into a single sequence by sharing the available processing resources of each encoder and decoder instance over time.
[0075] Fig.Figure 21 schematically illustrates examples of multiple layouts for the cube projection format: (a) raw cube projection layout; (b) an undesired cube projection layout; and (c) an example of a preferred cube projection layout. Conventional video compression technology exploits spatial correlation within an image. For example, when cube projection combines the projection sequences of the cube's top, bottom, front, back, left, and right faces into a single view sequence, different layouts of the cube faces can result in different compression efficiencies, even with the same 360-degree video data. In one or more implementations, the cube raw layout (a) does not produce discontinuous edge boundaries within the composite image, but compared to the other two layouts, the images are larger, and dummy data is present in the gray area.In one or more implementations, layout (b) is one of the undesirable layouts because it has a maximum number of discontinuous side boundaries (that is, 7) within the composite image. In one or more implementations, layout (c) is one of the desired cube projection layouts because it has a minimum, smaller number of discontinuous side boundaries (4). The side boundaries in layout (c) are... Fig. 22 refers to the horizontal page boundary between left and top and between right and back, and the vertical page boundary between top and bottom and between back and bottom. For some variations, experiments have shown that layout (c) performs roughly 3% better on average than layouts (a) and (b).
[0076] Therefore, in cube projections or other multi-sided projection formats, the layout should generally minimize spatial discontinuities (i.e., the number of discontinuous side boundaries) in the composite image to achieve better spatial prediction and thus better compression efficiency during video compression. For cube projection, for example, a preferred layout has four discontinuous side boundaries within a composite 360-degree video image. Further minimization of spatial discontinuities in the cube projection format can be achieved by introducing side rotations (i.e., 90, 180, 270 degrees) in the layout. In one or more instances, the layout information of the incoming 360-degree video is signaled in a higher-level video system message for compression and rendering.
[0077] Fig.Figure 22 schematically illustrates an example of unrestricted motion compensation. UMC (Unrestricted Motion Compensation) is a technique frequently used in video compression standards to achieve better compression efficiency. As shown in Fig. As shown in Figure 22, in the UMC, the reference block of a prediction unit may extend beyond the image boundaries. For a reference pixel ref [X] located outside the reference image p , Y p The nearest pixel on the image boundary is used. Here, the coordinates (X) are used. p , Y p ) for the reference pixels are determined based on the position of the current PU (Prediction Unit) and the motion vector; they can extend beyond the image boundaries.
[0078] If {refPic [X cp , Y cp ], X cp = 0.1, ...,inpuPicWidth - 1; Y cpIf the sampling matrix of the reference image is 0,1, ..., inpuPicHeight - 1}, then the UMC (as loading reference blocks) is defined as follows: refBlk[Xp,Yp]=refPic[clip3(0,inputPicWidth−1,Xp),clip3(0,inputPicHeight−1,Yp)] , where the clipping function clip3 (0, a, x) is defined as follows: int clip(0,a,x){if(x<0)return 0; else if(x>a)return a; else return x}
[0079] Fig.Figure 23 schematically illustrates examples of several layouts for 360-degree video projection formats. In some manifestations, the layouts for 360-degree video projection formats include, but are not limited to: (a) equirectangular projection layout; and (b) raw layout for cube projection. However, in a 360-degree video, the left and right image boundaries may belong to the same camera view. Similarly, the top and bottom image boundary lines may be adjacent in physical space, although they are widely separated in the image layouts for 360-degree videos. Fig.Figure 23 shows two examples. In layout example (a), both the left and right image borders belong to the Rear view. Although in layout example (b) the left and right image borders belong to different views (i.e., Left and Rear, respectively), these two columns of image borders are actually physically adjacent during video capture. Therefore, to achieve better compression efficiency for 360-degree videos, it makes sense to allow the reference block to load "all around" when reference pixels are located outside the image borders, rather than padding with the pixels closest to the image border, as defined in the current UMC.
[0080] In one or more implementations, the following syntax can enable the extended UMC at a high level. Table 2: Extended UMC Syntax syntax Bits semantics horizontal_UMC_warparound_enabled_flag 1 "1": Horizontal rotation of the reference block is enabled. "0": If reference pixels are located outside the image boundaries, the reference block is filled with padding. if (horizontal_UMC_warparound_enabled_flag)ΔW Difference in image size in the horizontal direction between the captured image and the coded image vertical_UMC_warparound_enabled_flag 1 "1": Vertical looping of the reference block is enabled. "0": If reference pixels are located outside the image boundaries, the reference block is filled with padding. if (horizontal_UMC_warparound_enabled_flag)ΔH Difference in image size in the vertical direction between the captured image and the coded image
[0081] It should be noted that the difference in image size (ΔW, ΔH) must be signaled because the encoded image size (inputPicWidth x inputPicHeight) must typically be a multiple of 8 or 16 in both directions, while this may not be the case for the captured image. This could result in a difference in image size between the captured and encoded images. The reference block is traversed along the image boundaries of the captured image, but NOT the encoded image.
[0082] In one or more implementations, the high-level syntax of Table 2 can be signaled either in a sequence header or in an image header, or in both, depending on the implementation.
[0083] Using the syntax defined above, the extended UMC can be defined as follows: refBlk[Xp,Yp]=refPic[Xcp,Ycp] , where (X cp , Y cp ) is calculated as follows: Xcp={clip3(0,inputPicWidth−1,Xp) if horizontal_UMC_warparound_enabled_flag=0warparound(0, inputPicWidth−1−ΔH,Yp) Otherwise Ycp={clip3(0,inputPicWidth−1,Yp) if vertical_UMC_warparound_enabled_flag=0warparound3(0, inputPicWidth−1−ΔH,Yp) Otherwise , wobei clip3 () in Gleichung 11 definiert ist und warparound3(0, a, x) wie folgt definiert ist: int wraparound3(0,a,x){while(x<0)x+=a; while(x>a) x−=a; return x;}
[0084] With current video compression standards, motion vectors with a wide margin can extend beyond the frame boundaries. This is why the "while" loop is included in Equation 13. To avoid the "while" loop, next-generation video compression standards may impose limits on how far a reference pixel (during motion compensation) can extend beyond the frame boundaries (for example, 64, 128, or 256 pixels, etc., depending on the maximum encoding block size defined in the next-generation standards). Once such a boundary condition has been established, `warparound3(0, a, x)` can be simplified as follows: int wraparound3(0,a,x){if(x<0)x+=a; if(x>a) x−=a; return x;}
[0085] Fig. Figure 24 schematically illustrates an example of extended unrestricted motion compensation 2400. In Fig.Option 24 enables the horizontal wrapping of the reference block. Instead of filling the part of the reference block that lies outside the left image boundary with image boundary pixels, the 'wrapping' part is loaded along the right (captured) image boundary. Adaptive projection format
[0086] Fig.Figure 25 illustrates an exemplary network environment 2500 in which the capture and playback of 360-degree videos can be implemented according to one or more implementations. Not all components shown may be used; however, one or more implementations may include additional components not shown in the figure. Variations in the arrangement and type of components are possible without derogating from the essence or scope of protection of the claims set forth in this document. Additional components, different components, or fewer components may be provided.
[0087] The exemplary network environment 2500 comprises a 360-degree video capture device 2502, a 360-degree video stitching device 2504, a video encoding device 2506, a video decoding device 2508, and a 360-degree video rendering device 2510. In one or more implementations, one or more of the devices 2502, 2504, 2506, 2508, and 2510 may be combined in the same physical device. For example, the 360-degree video capture device 2502, the 360-degree video joining device 2504 and the video encoding device 2506 can be combined in a single device, and the video decoding device 2508 and the 360-degree video rendering device 2510 can be combined in a single device.
[0088] The network environment 2500 may further include a projection format decision device for 360-degree videos (not shown), which can select the projection format before the video is stitched together by the stitching device 2504 and / or after video encoding by the video encoding device 2506. The network environment 2500 may also include a playback device for 360-degree videos (not shown), which plays back the content of the rendered 360-degree video. In one or more implementations, the video encoding device 2506 may be communicatively coupled to the video decoding device 2508 via a transmission link, such as a network.
[0089] In the claimed system, the 360-degree video assembly device 2504 can utilize the 360-degree video projection format decision device (not shown) on the 360-degree video acquisition / compression side to decide which projection format (for example, ERP (Equirectangular Projection), CMP (Cube Projection), ISP (Icosahedral Projection), etc.) is best suited for a current video segment (i.e., a group of images) or the current image to achieve the best possible compression efficiency. The decision can be based on encoding statistics (such as the bit rate distribution, intra / inter-modes across the segment or across the image, video quality measurements, etc.).), provided by the video coding device 2506, and / or raw data statistics (such as the distribution of spatial activity in relation to raw data, etc.) obtained from raw data of a 360-degree video camera by the 360-degree video acquisition device 2502. Once the projection format for the current segment or image has been selected, the 360-degree video merging device 2504 merges the video into the selected projection format and provides the merged 360-degree video to the video coding device 2506 for compression.
[0090] In the claimed system, the selected projection format and associated projection format parameters (for example, projection format ID, number of pages in the projection layout, page size, page coordinate offsets, page rotation angles, etc.) can be signaled by the video encoding device 2506 in a compressed 360-degree video bitstream via a suitable means, for example, in a SEI (Supplemental Enhancement Information) message, in a sequence header, in an image header, etc. The 360-degree video joining device 2504 can join the 360-degree video into different projection formats selected by the 360-degree video projection format decision device, instead of joining the 360-degree video into a single, fixed projection format (for example, ERP).
[0091] In the claimed system, the 360-degree video stitching device 2504 can use an additional coordinate system that allows more flexibility on the 360-degree video capture side when the captured 360-degree video is projected onto a coordinate system for 2D input images for storage or transmission. The 360-degree video stitching device 2504 can also support multiple projection formats for storing, compressing, transmitting, decoding, rendering, etc., 360-degree videos. For example, the video stitching device 2504 can remove overlapping areas captured by a camera mount and output, for example, six view sequences, each covering a 90° × 90° viewport.
[0092] The 2506 video encoding device can minimize spatial discontinuities (for example, the number of discontinuous side boundaries) in the composite image to achieve better spatial prediction and thus improved compression efficiency during video compression. For cube projection, for example, a preferred layout should have a minimized number of discontinuous side boundaries, such as four, within a composite 360-degree video image. To achieve better compression efficiency, the 2506 video encoding device can implement UMC.
[0093] On the playback page for 360-degree videos, a 360-degree video playback device (not shown) can receive the compressed 360-degree video bitstream and decompress it. The 360-degree video playback device can render the 360-degree video in different projection formats signaled in the 360-degree video bitstream, as opposed to rendering the video in a single, fixed projection format (for example, ERP). In this respect, the rendering of the 360-degree video is controlled not only by the viewing direction and field of view angle, but also by the projection format decoded from the compressed 360-degree video bitstream.
[0094] In the claimed system, a system message containing the presets for the (recommended) viewing direction (i.e., viewing direction angles), the field of view angles, and / or the rendering image size can be displayed along with the content of the 360-degree video. The 360-degree video playback device can leave the rendering image size unchanged but intentionally reduce the active rendering area to decrease rendering complexity and memory bandwidth requirements. The 360-degree video playback device can save the 360-degree video rendering settings (for example, the field of view angles, viewing direction angles, rendering image size, etc.) immediately before playback stops or a switch to another program channel, so that the saved rendering settings can be used when playback resumes on the same channel.The 360-degree video playback device can include a preview mode in which the viewing angles can be automatically changed every N frames to make it easier for viewers to select their desired viewing direction. The 360-degree video capture and playback device can calculate the projection image during processing (for example, block by block) to save memory bandwidth. In this case, the projection image cannot be loaded from the chip's external memory. With the claimed system, different views can be assigned different information to ensure fidelity.
[0095] In Fig.25. The video is captured using a camera mount and assembled into the equirectangular format. The video is then compressed into any suitable video compression format (for example, MPEG / ITU-T AVC / H.264, HEVC / H.265, VP9, etc.) and transmitted via the transmission link (for example, cable, satellite, terrestrial transmission, internet streaming, etc.). At the receiving end, the video is decoded and stored in the equirectangular format, then rendered and displayed according to the viewing angles and field of view. With this system, end users have control over the field of view and viewing angles to view the desired portion of the 360-degree video from their preferred angles.
[0096] With further reference to Fig.2. To utilize the existing video delivery infrastructure, which uses a single-layer video codec, the 360-degree video footage captured by multiple cameras at different angles is normally stitched together and assembled into a single video sequence, which is stored in a specific projection format, such as the widely used equirectangular projection format.
[0097] In addition to the equirectangular projection format (ERP), there are many other projection formats that can display a 360-degree video frame on a rectangular 2D image, such as the cube projection (CMP), the icosahedral projection (ISP), and the fisheye projection (FISHEYE).
[0098] Fig. Figure 26 schematically illustrates an example of a cube projection format. As in Fig.As shown in Figure 26, in CMP the surface of a sphere (for example, 2602) is projected onto six sides of a cube (that is, top, front, right, back, left and bottom) (for example, 2604), with each side covering a field of view of 90 × 90 degrees, and six cube sides are combined to form a single image. Fig. Figure 27 schematically illustrates an example of a fisheye projection format. In fisheye projection, the surface of a sphere (for example, 2702) is projected onto two cycles (for example, 2704, 2706); each cycle covers a field of view of 180 × 180 degrees. Fig. Figure 28 schematically illustrates an example of an icosahedral projection format. In ISP, a sphere is mapped onto a total of 20 triangles, which are then combined to form a single image. Fig. Figure 29 illustrates a schematic diagram of an example of a 360-degree video 2900 in several projection formats. For example, layout (a) shows Fig.29 the ERP projection format (for example 2902), layout (b) of Fig. Figure 29 shows the CMP projection format (for example, 2904), layout (c) by Fig. Figure 29 shows the ISP projection format (for example, 2906), and layout (d) of Fig. Figure 29 shows the FISHEYE projection format (for example, 2908).
[0099] For the same content in a 360-degree video, different projection formats can lead to different compression efficiencies after the video has been compressed, for example, using the MPEG / ITU AVC / H.264 or MPEG / ITU MPEG HEVC / H.265 video compression standards. Table 1 shows the difference in bitrate between ERP and CMP for twelve test sequences of 4K 360-degree videos. The difference in PSNR and bitrate is calculated in the CMP range, where negative numbers indicate better compression efficiency using ERP and positive numbers indicate better efficiency when compressed with CMP. Table 3 illustrates results from experiments regarding the compression efficiency of ERP compared to CMP, for example, using the reference software HM16.9 for MPEG / ITU MPEG HEVC / H.265. Table 3: Compression efficiency: ERP compared to CMP
[0100] As shown in Table 3, although the use of ERP generally leads to better compression efficiency, there are other cases (for example, GT_Sheriff-left, time-lapse_basejump) where the use of CMP yields more positive results. Therefore, it is desirable to change the projection format from time to time, if possible, so that it can be best adapted to the characteristics of the 360-degree video content to achieve the best possible compression efficiency.
[0101] Fig. Figure 30 illustrates a schematic diagram of an example of a system for capturing and playing back 360-degree videos with an adaptive projection format. To maximize compression efficiency for 360-degree videos, an adaptive method capable of compressing 360-degree video into mixed projection formats can be implemented. As shown in Fig.As shown in Figure 30, a projection format decision block (for example, 3002) on the capture / compression side of the 360-degree video can be used to decide which projection format (for example, ERP, CMP, ISP, etc.) is best suited for the current video segment (i.e., a group of frames) or the current image to achieve the best possible compression efficiency. The decision can be based on encoding statistics (such as the bit rate distribution, intra / inter-modes across a segment or image, video quality measurements, etc.) provided by the video encoder (for example, 2506), and / or raw data statistics (such as the distribution of spatial activity relative to raw data, etc.) obtained from the raw data of a 360-degree video camera.Once the projection format for the current segment or image has been selected, the 360-degree video join block (for example, 2504) joins the video to the selected projection format and provides the joined 360-degree video to the video encoder (for example, 2506) for compression.
[0102] The selected projection format and its associated parameters (such as projection format ID, number of pages in the projection layout, page size, page coordinate offsets, page rotation angle, etc.) are signaled in the compressed bitstream by any suitable means, for example, in a SEI (Supplemental Enhancement Information) message, in a sequence header, in an image header, etc. This differs from the above. Fig.In the illustrated system 25, the 360-degree video merging block (for example, 2504) is able to merge the 360-degree video into different projection formats selected by means of the projection format decision block (for example, 3002), instead of merging the video into a single, fixed projection format (for example, ERP).
[0103] On the playback side for 360-degree videos, the receiver (for example, 2508) receives the compressed 360-degree video bitstream and decompresses the bitstream. This differs from the process described in... Fig.In the system illustrated in Figure 25, the rendering block for the 360-degree video (for example, 2510) is capable of rendering 360-degree videos in different projection formats signaled in the bitstream, instead of rendering the video in a single, fixed projection format (for example, ERP). Thus, the rendering of the 360-degree video is controlled not only by the viewing direction and field of view angles, but also by the projection format information decoded from the bitstream.
[0104] Fig. Figure 31 illustrates a schematic diagram of another example of a system for capturing and playing back 360-degree videos with adjustments to the projection format from multiple sources across multiple channels. Fig.31 The 3100 system for capturing and playing back 360-degree videos is able, for example, to perform decoding and rendering of a 360-degree video on four channels in parallel, whereby the inputs for the 360-degree video can come from different sources (for example, 3102-1, 3102-2, 3102-3, 3102-4) and can be in different projection formats (for example, 3104-1, 3104-2, 3104-3, 3104-4) and compression formats (for example, 3106-1, 3106-2, 3106-3, 3106-4) and in different fidelity levels (for example, image resolution, frame rate, bit rate, etc.).
[0105] In one or more implementations, channel 0 may be live video in adaptive projection formats compressed with HEVC. In one or more implementations, channel 1 may be live video in a fixed ERP format compressed with MPEG-2. In one or more implementations, channel 2 may be 360-degree video content in adaptive projection formats compressed with VP9, but pre-stored on a server for streaming. In one or more implementations, channel 3 may be 360-degree video content in a fixed ERP format compressed with H.264, but pre-stored on a server for streaming.
[0106] For video decoding, the decoders (for example, 3108-1, 3108-2, 3108-3, 3108-4) are able to decode videos in different compression formats, and the decoders can be implemented in hardware (HW), using programmable processors (SW), or in a mixture of HW / SW.
[0107] For rendering 360-degree videos, the rendering engines (for example, 3110-1, 3110-2, 3110-3, 3110-4) are able to perform a viewport rendering from the 360-degree input video in different projection formats (for example, ERP, CMP, etc.) based on the viewing direction angle, the field of view angle, the image size of the 360-degree input video, the image size of the output viewport, the global rotation angles, etc.
[0108] Rendering engines can be implemented in hardware (HW), using programmable processors (such as GPUs), or in a hybrid of HW / SW. In some implementations, the same output from a video decoder can be fed into multiple rendering engines, allowing multiple viewports to be rendered for display from the same 360-degree input video.
[0109] Fig. Figure 32 illustrates a schematic diagram of an example of a projection format decision. While the playback of 360-degree videos (decompression plus rendering) is relatively fixed, there are several ways to generate a compressed video stream for a 360-degree video that has the greatest possible compression efficiency by adaptively changing the projection format from time to time. Fig.Figure 32 provides an example of how the projection format decision (for example, Figure 3002) can be made. In this example, 360-degree video content is provided in several projection formats, such as ERP, CMP, and SP (for example, Figure 3202). The videos in the different formats are compressed using the same type of video encoder (such as MPEG / ITU-T HEVC / H.265) (for example, Figure 3204), and the compressed bitstreams are stored (for example, Figure 3206). After encoding a segment or frame in different projection formats, the rate distortion costs (such as PSNR values for the same bit rate, or the bit rate for the same quality, or a combined metric) for the segment or frame in these projection formats can be measured.After the projection format decision (for example, 3002) has been made based on the rate distortion cost for the current segment or frame, the corresponding bitstream in the selected format is chosen (for example, 3208) and merged with the bitstream in the mixed projection format for the segment or frame (for example, 3210). It should be noted that with this system, the projection format can be changed from video segment to segment (group of frames) or from frame to frame.
[0110] Fig.Figure 33 schematically illustrates an exemplary project format transition 3300 without interprediction across the boundaries of the projection format transition. To support 360-degree videos with mixed projection formats using existing video compression standards, such as MPEG / ITU AVC / H.264, MPEG / ITU HEVC / H.265, and Google VP9, a limit must be imposed on how often the projection format can be changed. As shown in Fig. As shown in Figure 33, the projection format transition 3300, such as from ERP to CMP or from CMP to ERP in this example, can only occur at the RAPs (Random Access Points), so that interprediction does not need to extend beyond a boundary for the projection format transition. A RAP can be preceded by an IDR (Instantaneous Decoding Refresh) image or another type of image that provides the functionality for random access. Fig. 33 The projection format can only change from video segment to video segment, not from image to image, unless a segment consists of only one image.
[0111] Fig. Figure 34 schematically illustrates an example project format transition 3400 with interprediction across the boundaries of the projection format transition. The projection format can also change from frame to frame if the interprediction can extend beyond the boundaries of the projection format transition. With the same content in a 360-degree video, different projection formats can lead not only to different content in the image, but also to different image resolutions after the 360-degree video has been stitched together. As shown in Fig.As shown in Figure 34, the conversion of the projection format (for example, 3400) of reference images in the DPB (Decoded Picture Buffer) can be used to support interprediction beyond a projection format transition limit. During projection format conversion, a reference image of a 360-degree video in the DBP can be converted from a format (for example, ERP) to the projection format (for example, CMP), using the image resolution of the current image. The conversion can be implemented in an image-based manner, where reference images are converted to the projection format of the current image and pre-stored, or it can be implemented block-wise during processing based on the block size, position, and motion data of the current prediction block in the current image. Fig.34. The projection format can only change from image to image, but this requires a conversion of the projection format of the reference image; this is a tool that could be supported in future video compression standards. Suggested views
[0112] Several services, including YouTube and Facebook, have recently begun offering 360° video sequences. These services allow users to view the scene from all angles while the video is playing. Users can rotate the scene to focus on what they find interesting at any given time.
[0113] There are several formats used for 360° videos, but each requires some form of projection of a 3D surface (sphere, cube, octahedron, icosahedron, etc.) onto a 2D plane. The 2D projection is then encoded / decoded like any normal video sequence. In the decoder, a portion of this 360° view is rendered and displayed, depending on the user's viewing angle at any given time.
[0114] The end result is that the user is given the freedom to look around as they please, which significantly enhances the feeling of being "immersed" in the scene, making them feel as if they are actually there. Combined with spatial audio effects (rotating the surround sound to match the video), this effect can be quite captivating.
[0115] Fig.Figure 35 illustrates an exemplary network environment 3500 in which proposed views within a 360-degree video can be implemented according to one or more implementations. Not all components shown may be used; however, one or more implementations may include additional components not shown in the figure. Variations in the arrangement and type of components are possible without derogating from the nature or scope of protection of the claims set forth in this document. Additional components, different components, or fewer components may be provided.
[0116] The exemplary network environment 3500 comprises a 360-degree video capture device 3502, a 360-degree video stitching device 3504, a video encoding device 3506, a video decoding device 3508, and a 360-degree video rendering device 3510. In one or more implementations, one or more of the devices 3502, 3504, 3506, 3508, and 3510 may be combined in the same physical device. For example, the 360-degree video capture device 3502, the 360-degree video joining device 3504 and the video encoding device 3506 can be combined in a single device, and the video decoding device 3508 and the 360-degree video rendering device 3510 can be combined in a single device.In some embodiments, the video decoding device 3508 may include an audio decoding device (not shown), or in other embodiments, the video decoding device 3508 may be communicatively coupled with a separate audio decoding device to process an incoming or stored compressed 360-degree video bitstream.
[0117] On the playback side for 360-degree videos, the 3500 network environment may further include a demultiplexer device (not shown) that can demultiplex the incoming compressed 360-degree video bitstream and provide the demultiplexed bitstream to the 3508 video decoding device, the audio decoding device, and a view angle extraction device, respectively. In some manifestations, the demultiplexer device may be configured to decompress the 360-degree video bitstream. The network environment 3500 may further include a device for converting the projection format for 360-degree videos (not shown), which can perform a conversion of the projection format for 360-degree videos before video encoding by means of the video encoding device 3506 and / or after video decoding by means of the video decoding device 3508.The network environment 3500 may also include a 360-degree video playback device (not shown) that plays back the content of the rendered 360-degree video. In one or more implementations, the video encoding device 3506 may be communicatively coupled to the video decoding device 3508 via a transmission link, such as a network.
[0118] The 360-degree video playback device can save the rendering settings for the 360-degree video (for example, the field of view angles, viewing direction angles, rendering image size, etc.) immediately before playback stops or when switching to another program channel. This allows the saved rendering settings to be used when playback resumes on the same channel. The 360-degree video playback device can include a preview mode in which the viewing angles can be automatically changed every N frames to make it easier for viewers to select their preferred viewing direction. The 360-degree video capture and playback device can calculate the projection image during processing (for example, block by block) to save memory bandwidth. In this case, the projection image cannot be loaded from the chip's external memory.In the claimed system, different information regarding fidelity to the reproduction can be assigned to different views.
[0119] In the claimed system, content providers can offer a "suggested view" for a given 360-degree video. The suggested view can be a specific set of viewing angles for each frame of the 360-degree video, designed to provide a recommended experience for the user. If the user is not particularly interested in manually controlling the view at any given time, they can view (or play back) the suggested view and experience the perspective recommended by the content provider.
[0120] To enable the system used to record / play back the decompressed 360-degree video bitstream to store data in a specific view as it was originally viewed by a user during one or more specific viewing sessions, the yaw, pitch, and roll angles (viewing angle data) for each frame can be stored in the storage device. Combined with the already recorded, complete original 360-degree view data, a previously stored view can be recreated.
[0121] Depending on how the view angles were stored, a corresponding process for extracting them can be initiated. In one or more implementations where the view angles are stored within the compressed 360-degree video bitstream, the view angle extraction process can be used. For example, the 3508 video decoding device and / or the audio decoding device can extract view angles from the compressed 360-degree video bitstream (for example, from the SEI messages within an HEVC bitstream). In this respect, the view angles extracted by the 3508 video decoding device can be provided to the 3510 video rendering device.If the viewing angles are stored in a separate data stream (for example, MPEG-2 TS-PID), the demultiplexer device can extract this information and send it to the 3510 video rendering device as suggested viewing angles. In some examples, the demultiplexer feeds a separate viewing angle extraction device (not shown) to extract the suggested viewing angles. In this respect, the claimed system would have the ability to switch between the suggested view and the manually selected user view at any time.
[0122] Switching between the suggested view and the manual view may require the user to make a user-specific selection (for example, pressing a key) to enter or exit this mode. In one or more implementations, the system can perform the switching automatically. For example, if the user moves the view manually (with the mouse, a remote control, hand gestures, a headset, etc.), the view is updated to reflect the user's request. If the user has not made any manual adjustments for a specified period, the view can revert to the suggested view.
[0123] In one or more implementations, multiple suggested views can be provided where appropriate, and / or more than one suggested view can be rendered simultaneously. For example, in a soccer match, one view might track the goalkeeper, and other views might track the strikers. In the soccer example above, the user can have a split screen with four views running concurrently. Alternatively, different views can be used to track specific cars during a Formula 1 race. The user can select from these suggested views to personalize their experience without having to fully control the view at all times.
[0124] If a suggested view is unavailable or unsuitable for the entire scene, suggestions (or recommendations) can be provided to ensure the viewer doesn't miss any important action. A hint view (or preview) can be provided at the start of a new scene. This view can then be moved to examine the given angle and center the view on the main action. In one or more implementations, if the user wishes to act less directly (or independently), graphical arrows can be used on the screen to indicate that the user might be looking in the wrong direction and missing something interesting.
[0125] Two of the most commonly used types of projection are the equirectangular projection and the cubic projection. These map video from a sphere (equirectangular) or a cube (cubic) onto a flat 2D surface. Examples are in Fig. Figure 36 shows examples of an equirectangular projection (for example, 3602) and of a cube projection (for example, 3604).
[0126] Fig. Figure 37 schematically illustrates an example of a 360-degree video rendering. Currently, most 360-degree video content from streaming media services (YouTube, Facebook, etc.) is viewed on computers or smartphones. However, it is expected that 360° videos will be broadcast over standard cable / satellite networks in the near future. Sporting events, travelogues, extreme sports, and many other types of programs can be presented as 360° videos to capture and engage viewers.
[0127] While a 360-degree video application immerses the viewer in a scene in an entertaining way, it can often become tiring with longer programs if it's necessary to constantly control the view manually to track the most important relevant objects. For example, it might be entertaining to look around occasionally during a sporting event, but for most of the game, the user wants to focus solely on the main action.
[0128] To this end, content providers can offer a "suggested view." The suggested view could be a specific set of viewing angles for each frame, designed to provide a recommended user experience. If the user isn't particularly interested in manually controlling the view at any given time, they can simply view the suggested view and experience the perspective recommended by the content provider.
[0129] The concept of recording a user's view at any given time can be expressed in 3D space by three different angles. These are the so-called Euler angles. In aviation dynamics, these three angles are referred to as yaw, pitch, and roll. With further reference to Fig.10. The yaw angle, pitch angle, and roll angle can be encoded for each frame as a means of suggesting a view to the viewer. The decoder (for example, 3508) can extract and use these suggested view angles at any time if the user is not interested in controlling the view itself.
[0130] This viewing angle data can be stored in a number of ways. For example, the angles can be inserted into the video stream as user image data (AVC / HEVC SEI messages (Supplemental Enhancement Information)), or they can be carried as a separate data stream within the video sequence (another MPEG-2 TS-PID or MP4 data stream). Depending on how the viewing angles are stored, a corresponding mechanism for extracting them might be required. For example, if the angles are stored as SEI messages in the video bitstream, the video decoder can extract them. If they are stored in a separate MPEG-2 TS-PID, the demultiplexer can extract this information and send it to the rendering process.The 360-degree video rendering system (for example, 3500) would have the ability to switch between the suggested view and the manually selected user view at any time.
[0131] Fig.Figure 38 illustrates a schematic diagram for extraction and rendering with suggested viewpoints. Switching between the suggested view and the manual view can be so simple that only a single keystroke is required to enter or exit this mode. In some cases, the switching can be automated. For example, if the user manually moves the view (with a mouse, remote control, hand gestures, headset, etc.), the view updates to reflect the user's movement. If the user has not made any manual adjustments for a specified period, the view can revert to the suggested view.
[0132] Multiple suggested views can be provided where appropriate. For example, in a football match, one view might track the goalkeeper, and other views might track the strikers. In other scenarios, different views could be used to track specific cars during a Formula 1 race. The user can select from these suggested views to personalize their experience without having to fully control the view at all times.
[0133] More than one suggested view can be rendered simultaneously. In the football example above, the user can have a split screen with four views at once: one for the goalkeeper, one for the strikers, while manually controlling the others, and so on.
[0134] If a suggested view is unavailable or unsuitable for the entire scene, hints can be provided to ensure the viewer doesn't miss any important action. A hint view can be included at the start of a new scene. This view can then be moved to view the given angle, thus centering the view on the main action. In other presentations, where less direct control is desired, elements such as graphical arrows on the screen can be used to indicate that the user might be looking in the wrong direction and missing something interesting.
[0135] In most 360-degree viewing applications, the user cannot adjust the roll angle. The camera is typically fixed in a vertical orientation. The view can be rotated up / down and left / right, but it cannot be panned sideways. In this respect, a system that suggests only two viewing angles would be sufficient for most purposes.
[0136] It should be noted that not all 360-degree video streams actually cover the full 360° × 180° field of view. Some sequences may only allow viewing in the "front" direction (180° × 180°). In some cases, there may be limitations regarding viewing height (how high or low the viewer can see). These cases would all be covered by the concepts discussed in this document.
[0137] In one or more implementations, a system message can be signaled with the presets for the (recommended) viewing direction (i.e., the viewing direction angles), the field of view angles and / or the rendering image size, along with the content of the 360-degree video.
[0138] In one or more implementations, a 360-degree video playback system supports a scan mode in which the viewing angles automatically change every N frames to make it easier for viewers to select the desired viewing direction.For example, the vertical viewing angle y and the viewing angle ε along the z-axis are initially fixed at 0 degrees, the horizontal angle changes by one degree every N frames; after the viewer has selected the horizontal viewing angle θ, the horizontal viewing angle is fixed at the selected angle, and the viewing angle ε along the z-axis is still fixed at 0 degrees, the vertical viewing angle changes by one degree every N frames until the viewer selects the vertical viewing angle y; after both the horizontal and vertical viewing angles θ and y have been selected, both angles are fixed at the selected angles, the viewing angle ε along the z-axis begins to change by one degree every N frames until the viewer selects the viewing angle ε.The scan mode can traverse different viewing angles in parallel in some implementations, or it can traverse them sequentially in others. In some implementations, the viewing angles can be limited (or restricted) by a user profile or user type (for example, child, adult). In this example, 360-degree video content managed by parental controls can be restricted to a subset of viewing angles, as specified by the settings. In some implementations, the 360-degree video bitstream includes metadata that specifies the frame paths for the scan mode.
[0139] In one or more implementations, the multi-view layout format can also maintain the relevant viewpoints by specifying the importance of the view within the sequence. Views can be assigned with varying levels of fidelity (resolution, frame rate, bit rate, field of view angle, etc.). Information regarding the view's fidelity should be indicated via a system message.
[0140] Fig.Figure 39 schematically illustrates an electronic system 3900 with which one or more implementations of the claimed technology can be implemented. For example, the electronic system 3900 can be a network device, a media converter, a desktop computer, a laptop computer, a tablet computer, a server, a switch, a router, a base station, a receiver, a telephone, or generally any electronic device that transmits signals over a network. Such an electronic system 3900 has various types of computer-readable media and interfaces to various other types of computer-readable media.In one or more implementations, the electronic system 3900 may be or include one of the devices 102, 104, 106, 108, 110, the projection format conversion device for 360-degree videos, and / or the playback device for 360-degree videos. The electronic system 3900 comprises a bus 3908, one or more processing units 3912, a system memory 3904, a read-only memory (ROM) 3910, a permanent storage device 3902, an input device interface 3914, an output device interface 3906, and a network interface 3916, or subsets and variations thereof.
[0141] The bus 3908 encompasses all system buses, peripheral buses, and chipset buses that communicatively connect the numerous internal devices of the electronic system 3900. In one or more implementations, the bus 3908 communicatively connects the one or more processing unit(s) 3912 to the ROM 3910, the system memory 3904, and the permanent storage device 3902. From these various storage units, the one or more processing unit(s) 3912 retrieve instructions to be executed and data to be processed in order to carry out the processes of the claimed disclosure. The one or more processing unit(s) 3912 can be a single processor or a multi-core processor, depending on the implementation.
[0142] The ROM 3910 stores static data and instructions required by the one or more processing unit(s) 3912 and other modules of the electronic system. The permanent storage device 3902, on the other hand, is a read / write storage device. The permanent storage device 3902 is a non-volatile storage unit in which instructions and data are stored even when the electronic system 3900 is switched off. In one or more implementations of the claimed disclosure, a mass storage device (such as a magnetic disk or an optical disk and the corresponding disk drive) can be used as the permanent storage device 3902.
[0143] In other implementations, a removable storage device (such as a floppy disk, a flash drive, and the corresponding disk drive) is used as the permanent storage device 3902. Like the permanent storage device 3902, the system memory 3904 is a read / write storage device. However, unlike the permanent storage device 3902, the system memory 3904 is a volatile read / write memory, such as random access memory. Any instructions and data required by the one or more processing units 3912 at runtime are stored in the system memory 3904. In one or more implementations, the processes of the claimed disclosure are stored in the system memory 3904, in the permanent storage device 3902, and / or in the ROM 3910.From these various storage units, the one or more processing unit(s) retrieve 3912 instructions to be executed and data to be processed in order to execute the processes of one or more implementations.
[0144] Bus 3908 also provides a connection to the 3914 input device interface and the 3906 output device interface. The 3914 input device interface allows a user to transmit information and select commands to the electronic system. Input devices used with the 3914 input device interface include, for example, alphanumeric keyboards and pointing devices (also known as cursor control devices). The 3906 output device interface allows, for example, the display of images generated by the 3900 electronic system.Output devices used with the 3906 Output Device Interface include, for example, printers and display devices such as an LCD (Liquid Crystal Display), an LED (Light Emitting Diode), an OLED (Organic Light Emitting Diode), a flexible display, a flat panel display, a solid-state display, a projector, or any other device for outputting information. One or more implementations may include devices that function as both input and output devices, such as a touchscreen.In these implementations, the feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including auditory, speech, or tactile input.
[0145] Finally, as in Fig.Figure 39 shows that the bus 3908 also connects the electronic system 3900 to one or more networks (not shown) via one or more network interfaces 3916. In this way, the computer can be part of one or more networks of computers (such as a local area network, a wide area network, an intranet, or a network consisting of networks, such as the Internet). Any or all components of the electronic system 3900 can be used in connection with the claimed disclosure.
[0146] Implementations within the scope of protection of this disclosure may be carried out in whole or in part using a physical, computer-readable storage medium (or several physical, computer-readable storage media of one or more types) which encode one or more instructions. The physical, computer-readable storage medium may also be persistent by its nature.
[0147] The computer-readable storage medium can be any storage medium that can be read and written to, or otherwise accessed by a general-purpose or specialized computer device, including any processing electronics and / or processing circuitry capable of executing instructions. For example, the computer-readable medium can include, without limitation, any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. The computer-readable medium can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, Flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, Racetrack memory, FJG, and Millipede memory.
[0148] Furthermore, the computer-readable storage medium can comprise any non-semiconductor storage medium, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In some implementations, the physical, computer-readable storage medium can be directly coupled to a computer device, while in other implementations, the physical, computer-readable storage medium can be indirectly coupled to a computer device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.
[0149] Instructions can be directly executable or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or as instructions in a higher-level language that can be compiled to produce executable or non-executable machine code. Furthermore, instructions can be implemented as data or include data. Computer-executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, and so on. As experts in this field recognize, details, including but not limited to the number, structure, sequence, and organization of instructions, can vary considerably without altering the underlying logic, function, processing, and output.
[0150] While the above discussion mainly concerns microprocessors or multi-core processors that execute software, one or more implementations are executed by means of one or more integrated circuits, such as ASICs (Application Specific Integrated Circuits) or FPGAs (Field Programmable Gate Arrays). In one or more implementations, such integrated circuits execute instructions that are stored within the circuit itself.
[0151] Experts in this field would recognize that the various blocks, modules, elements, components, procedures, and algorithms described in this document for illustrative purposes can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various blocks, modules, elements, components, procedures, and algorithms used for illustrative purposes have been described above in general terms with regard to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints to which the overall system is subject. Experts in this field can implement the described functionality in different ways for each specific application.Different components and blocks may be arranged in other ways (for example, arranged in a different order or divided in other ways) without deviating from the scope of protection of the claimed technology.
[0152] It is understood that any specific order or hierarchy of blocks in the disclosed processes is merely an illustration of exemplary approaches. It is understood that the specific order or hierarchy of blocks in the processes can be rearranged based on design preferences, or that all illustrated blocks can be executed. Any of the blocks can be executed concurrently. In one or more implementations, multitasking and parallel processing may be advantageous. Furthermore, the separation of different system components in the embodiments described above should not be interpreted as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0153] As used in this patent specification and in any claims of this patent application, the terms "base station", "receiver", "computer", "server", "processor", and "memory" all refer to electronic or other technological devices. These terms do not refer to any person or group of persons. For the purposes of this patent specification, the terms "display" and "display" mean displaying on an electronic device.
[0154] In one or more implementations, the statement that a processor is configured to monitor and control an operation or component can also mean that the processor is programmed to monitor and control the operation, or that the processor is operational in such a way as to monitor and control the operation. Similarly, the statement that a processor is configured to execute code can be interpreted as meaning that a processor is programmed to execute code, or that it is operational in such a way as to execute code.
Claims
[1] System comprising the following: a video capture device configured to capture 360-degree videos; a device for assembly that is configured for the following: Combining the captured 360-degree video using an intermediate coordinate system between an input image coordinate system and a 360-degree video capture coordinate system, wherein the intermediate coordinate system is transformed from the 360-degree video capture coordinate system based on rotation angles along axes of the 360-degree video capture coordinate system; and a coding device configured for the following: Encoding the stitched 360-degree video into a 360-degree video bitstream; and Preparing the 360-degree video bitstream to be played back for transmission and storage. [2] System according to claim 1, wherein the assembly device is configured for the following: Calculating a normalized projection plane size using field-of-view angles; Calculating a coordinate in a normalized rendering coordinate system from a coordinate system of an output rendering image using the normalized projection plane size; Mapping the coordinate onto a viewing coordinate system using the normalized projection plane size; Converting the coordinate from the observation coordinate system to a recording coordinate system using a coordinate transformation matrix; Converting the coordinate from the acquisition coordinate system to the intermediate coordinate system using the coordinate transformation matrix; Converting the coordinate from the intermediate coordinate system into a normalized projection system; and Mapping the coordinate from the normalized projection system to the coordinate system for input images. [3] System according to claim 2, wherein the coordinate transformation matrix is pre-calculated using viewing direction angles and global rotation angles. [4] System according to claim 3, wherein the global rotation angles are signaled in a message contained in the 360-degree video bitstream. [5] System according to claim 1, wherein the coding device is further configured for the following: Encoding the stitched 360-degree video into a multitude of view sequences, with each of the multitude of view sequences corresponding to a different view region of the 360-degree video. [6] System according to claim 5, wherein at least two of the plurality of view sequences are encoded with a different format for the projection layout. [7] System according to claim 5, wherein the system further comprises: a rendering device configured to receive the multitude of view sequences as input and to render each of the multitude of view sequences using a rendering control input. [8] System according to claim 7, wherein the rendering device is further configured to select at least one of the plurality of view sequences for rendering and to exclude at least one of the plurality of view sequences from rendering.
Citation Information
Patent Citations
Method and apparatus for transmitting and receiving a panoramic video stream
EP2490179A1
Non-flat image processing apparatus, image processing method, recording medium, and computer program
US20040247173A1
Scalable video encoding in a multi-view camera system
US20100157016A1