Tiling in video encoding and decoding

A high-level syntax in video coding standards enables efficient extraction and rendering of tiled views by signaling multiview information, addressing the lack of standardization in view combination and display for 3D monitors.

JP7815403B2Active Publication Date: 2026-02-17DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024211900
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2007-04-20
Filing Date
2024-12-05
Publication Date
2026-02-17
Estimated Expiration
2028-04-11

AI Technical Summary

Technical Problem

Current video coding standards lack a standardized method for determining how multiple views are tiled and combined in a single frame, making it difficult for display devices to extract and render these views efficiently.

Method used

Introduce a high-level syntax, such as a Supplemental Enhancement Information (SEI) message, to signal multiview information in MPEG-4 AVC Standard-compliant bitstreams, allowing for easier extraction and rendering of tiled views on 3D monitors.

Benefits of technology

Facilitates efficient display of multiview video streams by providing a standardized method for determining view arrangement, improving coding efficiency and simplifying the implementation of multiview coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815403000007
    Figure 0007815403000007
  • Figure 0007815403000008
    Figure 0007815403000008
  • Figure 0007815403000009
    Figure 0007815403000009
Patent Text Reader

Abstract

To provide implementations that relate to view tiling in video encoding and decoding.SOLUTION: A decoding is performed by a step of accessing to a video and a picture, which include a plurality of pictures synthesized in a single picture; steps (806 and 808) of accessing to information indicating how the plurality of pictures in the accessed video and picture is synthesized; a step (824) of decoding the video and picture so as to supply a decoded expression of at least one of the plurality of pictures; and a step of supplying the accessed information and the decoded video and picture as an output. On an encoding side, the information indicating that how the plurality of pictures included in the single video and picture is synthesized is formatted or processed, and the encoded expression of the plurality of synthesized pictures is formatted and processed.SELECTED DRAWING: Figure 8A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of (1) U.S. Provisional Patent Application No. 60 / 923,014 (Attorney Docket No. PU070078), entitled "Multiview Information," filed April 12, 2007, and (2) U.S. Provisional Patent Application No. 60 / 925,400 (Attorney Docket No. PU0700103), entitled "View Tiling in MVC Coding," filed April 20, 2007. The entire contents of each of the foregoing two applications are incorporated herein by reference.

[0002] The present principles relate generally to video encoding and / or decoding. [Background technology]

[0003] Video display manufacturers may use a framework to arrange or tile separate views onto a single frame, in which case the views can be extracted and rendered from their respective positions. Summary of the Invention [Problem to be solved by the invention]

[0004] According to a general aspect, a video picture including multiple pictures combined into a single picture is accessed. Information indicating how the multiple pictures in the accessed video picture are combined is accessed. The video picture is decoded to provide a decoded representation of the combined multiple pictures. The accessed information and the decoded video picture are provided as an output. [Means for solving the problem]

[0005] According to another general aspect, information is generated that indicates how multiple pictures included in a video picture are combined into a single picture. The video picture is encoded to provide a coded representation of the combined multiple pictures. The generated information and the coded video picture are provided as output.

[0006] According to another general aspect, a signal or signal structure includes information indicating how multiple pictures included in a single video picture are combined into the single video picture, and the signal or signal structure also includes a coded representation of the combined multiple pictures.

[0007] According to another general aspect, a video picture including multiple pictures combined into a single picture is accessed. Information indicating how the multiple pictures in the accessed video picture are combined is accessed. The video picture is decoded to provide a decoded representation of at least one of the multiple pictures. The accessed information and the decoded representation are provided as an output.

[0008] According to another general aspect, a video picture including multiple pictures combined into a single picture is accessed. Information indicating how the multiple pictures in the accessed video picture are combined is accessed. The video picture is decoded to provide a decoded representation of the combined multiple pictures. User input selecting at least one of the multiple pictures for display is received. A decoded output of the selected at least one picture is provided, the decoded output being provided based on the accessed information, the decoded representation, and the user input. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 shows an example of four views tiled on a single frame. [Figure 2]FIG. 1 shows an example of four views flipped and tiled on a single frame. [Figure 3] 1 is a block diagram showing a video encoder to which the present principles may be applied, in accordance with an embodiment of the present principles; FIG. [Figure 4] 1 is a block diagram showing a video decoder to which the present principles may be applied, in accordance with an embodiment of the present principles; FIG. [Figure 5A] 1 is a flow diagram for a method for encoding multiple view pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 5B] 1 is a flow diagram for a method for encoding multiple view pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 6A] 1 is a flow diagram for a method for decoding multiple view pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 6B] 1 is a flow diagram for a method for decoding multiple view pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 7A] 1 is a flow diagram for a method for encoding multiple view and depth pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 7B] 1 is a flow diagram for a method for encoding multiple view and depth pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 8A] 1 is a flow diagram for a method for decoding multiple view and depth pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 8B] 1 is a flow diagram for a method for decoding multiple view and depth pictures using the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 9] 10 shows an example of a depth signal, in accordance with an embodiment of the present principles; [Figure 10] FIG. 10 shows an example of depth signals added as tiles, in accordance with an embodiment of the present principles; [Figure 11] FIG. 10 shows an example of five views tiled on a single frame, in accordance with an embodiment of the present principles. [Figure 12] 1 is a block diagram for an exemplary Multi-view Video Coding (MVC) encoder to which the present principles may be applied, in accordance with an embodiment of the present principles; [Figure 13] 1 is a block diagram for an exemplary Multi-view Video Coding (MVC) decoder to which the present principles may be applied, in accordance with an embodiment of the present principles; [Figure 14] 1 is a flow diagram for a method for processing pictures of multiple views in preparation for encoding the pictures using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 15A] 1 is a flow diagram for a method for encoding pictures for multiple views using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 15B] 1 is a flow diagram for a method for encoding pictures for multiple views using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 16] 1 is a flow diagram for a method for processing pictures of multiple views in preparation for decoding of the pictures using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 17A] 1 is a flow diagram for a method for decoding pictures of multiple views using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 17B] 1 is a flow diagram for a method for decoding pictures of multiple views using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 18]1 is a flow diagram for a method for processing pictures of multiple views and depths in preparation for encoding the pictures using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 19A] 1 is a flow diagram for a method for processing pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 19B] 1 is a flow diagram for a method for processing pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 20] 1 is a flow diagram for a method for processing pictures of multiple views and depths in preparation for decoding of those pictures using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; FIG. [Figure 21A] 1 is a flow diagram for a method for decoding pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 21B] 1 is a flow diagram for a method for decoding pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard, in accordance with an embodiment of the present principles; [Figure 22] FIG. 10 shows an example of tiling at pixel level, in accordance with an embodiment of the present principles; [Figure 23] 1 is a block diagram illustrating a video processing apparatus to which the present principles may be applied, in accordance with an embodiment of the present principles; DETAILED DESCRIPTION OF THE INVENTION

[0010] The details of one or more implementations are set forth in the accompanying drawings and the following specification. Although described in one particular way, an implementation may be configured or embodied in various ways. For example, an implementation may be performed as a method, or as a device configured to perform a set of operations, or as a device storing instructions for performing a set of operations, or embodied in a signal. Other aspects and features will become apparent from the following detailed description considered in conjunction with the accompanying drawings and claims. [Example]

[0011] Various implementations relate to methods and apparatus for tiling views in video encoding and decoding. Thus, those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the present application and are within its spirit and scope.

[0012] All examples and conditional language in the specification and claims are intended for teaching purposes to aid the reader in understanding the principles of the present application and the concepts contributed by the inventors of the present application to advance the art, and should be construed as without limitation to the foregoing specifically stated examples and conditions.

[0013] Moreover, all statements in the specification and claims reciting principles, aspects, and embodiments of the present application, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, such equivalents are intended to include both currently known equivalents as well as equivalents developed in the future (i.e., any components developed that perform the same function, regardless of structure).

[0014] Thus, for example, it will be appreciated by those skilled in the art that the block diagrams presented herein represent conceptual views of illustrative circuitry embodying the principles of the present application. Similarly, flowcharts, flow diagrams, state diagrams, pseudocode, etc., all of which are substantially represented in a computer-readable medium, represent various operations that may be performed by a computer or processor, whether or not such computer or processor is explicitly stated.

[0015] The functions of the various components shown in the figures may be provided through the use of dedicated hardware and hardware capable of executing software in association with appropriate software. If provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Furthermore, the explicit use of the terms "processor" or "controller" should not be construed as referring exclusively to hardware capable of executing software, but implicitly may include, without limitation, digital signal processor ("DSP") hardware, read-only memory ("ROM") for storing software, random access memory ("RAM"), and non-volatile storage.

[0016] Other hardware (general-purpose and / or custom) may also be included. Similarly, any switches shown in the figures are conceptual only. The foregoing functions may be performed by the operation of program logic, by dedicated logic, by the interplay of program control and dedicated logic, or manually, with the particular method being selectable by the implementer as more particularly indicated by the context.

[0017] In the claims, any element described as a means for performing a particular function is intended to encompass any means for performing that function (e.g., a) a combination of circuitry for performing that function, or b) software in any form, including firmware, microcode, etc., in combination with appropriate circuitry for executing the software to perform the function. The principle of the present application, as defined in the following claims, resides in the fact that the functionality provided by the various recited means is combined and gathered in the manner required by the claims. Thus, any means capable of providing the aforementioned functionality are deemed equivalent to those recited in the specification and claims of this application.

[0018] References herein to "one embodiment" (or "one implementation") or "an embodiment" (or "an implementation") of the present principles mean that the particular configuration, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment of the present principles. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment.

[0019] The use of the words "and / or" or "at least one of," for example, "A and / or B" or "at least one of A and B," is intended to encompass the selection of only the first listed alternative (A), the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, "A, B, and / or C" or "at least one of A, B, and C" is intended to encompass the selection of only the first listed alternative (A), the selection of only the second listed alternative (B), the selection of only the third listed alternative (C), the selection of only the first and second listed alternatives (A and B), the selection of only the first and third listed alternatives (A and C), the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). As will be readily apparent to those of ordinary skill in this and related arts, this can be expanded upon for any number of items listed.

[0020] Furthermore, although one or more embodiments of the present principles are described herein with respect to the MPEG-4 AVC standard, the present principles are not limited solely to this standard and, therefore, may be utilized with respect to other standards, recommendations, and extensions thereof (particularly video coding standards, recommendations, and extensions thereof, including extensions to the MPEG-4 AVC standard) while maintaining the spirit of the present principles.

[0021] Furthermore, although one or more other embodiments of the present principles are described herein with respect to a Multi-view Video Coding extension to the MPEG-4 AVC standard, the present principles are not limited to merely this extension and / or this standard and, thus, may be utilized with other video coding standards, recommendations, and their extensions related to Multi-view Video Coding (in particular, video coding standards, recommendations, and their extensions (extensions to the MPEG-4 AVC standard)) while maintaining the spirit of the present principles. Multi-view Video Coding (MVC) is a compression framework for coding multi-view sequences. A Multi-view Video Coding (MVC) sequence is a set of two or more video sequences capturing the same scene from different viewpoints.

[0022] Furthermore, although this specification and claims describe one or more other embodiments of the present principles that use depth information with respect to video content, the present principles are not limited to the foregoing embodiments, and thus other embodiments that do not use depth information can be implemented while maintaining the spirit of the present principles.

[0023] Additionally, as used herein, "high-level syntax" refers to syntax present in a bitstream hierarchically above the macroblock layer. For example, high-level syntax as used herein may refer to, but is not limited to, syntax at the slice header level, supplemental enhancement information (SEI) level, picture parameter set (PPS) level, sequence parameter set (SPS) level, view parameter set (VPS) level, and network abstraction layer (NAL) unit header level.

[0024] In current implementations of Multi-Video Coding (MVC) based on the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group-4 (MPEG-4) Part 10 Advanced Video Coding (AVC) Standard / International Telecommunications Union (ITU-T) H.264 Recommendation (hereinafter, the "MPEG-4 AVC Standard"), reference software achieves multi-view prediction by encoding each view with a single encoder and taking inter-view referencing into account. Each view is encoded by the encoder as a separate bitstream in its original resolution, and then all bitstreams are combined to form a single bitstream. This single bitstream is then decoded. Each view generates a separate YUV decoded output.

[0025] Another approach to multiview prediction involves grouping a set of views into pseudo-views. In one example of this approach, it is possible to tile pictures from N views out of a total of M views (sampled simultaneously) onto a larger frame, or super-frame, possibly with downsampling or other processing. Turning to Figure 1, an example of four views tiled onto a single frame is shown generally at 100. All four views are in their normal orientation.

[0026] Turning to Figure 2, an example of four views flipped and tiled on a single frame is shown generally at 200. The top left view is in its normal orientation. The top right view is flipped horizontally. The bottom left view is flipped vertically. The bottom right view is flipped horizontally and vertically. Thus, when there are four views, the pictures from each view are arranged in a tile-like super-frame. This results in a single uncoded input sequence at high resolution.

[0027] Alternatively, the image can be downsampled to yield a lower resolution. Thus, multiple sequences are created, each containing a different view tiled together. Each of these sequences then forms a pseudo-view, each containing N different tiled views. Figure 1 shows one pseudo-view, and Figure 2 shows another pseudo-view. These pseudo-views can then be coded using existing video coding standards, such as the ISO / IEC MPEG-2 standard or the MPEG-4 AVC standard.

[0028] Yet another approach to multi-view prediction involves using new standards to encode the different views separately and, after decoding, tiling the views as required by the player.

[0029] Additionally, in another approach, the views may be tiled pixel by pixel, for example in a superview containing four views, pixel (x,y) may be from view 0, while pixel (x+1,y) may be from view 1, pixel (x,y+1) may be from view 2, and pixel (x+1,y+1) may be from view 3.

[0030] Many display manufacturers use the above framework of placing or tiling separate views on a single frame and extracting and rendering the views from their respective positions. In such cases, there is no standard way to determine whether a bitstream has the above properties. Thus, if a system uses a method of tiling pictures of separate views on a large frame, the method of extracting the separate views is specific to that system.

[0031] However, there is no standard way to determine if a bitstream has the above properties. The inventors of this application propose a high-level syntax to make it easier for rendering devices or players to extract the above information to aid in display or other post-processing. Subpictures may have different resolutions, and some upsampling may be required to finally render the view. Users may want the upsampling to be indicated in the high-level syntax as well. Furthermore, it is also possible to send parameters to change the depth focus.

[0032] In one embodiment, we propose a new Supplemental Enhancement Information (SEI) message for signaling multiview information in MPEG-4 AVC Standard-compliant bitstreams, where each picture contains subpictures belonging to different views. The embodiment is intended for simple and convenient display of multiview video streams on, for example, three-dimensional (3D) monitors, which can use the aforementioned framework. This concept can be extended to other video coding standards and recommendations that use high-level syntax to signal such information.

[0033] Furthermore, one embodiment proposes a signaling method for arranging views before sending them to a multiview video encoder and / or decoder. Effectively, the above embodiment may lead to a simplified implementation of multiview coding, benefiting coding efficiency. Certain views can be aggregated to form a pseudo-view or super-flag, and the tiled super-view is treated as a normal view by a common multiview video encoder and / or decoder, for example, according to current MPEG-4 AVC standard-based implementations of multiview video coding. A new flag is proposed in the Sequence Parameter Set (SPS) of multiview video coding to signal the use of the pseudo-view technique. The embodiment is intended for easy and convenient display of multiview video streams on a 3D monitor that can use the aforementioned framework.

[0034] Encoding / Decoding Using Single-View Video Encoding / Decoding Standards / Recommendations In current implementations of multi-video coding (MVC) based on the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group-4 (MPEG-4) Part 10 Advanced Video Coding (AVC) Standard / International Telecommunications Union (ITU-T) H.264 Recommendation (hereinafter referred to as the “MPEG-4 AVC Standard”), reference software achieves multi-view prediction by encoding each view with a single encoder and taking into account inter-view references. Each view is encoded by the encoder as a separate bitstream in its original resolution, and then all bitstreams are combined to form a single bitstream. This single bitstream is then decoded. Each view produces a separate YUV decoded output.

[0035] Another approach to multiview prediction involves tiling pictures from each view (sampled simultaneously) onto a larger frame, or super-frame, possibly with a downsampling operation. Turning to FIG. 1, an example of four views tiled onto a single frame is shown generally by reference numeral 100. Turning to FIG. 2, an example of four views inverted and tiled onto a single frame is shown generally by reference numeral 200. Thus, when four views are present, the pictures from each view are arranged into a tile-like super-frame. This results in a single uncoded input sequence at high resolution. This signal can then be encoded using existing video coding standards, such as the ISO / IEC MPEG-2 standard or the MPEG-AVC standard.

[0036] Yet another approach to multi-view prediction involves using new standards to encode the different views separately and, after decoding, simply tiling the views as required by the player.

[0037] Many display manufacturers use the above framework of placing or tiling separate views on a single frame, and then extracting and rendering the views from their respective positions. In such cases, there is no standard way to determine whether a bitstream has the above properties. Therefore, if a system uses a method of tiling pictures of separate views on a large frame, the method of extracting the separate views is unique.

[0038] Turning to FIG. 3, a video encoder capable of performing video encoding in accordance with the MPEG-4 AVC standard is indicated generally by the reference numeral 300 .

[0039] The video encoder 300 includes a frame ordering buffer 310 having an output in signal communication with a non-inverting input of a combiner 385. The output of the combiner 385 is connected in signal communication with a first input of a transformer and quantizer 325. The output of the transformer and quantizer 325 is connected in signal communication with a first input of an entropy coder 345 and a first input of an inverse transformer and quantizer 350. The output of the entropy coder 345 is connected in signal communication with a first non-inverting input of a combiner 390. The output of the combiner 390 is connected in signal communication with a first input of an output buffer 335.

[0040] A first output of the encoder controller 305 is connected in signal communication with a second input of a frame ordering buffer 310, a second input of the inverse transformer and inverse quantizer 350, an input of a picture type decision module 315, an input of a macroblock type (MB type) decision module 320, a second input of the intra prediction module 360, a second input of the deblocking filter 365, a first input of the motion compensator 370, a first input of the motion compensator 375, and a second input of the reference picture buffer 380.

[0041] A second output of the encoder controller 305 is connected in signal communication with a first input of a supplemental enhancement information (SEI) inserter 330, a second input of the transformer and quantizer 325, a second input of the entropy encoder 345, a second input of the output buffer 335, and an input of a sequence parameter set (SPS) and picture parameter set (PPS) inserter 340.

[0042] A first output of the picture type decision module 315 is connected in signal communication with a third input of the frame ordering buffer 310. A second output of the picture type decision module 315 is connected in signal communication with a second input of a macroblock type decision module 320.

[0043] An output of the sequence parameter set (SPS) and picture parameter set (PPS) inserter 340 is connected in signal communication with a third non-inverting input of the combiner 390. An output of the SEI inserter 330 is connected in signal communication with a second non-inverting input of the combiner 390.

[0044] An output of the inverse quantizer and inverse transformer 350 is connected in signal communication with a first non-inverting input of the combiner 319. An output of the combiner 319 is connected in signal communication with a first input of the intra prediction module 360 ​​and a first input of the deblocking filter 365. An output of the deblocking filter 365 is connected in signal communication with a first input of the reference picture buffer 380. An output of the reference picture buffer 380 is connected in signal communication with a second input of the motion estimator 375 and a first input of the motion compensator 370. A first output of the motion estimator 375 is connected in signal communication with a second input of the motion compensator 370. A second output of the motion estimator 375 is connected in signal communication with a third input of the entropy coder 345.

[0045] An output of the motion compensator 370 is connected in signal communication with a first input of a switch 397. An output of the intra prediction module 360 ​​is connected in signal communication with a second input of the switch 397. An output of the macroblock type decision module 320 is connected in signal communication with a third input of the switch 397 for providing a control input to the switch 397. The third input of the switch 397 (compared to the control input, i.e., third input) determines whether the "data" input of the switch is provided by the motion compensator 370 or the intra prediction module 360. An output of the switch 397 is connected in signal communication with a second non-inverting input of the combiner 319 and an inverting input of the combiner 385.

[0046] The inputs of the frame ordering buffer 310 and the encoder controller 105 are available as inputs to the encoder 300 for receiving the input picture 301. Additionally, the input of the supplemental enhancement information (SEI) inserter 330 is available as an input to the encoder 300 for receiving metadata. The output of the output buffer 335 is available as an output of the encoder 300 for outputting a bitstream.

[0047] Turning to FIG. 4, a video decoder capable of performing video decoding in accordance with the MPEG-4 AVC standard is indicated generally by the reference numeral 400 .

[0048] The video decoder 400 includes an input buffer 410 having an output connected in signal communication with a first input of an entropy decoder 445. The first output of the entropy decoder 445 is connected in signal communication with a first input of an inverse transformer and inverse quantizer 450. The output of the inverse transformer and inverse quantizer 450 is connected in signal communication with a second non-inverting input of a combiner 425. The output of the combiner 425 is connected in signal communication with a second input of a deblocking filter 465 and a first input of the intra prediction module 250. The second output of the deblocking filter 465 is connected in signal communication with a first input of a reference picture buffer 480.

[0049] A second output of the entropy decoder 445 is connected in signal communication with a third input of the motion compensator 470 and a first input of the deblocking filter 465. A third output of the entropy decoder 445 is connected in signal communication with an input of a decoder controller 405. A first output of the decoder controller 405 is connected in signal communication with a second input of the entropy decoder 445. A second output of the decoder controller 405 is connected in signal communication with a second input of the inverse transformer and inverse quantizer 450. A third output of the decoder controller 405 is connected in signal communication with a third input of the deblocking filter 465. A fourth output of the decoder controller 405 is connected in signal communication with a second input of the intra prediction module 460, a first input of the motion compensator 470, and a second input of the reference picture buffer 480.

[0050] An output of the motion compensator 470 is connected in signal communication with a first input of a switch 497. An output of the intra prediction module 460 is connected in signal communication with a second input of the switch 497. An output of the switch 497 is connected in signal communication with a first non-inverting input of the combiner 425.

[0051] An input of the input buffer 410 is available as an input to the decoder 400 to receive the input bitstream, and a first output of the deblocking filter 465 is available as an output of the decoder 400 to output the output picture.

[0052] Turning to FIG. 5, an exemplary method for encoding pictures for multiple views using the MPEG-4 AVC Standard is indicated generally by the reference numeral 500 .

[0053] The method 500 includes a start block 502. The start block 502 passes control to a function block 504. The function block 504 arranges each view at a particular time point as a subpicture in a tiled format and passes control to a function block 506. The function block 506 sets the syntax element num_coded_views_minus1 and passes control to a function block 508. The function block 508 sets the syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 and passes control to a function block 510. The function block 510 sets a variable i equal to zero and passes control to a decision block 512. The decision block 512 determines whether the variable i is less than the number of views. If so, it passes control to a function block 514. Otherwise, it passes control to a function block 524.

[0054] The function block 514 sets the syntax element view_id[i] and passes control to a function block 516. The function block 516 sets the syntax element num_parts[view_id[i]] and passes control to a function block 518. The function block 518 sets the variable j to zero and passes control to a decision block 520. The decision block 520 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, it passes control to a function block 522. Otherwise, it passes control to a function block 528.

[0055] The function block 522 is depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and The frame_crop_bottom_offset[view_id[i]][j] syntax element is set, the variable j is incremented, and control is then returned to decision block 520 .

[0056] The function block 528 sets the syntax element upsample_view_flag[view_id[i]] and passes control to a decision block 530. The decision block 530 determines whether the current value of the syntax element upsample_view_flag[view_id[i]] is equal to 1. If so, then control is passed to a function block 532. Otherwise, control is passed to a decision block 534.

[0057] The function block 532 sets the syntax element upsample_filter[view_id[i]], and passes control to a decision block 534 .

[0058] The decision block 534 determines whether the current value of the syntax element upsample_filter[view_id[i]] is equal to 3. If so, then control is passed to a function block 536. Otherwise, control is passed to a function block 540.

[0059] Function block 536 is vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]] and passes control to a function block 538. The function block 538 sets the filter coefficients for each YUV component and passes control to a function block 540.

[0060] The function block 540 increments the variable i, and returns control to the decision block 512 .

[0061] The function block 524 writes the above syntax elements to at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a supplemental enhancement information (SEI) message, a network abstraction layer (NAL) unit header, and a slice header, and passes control to a function block 526. The function block 526 encodes each picture using the MPEG-4 AVC Standard or other single-view codec, and passes control to an end block 599.

[0062] Turning to FIG. 6, an exemplary method for decoding pictures for multiple views using the MPEG-4 AVC Standard is indicated generally by the reference numeral 600 .

[0063] The method 600 includes a start block 602. The start block 602 passes control to a function block 604. The function block 604 parses the following syntax elements from at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a supplemental enhancement information (SEI) message, a network abstraction layer (NAL) unit header, and a slice header and passes control to a function block 606. The function block 606 parses the syntax element num_coded_views_minus1 and passes control to a function block 608. The function block 608 parses the syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 and passes control to a function block 610. The function block 610 sets a variable i equal to zero and passes control to a decision block 612. The decision block 612 determines whether the variable i is less than the number of views. If so, then control is passed to a function block 614. Otherwise, control is passed to a function block 624.

[0064] The function block 614 parses the syntax element view_id[i] and passes control to a function block 616. The function block 616 parses the syntax element num_parts_minus1[view_id[i]] and passes control to a function block 618. The function block 618 sets a variable j equal to zero and passes control to a decision block 620. The decision block 620 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, it passes control to a function block 622. Otherwise, it passes control to a function block 628.

[0065] The function block 622 is depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j] parses the syntax elements of {circumflex over (x)}, increments the variable j, and then returns control to decision block 620.

[0066] The function block 628 parses the syntax element upsample_view_flag[view_id[i]] and passes control to a decision block 630. The decision block 630 determines whether the current value of the syntax element upsample_view_flag[view_id[i]] is equal to 1. If so, then control is passed to a function block 632. Otherwise, control is passed to a decision block 634.

[0067] The function block 632 parses the syntax element upsample_filter[view_id[i]], and passes control to a decision block 634 .

[0068] The decision block 634 determines whether the current value of the syntax element upsample_filter[view_id[i]] is equal to 3. If so, then control is passed to a function block 636. Otherwise, control is passed to a function block 640.

[0069] The function block 636 is vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]] and passes control to a function block 638.

[0070] The function block 638 analyzes the filter coefficients for each YUV component, and passes control to a function block 640 .

[0071] The function block 640 increments the variable i, and returns control to the decision block 612 .

[0072] The function block 624 decodes each picture using the MPEG-4 AVC Standard or other single-view codec, and passes control to a function block 626. The function block 626 separates each view from the picture using a high-level syntax, and passes control to an end block 699.

[0073] Turning to FIG. 7, an exemplary method for encoding multiple view and depth pictures using the MPEG-4 AVC Standard is indicated generally by the reference numeral 700 .

[0074] The method 700 includes a start block 702. The start block 702 passes control to a function block 704. The function block 704 arranges each view and corresponding depth at a particular time point as a subpicture in a tiled format and passes control to a function block 706. The function block 706 sets the syntax element num_coded_views_minus1 and passes control to a function block 708. The function block 708 sets the syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 and passes control to a function block 710. The function block 710 sets a variable i equal to zero and passes control to a decision block 712. The decision block 712 determines whether the variable i is less than the number of views. If so, it passes control to a function block 714. Otherwise, it passes control to a function block 724.

[0075] The function block 714 sets the syntax element view_id[i] and passes control to a function block 716. The function block 716 sets the syntax element num_parts[view_id[i]] and passes control to a function block 718. The function block 718 sets the variable j equal to zero and passes control to a decision block 720. The decision block 720 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, it passes control to a function block 722. Otherwise, it passes control to a function block 728.

[0076] The function block 722 is depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j] , increments the variable j, and then returns control to decision block 720.

[0077] The function block 728 sets the syntax element upsample_view_flag[view_id[i]] and passes control to a decision block 730. The decision block 730 determines whether the current value of the syntax element upsample_view_flag[view_id[i]] is equal to 1. If so, then control is passed to a function block 732. Otherwise, control is passed to a decision block 734.

[0078] The function block 732 sets the syntax element upsample_filter[view_id[i]], and passes control to a decision block 734 .

[0079] The decision block 734 determines whether the current value of the syntax element upsample_filter[view_id[i]] is equal to 3. If so, then control is passed to a function block 736. Otherwise, control is passed to a function block 740.

[0080] Function block 736 is vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]] and passes control to a function block 738.

[0081] The function block 738 sets the filter coefficients for each YUV component, and passes control to a function block 740 .

[0082] The function block 740 increments the variable i, and returns control to the decision block 712 .

[0083] The function block 724 writes the above syntax elements to at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a supplemental enhancement information (SEI) message, a network abstraction layer (NAL) unit header, and a slice header, and passes control to a function block 726. The function block 726 encodes each picture using the MPEG-4 AVC Standard or another single-view codec, and passes control to an end block 799.

[0084] Turning to FIG. 8, an exemplary method for decoding multiple view and depth pictures using the MPEG-4 AVC Standard is indicated generally by the reference numeral 800 .

[0085] The method 800 includes a start block 802. The start block 802 passes control to a function block 804. The function block 804 parses the following syntax elements from at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a supplemental enhancement information (SEI) message, a network abstraction layer (NAL) unit header, and a slice header and passes control to a function block 806. The function block 806 parses the syntax element num_coded_views_minus1 and passes control to a function block 808. The function block 808 parses the syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 and passes control to a function block 810. The function block 810 sets a variable i equal to zero and passes control to a decision block 812. The decision block 812 determines whether the variable i is less than the number of views. If so, then control is passed to a function block 814. Otherwise, control is passed to a function block 824.

[0086] The function block 814 parses the syntax element view_id[i] and passes control to a function block 816. The function block 816 parses the syntax element num_parts_minus1[view_id[i]] and passes control to a function block 818. The function block 818 sets a variable j equal to zero and passes control to a decision block 820. The decision block 820 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, it passes control to a function block 822. Otherwise, it passes control to a function block 828.

[0087] The function block 822 is depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j] , increments the variable j, and then returns control to decision block 820.

[0088] The function block 828 sets the syntax element upsample_view_flag[view_id[i]] and passes control to a decision block 830. The decision block 830 determines whether the current value of the syntax element upsample_view_flag[view_id[i]] is equal to 1. If so, then control is passed to a function block 832. Otherwise, control is passed to a decision block 834.

[0089] The function block 832 parses the syntax element upsample_filter[view_id[i]], and passes control to a decision block 834 .

[0090] The decision block 834 determines whether the current value of the syntax element upsample_filter[view_id[i]] is equal to 3. If so, then control is passed to a function block 836. Otherwise, control is passed to a function block 840.

[0091] The function block 836 is vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]] and passes control to a function block 838.

[0092] The function block 838 analyzes the filter coefficients for each YUV component, and passes control to a function block 840 .

[0093] The function block 840 increments the variable i, and returns control to the decision block 812 .

[0094] The function block 824 decodes each picture using the MPEG-4 AVC Standard or other single-view codec and passes control to a function block 826. The function block 826 separates each view and corresponding depth from the picture using a high-level syntax and passes control to a function block 827. The function block 827 potentially performs view synthesis using the extracted view and depth signals and passes control to an end block 899.

[0095] 7 and 8, FIG. 9 shows an example of a depth signal 900 (where the depth is provided as a pixel value for each corresponding location in an image (not shown)). Additionally, FIG. 10 shows an example of two depth signals contained in a tile 1000. The top right portion of the tile 1000 is a depth signal having a depth value corresponding to the top left image of the tile 1000. The bottom right portion of the tile 1000 is a depth signal having a depth value corresponding to the bottom left image of the tile 1000.

[0096] Turning to Figure 11, an example of five views tiled on a single frame is shown generally by the reference numeral 1100. The top four views are all in the normal orientation. The fifth view is also in the normal orientation, but is split into two parts along the bottom of the tile 1100. The left part of the fifth view shows the "top" of the fifth view, and the right part of the fifth view shows the "bottom" of the fifth view.

[0097] Encoding / Decoding using Multiview Video Encoding / Decoding Standards / Recommendations Turning to FIG. 12, an exemplary Multi-view Video Coding (MVC) encoder is indicated generally by the reference numeral 1200. The encoder 1200 includes a combiner 1205 having an output connected in signal communication with an input of a transformer 1210. The output of the transformer 1210 is connected in signal communication with an input of a quantizer 1215. The output of the quantizer 1215 is connected in signal communication with an input of an entropy coder 1220 and an input of an inverse quantizer 1225. The output of the inverse quantizer 1225 is connected in signal communication with an input of an inverse transformer 1230. The output of the inverse transformer 1230 is connected in signal communication with a first non-inverting input of a combiner 1235. The output of the combiner 1235 is connected in signal communication with an input of an intra predictor 1245 and an input of a deblocking filter 1250. The output of the deblocking filter 1250 is connected in signal communication with an input of a reference picture store (for view i) 1255. An output of the reference picture store 1255 is connected in signal communication with a first input of a motion compensator 1275 and a first input of a motion estimator 1280. An output of the motion estimator 1280 is connected in signal communication with a second input of the motion compensator 1275. An output of the reference picture store 1260 (for other views) is connected in signal communication with a first input of the disparity estimator 1270 and a first input of the disparity compensator 1265. An output of the disparity estimator 1270 is connected in signal communication with a second input of the disparity compensator 1265.

[0098] An output of the entropy decoder 1220 is available as an output of the encoder 1200. A non-inverting input of the combiner 1205 is available as an input of the encoder 1200 and is connected in signal communication with a second input of the disparity estimator 1270 and a second input of the motion estimator 1280. An output of the switch 1285 is connected in signal communication with a second non-inverting input of the combiner 1235 and an inverting input of the combiner 1205. The switch 1285 has a first input connected in signal communication with an output of the motion compensator 1275, a second input connected in signal communication with an output of the disparity compensator 1265, and a third input connected in signal communication with an output of the intra predictor 1245.

[0099] The mode decision module 1240 has an output connected to the switch 1285 to control which input is selected by the switch 1285 .

[0100] Turning to FIG. 13, an exemplary Multi-view Video Coding (MVC) decoder is indicated generally by the reference numeral 1300. The decoder 1300 includes an entropy decoder 1305 having an output connected in signal communication with an input of an inverse quantizer 1310. The output of the inverse quantizer is connected in signal communication with an input of an inverse transformer 1315. The output of the inverse transformer 1315 is connected in signal communication with a first non-inverting input of a combiner 1320. The output of the combiner 1320 is connected in signal communication with an input of a deblocking filter 1325 and an input of an intra predictor 1330. The output of the deblocking filter 1325 is connected in signal communication with an input of a reference picture store 1340 (for view i). The output of the reference picture store 1340 is connected in signal communication with a first input of a motion compensator 1335.

[0101] An output of the reference picture store (for other views) 1345 is connected in signal communication with a first input of a disparity compensator 1350 .

[0102] An input of the entropy coder 1305 is available as an input to the decoder 1300 to receive a residual bitstream. Additionally, an input of the entropy coder 1360 is also available as an input to the decoder 1300 to receive control syntax to control which input is selected by the switch 1355. Additionally, a second input of the motion compensator 1355 is available as an input to the decoder 1300 to receive a motion vector. Additionally, a second input of the disparity compensator 1350 is available as an input to the decoder 1300 to receive a disparity vector.

[0103] An output of the switch 1355 is connected in signal communication with a second non-inverting input of the combiner 1320. A first input of the switch 1355 is connected in signal communication with an output of the disparity compensator 1350. A second input of the switch 1355 is connected in signal communication with an output of the motion compensator 1335. A third input of the switch 1355 is connected in signal communication with an output of the intra predictor 1330. An output of the mode module 1360 is connected in signal communication with the switch 1355 for controlling which input is selected by the switch 1355. An output of the deblocking filter 1325 is available as an output of the decoder 1300.

[0104] Turning to FIG. 14, an exemplary method for processing pictures of multiple views in preparation for encoding the pictures using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 1400 .

[0105] The method 1400 includes a start block 1405. The start block 1405 passes control to a function block 1410. The function block 1410 arranges every N views out of a total of M views at a particular time point as a super picture in a tile format and passes control to a function block 1415. The function block 1415 sets a syntax element num_coded_views_minus1 and passes control to a function block 1420. The function block 1420 sets a syntax element view_id[i] for all (num_coded_views_minus1 + 1) views and passes control to a function block 1425. The function block 1425 sets inter-view reference dependency information of the anchor picture and passes control to a function block 1430. The function block 1430 sets inter-view reference dependency information of non-anchor pictures and passes control to a function block 1435. The function block 1435 sets the syntax element pseudo_view_present_flag and passes control to a decision block 1440. The decision block 1440 determines whether the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block 1445. Otherwise, control is passed to an end block 1499.

[0106] The function block 1445 is tiling_mode; org_pic_width_in_mbs_minus1 ; and org_pic_height_in_mbs_minus1 and passes control to a function block 1450. The function block 1450 calls the syntax element pseudo_view_info(view_id) for each coded view, and passes control to an end block 1499.

[0107] Turning to FIG. 15, an exemplary method for encoding pictures for multiple views using the Multi-view Video Coding (MVC) extension to the MPEG-4 AVC Standard is indicated generally by the reference numeral 1500 .

[0108] The method 1500 includes a start block 1502. The start block 1502 has an input parameter pseudo_view_id and passes control to a function block 1504. The function block 1504 sets a syntax element num_sub_views_minus1 and passes control to a function block 1506. The function block 1506 sets a variable i equal to zero and passes control to a decision block 1508. The decision block 1508 determines whether the variable i is less than the number of sub_views. If so, then control is passed to a function block 1510. Otherwise, control is passed to a function block 1520.

[0109] The function block 1510 sets the syntax element sub_view_id[i] and passes control to a function block 1512. The function block 1512 sets the syntax element num_parts_minus1[sub_view_id[i]] and passes control to a function block 1514. The function block 1514 sets a variable j equal to zero and passes control to a decision block 1516. The decision block 1516 determines whether the variable j is less than the syntax element num_parts[sub_view_id[i]]. If so, it passes control to a function block 1518. Otherwise, it passes control to a decision block 1522.

[0110] The function block 1518 is loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i][j]. , increments the variable j, and returns control to decision block 1516.

[0111] The function block 1520 encodes the current picture of the current view using Multiview Video Coding (MVC), and passes control to an end block 1599.

[0112] The decision block 1522 determines whether the syntax element tiling_mode is equal to zero. If so, then control is passed to a function block 1524. Otherwise, control is passed to a function block 1538.

[0113] The function block 1524 sets the syntax element flip_dir[sub_view_id[i]] and the syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block 1526. The decision block 1526 determines whether the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to 1. If so, then control is passed to a function block 1528. Otherwise, control is passed to a decision block 1530.

[0114] The function block 1528 sets the syntax element upsample_filter[sub_view_id[i]] and passes control to a decision block 1530. The decision block 1530 determines whether the value of the syntax element upsample_filter[sub_view_id[i]] is equal to 3. If so, then control is passed to a function block 1532. Otherwise, control is passed to a function block 1536.

[0115] The function block 1532 is vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]] and passes control to a function block 1534. The function block 1534 sets the filter coefficients for each YUV component and passes control to a function block 1536.

[0116] The function block 1536 increments the variable i, and returns control to the decision block 1508 .

[0117] The function block 1538 sets the syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]] and passes control to a function block 1540. The function block 1540 sets the variable j equal to zero and passes control to a decision block 1542. The decision block 1542 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, it passes control to a function block 1544. Otherwise, it passes control to a function block 1536.

[0118] The function block 1544 sets the syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]] and passes control to a function block 1546. The function block 1546 sets all the coefficients of the pixel tiling filter and passes control to the function block 1536.

[0119] Turning to FIG. 16, an exemplary method for processing pictures of multiple views in preparation for decoding of the pictures using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 1600 .

[0120] The method 1600 includes a start block 1605, which passes control to a function block 1615. The function block 1615 parses the syntax element num_coded_views_minus1 and passes control to a function block 1620. The function block 1620 parses the syntax elements view_id[i] of all (num_coded_views_minus1 + 1) views and passes control to a function block 1625. The function block 1625 parses the inter-view reference dependency information of the anchor picture and passes control to a function block 1630. The function block 1630 parses the inter-view reference dependency information of the non-anchor pictures and passes control to a function block 1635. The function block 1635 parses the syntax element pseudo_view_present_flag and passes control to a decision block 1640. The decision block 1640 determines whether the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block 1645. Otherwise, control is passed to an end block 1699.

[0121] The function block 1645 is tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1 and passes control to a function block 1650. The function block 1650 calls the syntax element pseudo_view_info(view_id) for each coded view, and passes control to an end block 1699.

[0122] Turning to FIG. 17, an exemplary method for decoding pictures of multiple views using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 1700 .

[0123] The method 1700 includes a start block 1702. The start block 1702 begins with an input parameter, pseudo_view_id, and passes control to a function block 1704. The function block 1704 parses the syntax element, num_sub_views_minus1, and passes control to a function block 1706. The function block 1706 sets a variable i equal to zero and passes control to a decision block 1708. The decision block 1708 determines whether the variable i is less than the number of sub_views. If so, then control is passed to a function block 1710. Otherwise, control is passed to a function block 1720.

[0124] The function block 1710 parses the syntax element sub_view_id[i] and passes control to a function block 1712. The function block 1712 parses the syntax element num_parts_minus1[sub_view_id[i]] and passes control to a function block 1714. The function block 1714 sets a variable j equal to zero and passes control to a decision block 1716. The decision block 1716 determines whether the variable j is less than the syntax element num_parts_minus1[sub_view_id[i]]]. If so, then control is passed to a function block 1718. Otherwise, control is passed to a decision block 1722.

[0125] Function block 1718 is loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i]][j] , increments the variable j, and returns control to decision block 1716.

[0126] The function block 1720 decodes the current picture for the current view using Multiview Video Coding (MVC) and passes control to a function block 1721. The function block 1721 separates each view from the picture using a high-level syntax and passes control to an end block 1799.

[0127] The separation of each view from the decoded picture is done using a high-level syntax indicated in the bitstream, which may indicate the possible orientations (and corresponding possible depths) and exact positions of the views present in the picture.

[0128] The decision block 1722 determines whether the syntax element tiling_mode is equal to zero. If so, then control is passed to a function block 1724. Otherwise, control is passed to a function block 1738.

[0129] The function block 1724 parses the syntax element flip_dir[sub_view_id[i]] and the syntax element upsample_view_flag[sub_view_id[i]] and passes control to a decision block 1726. The decision block 1726 determines whether the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to 1. If so, then control is passed to a function block 1728. Otherwise, control is passed to a decision block 1730.

[0130] The function block 1728 parses the syntax element upsample_filter[sub_view_id[i]] and passes control to a decision block 1730. The decision block 1730 determines whether the value of the syntax element upsample_filter[sub_view_id[i]] is equal to 3. If so, then control is passed to a function block 1732. Otherwise, control is passed to a function block 1736.

[0131] The function block 1732 is vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]] and passes control to a function block 1734. The function block 1734 parses the filter coefficients for each YUV component and passes control to a function block 1736.

[0132] The function block 1736 increments the variable i, and returns control to the decision block 1708 .

[0133] The function block 1738 parses the syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]] and passes control to a function block 1740. The function block 1740 sets the variable j equal to zero and passes control to a decision block 1742. The decision block 1742 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block 1744. Otherwise, control is passed to the function block 1736.

[0134] The function block 1744 parses the syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]] and passes control to the function block 1746. The function block 1776 parses all the coefficients of the pixel tiling filter and passes control to the function block 1736.

[0135] Turning to FIG. 18, an exemplary method for processing pictures of multiple views and depths in preparation for encoding the pictures using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 1800.

[0136] The method 1800 includes a start block 1805. The start block 1805 passes control to a function block 1810. The function block 1810 arranges, at a particular point in time, every N views and depth maps out of a total of M views and depth maps as a super-picture in tile format and passes control to a function block 1815. The function block 1815 sets the syntax element num_coded_views_minus1 and passes control to a function block 1820. The function block 1820 sets the syntax element view_id[i] for all (num_coded_views_minus1 + 1) depths corresponding to view_id[i] and passes control to a function block 1825. The function block 1825 sets inter-view reference dependency information for the anchor depth picture and passes control to a function block 1830. The function block 1830 sets inter-view reference dependency information for the non-anchor depth picture and passes control to a function block 1835. The function block 1835 sets the syntax element pseudo_view_present_flag and passes control to a decision block 1840. The decision block 1840 determines whether the current value of the syntax element pseudo_view_present_flag is equal to true. If so, it passes control to a function block 1845. Otherwise, it passes control to an end block 1899.

[0137] Function block 1845 is tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1 and passes control to a function block 1850. The function block 1850 calls the syntax element pseudo_view_info(view_id) for each coded view, and passes control to an end block 1899.

[0138] Turning to FIG. 19, an exemplary method for encoding pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 1900 .

[0139] The method 1900 includes a start block 1902. The start block 1902 passes control to a function block 1904. The function block 1904 sets the syntax element num_sub_views_minus1 and passes control to a function block 1906. The function block 1906 sets a variable i equal to zero and passes control to a decision block 1908. The decision block 1908 determines whether the variable i is less than the number of sub_views. If so, then control is passed to a function block 1910. Otherwise, control is passed to a function block 1920.

[0140] The function block 1910 sets the syntax element sub_view_id[i] and passes control to a function block 1912. The function block 1912 sets the syntax element num_parts_minus1[sub_view_id[i]] and passes control to a function block 1914. The function block 1914 sets a variable j equal to zero and passes control to a decision block 1916. The decision block 1916 determines whether the variable j is less than the syntax element num_parts_minus1[sub_view_id[i]]. If so, it passes control to a function block 1918. Otherwise, it passes control to a decision block 1922.

[0141] Function block 1918 is loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i]][j] , increments the variable j, and returns control to decision block 1916.

[0142] The function block 1920 encodes the current depth of the current view using multiview video coding (MVC) and passes control to an end block 1999. A depth signal can be encoded similarly to the way its corresponding video signal is encoded. For example, a view's depth signal can be included on a tile that contains other depth signals only, video signals only, or depth and video signals. The tile (pseudo-view) is then treated as a single view in MVC, and there may also be other tiles that are treated as other views in MVC.

[0143] The decision block 1922 determines whether the syntax element tiling_mode is equal to zero. If so, then control is passed to a function block 1924. Otherwise, control is passed to a function block 1938.

[0144] The function block 1924 sets the syntax element flip_dir[sub_view_id[i]] and the syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block 1926. The decision block 1926 determines whether the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to 1. If so, then control is passed to a function block 1928. Otherwise, control is passed to a decision block 1930.

[0145] The function block 1928 sets the syntax element upsample_filter[sub_view_id[i]] and passes control to a decision block 1930. The decision block 1930 determines whether the value of the syntax element upsample_filter[sub_view_id[i]] is equal to 3. If so, then control is passed to a function block 1932. Otherwise, control is passed to a function block 1936.

[0146] The function block 1932 is vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]] and passes control to a function block 1934. The function block 1934 sets the filter coefficients for each YUV component and passes control to a function block 1936.

[0147] The function block 1936 increments the variable i, and passes control to the decision block 1908 .

[0148] The function block 1938 sets the syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]] and passes control to a function block 1940. The function block 1940 sets the variable j to zero and passes control to a decision block 1942. The decision block 1942 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, it passes control to a function block 1944. Otherwise, it passes control to the function block 1936.

[0149] The function block 1944 sets the syntax element num_pixel_tiling_fιlter_coeffs_minus1[sub_view_id[i]] and passes control to a function block 1946. The function block 1946 sets all the coefficients of the pixel tiling filter and passes control to the function block 1936.

[0150] Turning to FIG. 20, an exemplary method for processing pictures of multiple views and depths in preparation for decoding the pictures using the Multiview Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 2000.

[0151] The method 2000 includes a start block 2005, which passes control to a function block 2015. The function block 2015 parses a syntax element num_coded_views_minus1 and passes control to a function block 2020. The function block 2020 parses syntax elements view_id[i] for all (num_coded_views_minus1 + 1) depths corresponding to view_id[i] and passes control to a function block 2025. The function block 2025 parses inter-view reference dependency information of an anchor depth picture and passes control to a function block 2030. The function block 2030 parses inter-view reference dependency information of a non-anchor depth picture and passes control to a function block 2035. The function block 2035 parses a syntax element pseudo_view_present_flag and passes control to a decision block 2040. The decision block 2040 determines whether the current value of the syntax element pseudo_view_present_flag is equal to true. If so, control is passed to a function block 2045. Otherwise, control is passed to a function block 2099.

[0152] Function block 2045 is tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1 and passes control to a function block 2050. The function block 2050 calls the syntax element pseudo_view_info(view_id) for each coded view, and passes control to an end block 2099.

[0153] Turning to FIG. 21, an exemplary method for decoding pictures of multiple views and depths using the Multi-view Video Coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral 2100 .

[0154] The method 2100 includes a start block 2102. The start block 2102 begins with an input parameter, pseudo_view_id, and passes control to a function block 2104. The function block 2104 parses a syntax element, num_sub_views_minus1, and passes control to a function block 2106. The function block 2106 sets a variable i equal to zero and passes control to a decision block 2108. The decision block 2108 determines whether the variable i is less than the number of sub_views. If so, then control is passed to a function block 2110. Otherwise, control is passed to a function block 2120.

[0155] The function block 2110 parses the syntax element sub_view_id[i] and passes control to a function block 2112. The function block 2112 parses the syntax element num_parts_minus1[sub_view_id[i]] and passes control to a function block 2114. The function block 2114 sets a variable j equal to zero and passes control to a decision block 2116. The decision block 2116 determines whether the variable j is less than the syntax element num_parts_minus1[sub_view_id[i]]. If so, it passes control to a function block 2118. Otherwise, it passes control to a decision block 2122.

[0156] The function block 2118 is loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i]][j] , increments the variable j, and returns control to decision block 2116.

[0157] The function block 2120 decodes the current picture using Multiview Video Coding (MVC) and passes control to a function block 2121. The function block 2121 separates each view from the picture using a high-level syntax and passes control to an end block 2199. The separation of each view using a high-level syntax is described above.

[0158] The decision block 2122 determines whether the syntax element tiling_mode is equal to zero. If so, then control is passed to a function block 2124. Otherwise, control is passed to a function block 2138.

[0159] The function block 2124 parses the syntax element flip_dir[sub_view_id[i]] and the syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block 2126. The decision block 2126 determines whether the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to 1. If so, then control is passed to a function block 2128. Otherwise, control is passed to a decision block 2130.

[0160] The function block 2128 parses the syntax element upsample_filter[sub_view_id[i]] and passes control to a decision block 2130. The decision block 2130 determines whether the value of the syntax element upsample_filter[sub_view_id[i]] is equal to 3. If so, then control is passed to a function block 2132. Otherwise, control is passed to a function block 2136.

[0161] The function block 2132 is vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]] and passes control to a function block 2134. The function block 2134 parses the filter coefficients for each YUV component and passes control to a function block 2136.

[0162] The function block 2136 increments the variable i, and returns control to the decision block 2108 .

[0163] The function block 2138 parses the syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]] and passes control to a function block 2140. The function block 2140 sets the variable j equal to zero and passes control to a decision block 2142. The decision block 2142 determines whether the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block 2144. Otherwise, control is passed to a function block 2136.

[0164] The function block 2144 parses the syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]] and passes control to a function block 2146. The function block 2146 parses all coefficients of the pixel tiling filter and passes control to the function block 2136.

[0165] Turning to Figure 22, an example of tiling at the pixel level is indicated generally by the reference numeral 2200. Figure 22 is further described below.

[0166] View tiling using MPEG-4 AVC or MVC One application of multi-view video coding is free-viewpoint TV (or FTV). Such applications require the user to be able to move freely between two or more views. To achieve this, it is necessary to interpolate or synthesize a "virtual" view between the two views. There are several techniques to perform view interpolation. One of these techniques uses depth for view interpolation / synthesis.

[0167] Each view may have an associated depth signal. Thus, depth may be considered as another form of video signal. FIG. 9 shows an example of a depth signal 900. To enable applications such as FTV, the depth signal is transmitted along with the video signal. In the proposed tiling framework, the depth signal may also be added as one of the tiles. FIG. 10 shows an example of a depth signal added as a tile. The depth signal / tile is shown on the right side of FIG. 10.

[0168] When depth is coded as tiles across a frame, the high-level syntax shall indicate which tiles are depth signals so that the renderer can use the depth signals appropriately.

[0169] If the input sequence (such as that shown in FIG. 1) was coded using an MPEG-4 AVC standard encoder (or an encoder corresponding to another video coding standard and / or recommendation), the proposed high-level syntax may be present, for example, in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), slice header, and / or Supplemental Enhancement Information (SEI) message. An example implementation of the proposed method is shown in Table 1, where the syntax is present in the Supplemental Enhancement Information (SEI) message.

[0170] If the pseudo-view input sequence (such as that shown in Figure 1) is coded using the Multiview Video Coding (MVC) extension of the MPEG-4 AVC Standard encoder (or a Multiview Video Coding Standard compliant encoder for another video coding standard and / or recommendation), the proposed high-level syntax may be present in the SPS, PPS, slice header, SEI message, or a specified profile.

[0171] One embodiment of the proposed method is shown in Table 1. Table 1 shows the syntax elements present in a Sequence Parameter Set (SPS) structure, including syntax elements proposed by an embodiment of the present principles.

[0172] [Table 1] Table 2 shows the syntax elements of the pseudo_view_info syntax element of Table 1, in accordance with an embodiment of the present principles.

[0173] [Table 2] Semantics of the syntax elements presented in Tables 1 and 2 pseudo_view_present_flag equal to true indicates that a particular view is a superview of multiple subviews.

[0174] A tiling_mode equal to 0 indicates that the subview is tiled at the picture level. A value of 1 indicates that tiling is done at the pixel level.

[0175] The new SEI message may use SEI payload type values ​​that are not used in the MPEG-4 AVC standard or in extensions to the MPEG-4 AVC standard. The new SEI message contains several syntax elements with the following semantics:

[0176] num_coded_views_minus1 plus 1 indicates the number of coded views supported by the bitstream. The value of num_coded_views_minus1 is in the range 0 to 1023, inclusive.

[0177] org_pic_width_in_mbs_minus1 plus 1 specifies the picture width in each view in macroblock units. The picture width variable, in macroblock units, is derived as follows:

[0178] PicWidthInMbs = org_pic_width_in_mbs_minus1 + 1 The picture width variable for the luma component is derived as follows:

[0179] PicWidthInSamplesL = PicWidthInMbs * 16 The picture width variables for the chroma components are derived as follows:

[0180] PicWidthInSamplesC = PicWidthInMbs * MbWidthC org_pic_height_in_mbs_minus1 plus 1 specifies the picture height in each view in macroblock units. The picture height variable, in macroblocks, is derived as follows:

[0181] PicHeightInMbs = org_pic_height_in_mbs_minus1 + 1 The picture height variable for the luma component is derived as follows:

[0182] PicHeightInSamplesL = PicHeightInMbs * 16 The picture height variables for the chroma components are derived as follows:

[0183] PicHeightInSamplesC = PicHeightInMbs * MbHeightC num_sub_yiews_minus1 plus 1 indicates the number of coded subviews included in the current view. The value of num_coded_views_minus1 is in the range of 0 to 1023, inclusive.

[0184] sub_view_id[i] specifies the sub_view_id of the subview in the decoding order indicated by i.

[0185] num_parts[sub_yiew_id[i]] specifies the number of parts into which the picture of sub_view_id[i] is divided.

[0186] loc_left_offset[sub_view_id[i]][j] and loc_top_offset[sub_view_id[i]][j] specify the location in left pixel offset and top pixel offset, respectively, where the current portion j is located in the final reconstructed picture of the view, and sub_view_id is equal to sub_view_id[i].

[0187] view_id[i] specifies the view id of the view, and the coding order is indicated by i.

[0188] frame_crop_left_offset[view_id[i]][j], frame_crop_right_offset[view_id[i]][j], frame_crop_top_offset[view_id[i]][j] and frame_crop_bottom_offset[view_id[i]][j] specify the sample of the picture in the coded video sequence that is part of num_part j and view_id i for the rectangular region specified in frame coordinates to output.

[0189] The variables CropUnitX and CropUnitY are derived as follows:

[0190] If chroma_format_idc is equal to 0, CropUnitX and CropUnitY are derived as follows:

[0191] CropUnitX = 1 CropUnitY = 2 - frame_mbs_only_flag Otherwise (if chroma_format_idc is equal to 1, 2 or 3), CropUnitX and CropUnitY are derived as follows:

[0192] CropUnitX = SubWidthC CropUnitY = SubHeightC * (2 -frame_mbs_only_flag) The frame crop rectangle is CropUnitX * frame_crop_left_offset〜PicWidthInSamplesL - (CropUnitX * Contains luma samples with horizontal frame coordinates of CropUnitY * frame_crop_right_offset + 1. Vertical frame coordinates are greater than or equal to CropUnitY * frame_crop_top_offset (16 * FrameHeightInMbs) - (CropUnitY * The value of frame_crop_left_offset must be in the range 0 to (PicWidthInSamplesL / CropUnitX) - (frame_crop_right_offset + 1), and the value of frame_crop_top_offset must be in the range 0 to (16 * FrameHeightInMbs / CropUnitY) - (frame_crop_bottom_offset + 1).

[0193] If chroma_format_idc is not equal to 0, the corresponding default samples of the two chroma arrays are samples with frame coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the frame coordinates of the default luma sample.

[0194] For a decoded field, the defined samples of the decoded field are those that fall within a rectangle defined in frame coordinates.

[0195] num_parts[view_id[i]] specifies the number of parts into which the picture of view_id[i] is divided.

[0196] depth_flag[view_id[i]] specifies whether the current part is a depth signal. If depth_flag is equal to 0, the current part is not a depth signal. If depth_flag is equal to 1, the current part is a depth signal associated with the view identified by view_id[i].

[0197] flip_dir[sub_view_id[i]][j] specifies the flip direction of the current part: flip_dir equal to 0 means no flip, flip_dir equal to 1 indicates a horizontal flip, flip_dir equal to 2 indicates a vertical flip, and flip_dir equal to 3 indicates a horizontal and vertical flip.

[0198] flip_dir[view_id[i]][j] specifies the flip direction of the current part, where flip_dir equal to 0 indicates no flip, flip_dir equal to 1 indicates a flip in the horizontal direction, flip_dir equal to 2 indicates a flip in the vertical direction, and flip_dir equal to 3 indicates a flip in both the horizontal and vertical directions.

[0199] loc_left_offset[view_id[i]][j] and loc_top_offset[view_id[i]][j] specify the location in pixel offsets where the current part j is located in the final reconstructed picture of the view, where view_id is equal to id[i].

[0200] upsample_view_flag[view_id[i]] indicates whether a picture belonging to the view specified by view_id[i] needs to be upsampled or not. upsample_view_flag[view_id[i]] equal to 0 specifies that a picture whose view_id is equal to view_id[i] will not be upsampled, and upsample_view_flag[view_id[i]] equal to 1 specifies that a picture whose view_id is equal to view_id[i] will be upsampled.

[0201] upsample_filter[view_id[i]] indicates the type of filter to use for upsampling. upsample_filter[view_id[i]] equal to 0 indicates that a 6-tap AVC filter should be used, upsample_filter[view_id[i]] equal to 1 indicates that a 4-tap SVC filter should be used, upsample_filter[view_id[i]] equal to 2 indicates that a bilinear filter should be used, and upsample_filter[view_id[i]] equal to 3 indicates that custom filter coefficients should be sent. upsample_filter[view_id[i]] is set to 0 if not present. In this example, a customized 2D filter is used. This can be easily extended to a 1D filter and to certain other nonlinear filters.

[0202] vert_dim[view_id[i]] specifies the vertical dimension of the custom 2D filter.

[0203] hor_dim[view_id[i]] specifies the horizontal dimension of the custom 2D filter.

[0204] quantizer[view_id[i]] specifies the quantization factor for each filter coefficient.

[0205] filter_coeffs[view_id[i]] [yuv][y][x] specifies the quantized filter coefficients, where yuv represents the component the filter coefficient falls into, with yuv equal to 0 specifying the Y component, yuv equal to 1 specifying the U component, and yuv equal to 2 specifying the V component.

[0206] pixel_dist_x[sub_view_id[i]] and pixel_dist_y[sub_view_id[i]] specify the horizontal and vertical distances in the final reconstructed pseudo-view between neighboring pixels in the view, respectively. sub_view_id is equal to sub_view_id[i].

[0207] num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i][j] plus one indicates the number of filter coefficients when tiling mode is set equal to 1.

[0208] pixel_tiling_filter_coeffs[sub_view_id[i][j] represents the filter coefficients needed to represent a filter that can be used to filter the tiled picture.

[0209] Pixel-level tiling example Turning to Figure 22, two examples illustrating the synthesis of pseudo-views by tiling pixels from four views are respectively designated by reference numerals 2210 and 2220. The four views together are designated by reference numeral 2250. The syntax values ​​for the first example in Figure 22 are set forth below in Table 3.

[0210] [Table 3] The syntax values ​​of the second example in FIG. 22 are all the same except for two syntax elements: loc_left_offset[3][0] equals 5 and loc_top_offset[3][0] equals 3.

[0211] The offset indicates that the pixels corresponding to a view should start at a particular offset location. This is shown in FIG. 22 (2220). This can be done, for example, when two views generate images of a common object that is shifted between views. For example, if a first camera and a second camera (representing the first and second views) take a picture of an object, the object may appear to be shifted five pixels to the right in the second view compared to the first view. This means that pixel (i-5, j) in the first view corresponds to pixel (i, j) in the second view. If the pixels of the two views are only tiled pixel-by-pixel, there may be little correlation between neighboring pixels within the tile, and the specific coding gain may be small. Conversely, shifting the tiling so that pixel (i-5, j) from view 1 is placed next to pixel (i, j) from view 2 can increase spatial correlation, and the spatial coding gain can also be increased. This is the case, for example, because corresponding pixels of the object in the first and second views are tiled next to each other.

[0212] Therefore, the presence of loc_left_offset and loc_top_offset can improve coding efficiency. The offset information can be obtained by external means. For example, the camera position information or the global disparity vector between views can be used to determine such offset information.

[0213] As a result of the offsetting, some of the pixels in the pseudo-views are not assigned pixel values ​​from either view. Continuing with the example above, if we tile pixel (i-5, j) from view 1 alongside pixel (i, j) from view 2, then for i=0 …For a value of 4, pixel (i-5, j) from view 1 of the tile object does not exist, so such pixel is empty in the tile. For pixels in the pseudo-view (tile) that are not assigned a pixel value from any view, at least one implementation uses an interpolation procedure similar to the sub-pixel interpolation procedure of motion compensation in AVC. That is, empty tile pixels can be interpolated from neighboring pixels. Such interpolation can result in greater spatial correlation in the tile and greater coding gain for the tile.

[0214] In video coding, it is possible to choose different coding types for each picture, such as I-pictures, P-pictures, and B-pictures. Furthermore, in the case of multi-view video coding, anchor and non-anchor pictures are defined. In one embodiment, we propose that grouping decisions can be made based on picture types. Such grouping information is signaled in the high-level syntax.

[0215] Turning to FIG. 11, an example of five views tiled onto a single frame is generally designated by the reference numeral 1100. In particular, a ballroom sequence is shown with five views tiled onto a single frame. Additionally, the fifth view is split into two parts so that it can be placed onto a rectangular frame. Here, each view is QVGA sized, with a total frame dimension of 640x640. Since 600 is not a multiple of 16, it is increased to 608.

[0216] For this example, a possible SEI message could be as shown in Table 4.

[0217] [Table 4-1] [Table 4-2] Table 5 shows the general syntax structure for transmitting the example multiview information shown in Table 4.

[0218] [Table 5] 23, there is shown a video processing device 2300. The video processing device 2300 may be, for example, a set-top box or other device that receives encoded video and provides decoded video, for example, for display to a user or for storage. Thus, the device 2300 may provide its output to a television set, a computer monitor, or a computer or other processing device.

[0219] The apparatus 2300 includes a decoder 2310 that receives a data signal 2320. The data signal 2320 may include, for example, an AVC or MVC compatible stream. The decoder 2310 decodes all or a portion of the received signal 2320 and provides a decoded video signal 2330 and tiling information 2340 as outputs. The decoded video 2330 and tiling information 2340 are provided to a selector 2350. The apparatus 2300 also includes a user interface that receives a user input 2370. The user interface 2360 provides a picture selection signal 2380 to the selector 2350 based on the user input 2370. The picture selection signal 2380 and the user input 2370 indicate which of a plurality of pictures the user wants to display. The selector 2350 provides the selected picture as an output 2390. Selector 2350 uses picture selection information 2380 to select which of the pictures in decoded video 2330 to provide as output 2390. Selector 2350 uses tiling information 2340 to locate the selected picture in decoded video 2330.

[0220] In various implementations, the selector 2350 includes a user interface 2360, while in other implementations, the user interface 2360 is not required because the selector 2350 receives user input 2370 directly without a separate interface function being performed. The selector 2350 can be implemented, for example, in software or as an integrated circuit. The selector 2350 can also incorporate the decoder 2310.

[0221] More generally, the decoders of various implementations described herein may provide a decoded output that includes an entire tile. Additionally, or alternatively, the decoders may provide a decoded output that includes only one or more selected pictures (e.g., image or depth signals) from the tile.

[0222] As mentioned above, the high-level syntax can be used to perform signaling in accordance with one or more embodiments of the present principles. For example, but not limited to, the high-level syntax can be used to signal any of the following: the number of coded views present in a larger frame; the original width and height of all of the views; for each coded view, a view identifier corresponding to the view; for each coded view, the number of portions the view's frame is divided into; the flip direction for each portion of the view (e.g., no flip, horizontal flip only, vertical flip only, or horizontal and vertical flip); for each portion of the view, the left position (in number of pixels or macroblocks) of the current portion in the last frame of the view; for each portion of the view, the top position (in number of pixels or macroblocks) of the current portion in the last frame of the view; for each portion of the view, the left position (in number of pixels or macroblocks) of the cropping window in the current larger decoded / coded frame; for each portion of the view, the top position (in number of pixels or macroblocks) of the cropping window in the current larger decoded / coded frame. The following information is provided for each view: the right position (in number of pixels or macroblocks) of the cropping window in the current large decoding / coding frame, for each portion of the view, the top position (in number of pixels or macroblocks) of the cropping window in the current large decoding / coding frame, and the bottom position (in number of pixels or macroblocks) of the cropping window in the current large decoding / coding frame for each portion of the view; and for each coded view, whether the view needs to be upsampled before output (if upsampling is required, a high-level syntax can be used to indicate the upsampling method, including but not limited to a 6-tap AVC filter, a 4-tap SVC filter, a bilinear filter, or a custom 1D, 2D linear or nonlinear filter).

[0223] The terms "encoder" and "decoder" refer to a general structure and are not limited to any particular function or configuration. For example, a decoder may receive a modulated carrier wave carrying an encoded bitstream, demodulate the encoded bitstream, and decode the bitstream.

[0224] Various methods are described. Many of the foregoing methods are described in detail to provide a thorough disclosure. However, variations are contemplated that may alter one or more of the specific features described in the foregoing methods. Furthermore, many of the features described are known in the art and, therefore, will not be described in detail.

[0225] Furthermore, some implementations refer to using a high-level syntax to send particular information, while other implementations use a lower-level syntax, or indeed an entirely different mechanism to provide the same information (or variations of that information) (e.g., sending the information as part of the encoded data).

[0226] Various implementations provide tiling and appropriate signaling to enable multiple views (generally, pictures) to be tiled into a single picture, encoded as a single picture, and transmitted as a single picture. The signaling information may enable a post-processor to separate the views / pictures. Furthermore, the tiled pictures may be views, but at least one of the pictures may be depth information. The above implementations may provide one or more advantages. For example, a user may want to display multiple views in a tiled manner, and the above various implementations provide an efficient way to encode, transmit, or store the views by encoding them in a tiled manner and tiling them before transmitting / storing.

[0227] Implementations of tiling multiple views in the context of AVC and / or MVC also provide additional benefits. Since AVC is ostensibly used for only a single view, no additional views are expected. However, the AVC-based implementations described above can provide multiple views in an AVC environment, since the tiled views can be arranged in such a way that the decoder knows that the tiled pictures belong to different views (e.g., the left and right pictures in the pseudo-view are view 1, the top right picture is view 2, etc.).

[0228] Furthermore, because MVC already includes multiple views, multiple views are not expected to be included in a single pseudo-view. Furthermore, MVC has a limit on the number of views that can be supported, and the aforementioned MVC-based implementation effectively increases the number of views that can be supported by allowing additional views to be tiled (as in AVC-based implementations). For example, each pseudo-view may correspond to one of MVC's supported views, and the decoder may know that each "supported view" actually includes four views in a pre-arranged tiling order. Thus, in the aforementioned implementation, the number of possible views is four times the number of "supported views."

[0229] The implementations described herein may be implemented, for example, as a method or process, an apparatus, or a software program. Even if described in the context of only one type of implementation (e.g., only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in an apparatus (e.g., a processor), which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processing devices also include communication devices (e.g., computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices) that facilitate communication of information between end users.

[0230] Implementations of the various processes and configurations described herein may be implemented in a variety of devices or applications, particularly those associated with, for example, data encoding and decoding. Examples of devices include video encoders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communications devices. Obviously, devices may be mobile and may be located in mobile vehicles.

[0231] Furthermore, the methods may be implemented by instructions executed by a processor, which may be stored on a processor-readable medium, such as an integrated circuit, software carrier, or other storage device, such as a hard disk, compact diskette, random access memory ("RAM"), or read-only memory ("ROM"). The instructions may constitute an application program tangibly embodied on the processor-readable medium. A processor may obviously include, for example, a processor-readable medium having instructions for performing a process. Such an application program may be uploaded to and executed by a machine having any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units ("CPUs"), random access memory ("RAM"), and input / output ("I / O") interfaces. The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be part of the application program or part of the microinstruction code (or a combination thereof), which may be executed by the CPU. In addition, various other peripheral devices may be connected to the computer platform such as an additional data storage device and a printing device.

[0232] As will be appreciated by those skilled in the art, implementations may also generate signals formatted to carry information that can be stored or transmitted, for example. The information may include, for example, data generated by one of the implementations or instructions for performing a method. The signals may be formatted, for example, as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as baseband signals. Formatting may include, for example, encoding a data stream, generating syntax, and modulating a carrier wave with the encoded data stream and syntax. The information carried by the signals may be, for example, analog or digital information. The signals may be transmitted over various wired or wireless links, as is known.

[0233] Additionally, because portions of the component systems and methods depicted in the accompanying figures are preferably implemented in software, the actual connections between the system components (or processing function blocks) may vary depending on the manner in which the present principles are programmed. Given the teachings herein, one of ordinary skill in the art will be able to contemplate these and similar implementations or configurations of the present principles.

[0234] Several implementations have been described. However, it will be understood that various modifications can be made. For example, elements of separate implementations can be combined, supplemented, modified, or removed to produce other implementations. Furthermore, those skilled in the art will understand that other structures and processes can be substituted for those disclosed, with the resulting implementations performing at least substantially the same functions in at least substantially the same manner to achieve at least substantially the same results as the disclosed implementations. In particular, although illustrative embodiments have been described herein with reference to the accompanying drawings, the principles of the present application are not limited to the precise embodiments described above, and various changes and modifications can be made by those skilled in the art without departing from the scope or spirit of the principles of the present application. Thus, the foregoing and other implementations are contemplated by this application and fall within the scope of the claims.

Claims

1. 1. An apparatus for generating an encoded video bitstream, the apparatus comprising: an input unit that receives a first picture and a second picture corresponding to a view of a multiview video, the first picture corresponding to a first view of the multiview video and the second picture corresponding to a second view of the multiview video; a processor; The processor: arranging the first picture and the second picture into a single picture at a pixel level; generating a supplemental enhancement information (SEI) message indicating that tiling is performed at a pixel level and that the second picture is shifted from the first picture by a certain number of pixels in the horizontal and vertical directions, the SEI message further indicating the number of pixels by which the second picture is shifted from the first picture in the horizontal and vertical directions; encoding the single picture and the SEI message to form an encoded video bitstream. Device.

2. The apparatus of claim 1 , wherein a common object appears to shift from the first view to the second view.

Citation Information

Patent Citations

  • JPP6825155B

  • JPP7116812B

  • JPP7357125B

  • JPP7605516B