Encoding and decoding method and device

By employing a forward projection method in multi-view video coding, the pixels of the reference view are projected onto the camera coordinate system of the current view to generate intermediate predicted images and store them in the reconstructed image buffer. This solves the problem of low compression efficiency of multi-view + depth content in existing technologies and achieves more efficient coding.

CN121887984APending Publication Date: 2026-04-17INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2020-11-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing multi-view video coding techniques fail to effectively utilize inter-view and depth information for efficient compression when processing multi-view + depth content captured by camera arrays, especially the redundancy between adjacent views is not fully utilized.

Method used

The forward projection method is used to project the pixels of the reference view from the camera coordinate system of the reference view to the camera coordinate system of the current view, generate an intermediate predicted image, and store it in the reconstructed image buffer of the current view so as to reconstruct the image of the current view.

Benefits of technology

By utilizing inter-view and depth information, the compression efficiency of multi-view video content is improved, the redundancy utilization between adjacent views is enhanced, and the encoding efficiency is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887984A_ABST
    Figure CN121887984A_ABST
Patent Text Reader

Abstract

The invention provides a method for decoding or encoding, the method comprising: obtaining (140) view parameters of a set of views comprising at least one reference view and a current view of multi-view video content, each view comprising a texture layer and a depth layer; generating (141), for at least one pair of reference views of the set of views and the current view, an intermediate prediction image whereby a forward projection method is applied to pixels of the reference views to project these pixels from a camera coordinate system of the reference view to a camera coordinate system of the current view, the prediction image comprising information allowing reconstruction of image data; storing (143) at least one final predicted image obtained from at least one intermediate predicted image in a reconstructed image buffer of the current view; a current image of the current view is reconstructed (144) from the image stored in the buffer, the buffer comprising the at least one final predicted image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application filed on November 30, 2020, with application number 202080088774.8 and invention title "Encoding and Decoding Method and Apparatus". Technical Field

[0002] At least one embodiment of the present invention relates in general to a method and apparatus for video encoding or decoding, and more specifically to a method and apparatus for video encoding or decoding MVD (Multi-View + Depth) data. Background Technology

[0003] Image and video experts have been studying the compression of multi-view images or video content (multiple views from multiple cameras) for years. Two types of content are generally considered: content called multi-view (MV) content, which consists of synchronized images, each corresponding to a different viewpoint on the same scene; and content called multi-view + depth content (MVD), where MV content is supplemented by depth information of the scene.

[0004] In 2015, to improve the encoding efficiency of multi-view content, two extensions of HEVC (ISO / IEC 23008-2-MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265) were adopted: ● MV-HEVC for music video content; ● 3D-HEVC for MVD content.

[0005] In MV-HEVC, in addition to HEVC spatial intra-frame image prediction (i.e., intra-frame prediction) and temporal inter-frame image prediction (i.e., inter-frame prediction), an inter-view prediction mode utilizing the similarity between views is introduced. A first view is selected as a reference, and at least a second view is encoded relative to this reference view using disparity-based motion prediction. Figure 11 This illustrates an example of the interdependencies between images within the MV content, both in terms of time and direction between views. (View) View0 This indicates that decoding is possible without any inter-view prediction to maintain a reference view that is backward compatible with HEVC. At time points... T0 View View1 and View2 Not encoded as all intra-frame images (I-images), but used in view View0 In T0 The reconstructed image at that location is used as a reference image for encoding / decoding for prediction. Not only the view... View1 The image is used as a reference image and the view is used. View0 In time T1 The image at the location is used to view View1 In time T1 The image at that location is encoded / decoded.

[0006] The MV-HEVC test content consisted only of stereoscopic or multi-view content, but only "3" views were captured by the pointed camera. However, after implementing this MV-HEVC method, inter-view prediction was based solely on disparity estimation between adjacent views, utilizing redundancy with neighboring views. This method is unsuitable for content captured by camera arrays.

[0007] In 3D-HEVC, the same approach as that used for MV-HEVC is employed, but the emission of dense depth information is also considered (i.e., one depth information per pixel per view). Since the content is identical, the same inter-view approach as for adjacent views is used. A more complex combination of inter-view predictions is introduced, including the use of additional depth information.

[0008] To utilize depth information in prediction mode selection, a basic view synthesis prediction mode (VSP) is introduced. The basic VSP mode uses a disparity motion vector corresponding to the depth information of the blocks adjacent to the current block for the current block. The depth information is used to obtain a textured block from a reference view, which serves as the predictor for the current block. Since the depth of the current block is decoded after texturing, the depth information used is from one of the already reconstructed neighboring blocks. The depth values ​​of the reconstructed neighboring blocks are generally considered to be suboptimal depth values ​​predicted between views.

[0009] We hope to propose a solution that allows for an improved VSP model. Summary of the Invention

[0010] In a first aspect, one or more embodiments of the present invention provide a method for decoding, the method comprising: Obtain view parameters for a set of views that include at least one reference view and the current view, comprising multi-view video content, wherein each view includes a texture layer and a depth layer; For at least one pair of reference views and the current view in the set of views, an intermediate predicted image is generated, thereby applying a forward projection method to the pixels of the reference views to project these pixels from the camera coordinate system of the reference views to the camera coordinate system of the current view. The predicted image includes information that allows for the reconstruction of image data. At least one final predicted image obtained from at least one intermediate predicted image is stored in the reconstructed image buffer of the current view; The current image of the current view is reconstructed from the image stored in the buffer, the buffer including the at least one final predicted image.

[0011] In a second aspect, one or more embodiments of the present invention provide a method for encoding, the method comprising: Obtain view parameters for a set of views that include at least one reference view and the current view, comprising multi-view video content, wherein each view includes a texture layer and a depth layer; For at least one pair of reference views and the current view in the set of views, an intermediate predicted image is generated, thereby applying a forward projection method to the pixels of the reference views to project these pixels from the camera coordinate system of the reference views to the camera coordinate system of the current view. The predicted image includes information that allows for the reconstruction of image data. At least one final predicted image obtained from at least one intermediate predicted image is stored in the reconstructed image buffer of the current view; and, The current image of the current view is reconstructed from the image stored in the buffer, the buffer including the at least one final predicted image.

[0012] In a third aspect, one or more embodiments of the present invention provide an apparatus for decoding, the apparatus comprising: A means for obtaining view parameters of at least one reference view and a set of views including multi-view video content and a current view, wherein each view includes a texture layer and a depth layer; A means for generating intermediate predicted images for at least one pair of reference views and a current view of a set of views, thereby applying a forward projection method to the pixels of the reference views to project these pixels from the camera coordinate system of the reference views to the camera coordinate system of the current view, the predicted images including information that allows for the reconstruction of image data. A means for storing at least one final predicted image obtained from at least one intermediate predicted image in a reconstructed image buffer of the current view; A means for reconstructing a current image of the current view from an image stored in the buffer, the buffer including the at least one final predicted image.

[0013] In a fourth aspect, one or more embodiments of the present invention provide an apparatus for encoding, the apparatus comprising: A means for obtaining view parameters of at least one reference view and a set of views including multi-view video content and a current view, wherein each view includes a texture layer and a depth layer; A means for generating intermediate predicted images for at least one pair of reference views and a current view of a set of views, thereby applying a forward projection method to the pixels of the reference views to project these pixels from the camera coordinate system of the reference views to the camera coordinate system of the current view, the predicted images including information that allows for the reconstruction of image data. A means for storing at least one final predicted image obtained from at least one intermediate predicted image in a reconstructed image buffer of the current view; A means for reconstructing a current image of the current view from an image stored in the buffer, the buffer including the at least one final predicted image.

[0014] In a fifth aspect, one or more embodiments of the present invention provide an apparatus comprising the device according to the third and / or fourth aspects.

[0015] In a sixth aspect, one or more embodiments of the present invention provide a signal comprising data generated by the method for encoding according to the second aspect or by the apparatus for encoding according to the fourth aspect.

[0016] In a seventh aspect, one or more embodiments of the present invention provide a computer program comprising program code instructions for implementing the method according to the first or second aspect.

[0017] In an eighth aspect, one or more embodiments of the present invention provide an information storage device that stores program code instructions for implementing the method according to the first or second aspect.

[0018] In the ninth aspect, one or more embodiments also provide a method and apparatus for transmitting or receiving signals according to the sixth aspect.

[0019] In a tenth aspect, one or more embodiments also provide a computer program product including instructions for performing at least a portion of any of the methods described above.

[0020] In any of the foregoing embodiments, the information that allows for the reconstruction of image data includes texture data and depth data.

[0021] In any of the embodiments of the foregoing aspects, the forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by deprojecting from the camera coordinate system of the reference view to the world coordinate system. This deprojection uses the pose matrix of the camera that acquires the reference view (called the reference camera), the inverse intrinsic matrix of the reference camera, and the depth value associated with the current pixel. The deprojected pixels are projected into the coordinate system of the current view using the intrinsic and extrinsic parameter matrices of the camera that acquires the current view (referred to as the current camera), respectively, to obtain the forward-projected pixels. Each matrix is ​​obtained from the view parameters; and, If the obtained forward-projected pixel does not correspond to a pixel on the current camera's pixel grid, then the pixel closest to the forward-projected pixel in the pixel grid is selected to obtain the corrected forward-projected pixel.

[0022] In any of the foregoing embodiments, the method includes filling in isolated missing pixels in each intermediate or final projected image.

[0023] In any of the foregoing embodiments, the information that allows the reconstruction of image data includes motion information.

[0024] In any of the embodiments of the foregoing aspects, the forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by deprojecting from the camera coordinate system of the reference view to the world coordinate system. This deprojection uses the pose matrix of the camera that acquires the reference view (called the reference camera), the inverse intrinsic matrix of the reference camera, and the depth value associated with the current pixel. The deprojected pixels are projected into the coordinate system of the current view using the intrinsic and extrinsic matrices of the camera that acquires the current view (referred to as the current camera) to obtain the forward-projected pixels. Each matrix is ​​obtained from the view parameters. If the obtained forward-projected pixel does not correspond to a pixel on the current camera's pixel grid, then the pixel closest to the forward-projected pixel in the pixel grid is selected to obtain the corrected forward-projected pixel. Calculate the motion vector representing the displacement between the forward-projected pixel or the corrected forward-projected pixel and the current pixel of the reference view.

[0025] In any of the foregoing embodiments, the method includes filling in isolated missing motion information in each intermediate or final projected image.

[0026] In any of the embodiments of the foregoing aspects, at least one final projected image is an intermediate projected image.

[0027] In any of the foregoing embodiments, at least one final projected image is generated from the aggregation of at least two intermediate predicted images.

[0028] In an implementation of any of the foregoing aspects, the method includes subsampling the depth layer of the reference views for at least one pair of reference views and the current view before applying the forward projection method to the pixels of the reference views.

[0029] In an implementation of any of the foregoing aspects, the method includes reconstructing a current block of the current image from a bidirectional predictor block calculated as a weighted sum of two unidirectional predictor blocks, each unidirectional predictor block being extracted from an image in a reconstructed image buffer stored in the current view, and at least one unidirectional predictor block being extracted from a final predicted image stored in the buffer.

[0030] In any of the implementations of the foregoing aspects, at least one weight used in the weighted sum is modified based on the confidence rate of the pixels in the unidirectional predictor block.

[0031] In any of the implementations of the foregoing aspects, view parameters for each view are provided by SEI messages.

[0032] In any of the foregoing embodiments, a syntax element representing information that allows the reconstruction of each final predicted image of the current view is included in the slice header or sequence header, or in the image header or at the synchronization point or image hierarchy.

[0033] In any of the foregoing embodiments, multi-view video content is encoded in or decoded from an encoded video stream, and wherein when the current block is encoded according to a prediction mode (referred to as VSP mode) used to generate a predictor block for the current block using the final predicted image, the encoding of the current block according to the VSP mode is explicitly signaled by a marker in the portion of the encoded video stream corresponding to the current block or implicitly signaled by a syntax element representing the index of the final predicted image in the list of reconstructed images stored in the reconstructed image buffer of the current view.

[0034] In any of the foregoing embodiments, the portion of the encoded video stream corresponding to the current block includes syntax elements representing motion information.

[0035] In any of the foregoing embodiments, motion information represents the index of the final predicted image in the final predicted image list stored in the reconstructed image buffer of the current view after the motion vector is refined and / or stored.

[0036] In any of the foregoing embodiments, when a current block encoded in merge or skip mode inherits its encoding parameters from a block encoded in VSP mode, the current block also inherits VSP parameters. Attached Figure Description

[0037] Figure 1 An example of a camera array suitable for acquiring MVD content is shown schematically; Figure 2 This schematically illustrates a processing module suitable for encoding MVD content provided by a camera array; Figure 3 This schematically illustrates a processing module suitable for decoding encoded video streams representing MVD content; Figure 4 An example of a hardware architecture for a processing module capable of implementing an encoding or decoding module is schematically shown, in which various aspects and implementation schemes are implemented; Figure 5 A block diagram of an example system in which various aspects and implementation schemes are implemented is shown; Figure 6 A method for partitioning an image is illustrated schematically; Figure 7 An example of a method for encoding a coded video stream representing a view is illustrated schematically; Figure 8 An example of a method for decoding an encoded video stream representing a view is illustrated schematically; Figure 9 An example of a method for encoding a coded video stream representing multi-view content is illustrated schematically; Figure 10 An example of a method for decoding an encoded video stream representing multi-view content is illustrated schematically; Figure 11 An example illustrating the dependencies between views in MV content; Figure 12 This schematically illustrates the transformation between the world coordinate system and the camera coordinate system; Figure 13A An example of the forward projection method used in the process of predicting image generation is illustrated schematically; Figure 13B schematically depicted Figure 13A Another representation of the example of the forward projection method depicted in the image; Figure 14A A first implementation scheme for the predictive image generation process is described; Figure 14B Details of a second implementation of the predictive image generation process are described; Figure 15 A third implementation scheme for the predictive image generation process is described; Figure 16 A fourth implementation scheme for the predictive image generation process is described; Figure 17 A basic implementation scheme for the syntax parsing process of a video compression method that does not use inter-view prediction is schematically depicted; Figure 18 A first implementation of the syntax parsing process for a video compression method using the new VSP mode is schematically depicted. Figure 19 A second implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated. Figure 20 A third implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated. Figure 21 A fourth implementation scheme of the syntax parsing process for a video compression method using the new VSP mode is schematically depicted. Figure 22A A fifth implementation scheme for the predictive image generation process is described; Figure 22B Details of a sixth implementation of the predictive image generation process are described; Figure 23 A seventh implementation scheme for the predictive image generation process is described; Figure 24 The typical coding structures and image dependencies of MV-HEVC and 3D-HEVC codecs are illustrated schematically; and, Figure 25 An eighth implementation of the predictive image generation process is described. Detailed Implementation

[0038] In the following description, some implementations use tools developed in the context of the international standard entitled "Universal Video Coding (VVC)" developed by the joint collaboration of ITU-T and ISO / IEC experts (known as the Joint Video Experts Group (JVET)) or in the context of HEVC, MV-HEVC, or 3D-HEVC. However, these implementations are not limited to video encoding / decoding methods corresponding to VVC, HEVC, MV-HEVC, or 3D-HEVC, and are applicable to other video encoding / decoding methods, as well as other image encoding / decoding methods suitable for MVD content.

[0039] The implementation scheme described below presents a new VSP model.

[0040] In the following text, Figure 24 , Figure 6 , Figure 7 and Figure 8 A basic implementation scheme that allows the introduction of some terms is described.

[0041] Figure 24 The typical coding structure and image dependencies of MV-HEVC and 3D-HEVC codecs are illustrated schematically.

[0042] MV and 3D-HEVC are known to employ a multi-layer approach, where layers are multiplexed into a single bitstream and can depend on each other. In MV and 3D-HEVC, layers can represent texture, depth, or other auxiliary information of a scene associated with a specific camera. All layers belonging to the same camera are represented as a view; while layers carrying the same type of information (e.g., texture or depth) are often referred to as components within the scope of 3D video.

[0043] Figure 24 This illustrates a typical encoding structure, which includes two views: view "0" (also known as the base view) 2409 and view... Figure 1 2410. Each view comprises two layers. View 0 2409 comprises a first layer consisting of texture images 2401 and 2405, and a second layer consisting of depth images 2403 and 2407. Figure 1 2410 includes a first layer consisting of texture images 2402 and 2406 and a second layer consisting of depth images 2404 and 2408.

[0044] Two successive timeframes are shown: through design selection, all images associated with the same capture or display time instance are contained within a single Access Unit (AU). Images 2401, 2402, 2403, and 2404 are in the same AU 0 2411. Images 2405, 2406, 2407, and 2408 are in the same AU 1 2412. It is typically required that the base layer conforms to the HEVC single-layer profile and is therefore the texture component of the base view.

[0045] In AutoCAD, layers following the base layer image are represented as enhancement layers, and views other than the base view are represented as enhancement views. In AutoCAD, all components are required to have the same view order. To facilitate composite encoding, 3D-HEVC further requires that the depth component of a particular view immediately follow its texture component. Figure 24 An overview of the dependencies between images in different layers and AUs is described in the text and discussed further below.

[0046] In MV-HEVC, in addition to the regular time-based inter-image prediction using the same view and components but in different AUs (by... Figure 24 In addition to the arrows associated with the abbreviation TIIP in the text, MV-HEVC allows predictions from images in the same AU and components but in different views, which is referred to below as inter-view prediction (by...). Figure 24 (The arrows associated with the abbreviation IVP are shown in the image). For inter-view prediction, decoded images from other views can be used as reference images for the current image.

[0047] The motion vector associated with the current block of the current image can be temporal (hereinafter referred to as TMV) when it is related to a temporal reference image of the same view, or disparity MV (hereinafter referred to as DMV) when it is related to a reference image between views. Existing block-level HEVC motion compensation modules can be used, which operate in the same way regardless of whether the MV is TMV or DMV.

[0048] To improve compression performance, 3D-HEVC extends MV-HEVC by allowing new types of inter-layer prediction. For example... Figure 24 The new forecast types are as follows: ● Combined time and view inter-prediction (by Figure 24 The arrows in the image with the abbreviation TII+IVP indicate that the images are in the same component but in different AUs and different views. ● Inter-component prediction (by Figure 24 The arrow in the image (indicated by the abbreviation ICP) refers to images in the same AU and view but in different components; ● Inter-component and inter-view predictions of the combination (by...) Figure 24 The arrow in the image (represented by the abbreviation ICIP) refers to an image that is in the same AU but in different views and components.

[0049] Another design change compared to MV-HEVC is that, in addition to sample and motion information, residual, disparity, and partitioning information can be predicted or inferred. A detailed overview of texture and depth coding tools is provided in the document "Multi-view and...". 3D Extended Overview ( Overview of the Multiview and 3D Extensions of High Efficiency Video Coding ), IEEE Video Technology Circuits and Systems Bulletin, Vol. 26 Volume, No. 1 Expect, 2016 Year 1 moon, G. Tech ; Y. Chen ; K. Müller ; JR. Ohm ; A. Vetro ; YK. Wang "middle.

[0050] Due to the similarity between HEVC and VVC, compression tools defined in the contexts of MV and 3D-HEVC should be adapted to the VVC context to obtain codecs capable of handling multi-view content (with or without depth information). As in the HEVC context, the base layer of multi-view content, which includes only texture information, should be fully compatible with VVC in the VVC context.

[0051] Figure 6, Figure 7 and Figure 8 This highlights some key features of basic compression methods that can be used to encode the base layer of multi-view content.

[0052] Figure 6 An example of partitioning is shown for the pixel image 11 of the original video 10. Here, a pixel is considered to consist of three components corresponding to the basic layer of the multi-view content: one luminance component and two chrominance components. The same partitioning can be applied to all layers of the multi-view content, i.e., the texture layer and the depth layer. Alternatively, the same partitioning can be applied to other numbers of components or layers, such as a texture layer and a depth layer that include four components (one luminance component, two chrominance components, and one alpha component).

[0053] The image is divided into multiple coded entities. First, as... Figure 6 As indicated by reference numeral 13, the image is divided into a grid of blocks called coding tree units (CTUs). A CTU is composed of... It consists of one luminance sample block and two corresponding chrominance sample blocks. N It is a power of two, for example, the maximum value is "128". Secondly, the image is divided into one or more tile rows and tile columns. A tile is a sequence of CTUs covering a rectangular area of ​​the image. In some video compression schemes, a tile can be divided into one or more bricks, each brick consisting of at least one CTU row within the tile. Above the concepts of tiles and bricks, there is another coded entity called a slice, which can contain at least one tile of an image or at least one brick of a tile.

[0054] exist Figure 6 In the example, as indicated by reference numeral 12, image 11 is divided into three slices S1, S2 and S3, each slice comprising multiple tiles (not shown).

[0055] like Figure 6 As indicated by reference numeral 14, a CTU can be partitioned into a hierarchical tree of one or more sub-blocks called coding units (CUs). The CTU is the root (i.e., the parent node) of the hierarchical tree and can be partitioned into multiple CUs (i.e., child nodes). If each CU is not further partitioned into smaller CUs, then each CU becomes a leaf of the hierarchical tree; or if each CU is further partitioned into smaller CUs (i.e., child nodes), then each CU becomes the parent node of the smaller CUs. Different types of hierarchical trees can be used: a quadtree, where the CTU or CU is divided into four square CUs of equal size; a binary tree, where the CTU (or CU) can be horizontally or vertically partitioned into "2" rectangular CUs of equal size; and a ternary tree, where the CTU (or CU) can be horizontally or vertically partitioned into "3" rectangular CUs.

[0056] exist Figure 6In the example, firstly, CTU 14 is partitioned into 4 square CUs using quadtree partitioning. The top-left CU is a leaf of the hierarchical tree because it is not further partitioned; that is, it is not the parent node of the other CUs. The top-right CU is further partitioned into 4 smaller square CUs using quadtree partitioning again. The bottom-right CU is vertically partitioned into 2 rectangular CUs using binary tree partitioning. The bottom-left CU is vertically partitioned into 3 rectangular CUs using ternary tree partitioning.

[0057] During image encoding, partitioning is adaptive, with each CTU being partitioned to optimize compression efficiency standards.

[0058] In some video compression schemes, the concepts of Prediction Unit (PU) and Transform Unit (TU) emerge. In practice, in this case, the coding entity used for prediction (i.e., PU) and the coding entity used for transformation (i.e., TU) can be sub-partitions of the CU. For example, as... Figure 6 The size is represented as The CU can be divided into sizes of or size PU 1411. Additionally, the CU can be divided into units of size [missing information]. The "4" TU 1412 or size is The "16" TUs.

[0059] In this application, the term "block" or "image block" may refer to any of CTU, CU, PU, ​​and TU. Additionally, the term "block" or "image block" may refer to macroblocks, partitions, and subblocks, and more generally to an array of samples of numerous sizes.

[0060] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image" and "picture." In the specific context of MVD data, [the following is a more general interpretation:] ... Figure 24 Similar to the AU in [the original text], a frame at time T is considered an entity comprising an image (texture and depth) corresponding to time T for each view. Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0061] Figure 7 A method for encoding a video stream, performed by an encoding module, is illustrated schematically. Variations of this encoding method are envisioned, but for clarity, they are described below. Figure 7The method for encoding is described, but not all anticipated variations are not described. Specifically, the method for encoding is applied to a base layer of multi-view content, where each pixel of the base layer comprises one luminance component and two chrominance components. Specific encoding tools suitable for encoding multi-view content, and particularly depth layers, are not further described.

[0062] The encoding of the current original image 501 during step 502 begins with a partition of the current original image 501, as per the relevant information. Figure 6 As described. Therefore, the current image is partitioned into 501 blocks: CTU, CU, PU, ​​TU, etc. For each block, the coding module determines the coding mode between intra-frame prediction and inter-frame prediction.

[0063] Intra-frame prediction includes, during step 503, predicting pixels of the current block from a prediction block derived from pixels of a reconstructed block located in the causal vicinity of the current block to be encoded, according to an intra-frame prediction method. The result of intra-frame prediction is a prediction direction indicating which pixels of the nearby blocks are used, and a residual block obtained by calculating the difference between the current block and the prediction block.

[0064] Inter-frame prediction involves predicting the pixels of the current block from pixel blocks (called reference blocks) in images preceding or following the current image (referred to as the reference image). During encoding of the current block according to the inter-frame prediction method, the closest block to the current block in the reference image is determined by motion estimation step 504 based on a similarity criterion. During step 504, a motion vector indicating the location of the reference block in the reference image, identified by an index, is determined. The motion vector and the index of the reference image are used during motion compensation step 505, during which a residual block is calculated as the difference between the current block and the reference block. It should be noted that only single-prediction inter-frame prediction is described here. Dual-prediction inter-frame prediction (or B-mode) also exists, for which the current block is associated with two motion vectors, thereby specifying two reference blocks in two different images (each reference block is specified by a reference image index), and the residual block of that block is then the average of the two residual blocks.

[0065] It should be noted that intra-frame prediction and inter-frame prediction are general terms that encompass many modes based on the general principles of spatial and temporal prediction.

[0066] During selection step 506, the encoding module selects a prediction mode from the tested prediction modes that optimizes compression performance according to the rate / distortion criterion. When a prediction mode is selected, the residual block is transformed during step 507 and quantized during step 509. Note that the encoding module may skip the transformation and directly apply quantization to the untransformed residual signal. When encoding the current block based on intra-frame prediction, the entropy encoder encodes the prediction direction and the transformed and quantized residual block during step 510. When encoding the current block based on inter-frame prediction, the motion vector of the block is predicted based on a set of prediction vectors selected from the reconstructed blocks located near the block to be encoded. Next, during step 510, the entropy encoder encodes the motion information (including the motion vector residual, the index of the motion vector predictor, and the index of the reference image) in the form of motion residuals and an index used to identify the prediction vectors. During step 510, the entropy encoder encodes the transformed and quantized residual block. It should be noted that the encoding module can bypass the transform and quantization processes; that is, entropy coding is applied to the residuals instead of the transform or quantization process. The result of the entropy coding is inserted into the encoded video stream 511.

[0067] After quantization step 509, the current block is reconstructed so that the pixels corresponding to that block are available for future prediction. This reconstruction stage is also called the prediction loop. Therefore, inverse quantization is applied to the transformed and quantized residual block during step 512, and inverse transform is applied during step 513. The prediction block of the block is reconstructed based on the prediction mode used for the block obtained during step 514. If the current block is encoded based on inter-frame prediction, the encoding module applies motion compensation using the motion vectors of the current block during step 516 to identify a reference block for the current block. If the current block is encoded based on intra-frame prediction, the reference block of the current block is reconstructed using the prediction direction corresponding to the current block during step 515. The reference block and the reconstructed residual block are added together to obtain the reconstructed current block.

[0068] After reconstruction, during step 517, an in-loop post-filter designed to reduce coding artifacts is applied to the reconstructed block. This post-filter is called an in-loop post-filter because it occurs in the prediction loop to obtain the same reference image as the decoder, thus avoiding drift between encoding and decoding. For example, in-loop post-filters include deblocking filtering and SAO (Sample Adaptive Shift) filtering. During entropy coding step 510, parameters representing the activation or deactivation of the in-loop deblocking filter and the characteristics of the in-loop deblocking filter when activated are introduced into the encoded video stream 511. When reconstructing a block, it is inserted into the reconstructed image stored in the reconstructed image memory 519 (also called a reference image memory, reference image buffer, or decoded image buffer (DPB)) during step 518. The reconstructed image thus stored can then be used as a reference image for other images to be encoded.

[0069] Figure 8 The schematic depiction illustrates the process performed by the decoding module for processing data based on... Figure 7 The described method is for decoding the encoded video stream 511. Variations of this decoding method are envisioned, but for clarity, the following description is preferred. Figure 8 The method used for decoding is described, but not all expected variations are described.

[0070] Decoding is performed block by block. For the current block, it begins with entropy decoding of the current block during step 610. Entropy decoding allows the acquisition of the block's predictive pattern.

[0071] If the block has already been encoded based on inter-frame prediction, entropy decoding allows for the acquisition of the prediction vector index, motion residual, index on the reference image, and residual block. During step 608, the prediction vector index and motion residual are used to reconstruct the motion vector of the current block.

[0072] If the block has already been encoded based on intra-frame prediction, entropy decoding allows for the acquisition of the prediction direction and residual block. Steps 612, 613, 614, 615, 616, and 617, implemented by the decoding module, are identical in all respects to steps 512, 513, 514, 515, 516, and 517, implemented by the encoding module. In step 618, the decoded block is stored in the decoded image, and the decoded image is stored in DPB 619. When the decoding module decodes a given image, the image stored in DPB 619 is identical to the image stored in DPB 519 by the encoding module during the encoding of the given image. The decoded image can also be output by the decoding module for, for example, display.

[0073] Figure 1 An example of a camera array suitable for acquiring MVD content is illustrated schematically.

[0074] Figure 1 Camera array 10 is defined as comprising 16 cameras 10A to 10P positioned on a 4×4 grid. Each camera in camera array 10 focuses on the same scene and is capable of acquiring an image, for example, where each pixel comprises one luminance component and two chrominance components. A computing or measuring device (not shown), connected to camera array 10, is used to generate a depth map for each image generated by the cameras of camera array 10. Figure 1 In the example, each depth map associated with an image has the same resolution as the image (i.e., the depth map includes the depth value for each pixel of the image). Therefore, camera array 10 generates MVD content comprising "16" texture layers and "16" depth layers. For such a camera array, overlap between the captured views is important. This is the goal of the following embodiments to improve the overall compression ratio achievable with such multi-view content.

[0075] Each camera in camera array 10 is associated with intrinsic and extrinsic camera parameters. These parameters are required by the decoder to create a predicted image, as described in this document. In this implementation, the intrinsic and extrinsic parameters are provided to the decoder in the form of an SEI (Supplemental Enhancement Information) message. SEI messages are defined in H.264 / AVC and HEVC to convey metadata.

[0076] Table TAB1 describes the syntax of SEI messages suitable for conveying the intrinsic and extrinsic parameters of a camera array. This syntax is identical to the syntax of the Multi-View Acquisition Information SEI message in HEVC (Section G.14.2.6). Table TAB1.

[0077] One objective of the implementation described below is to improve the prediction of a view based on at least one other view. Since we are targeting multi-view content captured by a camera array as described above, any camera can provide good predictions for some or all of the adjacent views. To create a predicted image of the current view, previously decoded views and their associated camera parameters, as well as camera parameters associated with the current view, are used. Note that texture layers or depth layers in the view can use a new vsp mode.

[0078] Consider a camera calibrated as a standard pinhole camera and the camera's intrinsic parameter matrix. : ● This represents the distance from the exit pupil to the sensor, expressed in pixels, and is often misused in the literature as "focal length." In Table TAB1, this information is described using the following set of parameters: ● The pixel coordinates represent the so-called "principal point," that is, the orthogonal projection of the pinhole onto the sensor. In Table 1, this information is described using the following set of parameters: ● α and γ represent the pixel aspect ratio and the sensor skew coefficient, respectively. Table 1 does not directly express the α value, but rather the αf value: ● Table 1 describes the γ value as follows: if Given the coordinates of a point in the camera's coordinate system (CS), then the coordinates of the projection of that point onto the image are... Given the following (in pixels): The symbol ≡ represents the equivalence relation between homogeneous vectors: make The pose matrix represents the camera, where and These represent the camera's orientation and position in the reference coordinate system (CS), respectively. The camera's external matrix is ​​defined by the following equation: For each camera, the R and T matrices in Table TAB1 are described as follows: if and Let represent the coordinates of the same point in the camera CS and the reference CS, respectively. and . Figure 12 This represents the projection from the reference CS to the camera CS and from the camera CS back to the reference CS.

[0079] Now, consider a given camera c Currently displaying the view. Camera c With intrinsic parameter matrix pose matrix Related. Let For the camera c Get the current pixel in the image of the current view, and z Its assumed depth. This is determined by the value corresponding to the current pixel. intrinsic parameter matrix and extrinsic parameter matrix Associated cameras The pixels of the image provided in the reference view Given from the following: Figure 9 An example of a method for encoding a encoded video stream representing MVD content is illustrated schematically.

[0080] Figure 9 The method allows encoding of the first view 501 and the second view 501B. In this example, as in... Figure 24 As shown in the diagram, the first view includes a base layer (layer "0") with texture data and a layer "1" with depth data. The second view includes a layer "2" with texture data and a layer "3" with depth data. For simplicity, only the encoding of the two views is shown, but it can be understood through... Figure 9 This method encodes more views. For example, it can be done through... Figure 9 The encoding method encodes the "16" views generated by the camera array 10.

[0081] In implementation (9a), the first view 501 is considered the root view from which all other views are predicted directly or indirectly. The first view 501 is encoded without any inter-view or inter-layer prediction. In one implementation, layers "0" and "1" are encoded in parallel or sequentially, respectively. In one implementation, using... Figure 7 The same steps 502, 503, 504, 505, 506, 507, 508, 509, 510, 512, 513, 514, 515, 516, 517, 518, and 519 are described to encode layer "0" and layer "1". In other words, using... Figure 7 The method encodes the texture and depth data of the first view 501 (which corresponds to...). Figure 24 (The arrow TIIP in the image).

[0082] In implementation scheme (9b), the following is used Figure 7 The method encodes layer "0", but the method is slightly modified for layer "1" to combine patterns defined in 3D HEVC to predict the depth layer of the view from the texture layer of the view (which corresponds to...). Figure 24 (The arrow in the image is ICP).

[0083] In implementation scheme (9c), the texture layer (layer "2") of the second view 501B is encoded by a process including steps 502B, 503B, 504B, 505B, 507B, 508B, 509B, 510B, 512B, 513B, 514, 515B, 516B, 517B, 518B and 519B, which are the same as steps 502, 503, 504, 505B, 507B, 508B, 509B, 510B, 512B, 513B, 515B, 516B, 517B, 518B and 519B, respectively.

[0084] In step 521, the processing module 20 generates a new predicted image and imports it into the DPB 519B. In step 522, the processing module 20 uses this new predicted image to determine the predictor of the current block of the current image of the texture layer of the second view 501B, referred to as... VSP Predictor. The predictions made by the VSP predictor correspond to the new predictions described below. VSP Pattern, also known as VSP model.

[0085] The VSP mode is very similar to the traditional inter-frame mode. In fact, when introduced in DPB 519B, the new prediction image generated during step 521 is considered a general reference image for temporal prediction (even if the prediction image generated during step 521 is temporally co-located with the current image of the texture layer of the second view 501B). Therefore, the new VSP mode can be considered an inter-frame mode using a specific reference image generated through inter-view prediction. Step 522 includes a motion estimation step and a motion compensation step. The block encoded using the VSP mode is encoded in the form of motion information and residuals, the motion information including the identifier of the prediction image generated during step 521.

[0086] During step 506B, processing module 20 performs a step that differs from step 506 only in that, in addition to the usual intra-frame and inter-frame predictors, it also considers the VSP predictor generated during step 522. Similarly, processing module 20 performs step 514B, which differs from step 514 only in that the new VSP mode belongs to a set of prediction modes that can potentially be applied to the current block. If a new VSP mode is selected for the current block during step 506B, then during step 523, processing module 20 reconstructs the corresponding VSP predictor.

[0087] In implementation scheme (9d), the same steps 502B, 503B, 504B, 505B, 506B, 507B, 508B, 509B, 510B, 512B, 513B, 514B, 515B, 516B, 517B, 518B, 519B, 521, 522, and 523 are used to encode the depth layer (layer "3") of the second view 501B. Therefore, the new VSP pattern is applied to the depth layer (layer "3") of the second view 501B. More generally, the VSP pattern can be applied to the depth layer of a view predicted from another view.

[0088] In implementation (9e), the encoding of layer "3" is combined with a pattern defined in 3D HEVC to predict the depth layer of the view from the texture layer of the view (which corresponds to...). Figure 24 (The arrow in the image is ICP).

[0089] exist Figure 9 In the example, the second view 501B is encoded from the first view 501, wherein at least texture layer "0" is encoded without any inter-view prediction. When via Figure 9 When encoding more than two views, any third view can be encoded from a view that encodes at least the texture layer without any inter-view prediction (e.g., from the first view 501) or from a view that encodes the texture layer with inter-view prediction (e.g., from the second view 501B).

[0090] As can be seen, Figure 9 The encoding method includes two encoding layers, each view Figure 1 One encoding layer. Of course, if more than two views are being encoded, then... Figure 9 The encoding method should include as many encoding layers as there are views. Figure 9 In the example, each encoding layer has its own DPB. In other words, each view is associated with its own DPB.

[0091] As will be described below, the images in the DPB are indexed by multiple reference indices: ● ref_idx The index of the reference image to be used in DPB; ● ref_idx_l0 : The index of the reference image to be used in the reference image list l0 stored in the DPB. This is for reference images within the frame. T View on i Decode it. ref_idx_l0 The list refers to views of different frames. i ; ● ref_idx_l1: The index of the reference image to be used in the reference image list l1 stored in the DPB. This is for reference images within the frame. T View on i Decode it. ref_idx_l1 The list refers to views of different frames. i ; ● ref_idx2 : The index of the reference image to be used, corresponding in time to the reference image of the current image. Index ref_idx2 This refers only to images generated through forward projection. (For the purposes of this section, the image must be understood within the frame.) T View on i Decode it. ref_ idx2 The list referred to includes those corresponding to frames. T Reference image.

[0092] Figure 10 An example of a method for decoding an encoded video stream representing multi-view content is illustrated schematically.

[0093] In implementation scheme (10a), corresponding to implementation scheme (9a), layer "0" and layer "1" are decoded in parallel or sequentially, respectively. In this implementation scheme, information about... Figure 8 The same steps 608, 610, 612, 613, 614, 615, 616, 617, 618, and 619 described above decode layer "0" and layer "1". In other words, using... Figure 8 The method decodes the texture and depth data of the first view 501 (which corresponds to...). Figure 24 (The arrow TIIP in the image).

[0094] In implementation scheme (10b) corresponding to implementation scheme (9b), using Figure 8 The method decodes layer "0", but the method is slightly modified for layer "1" to incorporate the patterns defined in 3D HEVC (which corresponds to...). Figure 24 (The arrow in the image is ICP).

[0095] In implementation scheme (10c) corresponding to implementation scheme (9c), the texture layer (layer "2") of the second view 501B is decoded by a process including steps 608B, 610B, 612B, 613B, 615B, 616B, 617B, 618B, and 619B, which are the same as steps 608, 610, 612, 613, 615, 616B, 617B, 618B, and 619, respectively. In step 621, processing module 20 generates a new predicted image that is the same as the image generated during step 521 and introduces this image into DPB 619B. Processing module 20 performs step 614B, which differs from step 614 only in that the new VSP mode belongs to a set of prediction modes that can be potentially applied to the current block. If a new VSP mode has been selected for the current block during step 506B, then during step 623, processing module 20 reconstructs the corresponding VSP predictor.

[0096] In implementation scheme (10d) corresponding to implementation scheme (9d), the depth layer (layer "3") of the second view 501B is decoded using the same steps 608B, 610B, 612B, 613B, 614B, 615B, 616B, 617B, 618B, 619B, 621 and 623.

[0097] As can be seen, Figure 10 The decoding method includes two encoding layers, each view Figure 1 One encoding layer. Of course, if decoding is required for more than two views, then... Figure 9 The decoding method will include as many decoding layers as there are views. Figure 9 In the example, each decoding layer has its own DPB. In other words, each view is associated with its own DPB.

[0098] Figure 2 This schematically illustrates a processing module suitable for encoding MVD content provided by a camera array.

[0099] exist Figure 2 The image shows a simplified representation of a camera array 10 consisting of only two cameras, 10A and 10B. Each camera in the camera array 10 communicates with the processing module 20 using a communication link, which can be wired or wireless. Figure 2 In the process, the processing module 20 encodes the multi-view content generated by the camera array 10 using the new VSP mode described below in the encoded video stream.

[0100] Figure 3 This schematically illustrates a processing module suitable for decoding encoded video streams representing MVD content.

[0101] exist Figure 3In this process, processing module 20 decodes the encoded video stream. Processing module 20 is connected to display device 26 via a communication link, which can be wired or wireless, and the display device can display the image generated from the decoding. The display device may be, for example, a virtual reality headset, a 3D TV, or a computer monitor.

[0102] Figure 4 An example of a hardware architecture for a processing module 20 capable of implementing an encoding or decoding module is schematically shown, which can implement the different embodiments described below. As a non-limiting example, the processing module 20 includes the following items connected by a communication bus 205: a processor or CPU (Central Processing Unit) 200 containing one or more microprocessors, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture; random access memory (RAM) 201; read-only memory (ROM) 202; a storage unit 203, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives and / or optical disk drives, or storage media readers such as SD (Secure Digital) card readers and / or hard disk drives (HDDs) and / or network-accessible storage devices; and at least one communication interface 204 for exchanging data with other modules, devices, or equipment. Communication interface 204 may include, but is not limited to, a transceiver configured to transmit and receive data via a communication channel. Communication interface 204 may include, but is not limited to, a modem or network interface card.

[0103] If the processing module 20 implements a decoding module, then the communication interface 204 enables, for example, the processing module 20 to receive encoded video streams and provide decoded video streams.

[0104] If the processing module implements the encoding module, then the communication interface 204 enables, for example, the processing module 20 to receive the raw image data to be encoded and provide the encoded video stream.

[0105] Processor 200 is capable of executing instructions loaded into RAM 201 from ROM 202, external memory (not shown), storage media, or a communication network. When processing module 20 is powered on, processor 200 is capable of reading instructions from RAM 201 and executing those instructions. These instructions form a computer program that causes, for example, processor 200 to implement... Figure 9 The described encoding method or about Figure 10 The decoding method described herein includes the various aspects and implementation schemes described below.

[0106] All or part of the algorithms and steps of the encoding or decoding method may be implemented in software by executing a set of instructions by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or in hardware by a machine or dedicated component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).

[0107] Figure 5 A block diagram illustrating an example of System 2 is shown, in which various aspects and embodiments are implemented. System 2 may be embodied as a device including the various components described below and configured to perform one or more aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, virtual reality headsets, and servers. Elements of System 2 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, System 2 includes a processing module 20 implementing a decoding module or an encoding module. However, in another embodiment, System 2 may include a processing module 20 implementing a decoding module and a processing module 20 implementing an encoding module, or a processing module 20 implementing both decoding and encoding modules. In various embodiments, System 2 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, System 2 is configured to implement one or more aspects and embodiments described in this document.

[0108] In one embodiment, system 2 includes at least one processing module 20, which is capable of implementing one or both of an encoding module or a decoding module.

[0109] Inputs to processing module 20 may be provided via various input modules as shown in box 22. Such input modules include, but are not limited to: (i) a radio frequency (RF) module that receives RF signals transmitted over the air, for example, by a broadcaster; (ii) a component (COMP) input module (or a set of COMP input modules); (iii) a universal serial bus (USB) input module; and / or (iv) a high-definition multimedia interface (HDMI) input module. Figure 5 Other examples not shown include composite video.

[0110] In various embodiments, the input module of block 22 has associated corresponding input processing elements as known in the art. For example, the RF module may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF module of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box implementation, the RF module and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various implementations rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various implementations, the RF module includes an antenna.

[0111] Additionally, the USB and / or HDMI modules may include corresponding interface processors for connecting System 2 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processing module 20. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processing module 20. Demodulation, error correction, and demultiplexing streams are provided to processing module 20.

[0112] Various components of System 2 can be housed within an integrated housing, where they can be interconnected and transmit data using suitable connection arrangements (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards). For example, in System 2, processing module 20 is interconnected with other components of System 2 via bus 205.

[0113] The communication interface 204 of the processing module 20 allows the system 2 to communicate over the communication channel 21. For example, the communication channel 21 can be implemented in a wired and / or wireless medium.

[0114] In various implementations, data is streamed or otherwise provided to System 2 using a wireless network such as Wi-Fi, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals in these implementations are received via a communication channel 21 and a communication interface 204 suitable for Wi-Fi communication. The communication channel 21 in these implementations is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cloud-based communications. Other implementations use a set-top box to provide streaming data to System 2, delivering data via an HDMI connection to input block 22. Still other implementations use an RF connection to input block 22 to provide streaming data to System 2. As mentioned above, various implementations provide data in a non-streaming manner. Furthermore, various implementations use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks. The data provided to System 2 includes, for example, MVD signals provided by camera array 10.

[0115] System 2 can provide output signals to various output devices, including to display 26 via display interface 23, to speaker 27 via audio interface 24, and to other peripheral devices 28 via interface 25. Display 26 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 26 can be used in televisions, tablets, laptops, cellular phones (mobile phones), smartphones, virtual reality headsets, or other devices. Display 26 can also be integrated with other components (e.g., as in a smartphone) or stand alone (e.g., as an external monitor for a laptop). In various examples of embodiments, other peripheral devices 28 include one or more of a standalone digital video disc (or digital versatile disc, both terms being DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 28 that provide functionality based on the output of System 2. For example, a disc player performs the function of playing the output of System 2.

[0116] In various embodiments, control signals are transmitted between system 2 and display 26, speaker 27, or other peripheral devices 28 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices are communicatively coupled to system 2 via dedicated connections through corresponding interfaces 23, 24, and 25. Alternatively, output devices can be connected to system 2 via communication interface 204 using communication channel 21. Display 26 and speaker 27 may be integrated into a single unit with other components of system 2 in electronic devices such as televisions. In various embodiments, display interface 23 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0117] For example, if the RF portion of input 22 is part of a separate set-top box, then display 26 and speaker 27 may optionally be separate from one or more other components. In various embodiments where display 26 and speaker 27 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0118] Various specific implementations involve decoding. As used in this application, "decoding" can encompass all or part of a process performed, for example, on a received encoded video stream, to produce a final output suitable for display. In various implementations, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and prediction. In various implementations, such a process also includes, or alternatively includes, processes performed by a decoder of the various specific implementations described in this application, such as for decoding a new VSP mode.

[0119] Whether the phrase “decoding process” specifically refers to a subset of operations or broadly refers to a wider decoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0120] Various specific implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass, for example, all or part of the process performed on an input video sequence to produce an encoded video stream. In various implementations, such processes include one or more processes typically performed by an encoder, such as partitioning, prediction, transform, quantization, and entropy coding. In various implementations, such processes also include, or alternatively include, processes performed by the encoder of the various specific implementations described herein, for example, encoding according to a new VSP mode.

[0121] Whether the phrase “encoding process” specifically refers to a subset of operations or broadly refers to a wider encoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0122] Note that, as used in this article, syntactic elements (e.g., tags) VSP and index ref_idx2 () are descriptive terms. Therefore, they do not preclude the use of other syntactic element names.

[0123] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0124] Various implementation schemes refer to rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate-distortion optimization is generally formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to solve the rate-distortion optimization problem. For example, these methods may be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce encoding complexity, particularly for the computation of approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0125] The specific embodiments and aspects described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of specific embodiment (e.g., discussed only as a method), specific embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented, for example, in a processor, which generally refers to a processing device. The processing device includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0126] The reference to "an implementation scheme" or "implementation scheme" or "a specific implementation" or "specific implementation," and other variations thereof, means that the specific features, structures, characteristics, etc., described in connection with the implementation scheme are included in at least one implementation scheme. Therefore, the appearance of the phrase "in an implementation scheme" or "in an implementation scheme" or "in a specific implementation" or "in a specific implementation," and any other variations appearing throughout this application, do not necessarily refer to the same implementation scheme.

[0127] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0128] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or more of these.

[0129] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or more. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0130] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” “at least one of A and B,” and “one or more of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one,” “one or more” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C,” “at least one of A, B, and C,” and “one or more of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many of the listed items as possible.

[0131] Moreover, as used herein, the term "signaling" refers to (among other things) instructing the corresponding decoder to do something. For example, in some implementations, the encoder signals information indicating a new VSP mode. Thus, in one implementation, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters and others, signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various implementations by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various implementations, information is signaled to the corresponding decoder using one or more syntax elements, tags, etc. Although the verb form of the term "signal" has been used above, the term "signal" can also be used as a noun here.

[0132] It will be apparent to those skilled in the art that the embodiments can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the embodiments. For example, a signal may be formatted to carry an encoded video stream of the embodiment. Such signals may be formatted as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the encoded video stream and modulating a carrier wave using the encoded video stream. The information carried by the signal may be, for example, analog or digital information. It is known that signals can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0133] Figure 13A An example of the forward projection method used in the process of predicting image generation is illustrated. Figure 13B yes Figure 13A Another representation of the forward projection method. Used during steps 521 and 621. Figure 13A and Figure 13B The forward projection process. Applying the forward projection process to the camera... m The pixels of the first view are acquired, and these pixels are projected from the camera coordinate system of the first view onto the coordinate system of the camera. n The camera coordinate system of the acquired second view. Each pixel is considered to include texture information and depth information.

[0134] In step 130, the processing module 20 applies deprojection to the current pixel P(u, v) of the first view to deproject from the camera coordinate system of the first view to the reference coordinate system (i.e., the world coordinate system) to obtain, Figure 13B The represented deprojected pixels To project using a camera. m pose matrix ,camera m inverse intrinsic parameter matrix And the depth value associated with the current pixel. The pose matrix is ​​defined using camera parameters obtained by the processing module 20, for example, from the SEI message described in Table TAB1. and intrinsic parameter matrix .

[0135] In step 131, the processing module 20 uses the camera. n intrinsic parameter matrix and extrinsic parameter matrix Deprojected pixels Projected into the coordinate system of the second view. Similarly, the intrinsic parameter matrix is ​​defined using camera parameters obtained by processing module 20, for example, from the SEI messages described in Table TAB1. and extrinsic parameter matrix When the projection does not fall into the camera n When the projection is within the area, it is rejected. When the projection falls into the camera's view... n When the projection falls within a certain area, it is most likely to fall between four pixels rather than on a real pixel.

[0136] In step 132, the processing module 20 selects the camera. n The closest pixel to the deprojected pixel in the pixel grid The projected pixel P'(u', v'). The nearest pixel, for example, is in the pixel grid of camera n such that it is the closest to the projected pixel. The distance to the projected pixel is minimized. This distance is calculated, for example, as the coordinates of the pixel in the grid and the distance to the deprojected pixel. The square root of the sum of the squared differences between the projections (or calculated as the sum of the absolute differences).

[0137] pass Figure 13A and Figure 13B The pixels P'(u', v') obtained through the forward projection process retain the texture and depth values ​​of the projected pixels P(u, v). This set of pixels P'(u', v') forms the projected image.

[0138] Figure 14A A first implementation scheme for the predictive image generation process is described.

[0139] exist Figure 14A In the implementation scheme (also known as implementation scheme (14A)), a view has been signaled to be used as a possible predictor for reconstructing the current view.

[0140] Figure 14A The described process includes Figure 9 During step 521 of the encoding method and Figure 10 Steps 140 to 143, performed during step 621 of the decoding method, are used to generate a reference image from the first view 501 to encode the current image of the second view 501B.

[0141] In step 140, processing module 20 obtains camera parameters (i.e., view parameters) of a reference view (e.g., first view 501) and a current view (e.g., second view 501B). This is performed during step 521. Figure 14A During the process, the processing module 20 obtains these parameters directly from the cameras of the camera array 10 or from the user. This is performed during step 621. Figure 14A During the process, the processing module 20 obtains these parameters from SEI messages (e.g., the SEI messages described in Table TAB1) or from the user.

[0142] In step 141, processing module 20 generates a predicted image G(k), which is then applied between a reference view (e.g., first view 501) and the current view (e.g., second view 501B). Figure 13A and Figure 13B The described forward projection method aims to introduce a predicted image G(k) into the DPB of the current view (e.g., DPB 519B (or 619B)) to become the k-th predicted image of the current view.

[0143] Each pixel in the predicted image G(k) generated from a successful prediction retains the texture and depth values ​​of the corresponding projected pixel in the reference view. After forward projection, isolated missing pixels can be preserved (because unsuccessful projections did not fall into the second view region).

[0144] In step 142, processing module 20 fills in isolated missing pixels. In one embodiment, isolated missing pixels are filled with the average value of adjacent pixel values. In another embodiment, isolated missing pixels are filled with the median value of adjacent pixel values. In yet another embodiment, isolated missing pixels are filled with a default value (typically 128 for values ​​encoded on 8 bits).

[0145] In step 143, the processing module 20 stores the predicted image G(k) in the DPB of the current view.

[0146] In step 144, the processing module 20 reconstructs the current image of the current view using a reference image included in the DPB of the current view, the DPB including the predicted image G(k).

[0147] when Figure 14A The process applied to Figure 9In the encoding method, steps 502B, 503B, 504B, 505B, 506B, 507B, 508B, 509B, 510B, 512B, 513B, 514B, 515B, 516B, 517B, 518B, 519B, 522, and 523 are generated.

[0148] when Figure 14A The process applied to Figure 10 In the decoding method, steps 608B, 610B, 612B, 613B, 614B, 615B, 616B, 617B, 618B, 619B, and 623 are generated.

[0149] Figure 14B Details of a second implementation of the predictive image generation process are described.

[0150] exist Figure 14B In the implementation scheme (also known as implementation scheme (14B)), several views can be used to predict the current view. For example, if returning to... Figure 9 Then at time T At this point, the images of the first view (texture and depth) and the second view have been encoded and reconstructed, and are ready to be used to encode the image of the third view using the predicted images generated from the reconstructed images of the first and second views. Similarly, if returning to... Figure 10 Then at time T At this point, the images of the first view (texture and depth) and the second view have been reconstructed and are ready to be used to decode the image of the third view using the predicted image generated from the reconstructed images of the first and second views.

[0151] In implementation (14B), multiple views are used to generate an aggregated prediction image to reconstruct the current view. More precisely, in implementation (14B), a prediction image is generated for each of the multiple views, and an aggregated prediction image is generated from the multiple prediction images.

[0152] In step 140, the processing module 20 obtains the camera parameters (i.e., view parameters) of each of the multiple views and the current view.

[0153] Compared with implementation scheme (14A), step 141 is replaced by steps 1411 to 1415.

[0154] In step 1411, the processing module 20 will change the variables j Initialize to "0". Variable j Used to enumerate all views in a set of views.

[0155] In step 1412, the processing module 20 generates the predicted image. Thus in the view j Apply between the current view Figure 13A and Figure 13B The described forward projection method. For example, the view. j It is either the first view 501 or the second view 501B, and the current view is the third view.

[0156] In step 1413, the processing module 20 will change the variables j The value and the number of views in multiple views Nb_views Compare. If j < Nb_views Then, step 1413 is followed by step 1414, where... j Increment by one unit.

[0157] Step 1414 is followed by step 1412, during which a new predicted image is generated. .

[0158] if j = Nb_views Then, after step 1413, step 1415 is performed, during which the predicted image is... Aggregate to generate an aggregated prediction image G(k) intended to be stored in the DPB of the current view.

[0159] In the implementation of the aggregation process, by maintaining the first predicted image among multiple predicted images... The pixel values ​​(texture and depth values) are used to aggregate and predict the image. The first predicted image is, for example, the predicted image. .

[0160] In the implementation of the aggregation process, by maintaining the view closest to the current view... j Generated Predicted Image The pixel values ​​(texture and depth values) are used to aggregate and predict the image. If several views are equidistant from the current view (i.e., there are several closest views), then the closest view among these closest views is randomly selected. For example, in camera array 10, assume that only the first view generated by camera 10A and the second view generated by camera 10C are available to predict the current view generated by camera 10B. Then, the first view and the second view are the closest views to the current view and are equidistant from it. One of the first view and the second view is randomly selected to provide the pixel value to the aggregated prediction image G(k).

[0161] In the implementation of the aggregation process, by maintaining the predicted image The pixel values ​​(texture and depth values) have the best quality for aggregating and predicting images. For example, information representing the quality of a pixel is the value of a quantization parameter applied to the transformed block including that pixel. (Predicted image) The quality of the pixels in the image determines the application of forward projection to obtain the predicted image. The quality of the image's pixels (i.e., the quantization parameters).

[0162] In the implementation of the aggregation process, by maintaining the predicted image The pixel values ​​(texture and depth values) are aggregated to predict the image using the closest depth value (z-buffer algorithm). .

[0163] In the implementation of the aggregation process, by using images already generated from the predicted images... Preserve the neighboring pixels of the aggregated prediction image G(k) when predicting the prediction image The pixel values ​​(texture and depth values) are used to aggregate and predict the image. .

[0164] In the implementation of the aggregation process, the predicted image is calculated. The average, weighted average, and median of pixel values ​​(texture and depth values) are used to aggregate and predict the image. .

[0165] Step 142 is performed after step 1415, during which processing module 20 fills in isolated missing pixels in the aggregated prediction image G(k).

[0166] In step 143, the aggregated predicted image G(k) is stored in the DPB of the current view.

[0167] In step 144, the processing module 20 reconstructs the current image of the current view using reference images included in the DPB of the current view, the DPB including the aggregated prediction image G(k).

[0168] Figure 15 A third implementation scheme for the predictive image generation process is described.

[0169] exist Figure 15 In the implementation scheme (also known as implementation scheme (15)), similar to implementation scheme (14B), several views can be used to predict the current view.

[0170] In implementation (15), a predicted image is generated for each of the multiple views. However, instead of generating an aggregated predicted image and inserting the aggregated predicted image into the DPB of the current view as in implementation (14B), in implementation (15), each generated predicted image is inserted into the DPB.

[0171] In step 140, the processing module 20 obtains the camera parameters (i.e., view parameters) of each of the multiple views and the current view.

[0172] In step 1501, the processing module 20 will change the variables j Initialize to "0". Variable j Used to enumerate all views in a set of views.

[0173] In step 1502, the processing module 20 generates a predicted image. Thus in the view j Apply between the current view Figure 13A and Figure 13B The described forward projection method.

[0174] In step 1503, the processing module 20 fills in the isolated missing pixels in the aggregated prediction image G(k).

[0175] In step 1504, the processing module 20 predicts the image. Stored in the DPB of the current view.

[0176] In step 1505, the processing module 20 will change the variables j The value and the number of views in multiple views Nb_views Compare. If j < Nb_views Then, after step 1505, proceed to step 1506, where... j Increment by one unit.

[0177] Step 1506 is followed by step 1502, during which a new predicted image is generated. .

[0178] if j = Nb_views Then, step 144 is performed after step 1505. In step 144, processing module 20 reconstructs the current image of the current view using reference images included in the DPB of the current view. The DPB includes multiple predicted images. .

[0179] In a variation of implementation scheme 15, in addition to predicting the image In addition, from the predicted image The generated aggregated prediction image and / or the image from the prediction image The aggregated predicted image generated from a subset of the data is inserted into the DPB of the current view.

[0180] In a variation of implementation scheme 15, instead of predicting the image... From the predicted image The generated aggregated prediction image and the image from the prediction image The aggregated predicted image generated from a subset is inserted into the DPB of the current view, or only from the predicted image. The aggregated predicted image generated from a subset of the data is inserted into the DPB of the current view.

[0181] Figure 16 A fourth implementation scheme for the predictive image generation process is described.

[0182] Figure 16 The purpose of the implementation scheme (also referred to as implementation scheme (16)) is to reduce the complexity of generating the predicted image. In implementation scheme (16), each image of the reference view used to generate the predicted image of the current view is divided into blocks. The depth layer of the reference view image is then subsampled so that each block retains only one depth value. Therefore, for forward projection, all pixels of a block use the same depth value. A strategy for selecting the depth value associated with a block is defined. This strategy may include one of the following methods: ● The depth value of a specific pixel in a block represents the block: for example, the top left block or the middle block; ● The average depth value or median depth value of a block (with an average location) represents the block; ● More frequent depth values ​​(with associated location or center location) indicate blocks; The implementation scheme (16) begins at step 140, during which the processing module 20 obtains the camera parameters (i.e., view parameters) of the reference view and the current view.

[0183] In step 161, the processing module 20 will change the variables n Initialize to "1".

[0184] In step 162, processing module 20 checks the variables. N_sub The value of determines whether to apply subsampling to the image of the reference view. If N_sub If the value is 1, then subsampling is not applied to the depth layer of the reference view. In this case, step 141 is performed after step 162, during which processing module 20 generates a prediction image G(k), thereby applying subsampling between the reference view and the current view. Figure 13A and Figure 13B The described forward projection method.

[0185] In step 143, the predicted image G(k) is stored in the DPB of the current view.

[0186] Step 142 is performed after step 143, during which time processing module 20 fills in isolated missing pixels.

[0187] In step 144, the processing module 20 reconstructs the current image of the current view using the reference images included in the DPB of the current view, where the DPB includes the predicted image G(k).

[0188] If N_sub > 1, subsampling is applied to the depth layers of the reference views. In one embodiment, when N_sub > 1, N_sub is a multiple of two. If the image of the reference view has a width w and a height h , then the image is divided, for example, into blocks of equal size.

[0189] After step 162, step 163 is performed, during which the processing module 20 uses the strategy defined for the selection to select the depth value of the block number n of the image of the reference view.

[0190] In step 164, for the block number n of the image of the reference view, the processing module 20 generates a predicted block Gblock(n,k), thereby applying the Figure 13A and Figure 13B described forward projection method between the reference view and the current view.

[0191] In step 165, the processing module 20 stores the predicted block Gblock(n, k) at a location co-located with the location of the block number n of the image of the reference view in the DPB of the current view.

[0192] In step 166, the processing module 20 compares the value of the variable n with the number of blocks NB_Blocks in the image of the reference view. If n < Nb_blocks, after step 166, step 167 is performed, during which the variable n is incremented by one unit. After step 167, step 162 is performed to apply the forward projection to the new block.

[0193] If n = Nb_blocks, after step 166, steps 142 and 144, which have been described, are performed. Note that at the end of the loop within NB_Blocks of the image of the reference view, the combination of the blocks Gblock(n, k) forms the predicted image G(k).

[0194] In a variant of embodiment (16), similar to embodiments (14B) and (15), embodiment (16) can be applied to the images of multiple reference views to obtain multiple predicted images.

[0195] In this variant implementation, the predicted image from the multiple predicted images is stored in the DPB of the current view.

[0196] In this variant implementation, prediction images from multiple prediction images are aggregated to form an aggregated prediction image, and the aggregated prediction image is stored in the DPB of the current view.

[0197] In this variant implementation, at least a subset of the predicted images from a plurality of predicted images are aggregated to form an aggregated predicted image, and each aggregated predicted image is stored in the DPB of the current view.

[0198] In this variant implementation, in addition to the predicted images in the multiple predicted images and in addition to the aggregated predicted image that aggregates all the predicted images in the multiple predicted images, at least a subset of the predicted images in the multiple predicted images are aggregated to form an aggregated predicted image, and each aggregated predicted image is stored in the DPB of the current view.

[0199] In a variation of implementation (16), the image of the reference view is divided into blocks of unequal size. For example, the image is divided into large blocks (128×128, 64×64, 32×32, 16×16, or 8×8) where the depth values ​​are uniform (e.g., in areas where the difference between the minimum and maximum depth values ​​does not exceed + or -10% of the minimum depth value) and small blocks (4×4 or 2×2) where the depth values ​​are non-uniform (e.g., in areas where the difference between the minimum and maximum depth values ​​exceeds + or -10% of the minimum depth value).

[0200] In an implementation scheme referred to as bidirectional, at least one predicted image of implementation schemes (14A), (14B), (15), and (16) is used to provide a reference block (i.e., a VSP predictor block) to the current block of the current image predicted using bidirectional prediction (i.e., bidirectional inter-frame prediction). In this case, the current block is associated with two motion information, thereby specifying two reference blocks in two different images, and then the residual block of that block is the average of the two residual blocks, i.e., the first reference block is weighted by... Weighting is applied, and the second reference block is weighted. Perform weighted calculations. The sample of the current block is obtained as follows. : in It is a sample of the first reference block, and It is a sample from the second reference block.

[0201] In what is called the implementation plan WPIn the implementation schemes, at least one predicted image of implementation schemes (14A), (14B), (15), and (16) is used to provide a reference block (i.e., a VSP predictor block) to the current block of the current image predicted using weighted prediction (WP). In this case, the current block is associated with two motion information, thereby specifying two reference blocks in two different images, and then the residual block of the block is a weighted average of the two residual blocks, with the first reference block being weighted by... Weighting is applied, and the second reference block is weighted. Perform weighted calculations. Similarly, obtain the sample of the current block as follows. : It should be noted that the implementation scheme WP can use weighted sampling to generalize to all patterns, such as the triangle pattern.

[0202] As seen above, forward prediction generates a predicted image that includes isolated missing pixels. To date, isolated missing pixels have been filled using a hole-filling process. However, the hole-filling process only provides an approximation of the true pixels.

[0203] In a variant of the bidirectional and WP implementation schemes (referred to as the implementation scheme with modified weighting), the weighting process is modified to consider the confidence rate values ​​representing samples (i.e., pixels) of the predicted image. In this variant, samples of the current block are obtained as follows: : in Depends on the sample The confidence rate in, and Depends on the sample The confidence rate in the data.

[0204] In the first variant with the modified weighted implementation, when the sample (or When obtained through hole filling, (or If it equals zero, then it equals "1"; otherwise, it equals "1". Then the processing module 20 will provide the default value. .

[0205] In the second variant with the modified weighted implementation, when the sample (or When obtained through hole filling, (or Set it to a low positive value (e.g., "1"), otherwise set it to a high positive value (e.g., "10000"). In other words, in the sample... (or The value obtained through hole filling (or (lower than in the sample) (or The value obtained directly through forward projection (or ).

[0206] In a third variation with the modified weighted implementation, the confidence rate of a sample depends on the similarity of the sample to samples in its neighborhood. For example, samples that are similar to samples in their neighborhood... (or ) and samples that are higher than or different from those in their neighborhood (or ) Related values (or The value of ) (or The difference between two samples is, for example, calculated as the square root of the difference between the values ​​in the two samples.

[0207] In a fourth variant of the implementation with the modified weighted scheme, the confidence rate of a sample depends on the similarity between the sample and samples in its neighborhood and on the process applied to obtain the sample (hole filling or direct forward projection).

[0208] In the fifth variant with the modified weighted implementation, the confidence rate of the predicted image samples is calculated at the block level (typically 4×4 size) instead of the pixel level. The values ​​associated with the samples in the block... (or The confidence level depends on the average confidence rate of the samples in the block.

[0209] In the sixth variant with the modified weighted implementation, the confidence rate of a sample of the predicted image depends on the consistency of its depth values ​​with other depth maps. Consider the view. n With depth Dn Sample location Pn To view m The forward projection on corresponds to having depth Dm Sample location Pm If it has depth Dm Sample location Pm To view n The forward projection on the sample location Pn Then at the sample location Pn depth at Dn It is considered a consistent depth. Otherwise, at the sample location... Pn depth at Dn Not considered as a view mConsistent. The same process is applied to other views, and then, the range between inconsistency and complete consistency at the sample location can be established. Pn depth at Dn The score. The confidence rate of the sample is proportional to the consistency of its depth.

[0210] In some cases, blocks predicted using bidirectional inter-frame prediction mode or weighted prediction can use a reference block from the predicted image and a reference block from the image not obtained through forward prediction. In the third, fourth, fifth, and sixth variations with modified weighted implementations, samples from the image not obtained through forward prediction are considered to have the highest possible confidence rate. For example, if (or The possible values ​​of ) are: ● “0” and “1”: Samples of images that did not pass positive prediction and values ​​equal to “1”. (or Related to; ● “1” and “10000”: Samples of images not obtained through positive prediction and values ​​equal to “10000”. (or Related to; ● Between "0" and "1", the samples of images not obtained through positive prediction and the value equal to "1". (or Related to; ● Between "1" and "10000", the samples of images not obtained through positive prediction and the value equal to "10000". (or (related to)

[0211] To reduce the burden on the decoder side and overcome the maximum complexity of compliant decoders that will use inter-view prediction, a signal is sent in the encoded video stream corresponding to the encoded MVD data (e.g., encoded video streams 511 and / or 511B) informing the relationship between the current view and the view used to generate the predicted image. From this signaling, the decoder can advantageously pre-compute the predicted image. The advantage of this approach is that it allows the use of conventional decoders with minimal changes, as only the reference picture buffer padding (i.e., DPB padding) is modified.

[0212] In the following text, it is referred to as view_parameter The syntax elements represent information that allows for the reconstruction or aggregation of predicted images and is suitable for the embodiments presented above. In one embodiment, the syntax elements are... view_ parameter Insert at the slice header level in the encoded video stream. In another implementation, the syntax element... view_ parameter Insert the syntax element into the sequence header (i.e., the sequence parameter set (SPS)), the header of an image or an image shared by multiple images (i.e., the picture parameter set (PPS)), or at a synchronization point or image level in the encoded video stream (e.g., in the header of an IDR (Instant Decode Refresh) image). Each time the decoder receives the syntax element, it can update its knowledge about the relationships between views. Table TAB2 .

[0213] Table TAB2 represents the syntax elements suitable for implementations in a DPB where only one predicted image or only one aggregated predicted image is inserted into the current view (typically implementations (14A), (14B), and (16) when only one predicted image or only one aggregated predicted image is generated). view_parameter The first form.

[0214] Syntax elements view_parameter The first type includes parameters view_id This parameter represents a unique identifier for the current view. If the current view is not the first view decoded for a frame, it is marked as such. vsp_flag Indicates whether the VSP mode is used for the current view. Parameter number_of_inter_view_predictor_used Indicates the maximum number of views (already decoded) used to decode the current view. Parameter predictor_id[view_id][i] An identifier is provided for each view to create a predicted image for the current view. In one implementation, the maximum number of views used for decoding the current view is fixed at "8". In this case, "3" bits are needed for the parameters. predictor_id[view_id] Encode it.

[0215] Of course, inter-view prediction between the first and second views is possible only if the camera parameters for both views are available on the decoder side, i.e., if the SEI message described in Table TAB1 is received and decoded by the decoder. Table TAB3 .

[0216] Table TAB3 represents the syntax elements suitable for implementations in a DPB where multiple predicted images and / or multiple aggregated predicted images are inserted into the current view (typically implementations (15) and (16) when multiple predicted images or multiple aggregated predicted images are generated). view_parameter The second form.

[0217] In this scenario, the current view can be associated with multiple reference views. (In the syntax element...) view_parameterIn this second form, the parameters number_inter_view_predictor_minus1 Specify for use by parameters view_id The number of predicted or aggregated predicted images between views of the current view. Parameter number_interview_ predictor_used_minus1 For each predicted image or aggregated predicted image, multiple reference views are specified for generating the predicted image or aggregated predicted image. In the case of predicted images, parameters... parameter number_inter_ view_predictor_used_minus1 Set to one. Parameter predictor_id Specify which view or views are used to generate the predicted image or aggregate the predicted image.

[0218] As can be seen from Tables TAB2 and TAB3, it can be determined by the markings. vsp_flag In syntax elements view_parameter Activate VSP mode at the slice header level.

[0219] If a block included in the slice might use VSP mode, a signal at the slice level allows the decoder to be notified. However, it does not specify which block in the slice actually uses VSP mode.

[0220] In one implementation, when activated at the slice level, the actual use of the VSP mode is signaled at the block level.

[0221] Figure 17 The basic implementation scheme of the syntax parsing process of a video compression method that does not use VSP mode is schematically depicted.

[0222] Figure 17 The basic implementation is based on the syntax of the blocks (also known as prediction units (PUs)) described in Table TAB4. This basic implementation is performed by the decoder when decoding the current block. However, the encoder encodes syntax that conforms to what the decoder can decode.

[0223] In step 1700, processing module 20 determines whether to encode the current block in skip mode. If so, processing module 20 assigns an identifier to the current block. merge_idx Decode. Identifier merge_idx The identifier identifies which candidate block in the neighborhood of the current block provides information for decoding the current block. (This is related to the identifier.) merge_idx After decoding, the current block is decoded using a decoding process suitable for skip mode.

[0224] If the current block is not encoded in skip mode, then in step 1701, processing module 20 determines whether the current block should be encoded in intra-frame mode. If so, then in step 1702, the current block is decoded using the intra-frame mode decoding process.

[0225] If the current block is not encoded in intra-frame mode, then in step 1703, the processing module determines whether to encode the current block in merge mode. If the current block is encoded in merge mode, then in step 1704, the processing module assigns an identifier to the current block. merge_idx Decode the identifier. merge_idx After decoding, the current block is decoded using a decoding process suitable for the merge mode.

[0226] If the current block is not encoded in merge mode, then in step 1705, the processing module 20 determines whether to encode the current block in bidirectional or unidirectional inter-frame prediction mode.

[0227] If the current block is encoded in unidirectional inter-frame prediction mode, then step 1712 is performed after step 1705, during which time the processing module 20 performs an index ( ) on a reference image list stored in the DPB. ref_idx_l0 or ref_ idx_l1 Decoding is performed. The index indicates which reference image provides the predictor block for the current block.

[0228] In step 1713, the processing module 20 refines the motion vector of the current block. mvd Decode it.

[0229] In step 1714, processing module 20 decodes the motion vector predictor index that indicates the motion vector predictor. Using this motion information, processing module 20 decodes the current block.

[0230] When encoding the current block in bidirectional prediction mode, step 1706 is performed after step 1705, during which processing module 20 performs processing on the first index (in the reference image list stored in the DPB). ref_idx_l0 Decode it.

[0231] In step 1707, the processing module 20 refines the first motion vector of the current block. mvd Decode it.

[0232] In step 1708, the processing module 20 decodes the first motion vector predictor index that indicates the first motion vector predictor.

[0233] In step 1709, the processing module 20 applies the second index ( ) to the reference image list stored in the DPB. ref_ idx_l1 Decode it.

[0234] In step 1710, the processing module 20 refines the second motion vector of the current block. mvdDecode it.

[0235] In step 1711, the processing module 20 decodes the second motion vector predictor index that indicates the second motion vector predictor.

[0236] Based on this motion information, the processing module 20 generates two predictors and uses these two predictors to decode the current block. Table TAB4 .

[0237] Figure 18 A first implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated schematically.

[0238] Figure 18 The implementation plan (hereinafter referred to as the implementation plan) 18 This is based on the syntax of the blocks described in Table TAB5. Differences between the syntaxes in Tables TAB4 and TAB5 are shown in bold. This implementation is performed by the decoder when decoding the current block. However, the encoder encodes syntax that conforms to what the decoder can decode. Table TAB5 .

[0239] As will be described below, in implementation scheme (18), by the mark VSP Signal the use of VSP mode at the block level. When marked VSP When =1, VSP mode is activated for the current block. Otherwise, it is deactivated. Furthermore, in implementation (18), when a block is encoded in VSP mode, the predictor block is co-localized with the current block. Therefore, the block predictor can be obtained from the reference image (which in this case is the prediction image or aggregated prediction image) without motion vectors. Additionally, as will be discussed later... Figure 18 As will be described, the combination of VSP mode and bidirectional inter-frame prediction is not possible when extracting two predictor blocks from the same predicted image or aggregated predicted image. In fact, since each predictor is co-located with the current block in VSP mode, the two predictor blocks are identical in the case of bidirectional inter-frame prediction.

[0240] When only one predicted image or aggregated predicted image is inserted into the DPB of the current layer, the syntax and parsing method of implementation (18) are suitable for implementations (14A), (14B) and (16).

[0241] In step 1800, processing module 20 determines whether to encode the current block in skip mode. If so, in step 1804, processing module 20 assigns an identifier to the current block.merge_idx Decode the identifier. merge_idx After decoding, the current block is decoded using a decoding process suitable for skip mode.

[0242] If the current block is not encoded in skip mode, then in step 1801, processing module 20 determines whether the current block should be encoded in intra-frame mode. If so, then in step 1802, the current block is decoded using the intra-frame mode decoding process.

[0243] If the current block is not encoded in intra-frame mode, then in step 1803, the processing module determines whether to encode the current block in merge mode. If the current block is encoded in merge mode, then in step 1806, the processing module assigns an identifier to the current block. merge_idx Decode the identifier. merge_idx After decoding, the current block is decoded using a decoding process suitable for the merge mode.

[0244] If the current block is not encoded in merge mode, then in step 1807, the processing module 20 determines whether to encode the current block in bidirectional or unidirectional inter-frame prediction mode.

[0245] If the current block is encoded in unidirectional inter-frame prediction mode, then step 1808 is performed after step 1807, during which time processing module 20 processes the markers. VSP Decoding is performed to determine whether the current block is encoded in VSP mode. If the current block is encoded in VSP mode, processing module 20 decodes the current block according to the VSP mode decoding process. In other words, the current block is predicted from blocks of predicted images (or aggregated predicted images) stored in the DPB of the current view. In this case, the position of the predicted images (or aggregated predicted images) in the DPB is implicit and known to the decoder (i.e., the predicted images (or aggregated predicted images) are systematically located at the same position in the DPB).

[0246] If the current block is not encoded in VSP mode, step 1810 is performed after step 1808, during which processing module 20 decodes an index (ref_idx_l0 or ref_idx_l1) in the list of reference images stored in DPB.

[0247] In step 1811, the processing module 20 refines the motion vector of the current block. mvd Decode it.

[0248] In step 1812, processing module 20 decodes the motion vector predictor index that indicates the motion vector predictor. Using this motion information, processing module 20 decodes the current block.

[0249] When encoding the current block in bidirectional prediction mode, step 1813 is performed after step 1807, during which processing module 20 decodes the marker VSP to determine whether the first predictor block of the current block is obtained from the prediction image (or from the aggregated prediction image). If the first predictor block of the current block is obtained from the prediction image (or from the aggregated prediction image), the first predictor block is obtained in step 1814, which is the same as step 1809. Step 1814 is followed by step 1819, during which processing module 20 decodes the index (ref_idx_l1) in the reference image list stored in the DPB.

[0250] In step 1820, the processing module 20 decodes the motion vector refinement mvd of the current block.

[0251] In step 1821, processing module 20 decodes the motion vector predictor index indicating the motion vector predictor. Using the motion information obtained in steps 1819, 1820, and 1821, processing module 20 determines the second predictor block. Using these two predictors, processing module 20 determines the bidirectional predictor to decode the current block.

[0252] If the first predictor block for the current block is not obtained from the predicted image (or from the aggregated predicted image), the processing module executes steps 1815, 1816, and 1817, which are the same as steps 1810, 1811, and 1812, respectively, to obtain the first predictor.

[0253] In step 1818, processing module 20 decodes the marked VSP to determine whether a second predictor block for the current block is obtained from the predicted image (or from the aggregated predicted image). If a second predictor block for the current block is obtained from the predicted image (or from the aggregated predicted image), then in step 1822, which is the same as step 1809, the second predictor block is obtained. Using the first and second predictor blocks, the processing module decodes the current block.

[0254] If the second predictor block for the current block is not obtained from the predicted image (or from the aggregated predicted image), then in step 1819, the processing module 20 adjusts the second index (in the reference image list stored in the DPB) ref_idx_l1 Decode it.

[0255] In step 1820, the processing module 20 refines the second motion vector of the current block. mvd Decode it.

[0256] In step 1821, the processing module 20 decodes the second motion vector predictor index that indicates the second motion vector predictor.

[0257] Using the motion information obtained in steps 1815, 1816, 1817, 1819, 1820, and 1821, the processing module 20 decodes the current block.

[0258] In a variant of implementation scheme (18), during step 1804, processing module 20 decodes the VSP tag of the current block. If VSP mode is activated for the current block, processing module 20 performs step 1805, which is the same as step 1809. If VSP mode is not activated for the current block, processing module 20 performs step 1806.

[0259] Figure 19 A second implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated schematically.

[0260] Figure 19 The implementation plan (hereinafter referred to as the implementation plan) 19 This is based on the syntax of the blocks described in Table TAB6. Differences between the syntaxes in Table TAB4 and Table TAB6 are shown in bold. This implementation is performed by the decoder when decoding the current block. However, the encoder encodes syntax that conforms to what the decoder can decode. Table TAB6 .

[0261] Implementation scheme (19) is very similar to implementation scheme (18). The difference between implementation scheme (19) and implementation scheme (18) is that the syntax of the blocks encoded in VSP mode includes the representation of motion vector differences. mvd The syntax elements. The result of this feature is that when extracting two predictor blocks from the same predicted image or aggregated predicted image, a combination of VSP mode and bidirectional inter-frame prediction is possible. In fact, in implementation (19), the motion vector difference... mvd The existence of this allows for the acquisition of two different predictor blocks.

[0262] When only one predicted image or aggregated predicted image is inserted into the DPB of the current layer, the syntax and parsing method of implementation (19) are suitable for implementations (14A), (14B) and (16).

[0263] The implementation scheme (19) includes steps 1900 to 1908, 1910 to 1813, and 1815 to 1821, which are the same as steps 1800 to 1808, 1810 to 1813, and 1815 to 1821, respectively.

[0264] When VSP mode is activated for the current block, step 1909 is performed after step 1908, during which the motion vector difference is calculated for the current block. mvdThe difference in motion vectors mvd This allows indicating predictor blocks within the predicted image or aggregated predicted image. The predictor is then used to decode the current block.

[0265] When the VSP flag specifies a first predictor for generating the current block from the predicted image or from the aggregated predicted image, step 1914 proceeds after step 1913, during which the motion vector difference is calculated for the current block. mvd The difference in motion vectors mvd Allows indicating the first predictor block in the predicted image or aggregated predicted image.

[0266] Step 1918 is performed after step 1914. Step 1922 is performed after step 1918 when the VSP flag specifies a second predictor for generating the current block from either the predicted image or the aggregated predicted image, during which the motion vector difference is calculated for the current block. mvd The difference in motion vectors mvd Allows indication of a second predictor block in the predicted image or aggregated predicted image. Decodes the current block from the first and second predictors, as in bidirectional prediction mode.

[0267] Note that steps 1905 and 1909 are the same.

[0268] Figure 20 A third implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated schematically.

[0269] Figure 20 The implementation plan (hereinafter referred to as the implementation plan) 20 This implementation is based on the syntax of the blocks described in Table TAB7. Differences between the syntaxes in Table TAB4 and Table TAB7 are shown in bold. This implementation is performed by the decoder when decoding the current block. However, the encoder encodes syntax that conforms to what the decoder can decode.

[0270] Implementation scheme (20) is very similar to implementation scheme (19). The difference between implementation scheme (20) and implementation scheme (19) is that the syntax of the blocks encoded in VSP mode no longer includes the representation of motion vector differences. mvd The syntax elements, but including at least one index from the list of reference images stored in the DPB ( ref_idx2_l0 or ref_idx2_l1 The index indicates which predicted image provides the predictor block for the current block. As a result of this feature, the combination of VSP mode and bidirectional inter-frame prediction is now possible. In fact, in implementation (20), the existence of two indices in the case where bidirectional inter-frame prediction indicates two different predicted images (or aggregated predicted images) allows for the acquisition of two different predictor blocks.

[0271] When multiple predicted images or aggregated predicted images are inserted into the DPB of the current layer, the syntax and parsing method of implementation (20) are adapted to implementations (15) and (16).

[0272] The implementation plan (20) includes steps 2000 to 2008, 2010 to 2013, and 2015 to 2021, which are the same as steps 1900 to 1908, 1910 to 1913, and 1915 to 1921, respectively.

[0273] Step 1909 of implementation scheme (19) is replaced by step 2009 in implementation scheme (20). In step 2009, the processing module (20) processes a list of reference images representing the predicted images or aggregated predicted images that temporally correspond to (i.e., in the same frame as) the image including the current block. l0 Syntax elements of indexes in (or l1) ref_idx2_l0 (or ref_ idx2_l1 Decoding is performed. Processing module 20 decodes from the index. ref_idx2_l0 (or ref_idx2_l1 The predictor block that is spatially colocalized with the current block is extracted from the predicted image (or aggregated predicted image) indicated by the current block. The processing module then uses the obtained predictor block to decode the current block.

[0274] Step 1914 of implementation scheme (19) is replaced by step 2014 in implementation scheme (20). In step 2014, the processing module (20) processes a first list of reference images to be used in the predicted images or aggregated predicted images that correspond temporally to (i.e., in the same frame as) the image including the current block. l0 Syntax elements of index in ref_idx2_l0 Decoding is performed. Processing module 20 extracts the first predictor block that is spatially co-located with the current block from the prediction image (or aggregated prediction image) indicated by index ref_idx2_l0.

[0275] Step 1922 of implementation scheme (19) is replaced by step 2022 in implementation scheme (20). In step 2022, the processing module (20) processes a second list of reference images to be used in the predicted images or aggregated predicted images that correspond temporally to (i.e., in the same frame as) the image including the current block. l1 Syntax elements of index in ref_idx2_l1 Decoding is performed. Processing module 20 retrieves the data from the index. ref_idx2_l1 The indicated prediction image (or aggregated prediction image) extracts a second predictor block that is spatially colocalized with the current block.

[0276] After step 2021 or 2022, the processing module 20 uses the first predictor and the second predictor to decode the current block, as in the dual-prediction inter-frame mode.

[0277] Note that steps 2005 and 2009 are the same. Table TAB .

[0278] ​ A fourth implementation of the syntax parsing process for a video compression method using the new VSP mode is illustrated.

[0279] ​ The implementation scheme (hereinafter referred to as implementation scheme (21)) is based on the syntax of the blocks described in Table TAB8. The differences between the syntaxes of Table TAB4 and Table TAB8 are indicated in bold. This implementation scheme is executed by the decoder when the current block is decoded. However, the encoder encodes the syntax that conforms to the content that the decoder can decode.

[0280] Implementation scheme (21) is very similar to implementation scheme (18). However, in implementation scheme (21), the VSP mode is used at the block level from the index of the reference image ( ​ or ​ Inference, rather than from the label ​ Explicitly specify.

[0281] When only one predicted image or aggregated image is inserted into the DPB of the current layer, the syntax and parsing method of implementation (21) are suitable for implementations (14A), (14B) and (16).

[0282] The implementation scheme (21) includes steps 2100 to 2103, 2105 to 2107, 2109 to 2112, 2114 to 2117, and 2119 to 2122, which are the same as steps 1800 to 1803, 1805 to 1807, 1809 to 1812, 1814 to 1817, and 1819 to 1822, respectively.

[0283] In step 2104, if the index of the reference image in the reference image list corresponds to the predicted image or the aggregated predicted image, then... ​ or index ​ Inherited free identifier ​ If the indicated candidate block is selected, the processing module 20 considers that the VSP mode should be activated for the current block.

[0284] In step 2108, if a list of reference images is indicated l0 Index of reference images in ​If the predicted image or aggregated predicted image is indicated, then the VSP mode is considered to be activated for the current block. For example, ​ =0 specifies a reference image corresponding to the predicted image or aggregated predicted image.

[0285] In step 2113, if a list of reference images is indicated l0 Index of reference images in ​ If the indicated predicted image or aggregated predicted image is provided, the processing module 20 considers that the first predictor of the current block is obtained from the predicted image or aggregated predicted image.

[0286] In step 2118, if a list of reference images is indicated l1 Index of reference images in ​ If the processing module 20 indicates a predicted image or an aggregated predicted image, it considers obtaining a second predictor for the current block from the predicted image or the aggregated predicted image. For example, ​ =0 specifies a reference image corresponding to the predicted image or aggregated predicted image. ​ .

[0287] In the syntax of Table TAB8, only in functions... ​ Difference of motion vectors when the return value is false ​ and motion vector predictor index ​ Perform decoding. This function... ​ Defined as: ​ ​ ● If referring to the index ​ If the frame is generated from a view within the same frame, then it returns true; ● Otherwise, return false.

[0288] A variation of implementation scheme (18) (hereinafter referred to as implementation scheme () ​ In the context of encoding the current block in merge mode or skip mode, no tags are specified for the current block. ​ Encoding is performed. In this case, processing module 20 first encodes the identifier. ​ Perform decoding and determine whether the identifier is in VSP mode. ​ ​ The indicated candidate block is encoded. If the candidate block is encoded in VSP mode, the current block inherits the VSP parameters from the candidate block, and these parameters are used to decode the current block. Otherwise, the normal merge mode decoding process is applied to decode the current block. This implementation (18bis) is based on the block syntax described in Table TAB9.

[0289] Implementation schemes (18), (18bis), (19), (20) and (21) can be combined to obtain additional implementation schemes.

[0290] For example, the syntax for encoding the current block in VSP mode may include motion vector differences. ​ And represents the first list of reference images to be used in the predicted images or aggregated predicted images that correspond to (i.e., in the same frame as) the image including the current block. l0 Neutralize / or second list l1 Syntax elements of index in ​ and / or ​ ​ This corresponds to the combination of implementation schemes (19) and (20).

[0291] In another example, the syntax for the current block encoded in VSP mode may include motion vector differences. ​ Furthermore, the use of VSP mode can be found in the syntax elements. ​ and / or ​ Inference, rather than by labels ​ Instructions. This corresponds to a combination of implementation schemes (19) and (21).

[0292] In another example, the syntax of the current block encoded in VSP mode may include syntax elements. ​ and / or ​ Furthermore, the use of VSP mode can be found in the syntax elements. ​ and / or ​ Inference, rather than by labels ​ Instructions. This corresponds to the combination of implementation schemes (20) and (21).

[0293] In other examples: ● Implementation scheme (18bis) can be combined with implementation schemes (19), (20) and (21); ● Implementation schemes (19), (20) and (21) can be combined; ● Implementation plans (19), (20), (21) and (22) can be combined; ● etc. ​ .

[0294] To date, it has been believed that projected images (or aggregated projected images) used for inter-view prediction include per-pixel texture data and depth data. This is referred to as being based on... ​ (Sports Information) ​In another embodiment of the implementation, the predicted image (and aggregated predicted image) is replaced with an image called an MI (motion information) predicted image (or MI aggregated predicted image), which includes motion information only for each pixel or subset of pixels.

[0295] In a MI-based VSP implementation plan, ​ The forward projection process includes an additional step 133. During step 133, the processing module 20 calculates a motion vector representing the displacement between the pixel P'(u', v') obtained through the forward projection in steps 130 to 132 and the projected pixel P(u, v). ​ The motion vector MV is intended to be stored in the MI prediction image.

[0296] The impact of the MI-based VSP implementation on implementations (14A), (14B), (15), and (16) is described below.

[0297] In the MI-based VSP implementation, implementation (14A) is modified and becomes implementation (14A_MI). Implementation (14A_MI) in ​ The Chinese side indicated that...

[0298] The first step of the implementation plan (14A_MI) is step 140, which has already been described with respect to the implementation plan (14A).

[0299] In step 141_MI, processing module 20 generates an MI prediction image MI(k), thereby applying the forward projection of steps 130 to 133 between a reference view (e.g., first view 501) and a current view (e.g., second view 501B). The aim is to introduce the MI prediction image MI(k) into the DPB (e.g., DPB 519B (or 619B)) of the current view to become the k-th prediction image of the current image of the current view.

[0300] In step 142_MI, processing module 20 fills in isolated missing motion information. In one embodiment, isolated missing motion information is filled with motion information of adjacent pixels. In another embodiment, isolated missing motion information is filled with a default value (typically motion vector = (0, 0)). In yet another embodiment, isolated missing motion information is considered invalid, and a flag indicating the validity of the motion information is associated with each piece of motion information.

[0301] In step 143_MI, the processing module 20 stores the MI prediction image MI(k) in the DPB of the current view.

[0302] In step 144_MI, processing module 20 reconstructs the current image of the current view using a reference image included in the DPB of the current view, the DPB including the MI prediction image MI(k). Processing module 20 uses the MI prediction image MI(k) to generate a prediction image G(k). In fact, the motion information included in the MI prediction image MI(k) is used to apply motion compensation to each pixel in the reference image specified by the motion information.

[0303] Optionally, implementation (14A_MI) includes step 220, which involves reducing the amount of motion information in the MI prediction image MI(k). In practice, motion information for each pixel location of the image represents a large amount of data. In one implementation, the MI prediction image MI(k) is divided into blocks of size N×M, where N and M are multiples of two and smaller than the width and height of the MI prediction image MI(k). Only one piece of motion information is retained for each N×M block. In other words, the motion information is subsampled by a factor of N×M. In the implementation where N=M=4, one piece of motion information retains no more than “16” pieces of motion information.

[0304] In the implementation scheme, subsampling involves selecting a specific motion information for each block from the N×M motion information.

[0305] In the implementation scheme, subsampling involves selecting the median of each block from the N×M motion information (the median is calculated using the norm of the motion vector).

[0306] In the implementation scheme, subsampling includes selecting the most frequently occurring motion information from N×M motion information.

[0307] In the implementation, subsampling includes selecting motion information of the minimum depth (z-buffer algorithm) in the view corresponding to the entire sub-block.

[0308] In the implementation scheme, subsampling includes retaining the first projection value in the N×M motion information.

[0309] In the MI-based VSP implementation, implementation (14B) is modified and becomes implementation (14B_MI). Implementation (14B_MI) in ​ The Chinese side indicated that...

[0310] Compared with implementation scheme (14B), in implementation scheme (14B_MI), step 1412 is replaced by step 1412_MI, and step 1415 is replaced by step 1415_MI.

[0311] In step 1412_MI, processing module 20 generates the MI prediction image. Thus in the view jApply the forward projection method from steps 130 to 133 between the view and the current view.

[0312] In step 1415_MI, the processing module 20 extracts data from the predicted image. Calculate the aggregated MI prediction image MI(k), which is intended to be stored in the DPB of the current view.

[0313] In an implementation of the aggregation process, the first MI prediction image among multiple MI prediction images is retained. Motion information is used to aggregate and predict images The first MI prediction image is, for example, the predicted image. .

[0314] In the implementation of the aggregation process, by maintaining the view closest to the current view... j Generated Predicted Image Motion information is used to aggregate MI prediction images If several views are equidistant from the current view (i.e., there are several closest views), then the closest view among these closest views is randomly selected. For example, in camera array 10, assume that only the first view generated by camera 10A and the second view generated by camera 10C are available to predict the current view generated by camera 10B. Then, the first view and the second view are the closest views to the current view and are equidistant from it. One of the first view and the second view is selected to provide motion information to the MI aggregation prediction image MI(k).

[0315] In the implementation of the aggregation process, by maintaining the predicted image Motion information of optimal quality is used to aggregate MI prediction images. For example, information representing the quality of a pixel is the value of a quantization parameter applied to the transformed block that includes the pixel.

[0316] In the implementation of the aggregation process, by maintaining the predicted image The motion information has the closest depth value (z-buffer algorithm) to aggregate MI prediction images. .

[0317] In the implementation of the aggregation process, by using images already generated from the predicted images... Preserve the neighboring pixels of the predicted image MI(k) when predicting aggregated MI. The pixel values ​​(texture and depth values) are used to aggregate the MI prediction image. .

[0318] In the implementation of the aggregation process, the MI prediction image is calculated. The average, weighted average, and median of motion information are used to aggregate MI prediction images. .

[0319] It should be noted that motion information includes information representing motion vectors and indices representing reference images in the list of reference images (e.g., ...). ​ , ​ , ​ , ​ (information).

[0320] As can be seen from the above, in implementation scheme (14B_MI), subsampling (step 220) is performed on the aggregated MI prediction image MI(k). In a variant of implementation scheme (14B_MI), in each MI prediction image... Perform subsampling (step 220).

[0321] As can be seen from the above, in implementation scheme (14B_MI), subsampling (step 220) and aggregation step 1415_MI are separate steps. In a variant of implementation scheme (14B_MI), subsampling is performed during the aggregation step.

[0322] In the MI-based VSP implementation, implementation (15) is modified and becomes implementation (15_MI). Implementation (15_MI) in ​ The Chinese side indicated that...

[0323] Compared with implementation scheme (15), in implementation scheme (15_MI), step 1502 is replaced by step 1502_MI, step 1503 is replaced by step 1503_MI, step 1504 is replaced by step 1504_MI, and step 144 is replaced by step 144_MI.

[0324] Step 1502_MI is the same as step 1412_MI.

[0325] Step 1503_MI is the same as step 142_MI, except that the hole filling process is applied to the MI prediction image. Instead of MI predicting the image MI(k).

[0326] During step 1504_MI, the predicted image Stored in the DPB of the current view.

[0327] Step 144_MI in implementation scheme (15_MI) is the same as step 144_MI in implementation scheme (14B_MI), the difference being that the DPB in the current view includes the quantity. ​ MI prediction image .

[0328] In a variation of the implementation scheme (15_MI), a subsampling step 220 is introduced between steps 1503_MI and 1504_MI.

[0329] In a variant of the implementation scheme (15_MI), in addition to predicting the image In addition, from the predicted image The generated aggregated prediction image and / or the image from the prediction image The aggregated predicted image generated from a subset of the data is inserted into the DPB of the current view.

[0330] In a variant of the implementation scheme (15_MI), instead of predicting the image... From the predicted image The generated aggregated prediction image and the image from the prediction image The aggregated predicted image generated from a subset is inserted into the DPB of the current view, or only from the predicted image. The aggregated predicted image generated from a subset of the data is inserted into the DPB of the current view.

[0331] In the MI-based VSP implementation, implementation (16) is modified and becomes implementation (16_MI). Implementation (16_MI) in ​ The Chinese side indicated that...

[0332] Compared with implementation scheme (16), in implementation scheme (16_MI), step 164 is replaced with step 164_MI, step 165 is replaced with step 165_MI, step 141 is replaced with step 141_MI, step 143 is replaced with step 143_MI, step 142 is replaced with step 142_MI, and step 144 is replaced with step 144_MI.

[0333] Step 141_MI in implementation scheme (16_MI) is the same as step 141_MI in implementation scheme (14A_MI).

[0334] Step 143_MI in implementation scheme (16_MI) is the same as step 143_MI in implementation scheme (14A_MI).

[0335] Step 142_MI in implementation scheme (16_MI) is the same as step 142_MI in implementation scheme (14A_MI).

[0336] Step 144_MI in implementation scheme (16_MI) is the same as step 144_MI in implementation scheme (14A_MI).

[0337] In step 164_MI, processing module 20 generates a predicted motion information block MIblock(n, k), thereby determining the block numbering of the reference view and the image of the reference view.n The forward projection method described in steps 130 to 133 is applied between the current views.

[0338] In step 165_MI, processing module 20 stores block MIblock(n, k) in the DPB of the current view with the block number of the image of the reference view. n The location is the same as the location of the coordinate system.

[0339] In a variant of implementation (16_MI), similar to implementations (14B) and (15), implementation (16) can be applied to images of multiple reference views to obtain multiple MI prediction images.

[0340] In this variant implementation, the MI prediction image among multiple MI prediction images is stored in the DPB of the current view.

[0341] In this variant implementation, MI prediction images from multiple MI prediction images are aggregated to form an aggregated MI prediction image, and the aggregated MI prediction image is stored in the DPB of the current view.

[0342] In this variant implementation, at least a subset of the MI prediction images from a plurality of MI prediction images are aggregated to form an aggregated MI prediction image, and each aggregated MI prediction image is stored in the DPB of the current view.

[0343] In this variant implementation, in addition to the MI prediction image and in addition to the aggregated MI prediction image which aggregates all the prediction images in the plurality of prediction images, at least a subset of the MI prediction images in the plurality of MI prediction images are aggregated to form an aggregated MI prediction image, and each aggregated MI prediction image is stored in the DPB of the current view.

[0344] The implementation schemes bidirectional, WP, and the modified weighted implementation schemes apply the same approach to all MI-based VSP implementation schemes (i.e., implementation schemes (14A), (14B), (15), and (16)).

[0345] To date, motion information has been considered to include information representing motion vectors and indices representing reference images in a list of reference images (e.g., ​ , ​ , ​ , ​ The information is as follows. In a variant of the MI-based VSP implementation, when the MI-predicted image MI(k) is divided into blocks of size N×M, the motion information associated with each N×M block includes parameters of the affine model of the motion, thereby allowing the pixels of the current block in the current view to be determined from the pixels of the N×M blocks in the reference view rather than information representing motion vectors.

[0346] When only one MI prediction image or an aggregated MI prediction image is inserted into the DPB of the current layer, the syntax and parsing methods of implementation schemes (18) and (18bis) are suitable for implementation schemes (14A_MI), (14B_MI) and (16_MI).

[0347] When only one MI prediction image or an aggregated MI prediction image is inserted into the DPB of the current layer, the syntax and parsing method of implementation (19) are suitable for implementations (14A_MI), (14B_MI) and (16_MI).

[0348] When multiple MI prediction images or aggregated MI prediction images are inserted into the DPB of the current layer, the syntax and parsing method of implementation (20) are suitable for implementations (15_MI) and (16_MI).

[0349] When only one MI prediction image or an aggregated MI prediction image is inserted into the DPB of the current layer, the syntax and parsing method of implementation (21) are suitable for implementations (14A_MI), (14B_MI) and (16_MI).

[0350] The implementation schemes featuring the combined implementation schemes (18), (18bis), (19), (20) and (21) are also applicable to MI-based VSP implementation schemes.

[0351] The foregoing describes several embodiments. Features of these embodiments may be provided individually or in any combination. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim classes and types: ● Includes a bitstream or signal of one or more syntax elements or their variants from the described syntax elements. ● To create and / or transmit and / or receive and / or decode bit streams or signals that include one or more of the described syntax elements or their variants. ● A television, set-top box, mobile phone, tablet computer or other electronic device that performs MVD encoding or decoding according to any of the described implementation schemes. ● A television, set-top box, mobile phone, tablet computer, or other electronic device that performs MVD decoding and displays the resulting image (e.g., using a monitor, screen, or other type of display) according to any of the described implementation schemes. ● A television, set-top box, mobile phone, tablet computer, or other electronic device that tunes (e.g., uses a tuner) a channel to receive signals including encoded video streams and performs multi-view decoding according to any of the described embodiments. ● A television, set-top box, mobile phone, tablet computer, or other electronic device that receives signals including encoded video streams over the air (e.g., using an antenna) and performs MVD decoding according to any of the described embodiments.

Claims

1. A method for decoding, comprising: Obtain first camera parameters associated with at least one reference view and second camera parameters associated with the current view of the multi-view video content, wherein each view includes a texture layer and a depth layer; An intermediate prediction image is generated by applying a forward projection method to the pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameters to the camera coordinate system of the current view defined by the second camera parameters. Each projected pixel of the intermediate prediction image is associated with a motion information value, which represents the displacement between the projected pixel and the pixel of the reference view from which it is projected. At least one final predicted image obtained from at least one intermediate predicted image is stored in the decoded image buffer of the reconstructed image for temporal prediction of the current view; as well as Reconstruct the current image of the current view from the image stored in the decoded image buffer.

2. The method of claim 1, wherein, The forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by projecting from the camera coordinate system of the reference view to the world coordinate system. The deprojection uses the pose matrix of the reference camera of the reference view, the inverse intrinsic parameter matrix of the reference camera, and the depth value associated with the current pixel. The pose matrix and inverse intrinsic parameter matrix of the reference camera are obtained according to the first camera parameters. The deprojected pixels are projected into the camera coordinate system of the current view using the intrinsic and extrinsic parameter matrices of the current camera to obtain forward-projected pixels. The intrinsic and extrinsic parameter matrices of the current camera are obtained according to the second camera parameters. If the projected pixel does not correspond to a pixel in the current camera's pixel grid, then the pixel in the pixel grid closest to the projected pixel is selected to obtain a corrected projected pixel; and Calculate a motion vector representing the displacement between the forward-projected pixel and the current pixel of the reference view.

3. The method according to claim 1, further comprising: In each intermediate or final predicted image, the motion information values ​​of isolated missing pixels are filled with the motion information values ​​of adjacent pixels or the default value.

4. The method according to claim 1, wherein, The at least one final predicted image is obtained by aggregating at least two intermediate predicted images.

5. A method for encoding, comprising: Obtain first camera parameters associated with at least one reference view and second camera parameters associated with the current view of the multi-view video content, wherein each view includes a texture layer and a depth layer; An intermediate prediction image is generated by applying a forward projection method to the pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameters to the camera coordinate system of the current view defined by the second camera parameters. Each projected pixel of the intermediate prediction image is associated with a motion information value, which represents the displacement between the projected pixel and the pixel of the reference view from which it is projected. At least one final predicted image obtained from at least one intermediate predicted image is stored in the decoded image buffer of the reconstructed image for temporal prediction of the current view; as well as Reconstruct the current image of the current view from the image stored in the decoded image buffer.

6. The method according to claim 5, wherein, The forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by projecting from the camera coordinate system of the reference view to the world coordinate system. The deprojection uses the pose matrix of the reference camera of the reference view, the inverse intrinsic parameter matrix of the reference camera, and the depth value associated with the current pixel. The pose matrix and inverse intrinsic parameter matrix of the reference camera are obtained according to the first camera parameters. The deprojected pixels are projected into the camera coordinate system of the current view using the intrinsic and extrinsic parameter matrices of the current camera to obtain forward-projected pixels. The intrinsic and extrinsic parameter matrices of the current camera are obtained according to the second camera parameters. If the projected pixel does not correspond to a pixel in the current camera's pixel grid, then the pixel in the pixel grid closest to the projected pixel is selected to obtain a corrected projected pixel; and Calculate a motion vector representing the displacement between the forward-projected pixel and the current pixel of the reference view.

7. The method according to claim 5, further comprising: In each intermediate or final predicted image, isolated missing pixel values ​​are filled with the average of neighboring pixel values, the median of neighboring pixel values, or a default value.

8. The method according to claim 5, wherein, The at least one final predicted image is obtained by aggregating at least two intermediate predicted images.

9. A device for decoding, comprising: processor; as well as A memory device operatively coupled to the processor, wherein the processor and the memory device are configured to: Obtain first camera parameters associated with at least one reference view and second camera parameters associated with the current view of the multi-view video content, wherein each view includes a texture layer and a depth layer; The configuration is to apply a forward projection method to the pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameters to the camera coordinate system of the current view defined by the second camera parameters to generate an intermediate prediction image, wherein each projected pixel of the intermediate prediction image is associated with a motion information value, the motion information value representing the displacement between the projected pixel and the pixel of the reference view from which it is projected; At least one final predicted image obtained from at least one intermediate predicted image is stored in the decoded image buffer of the reconstructed image for temporal prediction of the current view; as well as Reconstruct the current image of the current view from the image stored in the decoded image buffer.

10. The device according to claim 9, wherein, The forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by projecting from the camera coordinate system of the reference view to the world coordinate system. The deprojection uses the pose matrix of the reference camera of the reference view, the inverse intrinsic parameter matrix of the reference camera, and the depth value associated with the current pixel. The pose matrix and inverse intrinsic parameter matrix of the reference camera are obtained according to the first camera parameters. The deprojected pixels are projected into the camera coordinate system of the current view using the intrinsic and extrinsic parameter matrices of the current camera to obtain forward-projected pixels. The intrinsic and extrinsic parameter matrices of the current camera are obtained according to the second camera parameters. If the projected pixel does not correspond to a pixel in the current camera's pixel grid, then the pixel in the pixel grid closest to the projected pixel is selected to obtain a corrected projected pixel; and Calculate a motion vector representing the displacement between the forward-projected pixel and the current pixel of the reference view.

11. The device according to claim 9, wherein, The processor and memory are also configured to fill isolated missing pixel values ​​in each intermediate or final predicted image with the average of adjacent pixel values, the median of adjacent pixel values, or a default value.

12. The device according to claim 9, wherein, The at least one final predicted image is obtained by aggregating at least two intermediate predicted images.

13. A device for encoding, comprising: processor; as well as A memory device operatively coupled to the processor, wherein the processor and the memory device are configured to: Obtain first camera parameters associated with at least one reference view and second camera parameters associated with the current view of the multi-view video content, wherein each view includes a texture layer and a depth layer; The configuration is to apply a forward projection method to the pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameters to the camera coordinate system of the current view defined by the second camera parameters to generate an intermediate prediction image, wherein each projected pixel of the intermediate prediction image is associated with a motion information value, the motion information value representing the displacement between the projected pixel and the pixel of the reference view from which it is projected; At least one final predicted image obtained from at least one intermediate predicted image is stored in the decoded image buffer of the reconstructed image for temporal prediction of the current view; as well as Reconstruct the current image of the current view from the image stored in the decoded image buffer.

14. The device according to claim 13, wherein, The forward projection method includes: Deprojection is applied to the current pixel of the reference view to obtain a deprojected pixel by projecting from the camera coordinate system of the reference view to the world coordinate system. The deprojection uses the pose matrix of the reference camera of the reference view, the inverse intrinsic parameter matrix of the reference camera, and the depth value associated with the current pixel. The pose matrix and inverse intrinsic parameter matrix of the reference camera are obtained according to the first camera parameters. The deprojected pixels are projected into the camera coordinate system of the current view using the intrinsic and extrinsic parameter matrices of the current camera to obtain forward-projected pixels. The intrinsic and extrinsic parameter matrices of the current camera are obtained according to the second camera parameters. If the projected pixel does not correspond to a pixel in the current camera's pixel grid, then the pixel in the pixel grid closest to the projected pixel is selected to obtain a corrected projected pixel; and Calculate a motion vector representing the displacement between the forward-projected pixel and the current pixel of the reference view.

15. The device according to claim 13, wherein, The processor and memory are also configured to fill isolated missing pixel values ​​in each intermediate or final predicted image with the average of adjacent pixel values, the median of adjacent pixel values, or a default value.

16. The device according to claim 13, wherein, The at least one final predicted image is obtained by aggregating at least two intermediate predicted images.