Method and apparatus for video decoding, and storage medium

By extracting and decoding key image feature information in multi-view videos, feature-based multi-view coding is realized, solving the problem of inefficient multi-view video encoding in the prior art, and reducing the storage and transmission bandwidth requirements.

CN116261853BActive Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280006494.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-07-11
Filing Date
2022-07-12
Publication Date
2025-05-27
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

When existing video encoding technologies deal with multi-view videos, it is difficult to effectively utilize the similarity between views, resulting in low encoding efficiency and high storage/transmission bandwidth requirements.

Method used

By extracting the key pictures and feature information of each view, decoding and reconstruction of feature changes is performed using the multi-view code stream, and multi-view encoding based on feature is realized.

Benefits of technology

Improves the encoding efficiency of multi-view videos, reduces the bandwidth required for storage and transmission, while maintaining visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116261853B_ABST
    Figure CN116261853B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method, an apparatus, and a non-transitory computer-readable storage medium for video decoding. The apparatus includes processing circuitry configured to decode at least one first key picture of a plurality of pictures from a multi-view bitstream. The plurality of pictures correspond to different views. The at least one first key picture corresponds to at least a first view among the different views. The processing circuitry determines first feature information of content in the at least one first key picture. The processing circuitry decodes a first feature change of the first feature information based on the multi-view bitstream. The first feature change indicates a content change between one key picture and a first picture among the at least one first key picture. The processing circuitry reconstructs the first picture based on the decoded first feature change, the first feature information, and the key picture among the at least one first key picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority benefit of U.S. Patent Application No. 17 / 861,667, filed on July 11, 2022, entitled "FEATURE-BASED MULTI-VIEW REPRESENTATION AND CODING", which claims the priority of U.S. Provisional Application No. 63 / 221,351, filed on July 13, 2021, entitled "Features BasedMulti-View Representation and Delivery". The entire content of the prior application is incorporated herein by reference. Technical Field

[0002] Embodiments of the present application generally relate to video encoding / decoding, and in particular, to a method and apparatus for video decoding, and a storage medium. Background Art

[0003] The background description provided herein is for the purpose of presenting the disclosure generally. Certain work of the inventors (i.e., the work described in this background art section) and the content of the specification that is not prior art as of the filing date of the application, whether explicitly or implicitly, is not considered prior art with respect to the present disclosure.

[0004] Inter-picture prediction with motion compensation can be used to perform image and / or video encoding and decoding. Uncompressed digital images and / or videos can include a series of pictures, each picture having a spatial dimension of, for example, luminance samples of 1920x1080 and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bit rate requirements. For example, a 1080p60 4:2:0 video (1920x1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.

[0005] One purpose of image and / or video encoding and decoding can be to reduce redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, which in some cases can be reduced by two orders of magnitude or more than two orders of magnitude. Although the description herein uses video encoding / decoding as an illustrative example, the same techniques can be applied in a similar manner to image encoding / decoding without departing from the spirit of the present disclosure. Lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. Taking video as an example, lossy compression is widely used. The amount of tolerable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher admissible / acceptable distortion can result in a higher compression ratio.

[0006] Video encoders and video decoders can utilize a variety of broad categories of techniques, for example, including: motion compensation, transformation, quantization, and entropy coding.

[0007] Video codec techniques can include techniques referred to as intra-frame coding. In intra-frame coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are coded in intra-frame mode, the picture can be an intra-frame picture. Intra-frame pictures and their derived forms (such as independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and video session, or as a still image. The samples of an intra-frame block can be transformed, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are required to represent the entropy-coded block for a given quantization step size.

[0008] As known from, for example, MPEG-2 generation encoding techniques, traditional intra-frame coding does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to use, for example, surrounding sample data and / or metadata that are obtained during the encoding and / or decoding of data blocks that are spatially adjacent and prior in decoding order. Such techniques are hereafter referred to as "intra-frame prediction" techniques. It should be noted that, at least in some cases, intra-frame prediction only uses reference data from the currently being reconstructed picture and does not use reference data from reference pictures.

[0009] Intra prediction can have many different forms. When more than one such technique can be used in a given video coding technique, the technique in use can be coded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, and these sub - modes and / or parameters can be coded separately or included in the mode codeword. Which codeword is used for a given combination of mode, sub - mode, and / or parameter can affect the coding efficiency gain through intra prediction, and the entropy coding technique used to convert the codeword into a bitstream can also have an impact on it.

[0010] H.264 introduced certain intra prediction modes, which were improved in H.265 and further improved in new coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using the values of adjacent samples that belong to already available samples. The sample values of the adjacent samples are copied into the predictor block according to the direction. The reference to the direction used can be coded in the bitstream, or it can be predicted itself.

[0011] Reference Figure 1A , depicted in the lower - right is a subset of 9 predictor directions known from 33 possible predictor directions in H.265 (corresponding to 33 angular modes out of 35 intra modes). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction of the sample being predicted. For example, arrow (102) represents predicting sample (101) from one or more samples in the upper - right, at a 45 - degree angle to the horizontal line. Similarly, arrow (103) represents predicting sample (101) from one or more samples in the lower - left of sample (101), at a 22.5 - degree angle to the horizontal line.

[0012] Still referring Figure 1A , depicted in the upper - left is a square block (104) of 4×4 samples (represented by a bold dashed line). The square block (104) contains 16 samples, and each sample is labeled using "S" and its position in the Y - dimension (e.g., row index) and its position in the X - dimension (e.g., column index). For example, sample S21 is the second sample in the Y - dimension (starting from the top) and the first sample in the X - dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y - dimension and the X - dimension within block (104). Since the block is 4×4 samples, S44 is in the lower - right corner. Figure 1AReference samples are also shown therein, which follow a similar numbering scheme. The reference samples are labeled with an R and their Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus, negative values need not be used.

[0013] Intra picture prediction can work by copying the reference sample values from adjacent samples occupied by the signal-informed prediction direction. For example, assume that the encoded video bitstream includes signaling that indicates, for the block, a prediction direction consistent with arrow (102), that is, the samples are predicted from one or more prediction samples in the upper right corner at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted according to the same reference sample R05. Then, sample S44 is predicted according to reference sample R08.

[0014] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, in order to calculate a reference sample, especially when the direction is not divisible by 45 degrees.

[0015] With the development of video coding technology, the number of possible directions has increased. In H.264 (in 2003), nine different directions could be represented. This number increased to 33 in H.265 (in 2013), and at the time of the present disclosure, up to 65 directions are supported in JEM / VVC / BMS. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those possible directions with a small number of bits, accepting a certain cost for the less likely directions. Additionally, sometimes the direction itself can be predicted based on the neighboring directions used in the already decoded neighboring blocks.

[0016] Figure 1B is a schematic diagram (110) showing 65 intra prediction directions according to JEM, thus showing the increasing number of prediction directions over time.

[0017] The mapping of the intra prediction direction bits representing the direction in the encoded video bitstream can vary with different video coding technologies, and, for example, it can range from a simple direct mapping of the prediction direction to the intra prediction mode to a codeword, to a complex adaptive scheme involving the most likely modes and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in the video content compared to certain other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technology, those less likely directions will be represented by a larger number of bits compared to the likely directions.

[0018] Motion compensation can be a lossy compression technique and can involve the following techniques: Blocks of sample data from a previously reconstructed picture or a portion thereof (reference picture) are spatially offset in a direction indicated by a motion vector (hereinafter referred to as MV) and are then used to predict a newly reconstructed picture or picture portion. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture being used (which indirectly can be a temporal dimension).

[0019] In some video compression techniques, an MV applicable to a certain region of sample data can be predicted based on other MVs, for example, based on another MV related to another region of sample data that is spatially adjacent to the region being reconstructed and whose decoding order is prior to that of the MV. Doing so can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing the compression ratio. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is the following statistical possibility: A region larger than the region applicable to a single MV moves in a similar direction. Therefore, in some cases, a similar motion vector derived from the MVs of adjacent regions can be used to predict the larger region. This makes the MV found for a given region similar or identical to the MV predicted based on the surrounding MVs. Subsequently, after entropy coding, the MV found for the given region can be represented using fewer bits than the number of bits used when directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms provided by H.265, the technique described in this application is what is hereinafter referred to as "spatial merge".

[0021] Reference Figure 2, the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on a previous block of the same size that has generated a spatial offset. Additionally, the MV can be derived from metadata associated with one or more reference pictures instead of directly encoding the MV. For example, using the MV associated with any one of five surrounding samples A0, A1, and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV is derived from the metadata of the nearest reference picture (in decoding order). In H.265, MV prediction can use the prediction values of the same reference picture that adjacent blocks are also using. SUMMARY OF THE INVENTION

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding and decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry is configured to decode at least one first key picture among a plurality of pictures from a multi-view bitstream. The plurality of pictures correspond to different views. The at least one first key picture corresponds to at least one first view among the different views. First feature information of the content in the at least one first key picture can be determined. A first feature change of the first feature information can be decoded based on the multi-view bitstream. The first feature change can indicate a content change between one key picture and a first picture among the at least one first key picture. The first picture can be reconstructed based on the decoded first feature change, the first feature information, and the key picture among the at least one first key picture.

[0023] In one embodiment, the at least one first key picture corresponds to a first time instance. The at least one first key picture includes a plurality of first key pictures. At least one first view among different views includes a plurality of first views. The first feature information includes first 3D feature information indicated by the plurality of first views. The first 3D feature information at the first time instance can be determined based on a first predetermined 3D feature model and the plurality of first key pictures.

[0024] In one example, the first 3D feature information is used to decode pictures of each view among different views.

[0025] In one example, the plurality of first key pictures include each key picture at the first time instance.

[0026] In one example, second 3D feature information of the content in a plurality of second key pictures among the plurality of pictures at the first time instance is determined based on a second predetermined 3D feature model. The plurality of second key pictures correspond to the first time instance of a plurality of second views among different views.

[0027] In one example, the first feature information is associated with a first view among at least one first view. For each view that is not the first view among different views, the processing circuit may determine corresponding feature information based on the key picture of each view and another key picture of an adjacent view in the different views. The processing circuit may decode the feature change of the corresponding feature information based on the multi-view bitstream. The feature change corresponds to the corresponding picture of each view. Pictures of each view may be generated based on the corresponding feature change, the corresponding feature information, and the key picture of each view.

[0028] In one example, a first picture of one of the different views is at a second time instance.

[0029] In one example, the processing circuit may decode a subset of multiple pictures corresponding to different views. The subset of multiple pictures belongs to a first view among at least one first view corresponding to a respective time instance. The subset of pictures of the first view may include a key picture among at least one first key picture. The first picture belongs to a second view among different views. The first picture and the key picture among the at least one first key picture correspond to a first time instance. The first feature change indicates a feature change between the first picture of the second view at the first time instance and the key picture of the first view at the first time instance.

[0030] In one example, each picture of a first view among at least one first view is decoded as a key picture. Each picture of the first view corresponds to a respective time instance.

[0031] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for video decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0033] Figure 1A is a schematic diagram of an example subset of intra prediction modes.

[0034] Figure 1B is an illustration of an exemplary intra prediction direction.

[0035] Figure 2 shows a current block (201) and surrounding samples according to one embodiment.

[0036] Figure 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment.

[0037] Figure 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to an embodiment.

[0038] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.

[0039] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.

[0040] Figure 7 shows a block diagram of an encoder according to another embodiment.

[0041] Figure 8 shows a block diagram of a decoder according to another embodiment.

[0042] Figure 9A shows an exemplary arrangement of cameras in a multi-camera acquisition system according to an embodiment of the present disclosure.

[0043] Figure 9B shows an exemplary 1D parallel arrangement of cameras in a multi-camera acquisition system according to an embodiment of the present disclosure.

[0044] Figure 10 shows an example of spatial stitching according to an embodiment of the present disclosure.

[0045] Figure 11A shows an exemplary schematic diagram of feature-based video coding or a feature-based video coding process.

[0046] Figure 11B shows an exemplary schematic diagram of feature-based video coding or a feature-based video coding process.

[0047] Figure 12 shows an example of a picture in a bitstream according to an embodiment of the present disclosure.

[0048] Figure 13 shows exemplary key pictures of different views in a bitstream for determining feature information of key pictures according to an embodiment of the present disclosure.

[0049] Figure 14 shows an example subset of key pictures of different views in a bitstream for determining multiple 3D feature information according to an embodiment of the present disclosure.

[0050] Figure 15 shows an example subset of key pictures of different views in a bitstream for determining multiple feature information according to an embodiment of the present disclosure.

[0051] Figure 16Shows an exemplary encoding method for feature-based multi-view coding according to an embodiment of the present disclosure.

[0052] Figure 17 Shows an exemplary encoding method for feature-based multi-view coding according to an embodiment of the present disclosure.

[0053] Figure 18 Shows an exemplary encoding method for feature-based multi-view coding according to an embodiment of the present disclosure.

[0054] Figure 19 Shows an example of pictures of different views at different time instances according to an embodiment of the present disclosure.

[0055] Figure 20 Shows a flowchart outlining the encoding process according to an embodiment of the present disclosure.

[0056] Figure 21 Shows a flowchart outlining the decoding process according to an embodiment of the present disclosure.

[0057] Figure 22 Is a schematic diagram of a computer system according to an embodiment. Detailed Description

[0058] Figure 3 Shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected by a network (350). In Figure 3 the example, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can encode video data (such as a video picture stream captured by the terminal device (310)) for transmission through the network (350) to another terminal device (320). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video data, and display video pictures based on the recovered video data. Unidirectional data transmission is more common in applications such as media services.

[0059] In another example, a communication system (300) includes pairs of terminal devices (330) and (340) that perform two-way transmission of encoded video data, which may occur, for example, during a video conference. For two-way data transmission, each of the terminal devices (330) and terminal device (340) may encode video data (such as a video picture stream captured by the terminal device) for transmission over a network (350) to the other of the terminal devices (330) and terminal device (340). Each of the terminal devices (330) and terminal device (340) may also receive the encoded video data transmitted by the other of the terminal devices (330) and terminal device (340), may decode the encoded video data to recover the video pictures, and may display the video pictures on an accessible display device based on the recovered video data.

[0060] In Figure 3 the example, the terminal devices (310), terminal device (320), terminal devices (330) and terminal device (340) may be shown as servers, personal computers, and smart phones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that transfer encoded video data between the terminal devices (310), terminal device (320), terminal devices (330) and terminal device (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network (350) may be immaterial to the operations disclosed in this application.

[0061] As an example of an application of the disclosed subject matter, Figure 4 shows the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0062] A streaming system may include an acquisition subsystem (413), which may include a video source (401) such as a digital camera that creates an uncompressed video picture stream (402). In an embodiment, the video picture stream (402) includes samples taken by the digital camera. The video picture stream (402) is depicted as a thick line to emphasize the high data volume compared to the encoded video data (404) (or encoded video bitstream). The video picture stream (302) may be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume of the encoded video data (304) (or encoded video bitstream (304)) compared to the video picture stream (402), and it may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4 the client subsystem (406) and the client subsystem (408) in can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an output video picture stream (411) that can be presented on a display (412) (such as a display screen) or another rendering device (not depicted). In some streaming systems, the encoded video data (404), the video data (407), and the video data (409) (such as a video bitstream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, a video coding standard under development, informally referred to as Versatile Video Coding (VVC), the disclosed subject matter may be used in the context of VVC.

[0063] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0064] Figure 5is a block diagram of a video decoder (510) according to an embodiment disclosed in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 the video decoder (410) in the embodiment.

[0065] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (415) may be provided outside the video decoder (510) (not labeled). And in other cases, a buffer memory (not labeled) is provided outside the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) may be configured inside the video decoder (510) to, for example, handle the playback timing. And when the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (515), or the buffer memory may be made smaller. For use on a service packet network such as the Internet, a buffer memory (515) may also be required, which may be relatively large and may have an adaptive size, and may be implemented at least partially in the operating system or a similar element (not labeled) outside the video decoder (510).

[0066] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a rendering device such as a rendering device (512) (e.g., a display screen), which is not an integral part of the electronic device (530), but may be coupled to the electronic device (530), as Figure 5As shown. The control information for the display device can be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be based on video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0067] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0068] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) can involve multiple different units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and multiple units below are not described.

[0069] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.

[0070] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients as symbols (521) and control information from the parser (520), including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values, and the sample values can be input into the aggregator (555).

[0071] In some cases, the output samples of the scaler / inverse transform unit (551) can belong to intra-coded blocks; that is: blocks that do not use predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates surrounding blocks with the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (558). For example, the current picture buffer (558) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (555) adds the predictive information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.

[0072] In other cases, the output samples of the scaler / inverse transform unit (551) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (553) can access the reference picture memory (557) to extract samples for prediction. After motion compensation of the extracted samples according to the symbol (521), these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (which is called the residual sample or residual signal in this case), so as to generate output sample information. The motion compensation prediction unit (553) obtaining the prediction samples from the address in the reference picture memory (557) can be controlled by a motion vector, and the motion vector is in the form of the symbol (521) for use by the motion compensation prediction unit (553), and the symbol (421) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.

[0073] The output samples of the aggregator (555) can be adopted by various loop filtering techniques in the loop filter unit (556). Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters can be used in the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques can also respond to meta-information obtained during decoding of the previous (in decoding order) part of the encoded picture or the encoded video sequence, and respond to previously reconstructed and loop-filtered sample values.

[0074] The output of the loop filter unit (556) can be a sample stream, which can be output to the display device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.

[0075] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0076] The video decoder (510) can perform decoding operations according to predetermined video compression techniques, such as those in the ITU-T H.265 standard. In the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under that profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0077] In an embodiment, a receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by a video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, and the like.

[0078] Figure 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is disposed in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the example of

[0079] The video encoder (603) may receive video samples from a video source (601) (which is not Figure 6 part of the electronic device (620) in the example), and the video source may capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) is part of the electronic device (620).

[0080] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, and each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0081] According to an embodiment, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figures. Parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used for other suitable functions that relate to the video encoder (603) optimized for a certain system design.

[0082] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0083] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 5 the video decoder (510). Additionally, briefly referring to Figure 5 , however, when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0084] In one embodiment, any decoder technology other than parsing / entropy decoding that exists in the decoder must also exist in the corresponding encoder in substantially the same functional form. Accordingly, this application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is inverse to the decoder technology described comprehensively. In some aspects, more detailed descriptions will be provided below.

[0085] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures from a video sequence, which are designated as "reference pictures". In this way, the encoding engine (632) encodes the difference between a pixel block of the input picture and a pixel block of the reference picture, which can be selected as the prediction reference for the input picture.

[0086] The local video decoder (633) may decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data can be decoded on a video decoder ( Figure 6 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has the same content (without transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0087] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor (635) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (635), it can be determined that the input picture may have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (634).

[0088] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0089] The outputs of all the above functional units can be entropy encoded in an entropy encoder (645). The entropy encoder (645) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbol into an encoded video sequence.

[0090] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over a communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0091] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types:

[0092] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0093] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0094] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0095] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., each sample is a 4x4, 8x8, 4x8, or 16x16 block) and encoded on a block-by-block basis. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and that other block is determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the block can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be prediction-encoded with reference to a previously encoded reference picture by spatial prediction or by temporal prediction. Blocks of a B picture can be prediction-encoded with reference to one or two previously encoded reference pictures by spatial prediction or by temporal prediction.

[0096] The video encoder (603) can perform encoding operations according to a predetermined video encoding technique or standard such as, for example, the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including prediction encoding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.

[0097] In an embodiment, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0098] The captured video can be multiple source pictures (video pictures) in a time series. Intra picture prediction (usually abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension that identifies the reference picture.

[0099] In some embodiments, bidirectional prediction techniques can be used for inter - picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which are before the current picture in the video in the decoding order (but can be in the past and future respectively in the display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0100] In addition, the merge mode technique can be used for inter - picture prediction to improve the coding efficiency.

[0101] According to some embodiments disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, picture 0 in a video picture sequence is divided into coding tree units (CTUs) for compression. The CTUs in picture 0 have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter - prediction type or an intra - prediction type. In addition, depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction - block basis. Taking the luminance prediction block as an example of the prediction block, the prediction block includes a matrix of values (e.g., luminance values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0102] Figure 7 is a diagram of a video encoder (703) according to another embodiment disclosed in the present application. The video encoder (703) is used to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used to replace Figure 4The video encoder (403) in the embodiment.

[0103] In HEVC embodiments, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best encoded using an intra mode, an inter mode, or a dual prediction mode such as rate distortion optimization. When encoding the processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when encoding the processing block in the inter mode or the bi - prediction mode, the video encoder (703) may use inter prediction or bi - prediction techniques respectively to encode the processing block into an encoded picture. In some video coding techniques, the merge mode may be an inter - picture prediction sub - mode, where a motion vector is derived from one or more motion vector prediction values without resorting to encoded motion vector components external to the prediction values. In some other video coding techniques, there may be motion vector components applicable to the subject block. In an embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0104] In Figure 7 the embodiment of, the video encoder (703) includes Figure 7 an inter - frame encoder (730), an intra - frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown.

[0105] The inter - frame encoder (730) is configured to receive samples of a current block (e.g., the processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture), generate inter - frame prediction information (e.g., redundancy information description, motion vector, merge mode information according to inter - frame coding techniques), and calculate an inter - frame prediction result (e.g., the predicted block) based on the inter - frame prediction information using any suitable technique. In some embodiments, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0106] The intra - frame encoder (722) is configured to receive samples of the current block (e.g., the processing block), compare the block with blocks that have been encoded in the same picture in some cases, generate quantization coefficients after transformation, and also generate intra - frame prediction information (e.g., intra - frame prediction direction information according to one or more intra - frame coding techniques) in some cases. In an embodiment, the intra - frame encoder (722) also calculates an intra - frame prediction result (e.g., the predicted block) based on the intra - frame prediction information and reference blocks in the same picture.

[0107] The general controller (721) is used to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is the intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra prediction information and add the intra prediction information to the bitstream; and when the mode is the inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.

[0108] The residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is used to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is used to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and in some examples, the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0109] The entropy encoder (725) is used to format the bitstream to produce an encoded block. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, the selected prediction information (such as intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. It should be noted that according to the disclosed subject matter, there is no residual information when encoding a block in the merge submode of the inter mode or the bi - directional prediction mode.

[0110] Figure 8FIG. is for a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an embodiment, the video decoder (810) is configured to replace Figure 4 the video decoder (410) in the embodiment.

[0111] In Figure 8 the example of, the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together as shown in Figure 7 .

[0112] The entropy decoder (871) can be used to reconstruct certain symbols according to the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode used to encode the block (e.g., intra-frame mode, inter-frame mode, bi-prediction mode, merge sub-mode of the latter two, or another sub-mode), prediction information that can respectively identify certain samples or metadata for the intra-frame decoder (872) or the inter-frame decoder (880) to perform prediction (e.g., intra-frame prediction information or inter-frame prediction information), residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is an inter-frame or bi-prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information can be de-quantized and provided to the residual decoder (873).

[0113] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0114] The intra-frame decoder (872) is configured to receive the intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.

[0115] The residual decoder (873) is configured to perform de-quantization to extract the de-quantized transform coefficients and process the de-quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including the quantizer parameter QP), and this information can be provided by the entropy decoder (871) (the data path is not shown because this may be just low-data-volume control information).

[0116] The reconstruction module (874) is configured to combine, in a spatial domain, the residual output by the residual decoder (873) with a prediction result (which may be output by an inter-frame prediction module or an intra-frame prediction module) to form a reconstructed block, which may be part of a reconstructed picture, which in turn may be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations may be performed to improve the visual quality.

[0117] It should be noted that any suitable technology may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In an embodiment, one or more integrated circuits may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810).

[0118] According to embodiments of the present disclosure, a bitstream may include one or more coded video sequences (CVSs). A CVS may be encoded independently of other CVSs. Each CVS may include one or more layers, and each layer may be a representation of a video with a specific quality (e.g., spatial resolution), or a representation of a certain component interpretation attribute, e.g., as a depth map, a transparency map, or a perspective view. In the temporal dimension, each CVS may include one or more access units (AUs). Each AU may include one or more pictures corresponding to different layers at the same time instance. A coded layer video sequence (CLVS) is a layered CVS that may include a sequence of picture units in the same layer. If the bitstream has multiple layers, the CVSs in the bitstream may have one or more CLVSs for each layer.

[0119] In an embodiment, the CVS includes an AU sequence, where the AU sequence includes, in decoding order, an intra-random access point (IRAP) AU, followed by zero or more non-IRAP AUs. In one example, the zero or more AUs include all subsequent AUs until, but not including, any subsequent AU that is an IRAP AU. In one example, the CLVS includes picture sequences and associated non-video coding layer (VCL) network abstraction layer (NAL) units of the base layer of the CVS.

[0120] According to some aspects of the present disclosure, videos can be classified into single-view videos and multi-view videos. For example, a single-view video (e.g., a single-perspective video) is a two-dimensional medium that provides a single view of a scene to a viewer. A multi-view video can provide multiple viewpoints of a scene and can provide a sense of realism to the viewer. In an example, a 3D video can provide two views, such as a left view and a right view corresponding to a human viewer. The two views can be displayed (presented) simultaneously or almost simultaneously using different light polarizations, and the viewer can wear polarized glasses so that each eye of the viewer receives a corresponding one of the views.

[0121] The present disclosure includes embodiments related to efficient encoding and representation of multiple views. The present disclosure includes feature-based multi-view representation and transmission. In one embodiment, the content of each view (or each picture) can be extracted and represented using feature information indicating features and / or key points. The features of different views at the same time or the same time instance can be prioritized to achieve scalability of view access.

[0122] Multiple views can be used in video acquisition and encoding. To enrich the user's visual experience, for example, the scene of interest can be acquired using multiple cameras from different positions, as Figure 9A and Figure 9B shown.

[0123] Figure 9A Shows an exemplary arched arrangement of cameras in a multi-camera acquisition system according to an embodiment of the present disclosure. The cameras (e.g., indicated by i-1, i, i+1) are arranged around a one-dimensional (1D) arch. The distances between the multiple cameras and the scene of interest can be the same or different.

[0124] Figure 9BShows an exemplary 1D parallel arrangement of cameras in a multi-camera acquisition system according to an embodiment of the present disclosure. A plurality of cameras (e.g., 1-3) are arranged along a 1D axis (e.g., the camera horizontal axis). In one example, there is a distance (e.g., camera parallax) between adjacent cameras. The distance between the plurality of cameras and the scene of interest can be the same or different. The symbols v1, v2, and v3 represent the views corresponding to cameras 1, 2, and 3 respectively. The panoramic view Vpan can include views v1 - v3.

[0125] Applications of the plurality of cameras can include VR video or VR360, free viewpoint (e.g., freeviewpoint television (FTV)), light field video, etc. VR video can also be referred to as 360 VR or VR360. VR360 can refer to a video captured using an omnidirectional camera. An omnidirectional camera can capture 360 degrees or a part thereof simultaneously. In VR video, a user can look around the entire scene. Compared with conventional video, VR video can provide a more immersive and interactive experience.

[0126] FTV can include a system for viewing natural videos, allowing users to interactively control the viewpoint and generate new views of a dynamic scene according to 3D positions. With FTV, the focus of attention can be controlled by the viewer rather than the director, and each viewer can observe a unique perspective.

[0127] Light field video can be captured by a light field camera or a plenoptic camera. Some cameras only record the light intensity of the scene. A light field camera or a plenoptic camera can record the light field. Light field video can include information about the light field emitted from the scene, such as the light intensity in the scene, and the direction in which the light rays propagate in space. Light field video can include the intensity of light in the scene, and the direction in which the light rays propagate in space.

[0128] As Figure 9A and Figure 9B shown, the multi-camera array can capture VR360, FTV, light field video, etc.

[0129] As Figure 9A and Figure 9B described, a multi-view video can be created by simultaneously capturing a scene using a plurality of cameras, where the plurality of cameras are appropriately positioned such that each camera captures the scene from its respective viewpoint. The plurality of cameras can capture a plurality of video sequences corresponding to a plurality of viewpoints. To provide more views, more cameras can be used to generate a multi-view video with a large number of video sequences associated with the views. Multi-view video may require a large storage space for storage and / or a high bandwidth for transmission. Multi-view video coding techniques have been developed in this field to reduce the required storage space or transmission bandwidth.

[0130] To improve the efficiency of multi-view video coding, the similarity between views can be utilized. In some embodiments, one view among the multiple views, called the base view, is encoded like a monoscopic video. For example, during the encoding of the base view, intra (picture) and / or temporal inter (picture) prediction is used. A monoscopic decoder that performs intra (picture) prediction and inter (picture) prediction can be used to decode the base view. The other views in the multi-view video, besides the base view, can be called dependent views. To encode the dependent views, in addition to intra (picture) prediction and inter (picture) prediction, inter-view prediction with disparity compensation can also be used. In an example, in inter-view prediction, a reference block of samples from a picture of another view at the same time instance is used to predict the current block in the dependent view. The position of the reference block is indicated by a disparity vector. Inter-view prediction is similar to inter (picture) prediction, except that the motion vector is replaced by a disparity vector and the temporal reference picture is replaced by a reference picture from another view.

[0131] According to some aspects of the present disclosure, multi-view coding can adopt a multi-layer approach. The multi-layer approach can multiplex different encodings (e.g., HEVC encoding) of a video sequence, referred to as layers, into one bitstream. These layers can be dependent on each other. Inter-layer prediction can use the dependencies to improve the compression performance by exploiting the similarity between different layers. The layers can represent the texture, depth, or other auxiliary information of a scene related to a specific camera perspective. In some examples, all layers belonging to the same camera perspective are represented as views; and the layers carrying the same type of information (e.g., texture or depth) are called components within the multi-view video scope.

[0132] According to one aspect of the present disclosure, multi-view video coding can include the addition of a high-level syntax (HLS) (e.g., above the slice level) with an existing single-layer decoding kernel. In some examples, multi-view video coding does not change the syntax or decoding process required for single-layer encoding (e.g., HEVC) below the slice level. Reusing the existing implementation to build a multi-view video decoder can be allowed without significant changes. For example, a multi-view video decoder can be implemented based on video decoder (510) or video decoder (810).

[0133] In some examples, all pictures associated with the same capture or display time instance are included in an AU and have the same picture order count (POC). Multiview video coding may allow for inter-view prediction from pictures within the same AU. For example, decoded pictures from other views may be inserted into one or both of the reference picture lists for the current picture. Additionally, in some examples, when related to temporal reference pictures of the same view, the motion vectors may be actual temporal motion vectors, or when related to inter-view reference pictures, the motion vectors may be disparity vectors. A block-level motion compensation module (e.g., block-level encoding software or hardware, block-level decoding software or hardware) may be used, which operates in the same manner regardless of whether the motion vectors are temporal motion vectors or disparity vectors.

[0134] After being captured, the information in the multiviews may be processed, compressed, transmitted to a client, and / or stored. The video of each view may be treated as 2D video (e.g., single-viewpoint video), and the above video / image coding techniques (e.g., intra (picture) and / or inter-frame (picture) prediction), such as HEVC, VVC, etc., may be used to efficiently encode (e.g., compress). Certain applications (e.g., the above VR360, FTV, light field video) may impose relatively high bandwidth requirements due to the large number of views in the video. Various techniques may be used to alleviate the bandwidth burden.

[0135] As described above, the inter-view dependencies between different views may be exploited. A subset of views of all views may be encoded first. The content of the encoded view subset may be used as a reference for other views to be encoded, such as in inter-view prediction with disparity compensation. Compared to encoding other views independently, other views may be compressed more efficiently.

[0136] In one embodiment, a subset of views of all views may be selected for encoding / compression. Other views (referred to as non-encoded views) are not provided (e.g., sent) to reduce bandwidth requirements and do not require encoding / compression. At the client, the received bitstream only contains a portion (e.g., a subset of views) of all the captured views. If the content of a non-encoded view or other intermediate virtual (non-existent) views is to be accessed (e.g., consumed), the non-encoded view and / or the non-existent view may be rendered by using information from adjacent views. The non-encoded views may be captured by one or more cameras but are not encoded and not transmitted. The non-existent views are views that the cameras did not capture. In an example, depth information associated with the views (e.g., the distance between the scene of interest and the camera recording the scene of interest) is used for intermediate view rendering.

[0137] In one embodiment, spatial stitching may be performed, where selected views to be encoded are stitched together to form a larger video. Figure 10 An example of spatial stitching according to an embodiment of the present disclosure is shown. For example, six views to be encoded (e.g., views 0 - 5) may be stitched together spatially, for example, in a 3×2 setup, to form a video (1000). The resolution of the larger video (1000) is three times that of each of views 0 - 5 in the horizontal direction and two times that of each of views 0 - 5 in the vertical direction. A single video (1000) may be encoded (e.g., encoded and / or decoded) using a related video encoding method for 2D video (e.g., monoscopic video), which video encoding method may include, for example, the above-mentioned video / image encoding techniques (e.g., intra (picture) and / or temporal inter (picture) prediction), such as HEVC, VVC, etc. After the video (1000) is decoded, individual views (e.g., views 0 - 5) may be extracted from the larger video (1000).

[0138] In one embodiment, temporal stacking may be performed. One or more pictures may be used as one or more base views for only intra prediction, while other pictures may be predicted inter - reference to one or more encoded pictures, and pictures of different views corresponding to the same time instance are encoded sequentially. After encoding the views of the first time instance, a similar operation may be applied to the pictures of the views of different time instances. For example, all pictures up to a certain time instance are encoded sequentially before pictures of another time instance are processed. The above method may be referred to as a "time - first" encoding method. In one example, all pictures of the same time instance are processed before encoding pictures of another time instance.

[0139] In various examples, encoding multiple views using, for example, the above - mentioned encoding methods may be challenging due to the total bandwidth consumption. In some applications, a user can view only a small portion of all perspectives or viewpoints at a time. What a user can see in the reconstructed scene can be defined as a viewport. In one embodiment, a user may define a viewport. The viewport may have any suitable shape, such as a rectangular shape. The current viewport may be changed to another viewport, and the user can change the view. If the user does not switch to another viewport, data that is not relevant to the reconstruction of the current viewport may not be transmitted. In one example, the server only needs to transmit a portion of the multi - view video data corresponding to the current viewport. The server may not need to transmit data that does not correspond to the current viewport.

[0140] According to embodiments of the present disclosure, for example, when there is no significant change in the content between different pictures, feature-based video coding (or a feature-based video coding process) can be applied to selected applications. In some selected applications, such as in a video conferencing scenario, the content between different pictures (such as a face, a person's shoulders, and the background) does not change significantly. For example, when to some extent, the smoothness of playback and / or the subjective quality of the video are more important than high-fidelity with respect to the original content, feature-based video coding can be applied to selected applications. For example, in some video conferencing scenarios, the smoothness of playback and / or the subjective quality of the video are more important than the high-fidelity of the original content.

[0141] In the present disclosure, a key picture may refer to, for example, a picture whose samples (or pixels) are encoded (such as encoded and decoded) using video / image coding techniques such as HEVC and / or VVC, and the coding techniques include Figure 1A , Figure 1B and Figures 2 to 8 the intra (picture) and / or inter-temporal (picture) prediction as described.

[0142] A non-key picture refers to a picture whose samples are not directly encoded. For a non-key picture, the feature information or feature changes of the non-key picture relative to another picture (such as a key picture) can be encoded (such as encoded and decoded).

[0143] In a feature-based video coding process, feature information of the picture content can be determined (such as extracted) from the picture. The feature information can indicate features of the picture content, key points of the picture content, etc. The process of determining (such as extracting) the feature information (such as features and / or key points) can be referred to as a feature extraction process. The content of the picture (or the features of the content) can be represented or rendered by the feature information, for example, including the extracted features and / or key points of the picture. In one embodiment, the feature information indicates the features and / or key points that may change in the picture. For example, when there is no significant change in the content between different pictures of a view, the feature information may not include information about the parts of the picture that do not change significantly, so as to improve the coding efficiency without sacrificing the visual quality.

[0144] For example, in a video conferencing application, the picture includes a person (such as having a face and an upper body) and a background. For example, when the upper body does not change significantly, the feature information can represent the content related to the face and does not represent the content related to the background and the upper body.

[0145] A human face includes components common to different people in different pictures / videos, such as eyes, nose, mouth, chin, ears, etc. Differences in the shape, size, and structure of the components can distinguish one face from another. Feature information can include features such as the shape and size of the components and / or the arrangement of the components (e.g., the relative distance between two components or the position of the components on the human face). A picture can be determined based on the feature information of the human face and additional information (e.g., the background in the picture). The additional information can be determined based on another decoded picture (e.g., a decoded key picture).

[0146] Key points can refer to key positions in a picture and can be used to represent the structure of the content in the picture (e.g., a human face or components in a human face). In some examples, the features of components can be determined based on the key points of a face. In an example, in a 9-key point model of a human face, the key points describing the human face include two key points indicating the positions of two eyeballs, four key points indicating the near and far corners of two eyes, one key point indicating the midpoint of the nostrils, and two key points indicating the two corners of the mouth. For example, the 9-key point model can be adjusted to include additional key points (e.g., key points describing the positions of ears, eyebrows) to more accurately describe the features in the human face. In some examples, for example, a neural network is used to learn or otherwise determine a model including multiple key points and the positions of the key points.

[0147] In one example, such as in a video conference, the pictures in the video include a human face that changes from one picture to another and a background that remains relatively constant. Feature information can include key points, for example, including the key points in a 9-key point model, and / or features associated with the corresponding key points. For example, it includes two key points indicating the two corners of the mouth, and features indicating the shape / structure of the mouth, such as an open mouth, a closed mouth, can be extracted from the picture. In one example, only the key points are extracted. The features associated with certain key points can be determined based on (i) the key points and (ii) different pictures (e.g., key pictures) or the feature information of the different pictures.

[0148] The feature differences or feature changes in the current picture can indicate the changes or differences between the feature information of the current picture and the feature information of another picture (e.g., a key picture). Feature changes can involve changes in the position of feature information, changes in direction (e.g., in 3-D space), changes in feature size, etc. For example, the coordinates of one or more key points in a 9-key point model can change. The feature information of the current picture can be determined based on the feature changes in the current picture and the feature information of other pictures (e.g., key pictures).

[0149] Figure 11AShows a schematic diagram of a feature-based video coding or a feature-based video coding process (1100A). In an example, the video includes pictures of views. The feature-based video coding process (1100A) can be applied to encode pictures of a single view. The pictures can include key pictures (e.g., at time instance T0) and non-key pictures (e.g., original pictures at time instances T1 - T3). The video data includes data of key pictures and data of non-key pictures. On the encoder side (upper part), some parts of the video data (e.g., data of key pictures) can be compressed (e.g., encoded) using methods for single-view video, such as the video / image coding techniques described above (e.g., intra (picture) and / or inter-temporal (picture) prediction) (e.g., image / video coding techniques of HEVC and / or VVC). The remaining data (e.g., data of non-key pictures at T1 - T3) can be represented by the corresponding feature information at T1 - T3 respectively. The feature information at T1 - T3 can be encoded. In one example, the samples or pixels of the non-key pictures at T1 - T3 are not encoded using methods for single-view video, such as the video / image coding techniques described above (e.g., intra (picture) and / or inter-temporal (picture) prediction) (e.g., image / video coding techniques of HEVC and / or VVC).

[0150] Referring Figure 11A , the key picture at T0 is encoded as described above. Feature extraction processing can be performed on the key picture at T0 to determine the feature information of the key picture at T0. The feature information of the remaining part of the video data (e.g., data of non-key pictures at T1 - T3) can be determined. In an example, feature extraction processing can be performed on the non-key pictures at T1 - T3 relative to the key picture. For the non-key pictures at T1 - T3, the content of the original pictures at T1 - T3 can be represented or rendered by respectively using the feature information including the extraction of features and / or key points, and corresponding adjustments are made. For example, the feature information at T1, T2, and T3 corresponding to the original pictures at T1, T2, and T3 can be encoded into the bitstream. In one example, the encoded bitstream includes the encoded key picture at T0 and the encoded feature information at T1 - T3.

[0151] On the decoder side (lower part), the encoded key picture data at T0 can be decoded using methods for single-view video, such as the video / image coding techniques described above (e.g., intra (picture) and / or temporal inter (picture) prediction) (e.g., the image / video coding techniques of HEVC and / or VVC). In one example, the feature information (e.g., features and / or key points) of the key picture at T0 is extracted in the same manner as performed on the encoder side. For additional pictures after the key picture (e.g., pictures at T1 - T3), the encoded feature information indicating features and / or key points at T1 - T3 can be decoded. When the feature information (e.g., features and / or key points) of a specific picture (e.g., the picture at T1) is ready (e.g., being decoded), the corresponding picture data at the same time instance (e.g., T1) can be restored, reconstructed, or rendered based on the decoded feature information at the time instance (e.g., T1) and the already decoded key picture at T0. In one example, the picture at T1 is reconstructed by combining the decoded feature information at the time instance (e.g., T1) and the decoded key picture at T0.

[0152] As Figure 11B shown, when encoding (e.g., being encoded and / or decoded) feature changes or feature differences, the above description of the feature-based encoding process can be appropriately adapted. Referring to Figure 11B , the feature change of the current picture at a time instance (e.g., the original picture at T2) can indicate the difference between the feature information of another picture (e.g., the key picture or non-key picture at T0) and the feature information of the current picture at the time instance (e.g., T2). The feature change can indicate a change in the position, orientation (e.g., in 3D space), feature size, shape, etc. of features and / or key points from another picture (e.g., the key picture at T0) to the current picture. For example, the feature change includes, for example, a change in the coordinates of one or more key points in a 9-key point model due to a change in the facial expression (e.g., the mouth closing or opening).

[0153] According to an embodiment of the present disclosure, the feature changes of different pictures in a video can refer to the feature changes of different pictures relative to a single picture (e.g., a picture used as a feature reference) such as the key picture at T0. For example, the feature changes of the pictures at T1 - T3 are changes relative to the same key picture at T0.

[0154] According to an embodiment of the present disclosure, the feature change of different pictures in a video may refer to the feature change of different pictures relative to the corresponding pictures used as feature references. The feature references may include multiple pictures and may be different for different pictures. For example, the feature change at a time instance indicates the feature change between the current picture at that time instance and another picture at an adjacent time instance (or a non-adjacent time instance). In one example, the feature change of the picture at T1 is the feature change relative to the picture at T0 (e.g., the key picture), the feature change of the picture at T2 is relative to the picture at T1, and the feature change of the picture at T3 is relative to the picture at T2.

[0155] Figure 11B FIG. shows a schematic diagram of a feature-based video encoding or a feature-based video encoding process (1100B). As Figure 11A described, the video includes pictures of a view. The feature-based video encoding process (1100B) may be applied to encode pictures of a single view. The pictures may include key pictures (e.g., at T0) and non-key pictures (e.g., the original pictures at T1 - T3). The video data includes data of the key pictures and data of the non-key pictures. On the encoder side (upper part), as Figure 11A described, the data of the key picture at T0 may be compressed. The feature information of the key picture at T0 may be extracted from the key picture.

[0156] In Figure 11B the example shown, the feature change of different pictures refers to the feature change of different pictures relative to another picture (e.g., the key picture at T0). The feature change of the original picture at T1 may indicate the change between the key picture and the original picture at T1. In one example, the feature change at T1 is the feature change of the original picture at T1 relative to the key picture. For example, the feature change at T1 may be determined based on the feature information of the key picture and the original picture at T1, without determining the feature information of the original picture at T1. The feature change at T1 may be determined based on the feature information of the key picture and the original picture at T1. Similarly, for example, the feature change of the original pictures at T2 - T3 may be extracted based on the original pictures at T2 - T3 and the key picture.

[0157] Referring to Figure 11B , the change (or feature change) of the established feature (e.g., the feature indicated by the feature information of the key picture at T0) may be encoded (e.g., coded) into the bitstream. In one example, the encoded bitstream includes the encoded key picture at T0 and the encoded feature changes at T1 - T3.

[0158] On the decoder side (lower part), it may be through Figure 11AA method similar to the method described in is used to decode the key picture data. The encoded feature changes at T1 - T3 (e.g., the changes to the established features for the key picture at T0) can be decoded to recover the corresponding feature information indicating the features of the pictures at T1 - T3 and / or the features of the key points. In one example, the feature information of the key picture at T0 is decoded. The feature information indicating the features of the pictures at T1 - T3 and / or the features of the key points can be determined based on the corresponding decoded feature changes at T1 - T3 and the feature information of the key picture.

[0159] When decoding the feature information for a specific picture (e.g., the picture at T1), the corresponding picture data at the same time instance (e.g., T1) can be recovered, reconstructed, or rendered. The encoded picture can be decoded based on the decoded feature information of the picture and the decoded key picture at the same time instance.

[0160] A feature - based video encoding process, such as (1100A) or (1100B), may include a feature - based video encoding (or compression) process and a feature - based video decoding process (e.g., including feature - based rendering and / or reconstruction).

[0161] In one example, when certain conditions are met, the feature - based encoding process (1100A) or (1100B) may be more advantageous than the relevant video encoding techniques in a selected application. Two exemplary conditions are: (i) the bit - rate cost of encoding the features is less than the bit - rate cost of encoding the original video data; and (ii) the reconstructed visual quality of the feature - based video encoding is subjectively acceptable. In some examples, the pictures rendered from the features do not necessarily need to match the original pictures or have high subjective quality, e.g., as evaluated by peak signal - to - noise ratio (PSNR) and structural similarity index (SSIM).

[0162] In the feature - based encoding process, the feature information of the content in the picture can be extracted from the picture. The feature information can be used to represent the picture. The feature information or feature changes of the picture, rather than the samples (or pixels) in the picture, can be encoded.

[0163] Feature information need not include information about portions of the picture that do not change significantly. Thus, in some examples, when there is no significant change in the content between different pictures of a view, the feature information may include information about only a relatively small portion of the picture. Feature changes may include feature changes in only a relatively small portion of the picture. Information (or feature changes) about a relatively large portion of the picture, such as the background or other parts of the body, may not be included in the encoded feature information. Encoding efficiency can be improved without sacrificing visual quality.

[0164] Figure 11A and Figure 11B The above description in and can be applied to feature-based encoding of pictures of a single view. According to an embodiment of the present disclosure, feature-based video encoding can be used to encode a video having multiple views. The present disclosure includes embodiments for determining (e.g., defining) features in a multi-view context. The term "feature" in the present disclosure can be used in a method using key points (or key points) for video encoding and reconstruction.

[0165] Figure 12 An example of a picture in a bitstream (1200) according to an embodiment of the present disclosure is shown. The bitstream (1200) may include pictures of one or more views (e.g., views 0-(N-1)), where N is a positive integer indicating the number of views in the bitstream (1200). When N is 1, the bitstream (1200) is a single-view bitstream including pictures of a single view (e.g., view 0). When N is greater than 1, the bitstream (1200) is a multi-view bitstream including pictures of different views 0-(N-1). For each view, Figure 12 pictures at (M+1) time instances T0-TM are shown.

[0166] The pictures in the bitstream (1200) may be referred to as picture PIJ. I represents a view I, with values ranging from 0 to (N-1). Any appropriate number of views may be included in the bitstream (1200). J represents a time instance J, with values ranging from 0 to M. Any suitable number of pictures may be included in the views of the bitstream (1200). For example, picture P32 refers to the picture at Figure 3 T2 of view. In one example, the time instances T0-TM indicate the decoding order. For example, picture PI(J-1) (e.g., P31) is decoded before picture PIJ (e.g., P32).

[0167] In one example, pictures of one view are acquired by one camera. Pictures of another view are acquired by another camera. For example, N cameras may be used to acquire pictures of views 0-(N-1). Figure 9A and Figure 9B The multi-camera system shown in can be used to acquire pictures of views 0-(N-1).

[0168] The pictures of one view (e.g., view 0) can be encoded independently of the pictures of another view (e.g., view 1). The pictures of each view (e.g., P00 - P0M) can be encoded (e.g., encoded and / or decoded) based on a feature - based encoding process (1100A) or (1100B).

[0169] As described above, the inter - view dependencies between different views can be explored to obtain more efficient video / image encoding. The inter - view dependencies can be used to determine the feature information of the key pictures of different views. The inter - view dependencies can be utilized to encode non - key pictures based on the feature information of the key pictures of different views. The inter - view dependencies can be used to encode a non - key picture of a first view (e.g., view 1) based on another picture of a second view (e.g., view 1), where the first view is different from the second view.

[0170] According to an embodiment of the present disclosure, three - dimensional (3D) feature information can be determined based on the pictures of different views. Feature information indicating features / keypoints can be extracted based on the key pictures of multiple views at different angles, for example. In one example, a current feature model (e.g., a 3D feature model based on multiple views) is determined (e.g., established) based on the feature information. The same 3D feature model can be used to render pictures from different perspectives.

[0171] Figure 13 Exemplary key pictures of different views in a bitstream (1200) that can be used to determine feature information according to an embodiment of the present disclosure are shown. In Figure 12 is described Figure 13 the bitstream (1200), pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM in

[0172] In one embodiment, different pictures corresponding to different views at a time instance (e.g., T0) are used to determine 3D feature information or the current 3D feature model. The current 3D feature model can be used to encode and / or decode pictures from different views in the bitstream (1200). The 3D feature information (or 3D feature model) specific to the content (e.g., face) of a key picture can be determined based on a predetermined 3D feature model (e.g., a canonical 3D feature model or a general 3D feature model) and the key picture. In one example, the predetermined 3D feature model is adapted to the key picture to generate 3D feature information (or 3D feature model).

[0173] In one example, the entire set (1302) of pictures P00 - P(N - 1)0 at T0 is used to determine 3D feature information or the current 3D feature model. For example, the key pictures corresponding to all decoded views (e.g., views 0 - (N - 1)) are used to generate the current 3D feature model. The 3D feature information or the current 3D feature model generated based on the entire set (1302) can be referred to as unified 3D feature information or the current unified 3D feature model. The unified 3D feature information or the current unified 3D feature model can be used for encoding (e.g., encoding, decoding, or generating) a picture at a specific view (one of views 0 - (N - 1), e.g., view Figure 3 ) and a specific time instance (one of time instances T0 - TM, e.g., T2) (e.g., a non - key picture, e.g., P32).

[0174] As Figure 13 shown, a subset (1301) of pictures P00 - P(N - 1)0 at T0 is used to determine 3D feature information or the current 3D feature model. In one example, the subset (1301) includes pictures P00, P10, and P20 of views 0 - 2 at T0 respectively. The 3D feature information or the current 3D feature model generated based on the subset (1301) can be referred to as unified 3D feature information or the current unified 3D feature model, which can be used for encoding (e.g., encoding, decoding, or generating) a picture at a specific view (one of views 0 - (N - 1), e.g., view Figure 3 ) and a specific time instance (one of time instances T0 - TM, e.g., T2) (e.g., a non - key picture, e.g., P32).

[0175] Figure 14 shows a subset of exemplary key pictures of different views in a bitstream (1200) for determining multiple 3D feature information according to an embodiment of the present disclosure. The pictures (e.g., non - key pictures) in the bitstream (1200) can be determined based on the multiple 3D feature information determined or the corresponding different current 3D feature models. The bitstream (1200), pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM are described in Figure 12 . Figure 14

[0176] A first subset (e.g., subset (1401)) of a plurality of pictures (e.g., P00 - P(N - 1)0 at T0) at a time instance of a first view (e.g., view 0 - 2) can be used to determine first 3D feature information or a first current 3D feature model. In one example, the first subset (1401) includes pictures P00, P10, and P20 corresponding to view 0 - 2 at T0 respectively. The first 3D feature information or the first current 3D feature model can be used to encode (e.g., encode, decode, or generate) a picture (e.g., a non - key picture, e.g., P02) at one of the first views (one of views 0 - 2, e.g., view 0) and a specific time instance (one of time instances T0 - TM, e.g., T2).

[0177] A second subset (1403) of a plurality of pictures (e.g., P00 - P(N - 1)0 at T0) at a time instance can be used to determine second 3D feature information or a second current 3D feature model. In one example, the second subset (2103) includes pictures P20, P30, and P40 corresponding to... Figures 2 - 4 ...respectively. The second 3D feature information or the second current 3D feature model can be used to encode (e.g., encode, decode, or generate) a picture (e.g., a non - key picture, e.g., P32) at one of the second views (one of the views... Figures 2 - 4 ...e.g., one of the views... Figure 3 ...and a specific time instance (one of time instances T0 - TM, e.g., T2).

[0178] Referring back to... Figures 13 to 14 ...a feature change of the 3D feature information (e.g., obtained by a 3D transformation or 3D deformation between two adjacent feature information corresponding to two adjacent pictures) can be written into the bitstream. The feature change can be sent to the decoder in the bitstream (1200). On the decoder side, the corresponding 3D feature information or the corresponding 3D feature model (e.g., the unified 3D feature information or the current unified 3D feature model in... Figure 13 ...or the different 3D feature information or different 3D feature models in... Figure 14 ...can be applied to the written feature change at different perspectives to render a reconstructed picture of a specific view.

[0179] According to an embodiment of the present disclosure, a single feature information (or feature model) of each view or view subset can be determined (e.g., constructed) based on corresponding key pictures of multiple views. Each view can be associated with corresponding feature information indicating the corresponding features and / or key points of the view. In one example, key pictures of multiple views (e.g., multiple adjacent views) are used as inputs to generate single feature information (or feature model) of a specific view and a specific time instance.

[0180] Figure 15Shows a subset of exemplary key pictures of different views in a bitstream (1200) for determining multiple feature information according to an embodiment of the present disclosure. In Figure 12 is described in Figure 15 the bitstream (1200), pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM in

[0181] A first subset (e.g., subset (1501)) of pictures (e.g., P00 - P(N - 1)0 at T0) of a first view set (e.g., views 0 - 2) at a time instance can be used to determine first feature information or a first current feature model of a first view (e.g., view 0) in the first view set. In one example, the first subset (1501) includes pictures P00, P10, and P20 corresponding to views 0 - 2 at T0 respectively. The first feature information or the first current feature model can be used to encode (e.g., encode, decode, or generate) a picture (e.g., a non - key picture, e.g., P02) of the first view (e.g., view 0) and a specific time instance (one of time instances T0 - TM, e.g., T2). In one example, the first feature information or the first current feature model is first 3D feature information or a first current 3D feature model. In one example, the first feature information or the first current feature model is first 2D feature information or a first current 2D feature model.

[0182] A second subset (e.g., subset (1502)) of pictures (e.g., P00 - P(N - 1)0 at T0) of a second view set (e.g., views 1 - 3) at the same time instance can be used to determine second feature information or a second current feature model of a second view (e.g., view 1) in the second view set. In one example, the second subset (1502) includes pictures P10, P20, and P30 corresponding to views 1 - 3 at T0 respectively. The second feature information or the current second feature model can be used to encode (e.g., encode, decode, or generate) a picture (e.g., a non - key picture, e.g., P12) of the second view (e.g., view 1) and a specific time instance (one of time instances T0 - TM, e.g., T2).

[0183] In one example, a third subset (e.g., subset (1503)) of pictures (e.g., P00 - P(N - 1)0 at T0) of a third view set (e.g., view Figures 2 - 4 ) at the same time instance can be used to determine third feature information or a third current feature model of a third view (e.g., view Figure 2 ) in the third view set. In one example, the third subset (1503) includes at T0 view Figures 2 - 4Respectively corresponding pictures P20, P30, and P40. The third feature information or the current third feature model can be used for encoding (e.g., encoding, decoding, or generating) the third view (e.g., view Figure 2 ) and pictures (e.g., non-key pictures, such as P22) at a specific time instance (one of the time instances T0 - TM, e.g., T2).

[0184] Figure 16 An exemplary encoding method (1600) of feature-based multi-view encoding according to an embodiment of the present disclosure is shown. The pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM in the bitstream (1200) are described in Figure 12 . Figure 16 As described above, pictures across different views at one time instance (e.g., the first time instance of the entire sequence in the bitstream (1200), e.g., T0) can be encoded as key pictures. On the decoder side, the reconstructed key pictures of each view among different views (e.g., views 0 - (N - 1)) at the time instance (e.g., the first time instance, e.g., T0) can later be used as key pictures for feature-based rendering and / or reconstructing pictures of the same view but at different time instances (e.g., non-key pictures).

[0185] For example, the entire set (1302) of pictures at T0 (including P00 - P(N - 1)0) are key pictures. The key pictures P00 - P(N - 1)0 at T0 can be encoded (e.g., encoded and / or decoded) by using, for example, the method of single-view video, which is, for example, the video / image encoding technology described above (e.g., intra-frame (picture) and / or inter-temporal frame (picture) prediction) (e.g., the image / video encoding technology of HEVC and / or VVC). The encoded key pictures P00 - P(N - 1)0 can be sent to the decoder side. On the decoder side, during the feature-based rendering and / or reconstruction process, the reconstructed key picture (e.g., P20) of the first view (e.g., view

[0186] ) at T0 can later be used as a key picture to decode pictures (e.g., non-key pictures P20 - P2M) of the first view (e.g., view Figure 2 ) at different time instances (e.g., T1 - TM) respectively. Figure 2 ).

[0187] Figures 13 to 15 One or more embodiments in Figure 16 can be combined with the embodiments in Figure 16 . Referring to Figures 13 to 15, the feature information or feature model can be determined based on a subset (e.g., (1301), (1401), (1501)) or the entire set (e.g., (1302)) of all pictures at a first time instance (e.g., T0). The feature information or feature model can be Figure 13 the unified 3D feature information or unified 3D feature model applicable to all non-key pictures in the bitstream (1200), and can be Figure 14 one of the multiple 3D feature information or one of the different current 3D feature models applicable to a subset of non-key pictures of the corresponding view in the bitstream (1200), or can be Figure 15 the single feature information or single feature model applicable to a subset of non-key pictures of each view in the bitstream (1200).

[0188] Referring to Figure 16 , the feature change of the picture (e.g., non-key picture P22) of the first view (e.g., view Figure 2 ) can be generated based on the non-key picture P22 of the first view and the key picture (e.g., P20) at the first time instance. In one example, the feature change of the non-key picture P22 of the first view is generated based on the non-key picture P22 of the first view and the key picture of the first view (e.g., P20) at the first time instance, and the feature change indicates the feature change of the same view (e.g., the first view) in the time domain (e.g., from T0 to T2). The feature changes and key pictures (including P20) of all views at the first time instance can be encoded. On the decoder side, the non-key picture P22 can be decoded based on the decoded feature change and the decoded key picture P20. In one example, on the decoder side, the same method as on the encoder side can be used to obtain the feature information or feature model from the corresponding decoded key pictures (e.g., subset (1301), subset (1401), subset (1501), etc.). The non-key picture P22 can be decoded based on the feature change, feature information or feature model, and the decoded key picture P20.

[0189] Figure 17 shows an exemplary encoding method (1700) for feature-based multi-view coding according to an embodiment of the present disclosure. In Figure 12 it is described Figure 17 the pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM in the bitstream (1200) in

[0190] As described above, pictures of a first view (e.g., view 0) across different time instances (e.g., T0 - TM) can be encoded as key pictures. In an example, the key pictures include all pictures of the first view (e.g., P00 - P0M). On the decoder side, the key pictures reconstructed at each time instance at different time instances (e.g., T0 - TM) in the first view can later be used as key pictures for feature - based rendering and / or reconstruction of pictures of the same time instance but different views (e.g., non - key pictures).

[0191] For example, the entire set (1701) of pictures (including P00 - P0M) of a first view (e.g., view 0) are key pictures. The key pictures P00 - P0M of view 0 can be encoded (e.g., encoded and / or decoded) using, for example, a method for single - view video, which is, for example, the video / image encoding techniques described above (e.g., intra - frame (picture) and / or inter - frame (picture) prediction) (e.g., the image / video encoding techniques of HEVC and / or VVC). The encoded key pictures P00 - P0M can be sent to the decoder side. On the decoder side, during the feature - based rendering and / or reconstruction process, the reconstructed key picture (e.g., P01) of the time instance (e.g., T1) of view 0 can later be used as a key picture to separately decode pictures (e.g., non - key pictures P11 - P(N - 1)1) of the same time instance (e.g., T1) but different views (e.g., views 1 - (N - 1)).

[0192] In Figure 17 the example of Figure 2 , based on a non - key picture (e.g., P21) of a second view (e.g., view Figure 2 ) and a key picture (e.g., P01) of the first view (e.g., view 0) at the same time instance (e.g., T1), the feature change of the non - key picture (e.g., P21) of the second view (e.g., view Figure 2 ) relative to the key picture (e.g., P01) of the first view (e.g., view 0) at the same time instance (e.g., T1) is determined. The feature change indicates the feature change at the same time instance for multiple views (e.g., from the first view to the second view). The feature change can be encoded and transmitted to the decoder side. On the decoder side, a non - key picture (e.g., P21) can be generated based on the decoded feature change relative to the views (e.g., from the first view to the second view) and the key picture (e.g., P01) at the same time instance.

[0193] In one example, feature information of one or more key pictures (e.g., including P01) of a first view (e.g., view 0) is extracted at the encoder side and / or the decoder side. At the decoder side, non-key pictures (e.g., P21) can be generated based on the decoded feature changes with respect to multiple views (e.g., from the first view to the second view), the key pictures at the same time instance (e.g., P01), and the feature information of one or more key pictures (e.g., including P01) of the first view (e.g., view 0).

[0194] Encoding methods (1600) and (1700) can be referred to as feature-based multi-view encoding architectures. In one example, two feature-based multi-view encoding methods (1600) and (1700) can be combined. Key pictures can include the following combinations: (i) pictures of one view at different time instances and (ii) pictures of different views at one time instance. For example, key pictures can include the following combinations: (i) all pictures of one view at different time instances and (ii) all pictures of different views at one time instance.

[0195] Figure 18 An exemplary encoding method (1800) of feature-based multi-view encoding according to an embodiment of the present disclosure is shown. In Figure 12 is described Figure 18 the pictures P00 - P(N - 1)M, views 0 - (N - 1), and time instances T0 - TM in the bitstream (1200).

[0196] As described above, pictures of the first view (e.g., view 0) across different time instances (e.g., T0 - TM) can be encoded as key pictures. At one time instance (e.g., the first time instance (e.g., T0) of the entire sequence in the bitstream (1200)), pictures across different views can be encoded as key pictures. In one example, the key pictures include: (i) a first set including all pictures of the first view (e.g., P00 - P0M) and (ii) a second set including all pictures (e.g., P00 - P(N - 1)0) of all views (e.g., view 0 - (N - 1)) at the first time instance (e.g., T0). Since P00 is shared by the first set and the second set, the key pictures include P00 - P0M and P10 - P(N - 1)0. The key pictures P00 - P0M and P10 - P(N - 1)0 can be encoded (e.g., encoded and / or decoded) using, for example, a method for single - view video, which is, for example, the video / image encoding techniques described above (e.g., intra - (picture) and / or temporal inter - (picture) prediction) (e.g., the image / video encoding techniques of HEVC and / or VVC). The encoded key pictures P00 - P0M and P10 - P(N - 1)0 can be sent to the decoder side.

[0197] At the decoder side, a picture (e.g., non - key picture P22) of the second view (e.g., view ) at the second time instance (e.g., T2) can be reconstructed based on (i) a first reconstructed key picture (e.g., P02) of the first set (e.g., P00 - P0M) at the second time instance and (ii) a second reconstructed key picture (e.g., P20) of the second set (e.g., P00 - P(N - 1)0) at the second view (e.g., view Figure 2 ). Figure 2 ).

[0198] Different transmission or delivery priorities can be assigned to different views. In some examples, for instance when the bandwidth is limited, certain views (e.g., views with a lower priority than other views) can be discarded.

[0199] If associated with the corresponding view, multiple feature information and / or associated feature changes in different time instances can have the same priority as the associated view. In one example, the feature information of a view (e.g., view 0), such as the feature information of the (one or more) pictures of that view, can be assigned the same priority in transmission as that view (e.g., view 0). Returning to Figure 12, the first feature information and / or the first feature change associated with one or more pictures (e.g., one or more pictures among P00 - P0M) of the first view (e.g., view 0) may have the first priority as the first view (e.g., view 0). The second feature information and / or the second feature change associated with one or more pictures (e.g., one or more pictures among P10 - P1M) of the second view (e.g., view 1) may have the second priority as the second view (e.g., view 1).

[0200] Common feature information may be shared by multiple views (e.g., the first view (e.g., view 0) and the second view (e.g., view 1)). For example, the first feature information of the first view and the second feature information of the second view both indicate eyes, nose, and mouth. The first feature information indicates the left ear, and the second feature information indicates the right ear. The common feature information shared by the first view and the second view indicates the features and / or key points of eyes, nose, and mouth. The non - common feature information indicates the features and / or key points of the left ear and the right ear. According to an embodiment of the present disclosure, the common feature information (e.g., eyes, nose, and mouth) shared by multiple views (e.g., the first view and the second view) may have a higher priority than the non - common feature information associated with the same view.

[0201] If a unified feature information or a unified feature model is used to describe the features and / or key points of all views, the unified feature information or the unified feature model at different time instances may have the highest priority and will be transmitted to the decoder. In the example, the unified feature information or the unified feature model at different time instances is not discarded.

[0202] The priority of a view may be a separate set of priorities. For example, the priority of a view is different from the priority set for the temporal layer of the video bitstream.

[0203] Figure 19Shows an example of pictures at different time instances corresponding to different views according to an embodiment of the present disclosure. The time instances are indicated by T0 - T11. The views are indicated by S0 - S7. The pattern of the pictures may repeat along the time axis, for example, with a period of 8. The pattern of the pictures from T8 to T11 repeats the pattern of the pictures from T0 to T3. The picture represented by I0 can be independently encoded without referring to another picture. The picture indicated by P0 can be encoded based on another picture such as picture I0. The pictures indicated by B1, B2, and B3 can be encoded based on two other pictures. In one example, B1 is predicted based on I0(s) and / or P0(s). In one example, B2 is predicted based on I0(s), P0(s), and B1(s). In one example, B3 is predicted based on I0(s), P0(s), B1(s), and B2(s). In one example, B4 is predicted based on four other pictures including B1(s), B2(s), and B3(s).

[0204] The priorities of multiple views at a certain time instance (e.g., T0) can be arranged in descending order of priority as follows: the priority of view S0, the priority of view S2, the priority of view S4, the priority of view S6, the priority of view S7, and the priorities of views S1, S3, and S5. In one example, the priorities of views S1, S3, and S5 are the same. The priorities of multiple views at another time instance (e.g., T1) can be the same as the priorities in the above time instance T0.

[0205] The priorities of the time layers (e.g., T0 - T7) for a specific view (e.g., S0) can be arranged in descending order of priority as follows: the priority of time layer T0, the priority of time layer T4, the priorities of time layers T2 and T6, and the priorities of time layers T1, T3, T5, and T7. In one example, the priorities of time layers T2 and T6 are the same. In one example, the priorities of time layers T1, T3, T5, and T7 are the same. The priorities of the time layers in another view (e.g., S1) can be the same as the priorities in the above-described view S0.

[0206] The priorities of the views can be reused as the priorities set for the time layers of the video bitstream.

[0207] Figure 20A flowchart showing an overview of an encoding process (2000) according to an embodiment of the present disclosure is shown. In various embodiments, the process (2000) is executed by a processing circuit, such as the processing circuit in terminal devices (310), (320), (330), and (340), the processing circuit that executes a video encoder function (such as (403), (603), (703)), etc. In some embodiments, the process (2000) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (2000). The process starts at (S2001) and proceeds to (S2010).

[0208] At (S2010), first feature information or a first current feature model of the content in at least one first key picture among a plurality of pictures can be determined, as Figures 13 to 15 described. The plurality of pictures correspond to different views (e.g., view 0-(N-1)), and at least one first key picture corresponds to at least one first view among different views.

[0209] At (S2020), a first feature change of the first feature information can be determined. The first feature change can indicate a content change between one key picture in at least one first key picture and a first picture, such as as described above (e.g., in Figure 11B , Figure 16 and Figure 17 ).

[0210] At (S2030), the first feature change and at least one first key picture can be encoded. In one example, the first feature change and at least one first key picture are included in a multi-view bitstream, as Figure 11B , Figure 16 and Figure 17 described.

[0211] The process (2000) proceeds to (S2099) and stops.

[0212] The process (2000) can be suitably applied to various scenarios, and the steps in the process (2000) can be adjusted accordingly. One or more steps in the process (2000) can be modified, omitted, repeated, and / or combined. The process (2000) can be implemented in any suitable order. Additional steps can be added.

[0213] Figure 21A flowchart showing an overview decoding process (2100) according to an embodiment of the present disclosure is shown. In various embodiments, the process (2100) is executed by a processing circuit, such as the processing circuit in terminal devices (310), (320), (330), and (340), the processing circuit performing the function of video encoder (403), the processing circuit performing the function of video decoder (410), the processing circuit performing the function of video decoder (510), the processing circuit performing the function of video encoder (603), etc. In some embodiments, the process (2100) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (2100). The process starts at (S2101) and proceeds to (S2110).

[0214] At (S2110), at least one first key picture among a plurality of pictures from a multi-view bitstream can be decoded. The plurality of pictures may correspond to different views. The at least one first key picture may correspond to at least one first view among the different views.

[0215] At (S2120), first feature information or a first current feature model of the content in at least one first key picture among the plurality of pictures can be determined, as Figures 13 to 15 described.

[0216] In one example, the at least one first key picture corresponds to a first time instance. The at least one first key picture includes a plurality of first key pictures. At least one first view among different views includes a plurality of first views. The first feature information includes first 3D feature information or a first 3D current feature model indicated by the plurality of first views, as Figure 13 and Figure 14 described.

[0217] In one example, the first 3D feature information at the first time instance is determined based on a first predetermined 3D feature model and a plurality of first key pictures.

[0218] The first 3D feature information can be used to decode pictures of each view among different views.

[0219] The plurality of first key pictures includes each key picture at the first time instance. Referring to Figure 13 , the at least one first key picture includes P00 - P(N - 1)0.

[0220] In one example, the first picture belongs to one view among different views at a second time instance.

[0221] In one example, second 3D feature information of the content in multiple second key pictures in multiple pictures at a first time instance is determined based on a second predetermined 3D feature model, such as described in Figure 14 . The multiple second key pictures correspond to a first time instance of multiple second views in different views.

[0222] In (S2130), based on the multi-view bitstream, a first feature change of the first feature information can be decoded. The first feature change can indicate a content change between a key picture in at least one first key picture and a first picture.

[0223] In (S2140), the first picture can be reconstructed based on the decoded first feature change, the first feature information, and the key picture in at least one first key picture.

[0224] Process (2100) proceeds to (S2199) and stops.

[0225] Process (2100) can be suitably applied to various scenarios, and the steps in process (2100) can be adjusted accordingly. One or more steps in process (2100) can be modified, omitted, repeated, and / or combined. The process (2100) can be implemented in any suitable order. Additional steps can be added.

[0226] In one embodiment, the first feature information is associated with a first view in at least one first view, such as described in Figure 15 . For each view that is not the first view among the different views, corresponding feature information can be determined based on the key picture of each view and another key picture of an adjacent view in the different views. Based on the multi-view bitstream, a feature change of the corresponding feature information can be decoded, where the feature change corresponds to the corresponding picture of each view. Pictures of each view can be generated based on the corresponding feature change, the corresponding feature information, and the key picture of each view.

[0227] A subset of pictures of a first view in at least one first view corresponding to a corresponding time instance can be decoded, where the subset of pictures of the first view includes the key picture in at least one first key picture. The first picture belongs to a second view among the different views. The first picture and the key picture in at least one first key picture correspond to a first time instance. The first feature change indicates a feature change between the first picture of the second view at the first time instance and the key picture of the first view at the first time instance.

[0228] In one example, each picture of a first view among at least one first view can be decoded as a key picture, where each picture of the first view corresponds to a respective time instance.

[0229] Embodiments in the present disclosure can be applied to video sequences with multiple views or still pictures with multiple views. Applications include but are not limited to VR 360, free viewpoint systems, light field videos, etc.

[0230] Embodiments in the present disclosure can be used alone or in any order combination. Additionally, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0231] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 22 A computer system (2200) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0232] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpreted code, microcode, etc.

[0233] The instructions can be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0234] Figure 22 The components shown for the computer system (2200) are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one component or combination of components shown in the exemplary embodiments of the computer system (2200).

[0235] A computer system (2200) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, the following: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to collect certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0236] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (2201), mouse (2202), touchpad (2203), touch screen (2210), data glove (not shown), joystick (2205), microphone (2206), scanner (2207), camera (2208).

[0237] The computer system (2200) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen (2210), data glove (not shown), or joystick (2205), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (2209), headphones (not shown)), visual output devices (e.g., screens (2210) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input functionality, each with or without tactile feedback functionality - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0238] The computer system (2200) may also include human-accessible storage devices and their associated media: for example, optical media including CD / DVD ROM / RW (2220) with media such as CD / DVD (2221), thumb drives (2222), removable hard disk drives or solid state drives (2223), traditional magnetic media such as tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.

[0239] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover a transmission medium, a carrier wave, or other transitory signals.

[0240] The computer system (2200) may also include an interface (2254) to one or more communication networks (2255). The network can be, for example, a wireless network, a wired network, an optical network. The network can further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CANBus, and so on. Certain networks typically require an external network interface adapter (e.g., a USB port of the computer system (2200)) connected to certain common data ports or peripheral buses (2249); as described below, other network interfaces are typically integrated into the kernel of the computer system (2200) by connecting to the system bus (e.g., connecting an Ethernet interface in a PC computer system or connecting a cellular network interface in a smartphone computer system). The computer system (2200) can communicate with other entities using any of these networks. Such communication can be only one-way receiving (e.g., broadcast television), only one-way sending (e.g., CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0241] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be attached to the kernel (2240) of the computer system (2200).

[0242] The kernel (2240) may include one or more central processing units (CPUs) (2241), a graphics processing unit (GPU) (2242), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2243), hardware accelerators (2244) for certain tasks, a graphics adapter (2250), etc. These devices, as well as a read-only memory (ROM) (2245), a random access memory (2246), and internal mass storage such as an internal non-user-accessible hard disk drive, SSD, etc. (2247) may be connected via a system bus (2248). In some computer systems, the system bus (2248) may be accessible in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus of the kernel (2248) or connected to the system bus of the kernel (1848) via a peripheral bus (2249). In one example, a screen (2210) may be connected to the graphics adapter (2250). The architecture of the peripheral bus includes PCI, USB, etc.

[0243] The CPU (2241), GPU (2242), FPGA (2243), and accelerator (2244) may execute certain instructions, which may be combined to form the above-mentioned computer code. The computer code may be stored in the ROM (2245) or the RAM (2246). Transitional data may be stored in the RAM (2246), while permanent data may be stored in, for example, the internal mass storage (2247). Fast storage and retrieval of any storage device may be performed by using a cache, which may be closely associated with one or more CPUs (2241), GPUs (2242), mass storage (2247), ROM (2245), RAM (2246), etc.

[0244] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.

[0245] As a non-limiting example, a computer system having an architecture (2200), particularly a core (2240), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as certain non-transitory memories of the core (2240), such as on-core mass memory (2247) or ROM (2245). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2240). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (2240), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures (2246) stored in RAM and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., an accelerator (2244)), which can replace software or operate in conjunction with software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or including both. The present disclosure encompasses any suitable combination of hardware and software.

[0246] Appendix A: Abbreviations

[0247] JEM: Joint Exploration Model

[0248] VVC: Next Generation Video Coding

[0249] BMS: Benchmark Set

[0250] MV: Motion Vector

[0251] HEVC: High Efficiency Video Coding

[0252] SEI: Supplementary Enhancement Information

[0253] VUI: Video Usability Information

[0254] GOPs: Groups of Pictures

[0255] TUs: Transform Units

[0256] PUs: Prediction Units

[0257] CTUs: Coding Tree Units

[0258] CTBs: Coding Tree Blocks

[0259] PBs: Prediction Blocks

[0260] HRD: Hypothetical Reference Decoder

[0261] SNR: Signal-to-Noise Ratio

[0262] CPUs: Central Processing Units

[0263] GPUs: Graphics Processing Units

[0264] CRT: Cathode Ray Tube

[0265] LCD: Liquid Crystal Display

[0266] OLED: Organic Light-Emitting Diode

[0267] CD: Compact Disc

[0268] DVD: Digital Versatile Disc

[0269] ROM: Read-Only Memory

[0270] RAM: Random Access Memory

[0271] ASIC: Application-Specific Integrated Circuit

[0272] PLD: Programmable Logic Device

[0273] LAN: Local Area Network

[0274] GSM: Global System for Mobile Communications

[0275] LTE: Long-Term Evolution

[0276] CANBus: Controller Area Network Bus

[0277] USB: Universal Serial Bus

[0278] PCI: Peripheral Component Interconnect

[0279] FPGA: Field-Programmable Gate Array

[0280] SSD: Solid State Drive

[0281] IC: Integrated Circuit

[0282] CU: Coding Unit

[0283] R-D: Rate-Distortion

[0284] Although the present disclosure has described some exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it is to be understood that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A method for video decoding, characterized in that, the method comprises: decoding at least one first key picture among multiple pictures from a multi-view bitstream, the multiple pictures corresponding to different views, and the at least one first key picture corresponding to at least one first view among the different views; determining first feature information of the content in at least one first key picture among the multiple pictures; decoding a first feature change of the first feature information based on the multi-view bitstream, the first feature change indicating a position change, a direction change, a size change or a shape change between the first feature information of the content in one key picture among the at least one first key picture and the feature information of the corresponding content in a first picture, the first picture being one view among the different views, and the at least one first key picture corresponding to a first time instance; the first key pictures are multiple, the different views include multiple first views; the first feature information includes first 3D feature information indicated by the multiple first views, and the determining the first feature information includes determining the first 3D feature information at the first time instance based on a first predetermined 3D feature model and the multiple first key pictures; decoding a subset of the multiple pictures corresponding to different views, the subset of the multiple pictures belonging to one first view among the at least one first views corresponding to respective time instances, and the subset of the multiple pictures of the first view including the key pictures among the at least one first key picture; and reconstructing the first picture based on the decoded first feature change, the first feature information and the key pictures among the at least one first key picture.

2. The method according to claim 1, characterized in that, the first 3D feature information is used to decode pictures of each view among the different views.

3. The method according to claim 1, characterized in that, the multiple first key pictures include each key picture at the first time instance.

4. The method according to claim 1, characterized in that, the method comprises: determining second 3D feature information of the content in multiple second key pictures among the multiple pictures at the first time instance based on a second predetermined 3D feature model, the multiple second key pictures corresponding to a first time instance of multiple second views among the different views.

5. The method according to claim 1, characterized in that, the first feature information is associated with one first view among the at least one first views; for each view among the different views that is not the first view, the method further comprises: determining corresponding feature information based on the key pictures of each view and another key picture of an adjacent view among the different views; decoding a feature change of the corresponding feature information based on the multi-view bitstream, the feature change corresponding to the corresponding picture of each view; and generating a picture of each view based on the corresponding feature change, the corresponding feature information and the key pictures of each view.

6. The method according to claim 3, wherein, the first picture of one of the different views is at a second time instance.

7. The method according to claim 1, wherein, the method includes that the first picture belongs to a second view among the different views; the first picture and the key pictures among the at least one first key picture correspond to a first time instance; and a first feature change indicates a feature change between the first picture of the second view at the first time instance and the key picture of the first view at the first time instance.

8. The method according to claim 6, wherein, the method further comprises: decoding each picture of a first view among the at least one first view into a key picture, and each picture of the first view corresponds to a respective time instance.

9. A method for video coding, wherein, the method includes: determining at least one first key picture among a plurality of pictures, the plurality of pictures corresponding to different views, and the at least one first key picture corresponding to at least one first view among the different views; determining first feature information of the content in the at least one first key picture among the plurality of pictures; encoding a first feature change of the first feature information, the first feature change indicating a position change, a direction change, a size change or a shape change between the first feature information of the content in a key picture among the at least one first key picture and the feature information of the corresponding content in a first picture, the first picture being one of the different views, the at least one first key picture corresponding to a first time instance; there are a plurality of the first key pictures, the different views include a plurality of first views; the first feature information includes first 3D feature information indicated by the plurality of first views, and the determining the first feature information includes determining the first 3D feature information at the first time instance based on a first predetermined 3D feature model and the plurality of first key pictures; encoding a subset of the plurality of pictures corresponding to different views, the subset of the plurality of pictures belonging to a first view among the at least one first view corresponding to a respective time instance, and the subset of the plurality of pictures of the first view includes the key picture among the at least one first key picture; and encoding the first picture based on the encoded first feature change, the first feature information and the key picture among the at least one first key picture to obtain a multi-view bitstream.

10. An apparatus for video decoding, wherein, comprises: a memory for storing instructions; a processor for calling the instructions stored in the memory to implement the method according to any one of claims 1-8.

11. An apparatus for video coding, wherein, comprises: a memory for storing instructions; a processor for calling the instructions stored in the memory to implement the method according to claim 9.

12. A non-transitory computer-readable storage medium, wherein, It stores a program executed by at least one processor to perform the method according to any one of claims 1-9.

13. A method for processing a video stream, characterized in that the video stream is decoded according to the method according to any one of claims 1-8 or generated according to the method according to claim 9.