System and method for end-to-end lossless stereoscopic image compression
Through end-to-end lossless compression network, multi-scale codec structure and automatic encoder network, the problem of insufficient use of views when losslessly compressing stereo images in the prior art is solved, and efficient and simplified lossless compression effect is achieved.
Patent Information
- Application Number
- CN202380060154.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-17
- Filing Date
- 2023-08-17
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to fully utilize the correlation between views when compressing stereo images losslessly, resulting in poor encoding and decoding performance, and complex methods and requires multiple training stages.
The end-to-end lossless compression network is adopted to derive multi-scale auxiliary representation from the stereoscopic image through a multi-scale encoding and decoding structure, and establish hierarchical dependencies. The probability distribution is jointly estimated by using the automatic encoder network and the predictor, and combined with the inter-view interaction module and entropy model, bitstream generation is optimized.
It realizes more efficient lossless stereo image compression, makes full use of the correlation between views, simplifies the network architecture, reduces the training stage, and improves the encoding and codec performance.
Smart Images

Figure CN119948862A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This patent application claims the benefit of international application PCT / CN2022 / 112993 filed on August 17, 2022, which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the generation, storage and consumption of digital audio-visual media information in file format. Background Art
[0004] Digital video consumes the largest amount of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is likely to continue to grow. Summary of the invention
[0005] A first aspect relates to a method for processing video data, comprising: determining to apply an end-to-end lossless compression network to convert an input stereo image pair {x L ,x R}Compressed into a bit stream L ,b R}, where x represents the input stereoscopic image, b represents one of the bitstreams, L represents left, and R represents right; and conversion between visual media data and bitstream is performed based on an end-to-end lossless compression network.
[0006] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the end-to-end lossless compression network includes a multi-scale codec structure, the multi-scale codec structure is from {x L ,x R}Derive multi-scale auxiliary representation And establish {x L ,x R}and The hierarchical dependencies between are as follows:
[0007]
[0008] in,
[0009] and
[0010]
[0011] Where p represents the probability distribution, and where S represents the scale.
[0012] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that S is equal to 3.
[0013] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that S is a positive integer.
[0014] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the end-to-end lossless compression network includes an autoencoder network configured to be applied to each scale to estimate and
[0015] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the autoencoder network is configured to estimate The probability distribution of , and wherein the autoencoder network includes an encoder configured to be based on a non-quantized auxiliary representation of the previous scale To generate a non-quantized auxiliary representation
[0016] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the encoder includes N convolutional layers and M activation layers, wherein N and M are each positive integers.
[0017] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the encoder includes N residual blocks.
[0018] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the encoder includes a non-linear function configured to convert the input signal to a high-order domain.
[0019] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the autoencoder network includes a scale quantizer configured to Quantization as an auxiliary representation of quantization According to the information provided by the autoencoder network at scale s+2 and Compression is done via entropy codec.
[0020] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the scalar quantizer is configured to perform quantization using a rounding operation.
[0021] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the autoencoder network includes a predictor configured to jointly estimate the probability distribution and
[0022] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the predictor is configured to jointly estimate the probability distribution based on the intra-view prior and the inter-view prior and
[0023] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the predictor includes an inter-view interaction module and an entropy model.
[0024] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the inter-view interaction module includes a semi-coupled inter-view interaction module, and the semi-coupled inter-view interaction module is configured to extract view sharing information as valid inter-view information.
[0025] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the entropy model includes a joint conditional entropy model, which is configured to jointly estimate the distribution of the left view and the right view based on intra-view prior information and inter-view prior information.
[0026] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the semi-coupled view interaction module is configured to obtain the input stereo feature {f L ,f R} extracts inter-view prior information and combines the inter-view prior information with {f L ,f R}Merge to produce enhanced three-dimensional features
[0027] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the inter-view interaction module includes a semi-coupled extraction block, which is configured to extract view sharing information and generate a semi-coupled feature.
[0028] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the inter-view interaction module includes a parallax interaction transformer, which is configured to extract complementary information from the semi-coupled features and generate inter-view features.
[0029] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the inter-view interaction module includes a non-linear transformation block, which is configured to fuse the input features with the inter-view features and generate enhanced features, or a combination thereof.
[0030] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the semi-coupled extraction block is configured to extract the input stereo features {f L ,f RExtract semi-coupled features that preserve view-shared information while suppressing view-specific information And the semi-coupled extraction block includes a multi-stage extraction module, a semi-coupled depth-wise separable convolution and a fusion module.
[0031] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the semi-coupled extraction block adopts a multi-stage extraction strategy to progressively extract the view sharing information.
[0032] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that a semi-coupled depthwise separable convolution is applied to each stage c to extract a semi-coupled feature according to the following formula
[0033]
[0034] Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for left and right views respectively.
[0035] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the semi-coupled extraction block uses a fusion module to fuse the output semi-coupled features from each stage to generate a final semi-coupled feature, as shown below:
[0036]
[0037] Where {G L ,G R} is the aggregation block.
[0038] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the parallax interactive converter is configured to be based on the stereo feature {f L ,f R}From the semi-coupled feature Extract complementary information and generate inter-view features
[0039] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the parallax interactive converter includes an interactive structure, the interactive structure according to f L from Extract complementary information and according to f R from Extract complementary information.
[0040] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the parallax interaction transformer includes a parallax transformer configured to extract complementary information along a parallax direction.
[0041] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the parallax converter is based on f L from Extract complementary information.
[0042] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the query vector, the key vector, and the value vector are generated by a linear layer as follows: q L =Linear q (f L ),
[0043] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the squeeze operation is configured to squeeze the query vector, the key vector, and the value vector along a vertical direction as shown below: Where Squeeze(·): Includes extrusion operations.
[0044] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that a scaled dot product attention is performed along the disparity direction, followed by a feed-forward network generating an inter-view feature according to the following formula
[0045]
[0046] Where Softmax(·) represents the normalized exponential function (Softmax) operation, d k denotes a scaling factor, and FFN(·) denotes a feed-forward network.
[0047] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the parallax converter is configured to R from Extract complementary information and generate inter-view features or
[0048] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the nonlinear transformation block is configured to transform the input feature {f L ,f R} and view features Fusion, generating enhanced features according to the following formula
[0049]
[0050] where {H L ,HR} represents the nonlinear transformation implemented by the neural network.
[0051] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the joint conditional entropy model is configured to jointly estimate the probability distribution of the stereoscopic views.
[0052] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the stereoscopic views includes a probability dependency of the auxiliary representation and a corresponding entropy model.
[0053] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the stereoscopic view includes the probability dependency of the stereoscopic images and a corresponding entropy model.
[0054] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability dependency of the auxiliary representation and the corresponding entropy model are configured to be based on the prediction feature To estimate the left view auxiliary representation The probability distribution of Contains intra-view information and inter-view information.
[0055] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability dependency of the auxiliary representation and the corresponding entropy model are configured to be based on the estimated probability distribution of the auxiliary representation of the left view To assist in the left view Encoding and decoding.
[0056] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability dependency of the auxiliary representation and the corresponding entropy model are configured to use the decoded left view auxiliary representation To provide supplementary information to assist the right view Model the probability distribution of .
[0057] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability dependency of the auxiliary representation and the corresponding entropy model are configured to be based on the estimated probability distribution of the auxiliary representation of the right view To assist in the right view Encoding and decoding.
[0058] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the probability dependency and corresponding entropy model of the stereo image are configured to construct the stereo image {x L ,x R} interleaving dependencies between .
[0059] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that constructing the interlaced dependency between the stereo images comprises: L ,x R} is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3} and {x R,1 ,x R,2 ,x R,3}.
[0060] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that constructing the interlaced dependency between the stereo images includes estimating the sub-images {x L,1 ,x R,1} and compress the sub-image {x L,1 ,x R,1}Probability distribution of .
[0061] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the compressed sub-image includes a prediction feature based on the information including intra-view and inter-view information. To estimate the sub-image x L,1 The probability distribution of .
[0062] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the compressed sub-image includes an estimated probability distribution based on the auxiliary representation of the left view Come to x L,1 Encoding and decoding.
[0063] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the compressed sub-image includes using the decoded x L,1 To provide supplementary information for the sub-image x R,1 Model the probability distribution of .
[0064] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the probability distribution of the compressed sub-image includes a probability distribution based on the estimation To pair sub-image x R,1 Encoding and decoding.
[0065] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that constructing the interlaced dependency between the stereo images comprises decoding {x L,1 ,x R,1} is used as a condition to estimate the sub-image {x L,2 ,x R,2} and compress the sub-image {x L,2 ,xR,2}Probability distribution of .
[0066] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes based on the predicted features and {x L,1 ,x R,1} to estimate the sub-image x L,2 The probability distribution of .
[0067] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes estimating the probability distribution based on the auxiliary representation of the left view Come to x L,2 Encoding and decoding.
[0068] Optionally, in any of the preceding aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image comprises using the decoded x L,2 To provide supplementary information for the sub-image x R,2 Model the probability distribution of .
[0069] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes based on the estimated probability distribution To pair sub-image x R,2 Encoding and decoding.
[0070] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that constructing the interlaced dependency between the stereo images comprises decoding {x L,1 ,x R,1 ,x L,2 ,x R,2} is used as a condition to estimate the sub-image {x L,3 ,x R,3} and compress the sub-image {x L,3 ,x R,3}Probability distribution of .
[0071] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes based on the predicted features and {x L,1 ,x R,1 ,x L,2 ,x L,2} to estimate the sub-image x L,3 The probability distribution of .
[0072] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes estimating the probability distribution based on the auxiliary representation of the left view Come to x L,3 Encoding and decoding.
[0073] Optionally, in any of the preceding aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image comprises using the decoded x L,3 To provide supplementary information for the sub-image x R,3 Model the probability distribution of .
[0074] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the probability distribution of the sub-image includes based on the estimated probability distribution To pair sub-image x R,3 Encoding and decoding.
[0075] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further comprises using a plurality of encoders from {x L ,x R}Derive multi-scale auxiliary representation As shown below:
[0076]
[0077] in represents the encoder at scale s.
[0078] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further comprises: L ,x R}and A hierarchical dependency is established between them, as shown below:
[0079]
[0080] in,
[0081]
[0082] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method also includes estimating the factor distribution using multiple predictors.
[0083] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes providing a predictor that uses a neural network to estimate the distribution To provide predictive features The semi-coupled view interaction module (SI 2M) to extract inter-view information, and the process is formulated as follows:
[0084]
[0085] where P(·) indicates the predictor.
[0086] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further comprises estimating the distribution using a joint conditional entropy model (JCEM)
[0087] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes providing a semi-coupled view interaction module to extract the input stereo features {f L ,f R Extract semi-coupled features To generate enhanced features, semi-coupled features By preserving view-shared information while suppressing view-specific information, the enhanced features contain both intra-view information and inter-view information as effective priors.
[0088] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method also includes performing a progressive extraction comprising C stages.
[0089] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further comprises extracting the semi-coupled features using a semi-coupled depth-wise separable convolution applied to each stage c The process is formulated as:
[0090]
[0091] Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for left and right views respectively.
[0092] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the method further comprises using a fusion module to fuse the output semi-coupled features from each stage to generate a final semi-coupled feature, as shown below:
[0093]
[0094] Where {G L ,G R} is an aggregation block implemented by stacked 1×1 convolutional layers.
[0095] Optionally, in any one of the aforementioned aspects, another implementation of this aspect provides that the method further comprises using a parallax interactive transformer based on the stereo feature {f L ,f R}From the semi-coupled feature Extract complementary information and generate inter-view features
[0096] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides executing an interaction structure, the interaction structure according to f L from Extract complementary information and according to f R from Extract complementary information.
[0097] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides utilizing a disparity transformer that extracts complementary information along a disparity direction.
[0098] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the parallax converter is based on f L from Extract complementary information.
[0099] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the query vector, the key vector, and the value vector are generated by the linear layer according to the following manner:
[0100] q L =Linear Q (f L ),
[0101] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides squeezing the query vector, the key vector, and the value vector along the vertical direction according to the following manner:
[0102]
[0103] Where Squeeze(·): It is a squeeze operation.
[0104] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that a scaled dot product attention is further performed along the disparity direction, the disparity direction being the horizontal direction, and then the inter-view features are generated by the feed-forward network As shown below,
[0105]
[0106] Where Softmax(·) represents the Softmax operation, dk denotes a scaling factor, and FFN(·) denotes a feed-forward network.
[0107] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the parallax converter is based on f R from Extract complementary information and generate inter-view features or
[0108] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further comprises using a nonlinear transformation block, the nonlinear transformation block transforming the input feature {f L ,f R} and view features Fusion, generating enhanced features The process is formulated as follows:
[0109]
[0110] where {H L ,H R} represents the nonlinear transformation implemented by stacked 1×1 convolutional layers.
[0111] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes providing a joint conditional entropy model to jointly estimate the probability distribution of the stereoscopic view, thereby constructing an auxiliary representation The probability dependence of .
[0112] Optionally, in any one of the aforementioned aspects, another implementation of the aspect provides that the method further comprises: To estimate the left view auxiliary representation The probability distribution of .
[0113] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the distribution Perform parametric modeling as follows:
[0114]
[0115] in are the parameters corresponding to the logistic mixed model.
[0116] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, using a neural network to estimate As shown below:
[0117]
[0118] Among them, H L (·) denotes an estimator implemented by stacked 3×3 convolutional layers.
[0119] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides generating an estimated distribution
[0120] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution of the auxiliary representation of the left view To assist in the left view Encoding and decoding.
[0121] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that, using the decoded left view auxiliary representation To provide supplementary information to assist the right view Model the probability distribution of .
[0122] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the distribution Perform parametric modeling as follows:
[0123]
[0124] in are the parameters corresponding to the logistic mixed model.
[0125] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, using a neural network to estimate As shown below:
[0126]
[0127] Among them, H R (·) denotes an estimator implemented by stacked 3×3 convolutional layers, and Indicates channel-by-channel connection.
[0128] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides generating an estimated distribution
[0129] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution of the auxiliary representation of the right view To assist in the right view Encoding and decoding.
[0130] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that constructing an auxiliary representation {x L ,x R}probability dependence.
[0131] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that a stereoscopic image {x L ,x R} interleaving dependencies between .
[0132] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that {x L ,x R} is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3} and {x R,1 ,x R,2 ,x R,3}.
[0133] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that estimating the sub-image {x L,1 ,x R,1} and compress the sub-image {x L,1 ,x R,1}Probability distribution of .
[0134] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that based on the prediction features including intra-view and inter-view information To estimate the sub-image x L,1 The probability distribution of .
[0135] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution of the auxiliary representation of the left view Come to x L,1 Encoding and decoding.
[0136] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, using the decoded x L,1 To provide supplementary information for the sub-image x R,1 Model the probability distribution of .
[0137] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution For sub-image x R,1 Encoding and decoding.
[0138] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that, with the decoded {xL,1 ,x R,1} is used as a condition to estimate the sub-image {x L,2 ,x R,2} and compress the sub-image {x L,2 ,x R,2}Probability distribution of .
[0139] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that based on the prediction feature and {x L,1 ,x R,1} to estimate the sub-image x L,2 The probability distribution of .
[0140] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution of the auxiliary representation of the left view Come to x L,2 Encoding and decoding.
[0141] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, using the decoded x L,2 To provide supplementary information for the sub-image x R,2 Model the probability distribution of .
[0142] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that based on the estimated probability distribution To pair sub-image x R,2 Encoding and decoding.
[0143] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that, with the decoded {x L,1 ,x R,1 ,x L,2 ,x R,2} is used as a condition to estimate the sub-image {x L,3 ,x R,3} and compress the sub-image {x L,3 ,x R,3}Probability distribution of .
[0144] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that based on the prediction feature and {x L,1 ,x R,1 ,x L,2 ,x R,2} to estimate the sub-image x L,3 The probability distribution of .
[0145] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, based on the estimated probability distribution of the auxiliary representation of the left view Come to x L,3 Encoding and decoding.
[0146] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides, using the decoded x L,3 To provide supplementary information for the sub-image x R,3 Model the probability distribution of .
[0147] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that based on the estimated probability distribution To pair sub-image x R,3 Encoding and decoding.
[0148] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the aforementioned aspects.
[0149] A third aspect relates to a non-transitory computer-readable medium, which includes a computer program product for use by a video codec device, the computer program product including computer executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs any one of the methods of the aforementioned aspects.
[0150] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein the method includes any one of the disclosed methods.
[0151] A fifth aspect relates to a method for storing a bitstream of a video, comprising any one of the disclosed methods.
[0152] A sixth aspect relates to a method, device or system described in the present disclosure.
[0153] For clarity, any of the foregoing embodiments may be combined with any or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0154] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0155] For a more complete understanding of the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0156] Figure 1 An exemplary framework for the first exemplary embodiment is shown.
[0157] Figure 2 An exemplary framework for a second exemplary embodiment is shown.
[0158] Figure 3 An exemplary framework for a third exemplary embodiment is shown.
[0159] Figure 4 An exemplary architecture of a joint conditional entropy model is shown.
[0160] Figure 5 is a block diagram illustrating an exemplary video processing system.
[0161] Figure 6 is a block diagram of an exemplary video processing device.
[0162] Figure 7 is a flow chart of an exemplary method for video processing.
[0163] Figure 8 is a block diagram illustrating an exemplary video coding system.
[0164] Fig. 9 is a block diagram illustrating an exemplary encoder.
[0165] Fig.10 is a block diagram illustrating an exemplary decoder.
[0166] Fig.11 is a schematic diagram of an exemplary encoder. DETAILED DESCRIPTION
[0167] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.
[0168] The section headings used in this disclosure are for ease of understanding, and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the techniques described herein are applicable to other video codec protocols and designs.
[0169] 1. Preliminary Discussion
[0170] The present disclosure relates to the field of image compression, and in particular to an end-to-end stereoscopic image compression system and method.
[0171] Stereoscopic image compression is one of the key technologies in the field of digital image processing, and its purpose is to reduce the redundancy in stereoscopic images and compress them into a compact bitstream. In stereoscopic image compression, in addition to spatial redundancy, inter-view redundancy also needs to be addressed. Therefore, several compression methods, including codec-based methods and end-to-end methods, have been reported to improve codec efficiency. The history of traditional stereoscopic image compression can be traced back to the 1980s, and these methods generally adopt disparity compensation prediction (DCP) to reduce inter-view redundancy. Specifically, the exemplary stereoscopic image compression method is based on DCP, in which the left view is compressed independently, followed by block-by-block disparity estimation and compensation to produce inter-view predictions of the right view. For those blocks whose prediction error is less than a threshold, the prediction is assigned as reconstruction, and only the predicted disparity map is compressed and transmitted, while the remaining blocks in the right view are directly encoded. In another example, DCP is used to predict the right view, and a special codec adapted to the prediction residual characteristics is designed. In the example, a wavelet-based stereoscopic image compression scheme is used, in which a binocular compensation / suppression process based on the human visual system is designed to reduce inter-view redundancy. Inspired by the great success of end-to-end single image compression, several end-to-end stereo image compression methods have been proposed. In another example, Deep Stereo Image Compression (DSIC) is used, in which the parameter jump function is designed to share information from the left view with the right view codec branch. In another example, a stereo image compression network based on homography transformation, namely HESIC, is used to perform inter-view prediction using the homography matrix. In another example, an end-to-end stereo image compression method is used via a bidirectional codec, namely BCSIC-Net. The proposed BCSIC-Net consists of a bidirectional context transformation module and a bidirectional conditional entropy model to improve the codec efficiency. For lossless compression of stereo images, an example of an end-to-end lossless stereo image compression method called L3C-stereo is used, in which the left view and the right view are encoded respectively using two codec branches based on L3C. Specifically, the left view image is first compressed independently, and the decoded image is further warped to the right view by the estimated disparity as a conditional prior for the probability estimate in the right view codec branch.
[0172] 2. Technical problems solved by public technical solutions
[0173] The method for lossless stereoscopic image compression has the following problems.
[0174] Many compression methods are designed for lossy compression, where nonlinear transformations and quantization in these methods result in information loss. Therefore, these methods cannot be used to compress stereo images in a lossless manner.
[0175] Many compression methods designed for lossless compression do not fully exploit inter-view correlations. That is, the left view image is ineffective for conditioning on any inter-view priors, leading to suboptimal codec performance.
[0176] Many compression methods designed for lossless compression require multiple auxiliary disparity estimation modules, resulting in complex network architectures.
[0177] Many compression methods designed for lossless compression require multiple training phases.
[0178] 3. List of solutions and implementation examples
[0179] In order to solve the above problems, the following methods are disclosed. The embodiments should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these embodiments can be applied alone or in combination in any way. In order to solve the above problems, one or more of the following methods are disclosed.
[0180] Example 1
[0181] In this example, the end-to-end lossless compression network for stereo images takes as input a stereo image pair {x L ,x R}Compressed into a bit stream L ,b R In one example, the multi-scale codec structure is transformed from {x L ,x R}Derive multi-scale auxiliary representation And establish {x L ,x R}and The hierarchical dependencies between are as follows:
[0182]
[0183] in,
[0184]
[0185] In one example, S is equal to 3. In one example, S is a positive integer. In one example, an autoencoder network is applied to each scale to estimate and
[0186] Example 2
[0187] In one example, the autoencoder network from Example 1 estimates The probability distribution of , including: an encoder, which is based on the non-quantized auxiliary representation of the previous scale To generate a non-quantized auxiliary representation
[0188] In one example, the encoder may consist of N convolutional layers and M activation layers. In one example, the encoder may consist of N residual blocks. In one example, the encoder may be a nonlinear function that converts an input signal to a high-order domain.
[0189] In one example, the scaler quantizer will Quantization as an auxiliary representation of quantization According to the information provided by the autoencoder network at scale s+2 and Compression is performed by entropy coding and decoding. In one example, the quantizer can be implemented by a rounding operation.
[0190] In one example, the predictors jointly estimate the probability distribution and
[0191] Example 3
[0192] In one example, the predictor according to Example 2 jointly estimates based on the intra-view prior and the inter-view prior and The predictor includes an inter-view interaction module and an entropy model. In one example, the inter-view interaction module can be a semi-coupled inter-view interaction module that extracts view sharing information as effective inter-view prior information. In one example, the entropy model can be a joint conditional entropy model that jointly estimates the distribution of the left view and the right view under consideration of the intra-view prior information and the inter-view prior information.
[0193] Example 4
[0194] In one example, the semi-coupled view interaction module according to Example 3 is adapted to obtain the input stereo features {f L ,f R} extracts inter-view prior information and combines the inter-view prior information with {f L ,f R}Merge to generate enhanced stereo features In one example, the semi-coupled inter-view interaction module may include a semi-coupled extraction block to extract view sharing information and generate semi-coupled features. In one example, the semi-coupled inter-view interaction module may include a parallax interaction transformer to extract complementary information from the semi-coupled features and generate inter-view features. In one example, the nonlinear transformation block fuses the input features with the inter-view features and generates enhanced features.
[0195] Example 5
[0196] In one example, the semi-coupled extraction block according to Example 4 extracts from the input stereo features {f L ,g R Extract semi-coupled features that preserve view-shared information while suppressing view-specific information The semi-coupled extraction block includes a multi-stage extraction module, a semi-coupled depth-wise separable convolution and a fusion module. In one example, a multi-stage extraction strategy may be involved to progressively extract view-sharing information.
[0197] In one example, semi-coupled depthwise separable convolution is applied to each stage c to extract semi-coupled features The process can be formulated as:
[0198]
[0199] Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for left and right views respectively.
[0200] In one example, the fusion module fuses the output semi-coupled features from each stage to generate the final semi-coupled features as follows:
[0201]
[0202] Where {G L ,G R} is the aggregation block.
[0203] Example 6
[0204] In one example, the parallax interactive transformer according to Example 4 is based on the stereo feature {f L ,f R}From the semi-coupled feature Extract complementary information and generate inter-view features The parallax interactive transformer may include an interactive structure, the interactive structure according to f L from Extract complementary information and according to f R from Extracting complementary information. The parallax interactive converter may also include a parallax converter that extracts complementary information along the parallax direction. In one example, the parallax converter extracts complementary information according to f L from Extract complementary information.
[0205] The query vector, key vector, and value vector are generated by a linear layer as follows:
[0206] q L =Linear q(f L ),
[0207] The above three vectors are squeezed in the vertical direction as follows:
[0208]
[0209] Where Squeeze(·): It is a squeeze operation.
[0210] A scaled dot-product attention is further performed along the disparity direction (i.e., horizontal direction), followed by the feed-forward network to generate inter-view features Right now,
[0211]
[0212] Where Softmax(·) represents the Softmax operation, d k denotes a scaling factor, and FFN(·) denotes a feed-forward network.
[0213] In the example, the parallax transformer is based on f R from Extract complementary information and generate inter-view features This process is similar to resemblance.
[0214] Example 7
[0215] In one example, the nonlinear transformation block transforms the input features {f L ,f R} and view features Fusion, generating enhanced features The process can be formulated as follows:
[0216]
[0217] where {H L ,H R} represents the nonlinear transformation implemented by the neural network.
[0218] Example 8
[0219] In one example, the joint conditional entropy model of Example 3 jointly estimates the probability distribution of the stereoscopic view, including the probability dependency of the auxiliary representation and the corresponding entropy model; and / or the probability dependency of the stereoscopic image and the corresponding entropy model.
[0220] Example 9
[0221] In one example, the probabilistic dependency of the auxiliary representation according to Example 8 and the corresponding entropy model can be based on the prediction features containing intra-view information and inter-view information To estimate the left view auxiliary representation The probability dependence of the auxiliary representation and the corresponding entropy model can also be based on the estimated probability distribution of the auxiliary representation of the left view To assist in the left view The probabilistic dependency of the auxiliary representation and the corresponding entropy model can also be used to decode the left view auxiliary representation To provide supplementary information to assist the right view The probability dependence of the auxiliary representation and the corresponding entropy model can also be based on the estimated probability distribution of the auxiliary representation of the right view To assist in the right view Encoding and decoding.
[0222] Example 10
[0223] In one example, the probability dependency of the stereoscopic image according to Example 8 and the corresponding entropy model can be constructed by, for example, performing any of the following: L ,x R}Interleaved dependencies between:
[0224] {x L ,x R} is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3} and {x R,1 ,x R,2 ,x R,3}.
[0225] Estimate the sub-image {x L,1 ,x R,1}, and then compress the sub-image {x L,1 ,x R,1}: Based on the prediction features containing intra-view information and inter-view information To estimate the sub-image x L,1 Probability distribution of; estimated probability distribution based on auxiliary representation of left view Come to x L,1 Encode and decode; use the decoded x L,1 To provide supplementary information for the sub-image x R,1 modeling the probability distribution of; and / or based on the estimated probability distribution For sub-image x R,1 Encoding and decoding.
[0226] With decoded {x L,1 ,x R,1} is used as a condition to estimate the sub-image {x L,2 ,x R,2}, and then perform the following operations on the sub-image {x L,2 ,x R,2} probability distribution is compressed: based on the predicted features and {x L,1 ,x R,1}Estimated sub-image x L,2 Probability distribution of; estimated probability distribution based on auxiliary representation of left view x L,2 Encode and decode; use the decoded x L,2 To provide supplementary information for the sub-image x R,2 modeling the probability distribution of; and / or based on the estimated probability distribution For sub-image x R,2 Encoding and decoding.
[0227] With decoded {x L,1 ,x R,1 ,x L,2 ,x R,2} is used as a condition to estimate the sub-image {x L,3 ,x R,3}, and then perform the following operations on the sub-image {x L,3 ,x R,3} probability distribution is compressed: based on the predicted features and {x L,1 ,x R,1 ,x L,2 ,x R,2}Estimated sub-image x L,3 Probability distribution of; estimated probability distribution based on auxiliary representation of left view x L,3 Encode and decode; use the decoded x L,3 To provide supplementary information for the sub-image x R,3 modeling the probability distribution of; and / or based on the estimated probability distribution For sub-image x R,3 Encoding and decoding.
[0228] 5. Exemplary Embodiments
[0229] In order to make the objectives, technical solutions and advantages of the present application more clear, the embodiments of the present disclosure are further described in detail below.
[0230] A first exemplary embodiment will now be described. Figure 1An exemplary framework for a first exemplary embodiment is shown. The embodiments of the present disclosure provide an end-to-end lossless compression network for stereoscopic images, such as Figure 1 As shown, the following steps are included:
[0231] Using multiple encoders from {x L ,x R}Derive multi-scale auxiliary representation As shown below:
[0232]
[0233] in represents the encoder at scale s.
[0234] In {x L ,x R}and A hierarchical dependency is established between them, as shown below:
[0235]
[0236] in,
[0237] and
[0238]
[0239] In embodiment #1, multiple predictors are used to estimate the factor distribution.
[0240] A second exemplary embodiment will now be described. Figure 2 An exemplary framework for the second exemplary embodiment is shown. The embodiment of the present disclosure provides an estimation distribution The predictor of Figure 2 As shown, the following steps are included:
[0241] Using neural networks to provide predictive features The semi-coupled view interaction module (SI2M) is introduced to extract the inter-view information. The process can be formulated as follows:
[0242]
[0243] where P(·) indicates the predictor.
[0244] Use the Joint Conditional Entropy Model (JCEM) to estimate the distribution
[0245] A third exemplary embodiment will now be described. Figure 3 An exemplary framework for a third exemplary embodiment is shown.
[0246] The embodiments of the present disclosure provide a semi-coupled inter-view interaction module to generate enhanced features that contain intra-view information and inter-view information as effective priors, such as Figure 3 As shown, the following steps are included:
[0247] The semi-coupled extraction (SE) block is used to extract the input stereo features {f L ,f R Extract semi-coupled features These semi-coupled features retain the information shared by views while suppressing view-specific information, and include the following steps: Perform a progressive extraction consisting of C stages. The semi-coupled features are extracted using a semi-coupled depthwise separable convolution applied to each stage c. The process can be formulated as:
[0248]
[0249] Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for the left view and the right view, respectively. The fusion module is used to fuse the output semi-coupled features from each stage to generate the final semi-coupled features as follows:
[0250]
[0251] Where {G L ,G R} is an aggregation block implemented by stacked 1×1 convolutional layers.
[0252] Using the parallax interactive transformer based on stereo features {f L ,f R}From the semi-coupled feature Extract complementary information and generate inter-view features The following steps are included: executing the interaction structure, which is based on f L from Extract complementary information and according to f R from Extract complementary information. Use a parallax converter to extract complementary information along the parallax direction. In one example, the parallax converter is based on f L from Extract complementary information. The query vector, key vector, and value vector are generated by a linear layer as follows:
[0253] q L =Linear q (f L ),
[0254] The above three vectors are squeezed in the vertical direction as follows:
[0255]
[0256] Where Squeeze(·): is a squeezing operation. A scaled dot product attention is further performed along the disparity direction (i.e., horizontal direction), followed by the feed-forward network to generate inter-view features Right now,
[0257]
[0258] Where Softmax(·) represents the Softmax operation, d k represents a scaling factor, and FFN(·) represents a feed-forward network. In the example, the disparity transformer is based on f R from Extract complementary information and generate inter-view features This process is similar to resemblance.
[0259] According to Example 4, a nonlinear transformation block is used to transform the input feature {f L ,f R} and view features Fusion, enhanced features The process can be formulated as follows:
[0260]
[0261] where {H L ,H R} represents the nonlinear transformation implemented by stacked 1×1 convolutional layers.
[0262] A fourth exemplary embodiment will now be described. Figure 4 An exemplary architecture of a joint conditional entropy model is shown.
[0263] The embodiments of the present disclosure provide a joint conditional entropy model to jointly estimate the probability distribution of stereoscopic views, such as Figure 4 As shown, the following steps are included:
[0264] Constructing auxiliary representations The probability dependence of , including the following steps: based on the prediction features containing intra-view information and inter-view information To estimate the left view auxiliary representation The probability distribution of includes the following steps: using the logistic mixture model to distribute Perform parametric modeling as follows:
[0265]
[0266] in are the parameters corresponding to the logistic mixture model. A neural network is used to estimate As shown below:
[0267]
[0268] Among them, H L (*) denotes an estimator implemented by stacking 3×3 convolutional layers. Produces the estimated distribution Estimated probability distribution based on auxiliary representation of left view To assist in the left view Encode and decode. Use the decoded left view to assist in the representation To provide supplementary information to assist the right view Use a logistic mixture model to model the probability distribution of Perform parametric modeling as follows:
[0269]
[0270] in are the parameters corresponding to the logistic mixture model. A neural network is used to estimate As shown below:
[0271]
[0272] Among them, H R (·) denotes an estimator implemented by stacked 3×3 convolutional layers, and represents the channel-by-channel connection. Produces the estimated distribution Estimated probability distribution based on right view auxiliary representation To assist in the right view Encoding and decoding.
[0273] Construct auxiliary representation {x L ,x R}, including the following steps: constructing a stereo image {x L ,x R}, including the following steps: L ,x R} is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3} and {x R,1 ,x R,2 ,x R,3}. Estimate the sub-image {x L,1 ,x R,1}, and then for the sub-image {x L,1 ,x R,1}, including: based on the prediction features containing intra-view information and inter-view information To estimate the sub-image x L,1 The estimated probability distribution based on the auxiliary representation of the left view Come to x L,1 Encode and decode. Using the decoded x L,1 To provide supplementary information for the sub-image x R,1 Based on the estimated probability distribution To pair sub-image x R,1 Encode and decode. L,1 ,x R,1} is used as a condition to estimate the sub-image {x L,2 ,x R,2}, and then for the sub-image {x L,2 ,x R,2} is compressed based on the probability distribution of the predicted features. and {x L,1 ,x R,1} to estimate the sub-image x L,2 The estimated probability distribution based on the auxiliary representation of the left view Come to x L,2 Encode and decode. Using the decoded x L,2 To provide supplementary information for the sub-image x R,2 Based on the estimated probability distribution To pair sub-image x R,2 Encode and decode. L,1 ,x R,1 ,x L,2 ,x R,2} is used as a condition to estimate the sub-image {x L,3 ,x R,3}, and then for the sub-image {x L,3 ,x R,3} is compressed based on the probability distribution of the predicted features. and {x L,1 ,x R,1 ,x L,2 ,x 3,2} to estimate the sub-image x L,3 The estimated probability distribution based on the auxiliary representation of the left view Come to x L,3 Encode and decode. Using the decoded x L,3 To provide supplementary information for the sub-image x R,3Based on the estimated probability distribution To pair sub-image x R,3 Encoding and decoding.
[0274] Figure 5 4000 . 4002 . 4004 . 4005 . 4006 . 4007 . 4008 . 4009 . 4010 . 4011 . 4012 . 4013 . 4014 . 4015 . 4016 . 4017 . 4018 . 4019 . 4020 . 4021 . 4022 . 4023 . 4024 . 4025 . 4026 . 4027 . 4028 . 4029 . 4030 . 4031 . 4032 . 4033 . 4034 . 4035 . 4036 . 4037 . 4038 . 4039 . 4100 . 4101 . 4102 . 4103 . 4104 . 4105 . 4106 . 4107 . 4108 . 4109 . 4201 . 4202 . 4203 . 4204 . 4205
[0275] System 4000 may include codec component 4004, which can implement various codecs or coding methods described in this disclosure. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to generate the codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video code conversion technology. The output of codec component 4004 can be stored, or transmitted (as represented by component 4006) via connected communication. Component 4008 can use the bit stream (or codec) representation of the storage or communication of the video received at input 4002 to generate pixel values or displayable video sent to display interface 4010. The process of generating user-visible video from bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at encoders, and corresponding decoding tools or operations opposite to the codec results will be performed by decoders.
[0276] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Display Port, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in the present disclosure may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0277] Figure 64106 is a block diagram of an exemplary video processing device 4100. Device 4100 may be used to implement one or more of the methods described herein. Device 4100 may be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. (Multiple) processors 4102 may be configured to implement one or more methods described in the present disclosure. (Multiple) memories 4104 may be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 may be used to implement some of the techniques described in the present disclosure in hardware circuits. In some embodiments, video processing circuitry 4106 may be at least partially included in processor 4102, such as a graphics coprocessor.
[0278] Figure 7 4200 is a flow chart of an exemplary method 4200 for video processing. The method 4200 includes, at step 4202, determining to apply an end-to-end lossless compression network to convert an input stereo image pair {x L ,x R}Compressed into a bit stream L ,b R At step 4204, conversion between the visual media data and the bitstream is performed based on the end-to-end lossless compression network. Depending on the example, the conversion of step 4204 may include encoding at the encoder or decoding at the decoder.
[0279] It should be noted that method 4200 may be implemented in an apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In such cases, the instructions, after being executed by the processor, cause the processor to perform method 4200. In addition, method 4200 may be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs method 4200.
[0280] Figure 8 43 is a block diagram illustrating an exemplary video codec system 4300 that can use the techniques of the present disclosure. The video codec system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data and may be referred to as a video encoding device. The target device 4320 may decode the encoded video data generated by the source device 4310 and may be referred to as a video decoding device.
[0281] Source device 4310 may include video source 4312, video encoder 4314 and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers and / or computer graphics systems for generating video data, or a combination of such sources. Video data may include one or more pictures. Video encoder 4314 encodes video data from video source 4312 to generate a bit stream. The bit stream may include a bit sequence that forms a codec representation of video data. The bit stream may include a coded picture and associated data. The coded picture is a coded representation of the picture. Associated data may include sequence parameter sets, image parameter sets and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be sent directly to target device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0282] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or may be external to target device 4320, and target device 4320 may be configured to interface with an external display device.
[0283] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or further standards.
[0284] Fig. 9 is a block diagram showing an example of a video encoder 4400, which may be Figure 8 Video encoder 4314 in system 4300 shown. Video encoder 4400 can be configured to perform any or all of the techniques of the present disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0285] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra-frame prediction unit 4406, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413 and an entropy coding unit 4414.
[0286] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In an example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.
[0287] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated but are represented separately in the example of the video encoder 4400 for purposes of explanation.
[0288] The segmentation unit 4401 may segment the picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0289] The mode selection unit 4403 may select one of the coding modes (intra or inter) based on the error result, for example, and provide the resulting intra or inter coded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 may select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0290] In order to perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information of the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block of the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 4413 and decoded samples.
[0291] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0292] In some examples, the motion estimation unit 4404 may perform unidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block of the current video block in the reference picture of list 0 or list 1. The motion estimation unit 4404 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0293] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. The motion estimation unit 4404 may then generate a reference index and a motion vector, the reference index indicating the reference pictures in list 0 and list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0294] In some examples, motion estimation unit 4404 may output a full set of motion information for a decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a full set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal motion information of the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0295] In one example, the motion estimation unit 4404 may indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0296] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0297] As described above, the video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0298] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.
[0299] The residual generating unit 4407 may generate residual data of the current video block by subtracting the predicted video block of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample point components of the samples in the current video block.
[0300] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 4407 may not perform a subtraction operation.
[0301] Transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0302] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0303] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block to be stored in the buffer 4413.
[0304] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0305] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives the data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data, and output a bitstream including the entropy coded data.
[0306] Fig.10 is a block diagram showing an example of a video decoder 4500, which may be Figure 8 Video decoder 4324 in system 4300 shown. Video decoder 4500 can be configured to perform any or all of the techniques of the present disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video decoder 4500. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0307] In the illustrated example, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 may perform a decoding process that is generally reciprocal to the encoding process described with reference to the video encoder 4400.
[0308] The entropy decoding unit 4501 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 may decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 4502 may determine such information by performing AMVP and Merge modes.
[0309] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0310] The motion compensation unit 4502 may calculate interpolated values of sub-integer pixels of the reference block using interpolation filters as used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information and use the interpolation filters to generate a prediction block.
[0311] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the encoded video sequence.
[0312] The intra prediction unit 4503 may form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0313] The reconstruction unit 4506 may add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for presentation on a display device.
[0314] Fig.11 4600 is a schematic diagram of an exemplary encoder 4600. Encoder 4600 is suitable for implementing techniques for VVC. Encoder 4600 includes three loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses a predefined filter, SAO 4604 and ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, where the auxiliary information of the codec signals the offset and filter coefficients. ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool that attempts to capture and repair artifacts produced by previous stages.
[0315] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using a reference picture obtained from a reference picture buffer 4612. The residual block from the inter prediction or intra prediction is fed to a transform (T) component 4614 and a quantization (Q) component 4616 to produce quantized residual transform coefficients, which are fed to an entropy coding component 4618. The entropy coding component 4618 entropy codes and decodes the prediction result and the quantized transform coefficients and sends them to a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before the images are stored in the reference picture buffer 4612 .
[0316] A list of some example preferred solutions is provided next.
[0317] The following solutions show examples of the techniques discussed in this article.
[0318] 1. A method for processing video data (e.g., Figure 7 The method 4200 depicted in FIG. 4 includes: determining (4202) applying an end-to-end lossless compression network to convert an input stereo image pair {x L ,x R}Compressed into a bit stream L ,b R}; and performing (4204) conversion between visual media data and bitstream based on an end-to-end lossless compression network.
[0319] 2. The method according to solution 1, wherein the end-to-end lossless compression network includes a multi-scale codec structure, the multi-scale codec structure is from {x L ,x R}Export multi-scale auxiliary representation And establish {x L ,x R}and The hierarchical dependencies between are as follows: in, and
[0320] 3. A method according to any one of solutions 1-2, wherein the end-to-end lossless compression network comprises an autoencoder network configured to be applied at each scale to estimate and
[0321] 4. A method according to any one of solutions 1-3, wherein the autoencoder network is configured to estimate The probability distribution of , and wherein the autoencoder network comprises: an encoder configured to be based on a non-quantized auxiliary representation of the previous scale To generate a non-quantized auxiliary representation The scaler is configured to Quantization as an auxiliary representation of quantization According to the information provided by the autoencoder network at scale s+2 and compression by entropy coding; and a predictor configured to jointly estimate the probability distribution and
[0322] 5. A method according to any of solutions 1-4, wherein the predictor jointly estimates based on intra-view and inter-view priors and
[0323] 6. A method according to any one of solutions 1-5, wherein the predictor includes an inter-view interaction module and an entropy model, wherein the inter-view interaction module is a semi-coupled inter-view interaction module configured to extract view sharing information as effective inter-view information, and wherein the entropy model is a joint conditional entropy model configured to jointly estimate the distribution of left views and right views taking into account intra-view prior information and inter-view prior information.
[0324] 7. The method according to any one of solutions 1-6, wherein the inter-view interaction module is configured to extract the stereo features {f L ,f R} extracts inter-view prior information and combines the inter-view prior information with {f L ,f R}Merge to produce enhanced three-dimensional features
[0325] 8. A method according to any one of solutions 1-7, wherein the inter-view interaction module includes a semi-coupled extraction block configured to extract view sharing information and generate semi-coupled features, a disparity interaction transformer configured to extract complementary information from the semi-coupled features and generate inter-view features, and a nonlinear transformation block configured to fuse input features with inter-view features and generate enhanced features or a combination thereof.
[0326] 9. The method according to any one of solutions 1-8, wherein the semi-coupled extraction block is configured to extract the input stereo features {f L ,fR Extract semi-coupled features Semi-coupled features View-shared information is preserved while view-specific information is suppressed, where the semi-coupled extraction block includes a multi-stage extraction module, a semi-coupled depth-wise separable convolution, and a fusion module.
[0327] 10. A method according to any one of solutions 1-9, wherein the semi-coupled extraction block adopts a multi-stage extraction strategy to progressively extract view-sharing information, wherein a semi-coupled depth-wise separable convolution is applied to each stage c to extract the semi-coupled features according to the following formula Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for left and right views respectively.
[0328] 11. The method according to any one of solutions 1-10, wherein the semi-coupled extraction block employs a fusion module to fuse the output semi-coupled features from each stage to produce the final semi-coupled features as follows: Where {G L ,G R} is the aggregation block.
[0329] 12. A method according to any one of solutions 1-11, wherein the parallax interactive transformer is configured to be based on the stereo feature {f L ,f R}From the semi-coupled feature Extract complementary information and generate inter-view features The parallax interactive transformer includes an interactive structure, which is based on f L from Extract complementary information and according to f R from Extracting complementary information, wherein the parallax interactive transformer includes a parallax transformer configured to extract complementary information along the parallax direction, wherein the parallax transformer extracts complementary information according to f L from Extract complementary information, where the query vector, key vector, and value vector are produced by a linear layer as follows: L =Linear q (f L ), The query vector, key vector, and value vector are squeezed vertically as follows: in is a squeezing operation, where a scaled dot-product attention is performed along the disparity direction, followed by a feed-forward network that generates inter-view features according to Where Softmax(·) represents the Softmax operation, d k represents a scaling factor, and FFN(·) represents a feed-forward network, and wherein the disparity transformer is configured to be based on f R from Extract complementary information and generate inter-view features
[0330] 13. A method according to any one of solutions 1-12, wherein the nonlinear transformation block is configured to transform the input feature {f L ,f R} and view features Fusion, the enhanced features are obtained according to the following formula where {H L ,H R} represents the nonlinear transformation implemented by the neural network.
[0331] 14. A method according to any of solutions 1-13, wherein the joint conditional entropy model is configured to jointly estimate the probability distribution of the stereoscopic view, including the probability dependencies and corresponding entropy models of the auxiliary representation and the probability dependencies and corresponding entropy models of the stereoscopic image.
[0332] 15. The method according to any one of solutions 1-14, wherein the probabilistic dependency of the auxiliary representation and the corresponding entropy model are configured to: To estimate the left view auxiliary representation Probability distribution of; estimated probability distribution based on auxiliary representation of left view To assist in the left view Encode and decode; use the decoded left view as an auxiliary representation To provide supplementary information to assist the right view The probability distribution of is modeled; and the estimated probability distribution based on the auxiliary representation of the right view To assist in the right view Encoding and decoding.
[0333] 16. The method according to any one of solutions 1-15, wherein the probabilistic dependency of the stereoscopic images and the corresponding entropy model are configured to construct the stereoscopic images {x L ,x R}: Interleaving dependencies between {x L ,x R} is divided into sub-images along the channel {xL,1 ,x L,2 ,x L,3} and {x R,1 ,x R,2 ,x R,3}; Estimate sub-image {x L,1 ,x R,1}, and then perform the following operations to transform the sub-image {x L,1 ,x R,1} is compressed based on the prediction features containing intra-view information and inter-view information To estimate the sub-image x L,1 Probability distribution of; estimated probability distribution based on auxiliary representation of left view Come to x L,1 Encode and decode; use the decoded x L,1 To provide supplementary information for the sub-image x R,1 Modeling the probability distribution of ; and based on the estimated probability distribution To pair sub-image x R,1 Encode and decode; decoded {x L,1 ,x R,1} is used as a condition to estimate the sub-image {x L,2 ,x R,2}, and then perform the following operations on the sub-image {x L,2 ,x R,2} probability distribution is compressed: based on the predicted features and {x L,1 ,x R,1} to estimate the sub-image x L,2 Probability distribution of; estimated probability distribution based on auxiliary representation of left view Come to x L,2 Encode and decode; use the decoded x L,2 To provide supplementary information for the sub-image x R,2 Modeling the probability distribution of ; and based on the estimated probability distribution To pair sub-image x R,2 Encode and decode; and decode {x L,1 ,x R,1 ,x L,2 ,x R,2} is used as a condition to estimate the sub-image {x L,3 ,x R,3}, and then perform the following operations on the sub-image {x L,3 ,x R,3} probability distribution is compressed: based on the predicted features and {x L,1 ,x R,1 ,xL,2 ,x R,2} to estimate the sub-image x L,3 Probability distribution of; estimated probability distribution based on auxiliary representation of left view Come to x L,3 Encode and decode; use the decoded x L,3 To provide supplementary information for the sub-image x R,3 Modeling the probability distribution of ; and based on the estimated probability distribution To pair sub-image x R,3 Encoding and decoding.
[0334] 17. An apparatus for processing video data comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-16.
[0335] 18. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer executable instructions are executed by a processor, the video codec device performs the method of any one of solutions 1-16.
[0336] 19. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein the method comprises: determining to apply an end-to-end lossless compression network to convert an input stereoscopic image pair {x L ,x R}Compressed into a bit stream L ,b R}; and generating a bitstream based on the determination.
[0337] 20. A method for storing a bitstream of a video, comprising: determining to apply an end-to-end lossless compression network to convert an input stereo image pair {x L ,x R}Compressed into a bit stream L ,b R}; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0338] 21. A method, apparatus or system as described in the present disclosure.
[0339] In the solution described herein, an encoder may comply with the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder may parse syntax elements in the codec representation using the format rules, understand the presence and absence of syntax elements according to the format rules, and generate decoded video.
[0340] In the present disclosure, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, or vice versa. As defined by the syntax, the bitstream representation of the current video block may correspond, for example, to bits that are co-located or dispersed at different positions within the bitstream. For example, a macroblock may be encoded based on a transformed and encoded error residual value, and also using bits in a header and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not exist, as described in the above solution. Similarly, the encoder may determine whether to include or not include certain syntax fields, and generate a codec representation accordingly by including or excluding syntax fields in the codec representation.
[0341] The disclosed and other solutions, examples, embodiments, modules and functional operations described in the present disclosure may be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the structures disclosed in the present disclosure and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material combination that implements a machine-readable propagation signal, or a combination of one or more of them. The term "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. In addition to hardware, the device may include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagation signal is an artificially generated signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0342] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing portions of one or more modules, subroutines, or code). A computer program may be deployed to execute on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0343] The processes and logic flows described in the present disclosure may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by special purpose logic circuits, and the apparatus may also be implemented as special purpose logic circuits, such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs).
[0344] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be coupled in an operational manner to receive data from one or more mass storage devices or to transmit data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read-only memory (CD ROM) and digital versatile disc read-only memory (DVD-ROM) disks. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit.
[0345] Although the present disclosure includes many details, these should not be understood as limitations on any subject matter or the scope that may be claimed, but rather as descriptions of features peculiar to specific embodiments of specific technologies. Certain features described in the context of separate embodiments of the present disclosure may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may involve a sub-combination or a variant of a sub-combination.
[0346] Similarly, although operations are depicted in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this disclosure should not be understood as requiring such separation in all embodiments.
[0347] Only a few implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this disclosure.
[0348] A first component is directly coupled to a second component when there are no intermediate components other than a line, trace, or another medium between the first and second components. A first component is indirectly coupled to a second component when there are intermediate components other than a line, trace, or another medium between the first and second components. The term "coupled" and variations thereof include both direct and indirect couplings. Unless otherwise indicated, use of the term "about" is intended to include a range of ±10% of the subsequent numerical value.
[0349] Although several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are considered to be illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0350] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and shown as discrete or independent in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as coupled may be directly connected, or may be indirectly coupled or communicated electrically, mechanically, or otherwise through some interface, device, or intermediate component. Other examples of changes, substitutions, and alterations may be determined by those skilled in the art, and these changes, substitutions, and alterations may be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: Determine to apply an end-to-end lossless compression network to transform the input stereo image pair {x L ,x R }Compressed into a bit stream L ,b R }, wherein x represents an input stereoscopic image, b represents one of the bit streams, L represents left, and R represents right; as well as Conversion between visual media data and the bitstream is performed based on the end-to-end lossless compression network.
2. The method according to claim 1, wherein: The end-to-end lossless compression network includes a multi-scale codec structure, wherein the multi-scale codec structure is from {x L ,x R }Derive multi-scale auxiliary representation And establish {x L ,x R }and The hierarchical dependencies between are as follows: in, and Where p represents the probability distribution, and where S represents the scale.
3. The method according to claim 1 or 2, wherein S is equal to 3.
4. The method according to claim 1 or 2, wherein S is a positive integer.
5. The method according to any one of claims 1 to 4, wherein: The end-to-end lossless compression network includes an autoencoder network configured to apply each scale to estimate and 6. The method according to any one of claims 1 to 5, wherein: The autoencoder network is configured to estimate The probability distribution of , and wherein the autoencoder network includes an encoder configured to be based on a non-quantized auxiliary representation of the previous scale To generate a non-quantized auxiliary representation 7. The method according to claim 6, wherein: The encoder includes N convolutional layers and M activation layers, where N and M are each positive integers.
8. The method according to claim 6 or 7, wherein: The encoder includes N residual blocks.
9. The method according to any one of claims 6 to 8, wherein: The encoder includes a non-linear function configured to convert an input signal to a higher order domain.
10. The method according to any one of claims 1 to 9, wherein: The autoencoder network includes a scaler quantizer configured to Quantization as an auxiliary representation of quantization According to the autoencoder network at scale s+2 and Compression is done via entropy codec.
11. The method according to claim 10, wherein: The scalar quantizer is configured to perform quantization using a rounding operation.
12. The method according to any one of claims 1 to 5, wherein: The autoencoder network includes a predictor configured to jointly estimate a probability distribution and 13. The method of claim 12, wherein the predictor is configured to jointly estimate the probability distribution based on an intra-view prior and an inter-view prior and 14. The method according to claim 13, wherein: The predictor includes an inter-view interaction module and an entropy model.
15. The method according to claim 14, wherein: The inter-view interaction module includes a semi-coupled inter-view interaction module, and the semi-coupled inter-view interaction module is configured to extract view sharing information as valid inter-view information.
16. The method according to claim 14 or 15, wherein: The entropy model comprises a joint conditional entropy model configured to jointly estimate distributions of left views and right views according to the intra-view prior information and the inter-view prior information.
17. The method according to claim 15 or 16, wherein: The semi-coupled view interaction module is configured to extract the stereo features {f L ,f R } extract the inter-view prior information, and by combining the inter-view prior information with {f L ,f R }Merge to produce enhanced three-dimensional features 18. The method according to claim 17, wherein: The inter-view interaction module includes a semi-coupled extraction block configured to extract view sharing information and generate a semi-coupled feature.
19. The method according to claim 17 or 18, wherein: The inter-view interaction module includes a disparity interaction transformer configured to extract complementary information from the semi-coupled features and generate inter-view features.
20. The method according to any one of claims 17 to 19, wherein: The inter-view interaction module includes a non-linear transformation block configured to fuse input features with inter-view features and generate enhanced features, or a combination thereof.
21. The method according to any one of claims 18 to 20, wherein: The semi-coupled extraction block is configured to extract the input stereo features {f L ,f R Extract semi-coupled features that preserve view-shared information while suppressing view-specific information And wherein the semi-coupled extraction block includes a multi-stage extraction module, a semi-coupled depth-wise separable convolution and a fusion module.
22. The method according to claim 21, wherein: The semi-coupled extraction block adopts a multi-stage extraction strategy to progressively extract the view sharing information.
23. The method according to claim 22, wherein: The semi-coupled depthwise separable convolution is applied to each stage c to extract the semi-coupled features according to Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for the left view and right view, respectively.
24. The method according to claim 21 or 22, wherein: The semi-coupled extraction block adopts a fusion module to fuse the output semi-coupled features from each stage to generate the final semi-coupled features as shown below: Where {G L ,G R } is the aggregation block.
25. The method according to any one of claims 19 to 24, wherein: The parallax interactive transformer is configured to be based on the stereo feature {f L ,f R }From the semi-coupled feature Extract complementary information and generate inter-view features 26. The method according to claim 25, wherein: The parallax interactive converter includes an interactive structure, wherein the interactive structure is based on f L from Extract complementary information and according to f R from Extract complementary information.
27. The method according to claim 25 or 26, wherein: The parallax interactive transformer includes a parallax transformer configured to extract complementary information along a parallax direction.
28. The method according to any one of claims 25 to 27, wherein: The parallax converter is based on f L from Extract complementary information.
29. The method according to claim 28, wherein: The query vector, key vector, and value vector are generated by a linear layer as follows: L =Linear q (f L ), 30. The method of claim 29, wherein: The squeeze operation is configured to squeeze the query vector, the key vector, and the value vector along a vertical direction as follows: in Including the squeeze operation.
31. The method according to claim 30, wherein: A scaled dot-product attention is performed along the disparity direction, and then the feed-forward network generates the inter-view features according to the following formula Where Softmax(·) represents the normalized exponential function Softmax operation, d k represents a scaling factor, and FFN(·) represents the feed-forward network.
32. The method according to any one of claims 25 to 27, wherein: The parallax converter is configured to R from Extract complementary information and generate inter-view features or 33. The method according to any one of claims 20 to 32, wherein: The nonlinear transformation block is configured to transform the input features {f L ,f R } and view features Fusion, generating enhanced features according to the following formula where {H L ,H R } represents the nonlinear transformation implemented by the neural network.
34. The method according to any one of claims 16 to 33, wherein: The joint conditional entropy model is configured to jointly estimate a probability distribution of stereo views.
35. The method of claim 34, wherein: The probability distribution of the stereoscopic views includes the probability dependency of the auxiliary representation and the corresponding entropy model.
36. The method of claim 34, wherein: The probability distribution of the stereoscopic view includes the probability dependency of the stereoscopic images and the corresponding entropy model.
37. The method according to any one of claims 34 to 36, wherein: The probabilistic dependency of the auxiliary representation and the corresponding entropy model are configured to predict features based on To estimate the left view auxiliary representation The probability distribution of the predicted features Contains intra-view information and inter-view information.
38. The method according to any one of claims 34 to 37, wherein: The probability dependence of the auxiliary representation and the corresponding entropy model are configured to be based on the estimated probability distribution of the auxiliary representation of the left view To assist in the left view Encoding and decoding.
39. The method according to any one of claims 34 to 38, wherein: The probabilistic dependencies of the auxiliary representation and the corresponding entropy model are configured to use the decoded left view auxiliary representation To provide supplementary information to assist the right view Model the probability distribution of .
40. The method according to any one of claims 34 to 39, wherein: The probability dependence of the auxiliary representation and the corresponding entropy model are configured to be based on the estimated probability distribution of the auxiliary representation of the right view To assist in the right view Encoding and decoding.
41. The method according to any one of claims 36 to 40, wherein: The probabilistic dependency of the stereo image and the corresponding entropy model are configured to construct a stereo image {x L ,x R } interleaving dependencies between .
42. The method according to claim 41, wherein: Constructing the interlaced dependency between the stereo images includes converting {x L ,x R } is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3 } and {x R,1 ,x R,2 ,x R,3 }.
43. The method according to claim 41 or 42, wherein: Constructing the interlaced dependency between the stereo images includes estimating the sub-images {x L,1 ,x R,1 }, and compress the sub-image {x L,1 ,x R,1 }Probability distribution of .
44. The method of claim 43, wherein: The probability distribution of compressing the sub-image includes prediction features based on intra-view and inter-view information To estimate the sub-image x L,1 The probability distribution of .
45. The method of claim 44, wherein: The probability distribution of compressing the sub-image includes an estimated probability distribution based on the auxiliary representation of the left view Come to x L,1 Encoding and decoding.
46. The method according to claim 44 or 45, wherein: Compressing the probability distribution of the sub-image comprises using the decoded x L,1 To provide supplementary information for the sub-image x R,1 Model the probability distribution of .
47. The method according to any one of claims 44 to 46, wherein: Compressing the probability distribution of the sub-image includes based on the estimated probability distribution To the sub-image x R,1 Encoding and decoding.
48. The method according to any one of claims 41 to 47, wherein: Constructing the interlaced dependency between the stereo images includes decoding {x L,1 ,x R,1 } is used as a condition to estimate the sub-image {x L,2 ,x R,2 }, and compress the sub-image {x L,2 ,x R,2 }Probability distribution of .
49. The method of claim 48, wherein: Estimating the probability distribution of the sub-image includes based on the predicted features and {x L,1 ,x R,1 } to estimate the sub-image x L,2 The probability distribution of .
50. The method of claim 49, wherein: Estimating the probability distribution of the sub-image includes estimating the probability distribution based on the auxiliary representation of the left view Come to x L,2 Encoding and decoding.
51. The method of claim 49 or 50, wherein: Estimating the probability distribution of the sub-image comprises using the decoded x L,2 To provide supplementary information for the sub-image x R,2 Model the probability distribution of .
52. The method according to any one of claims 49 to 51, wherein: Estimating the probability distribution of the sub-image includes based on the estimated probability distribution To the sub-image x R,2 Encoding and decoding.
53. The method according to any one of claims 41 to 52, wherein: Constructing the interlaced dependency between the stereo images includes decoding {x L,1 ,x R,1 ,x L,2 ,x R,2 } is used as a condition to estimate the sub-image {x L,3 ,x R,3 }, and compress the sub-image {x L,3 ,x R,3 }Probability distribution of .
54. The method of claim 53, wherein: Estimating the probability distribution of the sub-image includes based on the predicted features and {x L,1 ,x R,1 ,x L,2 ,x R,2 } to estimate the sub-image x L,3 The probability distribution of .
55. The method of claim 54, wherein: Estimating the probability distribution of the sub-image includes estimating the probability distribution based on the auxiliary representation of the left view Come to x L,3 Encoding and decoding.
56. The method of claim 54 or 55, wherein: Estimating the probability distribution of the sub-image comprises using the decoded x L,3 To provide supplementary information for the sub-image x R,3 Model the probability distribution of .
57. The method according to any one of claims 54 to 56, wherein: Estimating the probability distribution of the sub-image includes based on the estimated probability distribution To the sub-image x R,3 Encoding and decoding.
58. The method of claim 1, wherein: The method also includes using a plurality of encoders to L ,x R }Derive multi-scale auxiliary representation As shown below: in represents the encoder at scale s.
59. The method of claim 58, wherein: The method further comprises: L ,x R }and A hierarchical dependency is established between them, as shown below: in, 60. The method of claim 59, wherein: The method also includes estimating the factor distribution using a plurality of predictors.
61. The method of claim 1, wherein: The method also includes providing a predictor that uses a neural network to estimate the distribution To provide predictive features The semi-coupled view interaction module SI is introduced 2 M is used to extract inter-view information, and the process is formulated as follows: where P(·) indicates the predictor.
62. The method of claim 61, wherein: The method also includes estimating the distribution using the joint conditional entropy model JCEM 63. The method of claim 1, wherein: The method further includes providing a semi-coupled inter-view interaction module to extract the input stereo features {f L ,f R Extract semi-coupled features To generate enhanced features, the semi-coupled features The view-shared information is preserved while the view-specific information is suppressed, and the enhanced features contain both intra-view information and inter-view information as effective priors.
64. The method of claim 63, wherein: The method further comprises performing a progressive extraction comprising C stages.
65. The method of claim 63 or 64, wherein: The method further comprises extracting semi-coupled features using a semi-coupled depthwise separable convolution applied to each stage c The process is formulated as: Where * is the depth-wise convolution operation, is the view-shared convolution kernel, and are view-specific convolution kernels for the left view and right view, respectively.
66. The method according to any one of claims 63 to 65, wherein: The method also includes fusing the output semi-coupled features from each stage using a fusion module to generate a final semi-coupled feature as shown below: Where {G L ,G R } is an aggregation block implemented by stacked 1×1 convolutional layers.
67. The method according to any one of claims 63 to 66, wherein: The method further comprises using a parallax interactive transformer based on stereo features {f L ,f R }From the semi-coupled feature Extract complementary information and generate inter-view features 68. The method of claim 67 further comprising executing an interaction structure, the interaction structure being based on from Extract complementary information and according to f R from Extract complementary information.
69. The method of claim 68, further comprising utilizing a disparity transformer that extracts complementary information along the disparity direction.
70. The method of claim 69, wherein: The parallax converter is based on f L from Extract complementary information.
71. The method of claim 70, wherein: The query vector, the key vector, and the value vector are generated by the linear layer as follows: q L =Linear q (f L ), 72. The method of claim 71, wherein: The query vector, the key vector, and the value vector are squeezed along the vertical direction according to: in is the squeezing operation.
73. The method according to any one of claims 70 to 72, wherein: The scaled dot product attention is further performed along the disparity direction, the disparity direction being the horizontal direction, and then the inter-view features are generated by the feed-forward network As shown below: Where Softmax(·) represents the normalized exponential function Softmax operation, d k denotes the scaling factor, and FFN(·) denotes the feedforward network.
74. The method of claim 69, wherein: The parallax converter is based on f R from Extract complementary information and generate inter-view features or 75. The method of any one of claims 63-74, wherein: The method further comprises using a nonlinear transformation block, wherein the nonlinear transformation block transforms the input feature {f L ,f R } and view features Fusion, generating enhanced features This process is formulated as follows: where {H L ,H R } represents the nonlinear transformation implemented by stacked 1×1 convolutional layers.
76. The method of claim 1, wherein: The method also includes providing a joint conditional entropy model to jointly estimate the probability distribution of the stereoscopic views, thereby constructing an auxiliary representation The probability dependence of .
77. The method of claim 76, wherein: The method further includes predicting features based on intra-view and inter-view information To estimate the left view auxiliary representation The probability distribution of .
78. The method of claim 77, further comprising using a logistic mixture model to Perform parametric modeling as follows: in are the parameters corresponding to the logistic mixed model.
79. The method of claim 77 or 78, further comprising using a neural network to estimate As shown below: Among them, H L (·) denotes an estimator implemented by stacked 3×3 convolutional layers.
80. The method of any one of claims 77-79, further comprising generating an estimated distribution 81. The method of any one of claims 77-80, wherein: The method further comprises an estimated probability distribution based on the auxiliary representation of the left view To assist in the left view Encoding and decoding.
82. The method of any one of claims 77-81, wherein: The method further comprises using the decoded left view auxiliary representation To provide supplementary information to assist the right view Model the probability distribution of .
83. The method of claim 82, wherein: The method also includes using the logistic mixture model to Perform parametric modeling as follows: in are the parameters corresponding to the logistic mixed model.
84. The method of any one of claims 77-83, wherein: The method also includes using a neural network to estimate As shown below: Among them, H R (·) denotes the estimator implemented by stacked 3×3 convolutional layers, and Indicates channel-by-channel connection.
85. The method of any one of claims 77-84, wherein: The method also includes generating an estimated distribution 86. The method of any one of claims 77-85, wherein: The method further includes an estimated probability distribution based on the auxiliary representation of the right view To assist in the right view Encoding and decoding.
87. The method of any one of claims 77-86, wherein: The method further comprises constructing an auxiliary representation {x L ,x R }probability dependence.
88. The method according to claim 87 further comprises constructing a stereoscopic image {x L ,x R } interleaving dependencies between .
89. The method according to claim 88 further comprises: L ,x R } is divided into sub-images along the channel {x L,1 ,x L,2 ,x L,3 } and {x R,1 ,x R,2 ,x R,3 }.
90. The method according to claim 88 or 89, further comprising estimating the sub-image {x L,1 ,x R,1 }, and compress the sub-image {x L,1 ,x R,1 }Probability distribution of .
91. The method according to claim 90 further comprises predicting features based on the information including intra-view and inter-view information. To estimate the sub-image x L,1 The probability distribution of .
92. The method according to claim 91 further comprises an estimated probability distribution based on the auxiliary representation of the left view Come to x L,1 Encoding and decoding.
93. The method according to claim 91 or 92, further comprising using the decoded x L,1 To provide supplementary information for the sub-image x R,1 Model the probability distribution of .
94. The method according to any one of claims 91-93, further comprising: To the sub-image x R,1 Encoding and decoding.
95. The method of any one of claims 88-94, wherein: The method further comprises decoding {x L,1 ,x R,1 } is used as a condition to estimate the sub-image {x L,2 ,x R,2 }, and compress the sub-image {x L,2 ,x R,2 }Probability distribution of .
96. The method of claim 95 further comprising: and {x L,1 ,x R,1 } to estimate the sub-image x L,2 The probability distribution of .
97. The method of claim 96 further comprising: Come to x L,2 Encoding and decoding.
98. The method according to claim 96 or 97, further comprising using the decoded x L,2 To provide supplementary information for the sub-image x R,2 Model the probability distribution of .
99. The method according to any one of claims 96-98, further comprising: To the sub-image x R,2 Encoding and decoding.
100. The method of any one of claims 88-99, wherein: The method further comprises decoding {x L,1 ,x R,1 ,x L,2 ,x R,2 } is used as a condition to estimate the sub-image {x L,3 ,x R,3 }, and compress the sub-image {x L,3 ,x R,3 }Probability distribution of .
101. The method according to claim 100, further comprising: and {x L,1 ,x R,1 ,x L,2 ,x R,2 } to estimate the sub-image x L,3 The probability distribution of .
102. The method according to claim 101 further comprises an estimated probability distribution based on the auxiliary representation of the left view Come to x L,3 Encoding and decoding.
103. The method according to claim 101 or 102, further comprising using the decoded x L,3 To provide supplementary information for the sub-image x R,3 Model the probability distribution of .
104. The method according to claim 101 or 102, further comprising: To the sub-image x R,3 Encoding and decoding.
105. A device for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-104.
106. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when the computer executable instructions are executed by a processor, the video codec device performs the method according to any one of claims 1-104.
107. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein: The method comprises the method according to any one of claims 1-104.
108. A method of storing a bitstream of a video, comprising the method according to any one of claims 1-104.
109. A method, apparatus or system as described in the present disclosure.