Cross-component prediction method and system

By adopting a DNN-based cross component prediction model in video encoding, the problem of insufficient compression performance in intra prediction by traditional methods is solved, and higher video encoding compression quality and efficiency are achieved.

CN116601945BActive Publication Date: 2025-05-13TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280008144.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-20
Filing Date
2022-05-31
Publication Date
2025-05-13
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The cross component linear prediction mode of traditional video encoding methods cannot work effectively in intra prediction, resulting in insufficient compression performance.

Method used

A cross-component prediction model based on deep neural network (DNN) is used to predict the chrominance component using the information provided by the encoder such as brightness components, quantization parameters and block depth, thereby achieving better compression performance.

Benefits of technology

Through the use of DNN models, nonlinear and non-local spatiotemporal correlations can be explored more effectively, and the compression quality and efficiency of video encoding can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116601945B_ABST
    Figure CN116601945B_ABST
Patent Text Reader

Abstract

Systems and methods for cross-component prediction based on deep neural networks (DNNs) are provided. A method includes: inputting a reconstructed luminance block of an image or video into a DNN; and predicting a reconstructed chrominance block of the image or video by the DNN based on the input reconstructed luminance block. Luminance and chrominance reference information and auxiliary information may also be input into the DNN to predict the reconstructed chrominance block. Various inputs may also be generated using processes such as downsampling and transformation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 210,741 filed on June 15, 2021 and U.S. Application No. 17 / 749,730 filed on May 20, 2022, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] Embodiments of the present disclosure relate to methods and systems for cross-component prediction based on DNN. Background Art

[0004] Conventional video coding standards (e.g., H.264 / Advanced Video Coding (H.264 / AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC)) are all designed on a similar (recursive) block-based hybrid prediction / transform framework, where various coding tools (e.g., intra / inter prediction, integer transform, and context-adaptive entropy coding) are carefully crafted to optimize the overall efficiency. Basically, the spatiotemporal pixel neighborhood is used to predict the signal construction to obtain the corresponding residual for subsequent transform, quantization, and entropy coding. On the other hand, the essence of deep neural networks (DNNs) is to extract different levels of spatiotemporal stimulation by analyzing the spatiotemporal information of the receptive fields from neighboring pixels. The ability to explore highly nonlinear and non-local spatiotemporal correlations provides a promising opportunity to greatly improve the compression quality.

[0005] One goal of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression helps reduce the bandwidth or storage space requirements mentioned above, in some cases by two orders of magnitude or more. Lossless and lossy compression and combinations thereof can be employed. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely employed. The amount of distortion allowed depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of TV contribution applications. The achievable compression ratio may reflect that higher allowed / tolerable distortion can produce higher compression ratios. Summary of the invention

[0006] Using information from different components and other auxiliary information, traditional encoders can predict other components to achieve better compression performance. However, compared with DNN-based methods, the cross-component linear prediction mode in intra-frame prediction does not work well. The essence of DNN is to extract different high-level stimuli, and the ability to explore highly nonlinear and non-local correlations provides promising opportunities for high compression quality. The embodiments of the present disclosure use a DNN-based model to process arbitrarily shaped luminance components, reference components, and auxiliary information to predict the reconstructed chrominance components, thereby achieving better compression performance.

[0007] The embodiments of the present disclosure provide a cross component prediction (CCP) model as a new mode in intra prediction by using a deep neural network (DNN). The model uses information provided by the encoder, such as the luminance component, the quantization parameter (QP) value, the block depth, etc., to predict the chrominance component to achieve better compression performance. Previous NN-based intra prediction methods only predict the luminance component, or generate predictions for all three channels without considering the correlation between the chrominance components and other additional information.

[0008] According to an embodiment, a cross-component prediction method performed by at least one processor is provided. The method includes: obtaining a reconstructed luminance block of an image or video; inputting the reconstructed luminance block into a DNN; obtaining a reference component and auxiliary information associated with the reconstructed luminance block; inputting the reference component and auxiliary information into the DNN; and predicting a reconstructed chrominance block of the image or video by the DNN based on the reconstructed luminance block, the reference component and the auxiliary information.

[0009] According to an embodiment, a cross-component prediction system is provided. The system includes: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate according to the instructions of the computer program code. The computer program code includes: an input code configured to enable at least one processor to input a reconstructed luminance block of an image or video, a reference component associated with the reconstructed luminance block, and auxiliary information into a deep neural network DNN implemented by at least one processor; and a prediction code configured to enable at least one processor to predict a reconstructed chrominance block of an image or video based on the reconstructed luminance block, the reference component, and the auxiliary information by the DNN.

[0010] According to an embodiment, a non-transitory computer-readable medium storing computer code is provided. The computer code is configured to, when executed by at least one processor, cause the at least one processor to: implement a DNN; input a reconstructed luminance block of an image or video, a reference component associated with the reconstructed luminance block, and auxiliary information into the DNN; and predict a reconstructed chrominance block of the image or video based on the input reconstructed luminance block, the reference component, and the auxiliary information by the DNN. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0012] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment;

[0013] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to one embodiment;

[0014] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment;

[0015] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment;

[0016] Figure 5 is a schematic diagram of a simplified block diagram of an input generation process according to one embodiment;

[0017] Figure 6 is a schematic diagram of a simplified block diagram of a cross-component prediction process according to one embodiment;

[0018] Figure 7 is a block diagram of computer code according to an embodiment;

[0019] Figure 8 is a schematic diagram of a computer system suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 110, 120 interconnected via a network 150. For unidirectional transmission of data, the first terminal 110 may encode video data at a local location for transmission to another terminal 120 via the network 150. The second terminal 120 may receive the encoded video data of the other terminal from the network 150, decode the encoded data and display the recovered video data. Unidirectional data transmission is common in media service applications and the like.

[0021] Figure 1 A second pair of terminals 130, 140 is shown, provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal 130, 140 may encode video data captured at a local location for transmission to the other terminal via the network 150. Each terminal 130, 140 may also receive encoded video data transmitted by the other terminal, may decode the encoded data, and may display the recovered video data on a local display device.

[0022] exist Figure 1 In the embodiment of the present invention, terminals 110-140 can be shown as servers, personal computers and smart phones and / or any other type of terminals. For example, terminals 110-140 can be laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network 150 represents any number of networks that transmit encoded video data between terminals 110-140, including, for example, wired and / or wireless communication networks. Communication network 150 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet. For the purpose of current discussion, the architecture and topology of network 150 may be unimportant to the operation of the present disclosure unless explained below.

[0023] As examples of applications of the disclosed subject matter, Figure 2 The placement of the video encoder and decoder in a streaming environment is shown. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., etc.

[0024] like Figure 2 As shown, the streaming system 200 may include a capture subsystem 213, which may include a video source 201 and an encoder 203. The video source 201 may be, for example, a digital camera, and may be configured to create an uncompressed video sample stream 202. The uncompressed video sample stream 202 may provide a high data volume compared to an encoded video bitstream, and may be processed by an encoder 203 coupled to the video source 201. The encoder 203 may include hardware, software, or a combination thereof to implement or implement various aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may include a lower data volume compared to a sample stream, and may be stored on a streaming server 205 for future use. One or more streaming clients 206 may access the streaming server 205 to retrieve a video bitstream 209, which may be a copy of the encoded video bitstream 204.

[0025] In an embodiment, the streaming server 205 may also function as a media aware network element (MANE). For example, the streaming server 205 may be configured to prune the encoded video bitstream 204 to customize a potentially different bitstream for one or more streaming clients 206. In an embodiment, the MANE may be provided separately from the streaming server 205 in the streaming system 200.

[0026] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may, for example, decode a video bitstream 209, which is an input copy of the encoded video bitstream 204, and create an output video sample stream 211 that can be presented on a display 212 or another rendering device (not shown). In some streaming systems, the video bitstreams 204, 209 may be encoded according to a specific video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. Under development is a video coding standard informally referred to as Versatile Video Coding (VVC). Embodiments of the present disclosure may be used in the context of VVC.

[0027] Figure 3 An example functional block diagram of a video decoder 210 connected to a display 212 is shown according to an embodiment of the present disclosure.

[0028] The video decoder 210 may include a channel 312, a receiver 310, a buffer memory 315, an entropy decoder / parser 320, a scaler / inverse transform unit 351, an intra-frame picture prediction unit 352, a motion compensation prediction unit 353, an aggregator 355, a loop filter unit 356, a reference picture memory 357, and a current picture memory. In at least one embodiment, the video decoder 210 may include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder 210 may also be partially or fully embodied in software running on one or more CPUs with associated memory.

[0029] In this and other embodiments, the receiver 310 may receive one or more codec video sequences to be decoded by the decoder 210, one at a time, wherein the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequence may be received from a channel 312, which may be a hardware / software link to a storage device storing the coded video data. The receiver 310 may receive the coded video data and other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective use entities (not shown). The receiver 310 may separate the coded video sequence from the other data. To combat network jitter, a buffer memory 315 may be coupled between the receiver 310 and the entropy decoder / parser 320 (hereinafter referred to as "parser"). When the receiver 310 receives data from a storage / forward device with sufficient bandwidth and controllability or from a synchronous network, the buffer memory 315 may not be needed, or may be small. In order to strive for use on a packet network (e.g., the Internet), a buffer memory 315 may be required, which may be relatively large and may advantageously have an adaptive size.

[0030] The video decoder 210 may include a parser 320 to reconstruct symbols 321 from the entropy coded video sequence. The categories of these symbols include, for example, information for managing the operation of the decoder 210 and potentially controlling a rendering device (e.g., display 212) that may be coupled to the decoder, such as Figure 2 As shown. The control information for the rendering device may be in the form of, for example, supplemental enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser 320 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be based on a video coding technique or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 320 may extract a set of subgroup parameters of at least one pixel subgroup in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser 320 may also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0031] Parser 320 may perform entropy decoding / parsing operations on the video sequence received from buffer memory 315 , thereby creating symbols 321 .

[0032] Depending on the type of coded video pictures or parts thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors, the reconstruction of symbol 321 may involve a number of different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 320. For clarity, such subgroup control information flow between parser 320 and the following multiple units is not described.

[0033] In addition to the functional blocks already mentioned, decoder 210 can be conceptually subdivided into a plurality of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0034] One unit may be a scaler / inverse transform unit 351. The scaler / inverse transform unit 351 may receive quantized transform coefficients and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc., as symbols 321 from the parser 320. The scaler / inverse transform unit 351 may output blocks including sample values, which may be input into the aggregator 355.

[0035] In some cases, the output samples of the scalar / inverse transform unit 351 may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed image, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra picture prediction unit 352. In some cases, the intra picture prediction unit 352 generates a block of the same size and shape as the block being reconstructed, using surrounding already-reconstructed information obtained from the current (partially reconstructed) picture from the current picture memory 358. In some cases, the aggregator 355 adds the prediction information already generated by the intra picture prediction unit 352 to the output sample information provided by the scalar / inverse transform unit 351 on a per-sample basis.

[0036] In other cases, the output samples of the scalar / inverse transform unit 351 may belong to an inter-coded and possibly motion compensated block. In this case, the motion compensated prediction unit 353 may access the reference picture memory 357 to obtain samples for prediction. After the extracted samples are motion compensated according to the symbol 321 associated with the block, these samples may be added by the aggregator 355 to the output of the scalar / inverse transform unit 351 (in this case referred to as residual samples or residual signals) to generate output sample information. The address within the reference picture memory 357 from which the motion compensated prediction unit 353 obtains the predicted samples may be controlled by motion vectors. The motion compensated prediction unit 353 may obtain these motion vectors in the form of symbols 321, which may have, for example, X, Y, and reference picture components. When sub-sampled accurate motion vectors are used, motion compensation may also include interpolation of sample values ​​obtained from the reference picture memory 357, motion vector prediction mechanisms, and the like.

[0037] The output samples of the aggregator 355 may be subjected to various loop filtering techniques in a loop filter unit 356. The video compression techniques may include loop filtering techniques which are controlled by parameters contained in the coded video bitstream and available to the loop filtering unit 356 as symbols 321 from the parser 320, but may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of a coded picture or coded video sequence and to previously reconstructed and loop filtered sample values.

[0038] The output of the loop filter unit 356 may be a sample stream that may be output to a rendering device (eg, the display 212 ) as well as stored in the reference picture memory 357 for use in future inter-picture prediction.

[0039] Once fully reconstructed, certain coded pictures may be used as reference pictures for future predictions. Once a coded picture is fully reconstructed, and the coded picture has been identified as a reference picture (e.g., by parser 320), the current reference picture may become part of reference picture buffer memory 357, and a new current picture memory may be reallocated before starting reconstruction of the next coded picture.

[0040] The video decoder 210 may perform decoding operations according to a predetermined video compression technique, which may be described in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in a sense, conforming to the syntax of the video compression technique or standard, as specified in the video compression technology document or standard, particularly in the profile document therein. In addition, to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the metadata of the hypothetical reference decoder (HRD) specification and HRD buffer management signaled in the encoded video sequence.

[0041] In one embodiment, receiver 310 may receive additional (redundant) data with the encoded video. Additional data may be included as part of the encoded video sequence. Video decoder 210 may use the additional data to correctly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, time, space or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0042] Figure 4 An example functional block diagram of a video encoder 203 associated with a video source 201 according to an embodiment of the present disclosure is shown.

[0043] The video encoder 203 may include, for example, an encoder as a source encoder 430 , an encoding engine 432 , a (local) decoder 433 , a reference picture memory 434 , a predictor 435 , a transmitter 440 , an entropy encoder 445 , a controller 450 , and a channel 460 .

[0044] Encoder 203 may receive video samples from a video source 201 (not part of the encoder), which may capture video images to be encoded by encoder 203 .

[0045] The video source 201 may provide a source video sequence to be encoded by the encoder 203 in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source 201 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 203 may be a camera that captures local image information as a video sequence. Video data may be provided as multiple separate pictures that impart motion when viewed sequentially. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. in use. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0046] According to one embodiment, the video encoder 203 can encode and compress the pictures of the source video sequence into the encoded video sequence 443 in real time or under any other time constraints required by the application. It is a function of the controller 450 to implement the appropriate encoding speed. The controller 450 also controls other functional units as described below and can be coupled to these units in function. For the sake of clarity, the coupling is not described. The parameters set by the controller 450 may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 450 because they may be related to the video encoder 203 optimized for a specific system design.

[0047] Some video encoders operate in what is readily recognizable to those skilled in the art as a "coding loop". As an overly simplified description, the coding loop may consist of the encoding portion of a source encoder 430 (responsible for creating symbols based on the input picture to be encoded and the reference picture) and a (local) decoder 433 embedded in the encoder 203, which reconstructs the symbols to create sample data that the (remote) decoder will also create when the compression between the symbols and the encoded video bitstream is lossless in a particular video compression technique. This reconstructed sample stream may be input to a reference picture memory 434. Since the decoding of the symbol stream results in bit-accurate results that are independent of the decoder location (local or remote), the reference picture memory contents are also bit-accurate between the local encoder and the remote encoder. In other words, when prediction is used during decoding, the prediction portion of the encoder "sees" as reference picture samples exactly the same sample values ​​as the sample values ​​"seen" by the decoder. The basic principles of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) are well known to those skilled in the art.

[0048] The operation of the "local" decoder 433 may be identical to the operation of the "remote" decoder 210, which has been described above in connection with Figure 3 However, since the symbols are available and the encoding / decoding of the symbols of the encoded video sequence by the entropy encoder 445 and the parser 320 can be lossless, the entropy decoding part of the decoder 210 (including the channel 312, the receiver 310, the buffer memory 315 and the parser 320) may not be fully implemented in the local decoder 433.

[0049] It can be observed at this point that, except for the parsing / entropy decoding present in the decoder, any decoder techniques need to be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. The description of the encoder techniques can be simplified because these techniques can be the inverse of the fully described decoder techniques. A more detailed description is only required in certain areas and is provided below.

[0050] As part of its operation, source encoder 430 may perform motion compensated predictive coding, which predictively encodes an input frame with reference to one or more previously encoded frames from a video sequence that are designated as “reference frames.” In this manner, encoding engine 432 encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame.

[0051] The local decoder 433 can decode the encoded video data of the frame that can be designated as the reference frame based on the symbols created by the source encoder 430. The operation of the encoding engine 432 can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local decoder 433 replicates the decoding process that may be performed by the video decoder on the reference frame, and may cause the reconstructed reference frame to be stored in the reference picture memory 434. In this way, the encoder 203 may locally store copies of the reconstructed reference frames that have the same content (absent transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.

[0052] The predictor 435 may perform a prediction search on the encoding engine 432. That is, for a new frame to be encoded, the predictor 435 may search the reference picture memory 434 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor 435 may operate on a sample block-pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor 435, the input picture may have a prediction reference extracted from multiple reference pictures stored in the reference picture memory 434.

[0053] The controller 450 may manage encoding operations of the source encoder 430 , including, for example, setting of parameters and sub-group parameters for encoding video data.

[0054] The outputs of all the aforementioned functional units may undergo entropy coding in an entropy encoder 445. The entropy encoder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0055] The transmitter 440 may buffer the encoded video sequence created by the entropy encoder 445 in preparation for transmission via the communication channel 460, which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter 440 may combine the encoded video data from the source encoder 430 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0056] The controller 450 may manage the operation of the encoder 203. During encoding, the controller 450 may assign a specific encoding picture type to each encoded picture, which may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may generally be designated as an intra picture (I picture), a predicted picture (P picture), or a bidirectionally predicted picture (B picture).

[0057] An intra picture (I picture) may be a picture that is encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are aware of those variants of I pictures and their corresponding applications and features.

[0058] A prediction picture (P picture) may be a picture that is encoded and decoded using intra prediction or inter prediction by predicting sample values ​​of each block using at most one motion vector and a reference index.

[0059] Bidirectional prediction pictures (B pictures) can be pictures that use up to two motion vectors and reference indices to predict the sample values ​​of each block, and are encoded and decoded using intra-frame prediction or inter-frame prediction. Similarly, multi-prediction pictures can use more than two reference pictures and related metadata to reconstruct a single block.

[0060] A source picture may typically be spatially subdivided into a plurality of blocks of samples (e.g., 4×4, 8×8, 4×8, or 16×16 blocks of samples each), and encoded on a block-by-block basis. Blocks may be predictively encoded with reference to other (already encoded) blocks determined by the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively encoded, or may be predictively encoded (spatial prediction or intra prediction) with reference to already encoded blocks of the same picture. Blocks of pixels of a P picture may be non-predictively encoded, either via spatial prediction or via temporal prediction, with reference to one previously encoded reference picture. Blocks of a B picture may be predictively encoded, either via spatial prediction or via temporal prediction, with reference to one or two previously encoded reference pictures.

[0061] The video encoder 203 can perform encoding operations according to a predetermined video encoding technology or standard (e.g., ITU-T Rec. H.265). In its operation, the video encoder 203 can perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video encoding technology or standard being used.

[0062] In one embodiment, transmitter 440 may transmit additional data along with the encoded video. Source encoder 430 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (e.g., redundant pictures and slices), supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, etc.

[0063] The embodiments of the present disclosure provide cross-component prediction based on DNN. Figure 5-6 Example embodiments are described.

[0064] According to an embodiment of the present disclosure, a video compression framework is described as follows. Assume that an input video includes a plurality of image frames equal to the total number of frames in the video. The frame is divided into spatial blocks, and each block can be iteratively divided into smaller blocks. The block contains a luminance component 510y and a chrominance component 520 including chrominance channels 520u and 520t. During the intra-frame prediction process, the luminance component 510y can be predicted first, and then the two chrominance channels 520u and 520t can be predicted later. The prediction of the chrominance channels 520u and 520t can be performed jointly or separately.

[0065] In one embodiment of the present disclosure, the reconstructed chroma component 520 is generated by a DNN-based model in the encoder and decoder, or is generated only in the decoder. Two chroma channels 520u and 520t can be generated together with a single network, or separately with different networks. For each chroma channel, a different network can be used to generate the chroma channel based on the block size. In the cross-component prediction based on DNN, one or more processes including signal processing, spatial or temporal filtering, scaling, weighted averaging, up / down sampling, aggregation, recursive processing with memory, linear system processing, nonlinear system processing, neural network processing, deep learning-based processing, AI processing, pre-trained network processing, machine learning-based processing or a combination thereof can be used as a module in an embodiment of the present disclosure. In order to process the reconstructed chroma component 520, a reconstructed chroma channel (e.g., one of the chroma channels 520u and 520t) can be used to generate another reconstructed chroma channel (e.g., the other of the chroma channels 520u and 520t).

[0066] According to an embodiment of the present disclosure, a cross-component prediction model based on a DNN may be provided, which enhances the compression performance of the reconstructed chroma channels 520u and 520t of the block based on the reconstructed luminance component 510y of the block, the reference component, and other auxiliary information provided by the encoder. According to an embodiment, 4:2:0 may be used to subsample the chroma channels 520u and 520t. Therefore, the chroma channels 520u and 520t may have a lower resolution than the luminance component 510y.

[0067] refer to Figure 5 , the process 500 is described below. The process 500 includes a workflow for generating input samples 580 for training and / or prediction in a general hybrid video coding system according to an embodiment of the present disclosure.

[0068] The reconstructed luma component 510y may be a luma block that is a 2N×2M block, where 2N is the width of the luma block and 2M is the height of the luma block. According to an embodiment, a first luma reference 512y as a 2N×2K block and a second luma reference 514y as a 2K×2M block may also be provided, where 2K represents the number of rows or columns in the luma reference. In order to make the luma size the same as the predicted output size, a downsampling process 591 is applied to the luma component 510y, the first luma reference 512y, and the second luma reference 514y. The downsampling process 530 may be a conventional method (e.g., bicubic and bilinear), or may be an NN-based downsampling method. After downsampling, the luma component 510y may become a downsampled luma component 530y with a block size of N×M, the first luma reference 512y may become a downsampled first luma reference 532y with a block size of N×K, and the second luma reference 514y may become a downsampled second luma reference 534y with a block size of K×M. The downsampled first luma reference 532y and the downsampled second luma reference 534y may be transformed (at step 592) to become a first transformed luma reference 552y and a second transformed luma reference 554y, respectively, that match the size of the downsampled luma component 530y (also referred to as a luma block), and the first transformed luma reference 552y, the second transformed luma reference 554y, and the downsampled luma component 530y may be connected together (at step 592). For example, the transformation may be performed by copying the values ​​of the downsampled first luma reference 532y and the downsampled second luma reference 534y several times until their size is the same as the output block size (e.g., the size of the downsampled luma component 530y).

[0069] In order to predict the chroma component 520, neighboring references of the chroma component 520 (eg, the first chroma reference 522 and the second chroma reference 524) may also be added as optional references for generating better chroma components. Figure 5, the chroma component 520 may be a block of size N×M, which is a reconstructed chroma block that may be generated / predicted in an embodiment of the present disclosure. The chroma component 520 has two chroma channels 520u and 520t, and the two channels 520u and 520t may be used jointly. A first chroma reference 522 and a second chroma reference 524, which may have block sizes of N×K and K×M, respectively, may be obtained (at step 593). According to an embodiment, the first chroma reference 522 and the second chroma reference 524 may be obtained twice to correspond to the two chroma channels 520u and 520t. The first chroma reference 522 and the second chroma reference 524 may be transformed (at step 594) into a first transformed chroma reference 542 and a second transformed chroma reference 544, respectively, that match the N×M size. All image-based information (e.g., downsampled luma component 530y, first transformed luma reference 552y, second transformed luma reference 554y, first transformed chroma reference 542, and second transformed chroma reference 544) may be concatenated together (at step 595) to obtain input samples 580 for training the DNN and / or prediction using the DNN. In addition to luma and chroma components, auxiliary information may be added to the input for training the neural network and / or prediction. For example, QP values ​​and block segmentation depth information may be used to generate a feature map of size N×M and may be concatenated together (at step 595) with the image-based feature map (e.g., downsampled luma component 530y, first transformed luma reference 552y, second transformed luma reference 554y, first transformed chroma reference 542, and second transformed chroma reference 544) to generate input samples 580 for training and / or prediction.

[0070] Reference below Figure 6 The workflow of process 600 in a general hybrid video coding system is described.

[0071] A set of reconstructed luma blocks 610 (also referred to as luma components), auxiliary information 612, adjacent luma references 614 of the luma block 610, and adjacent chroma references 616 of the chroma block to be reconstructed can be used as inputs to the DNN 620 so that the model of the embodiment of the present disclosure can perform training and prediction. The output 630 of the DNN 620 can be a predicted chroma component, and different DNN models or the same DNN model can be used to predict the two chroma channels.

[0072] According to an embodiment, the input of DNN 620 may be a reference Figure 5580 is described. For example, the reconstructed luma block 610 may be the downsampled luma component 530y, the neighboring luma reference 614 may be one or more of the first transformed luma reference 552y and the second transformed luma reference 554y, and the neighboring chroma reference 616 may be one or more of the first transformed chroma reference 542 and the second transformed chroma reference 544 (for one or both of the chroma channels 520u and 520t). According to an embodiment, the auxiliary information may include, for example, a QP value and block partition depth information.

[0073] The combination, connection or order of how to use the reconstructed luma block 610, the auxiliary information 612, the adjacent luma reference 614 and the adjacent chroma reference 616 as inputs may be changed differently. According to an embodiment, based on the decision of the encoding system of an embodiment of the present disclosure, the auxiliary information 612, the adjacent luma reference 614 and / or the adjacent chroma reference 616 may be optional inputs of the DNN 620.

[0074] According to an embodiment, the encoding system of the present disclosure may calculate the reconstruction quality (step 640) by, for example, comparing the output 630 (e.g., predicted chroma components) of the DNN 620 with the original chroma block 660, and comparing one or more chroma blocks (step 650) from other prediction modes with the original chroma block 660. Based on determining that one of the output 630 (e.g., predicted chroma components) and one or more chroma blocks (step 650) from other prediction modes has the highest reconstruction quality (e.g., closest to the original chroma block 660), the encoding system may select such a block (or mode) as the reconstructed chroma block 670.

[0075] According to an embodiment, at least one processor and a memory storing computer program instructions may be provided. When executed by at least one processor, the computer program instructions may implement a system that performs any number of functions described in the present disclosure. For example, referring to Figure 7 , at least one processor can implement system 700. System 700 can include a DNN and at least one model thereof. Computer program instructions can include, for example, DNN code 710, input generation code 720, input code 730, prediction code 740, reconstruction quality code 750, and image acquisition code 760.

[0076] According to an embodiment of the present disclosure, the DNN code 710 may be configured to enable at least one processor to implement a DNN (and its model).

[0077] According to the embodiments of the present disclosure (for example, referring to Figure 5 ), the input generation code 720 may be configured to cause at least one processor to generate input for the DNN. For example, the input generation code 720 may cause the execution of the reference Figure 5Describe the process.

[0078] According to an embodiment of the present disclosure, the input code 730 may be configured to cause at least one processor to input the input into the DNN (e.g., referring to Figure 6 For example, see Figure 6 , the input may include a reconstructed luma block 610 , auxiliary information 612 , a luma reference 614 and / or a chroma reference 616 .

[0079] According to an embodiment of the present disclosure, the prediction code 740 may be configured to enable at least one processor to predict the reconstructed chrominance block (e.g., reference Figure 6 630 shown in the description).

[0080] According to an embodiment of the present disclosure, the reconstruction quality code 750 may be configured to cause at least one processor to calculate the reconstruction quality of a reconstructed chroma block predicted by the DNN and the reconstruction quality of another reconstructed chroma block predicted using a different prediction mode (e.g., reference Figure 6 640 and 650 shown in FIG. 1 ).

[0081] According to an embodiment of the present disclosure, the image acquisition code 760 may be configured to enable at least one processor to obtain an image (e.g., referring to FIG. 1 ) using a reconstructed chroma block predicted by a DNN or another reconstructed chroma block predicted using a different prediction mode. Figure 6 ). For example, the image acquisition code 760 may be configured to cause at least one processor to select one of the reconstructed chroma blocks and the other reconstructed chroma block based on the one with the highest calculated reconstruction quality, and use such a reconstructed chroma block to obtain an image. According to an embodiment, the image acquisition code 760 may be configured to cause at least one processor to obtain an image using the reconstructed chroma block predicted by the DNN without calculating the reconstruction quality and / or for selecting between the reconstructed chroma blocks. According to an embodiment, the reconstructed luminance block may also be used to obtain an image.

[0082] Compared with the existing cross-component prediction method in intra-frame prediction mode, the embodiments of the present disclosure provide multiple benefits. For example, the embodiments of the present disclosure provide a flexible and general framework that adapts to reconstruction blocks of various shapes. In addition, the embodiments of the present disclosure include aspects of utilizing a transformation mechanism with various input information, thereby optimizing the learning ability of the DNN model to improve coding efficiency. In addition, auxiliary information can be used with the DNN to improve the prediction results.

[0083] The techniques of the embodiments of the present disclosure described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Figure 8 A computer system 900 suitable for implementing embodiments of the disclosed subject matter is shown.

[0084] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.

[0085] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, IoT devices, etc.

[0086] Figure 8 The components of the computer system 900 shown in the figure are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependency or requirement on any one component or combination of components shown in the exemplary embodiment of the computer system 900.

[0087] Computer system 900 may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, for example, audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0088] Input human interface devices may include one or more of the following (only one of each is shown): keyboard 901 , mouse 902 , trackpad 903 , touch screen 910 , data gloves, joystick 905 , microphone 906 , scanner 907 , and camera 908 .

[0089] The computer system 900 may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen 910, a data glove, or a joystick 905, but there may also be tactile feedback devices that are not used as input devices). For example, such devices may be audio output devices (e.g., speakers 909, headphones (not shown)), visual output devices (e.g., screens 910, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or more than three-dimensional outputs by means such as stereo output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0090] The computer system 900 may also include human-accessible storage devices and their associated media, such as optical media 921 including CD / DVD ROM / RW 920 with CD / DVD or similar media 921, a thumb drive 922, a removable hard drive or solid state drive 923, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), and the like.

[0091] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0092] The computer system 900 may also include an interface to one or more communication networks. The network may be, for example, wireless, wired, optical. The network may also be local, wide, urban, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks, such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter connected to some universal data port or peripheral bus 949 (e.g., a USB port of the computer system 900); others are typically integrated into the core of the computer system 900 by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smart phone computer system). Using any of these networks, the computer system 900 can communicate with other entities. Such communications may be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., CANbus to certain CANbus devices), or bidirectional, for example, to other computer systems using a local area network or wide area network. Such communications may include communications to a cloud computing environment 955. As described above, certain protocols and protocol stacks may be used on each of these networks and network interfaces.

[0093] The aforementioned human interface devices, human-accessible storage devices, and network interface 954 may be connected to the core 940 of the computer system 900 .

[0094] The core 940 may include one or more central processing units (CPUs) 941, graphics processing units (GPUs) 942, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 943, hardware accelerators 944 for specific tasks, and the like. These devices, along with read-only memory (ROM) 945, random access memory 946, and internal mass storage 947 such as internal non-user accessible hard drives, SSDs, and the like, may be connected via a system bus 948. In some computer systems, the system bus 948 may be accessed in the form of one or more physical plugs to allow expansion of additional CPUs, GPUs, and the like. Peripheral devices may be connected directly to the core's system bus 948, or may be connected via a peripheral bus 949. The architecture of the peripheral bus includes PCI, USB, and the like. A graphics adapter 950 may be included in the core 940.

[0095] The CPU 941, GPU 942, FPGA 943, and accelerator 944 may execute certain instructions, which in combination may constitute the above-mentioned computer code. The computer code may be stored in ROM 945 or RAM 946. Transitional data may also be stored in RAM 946, while permanent data may be stored in, for example, internal mass storage 947. Fast storage and retrieval of any storage device may be achieved by using a cache memory, which may be closely associated with one or more CPUs 941, GPUs 942, mass storage 947, ROM 945, RAM 946, etc.

[0096] The computer readable medium may have computer codes for performing various computer-implemented operations. The media and computer codes may be specially designed and constructed for the purposes of the present disclosure, or may be of a type well known and available to those skilled in the art of computer software.

[0097] As an example and not a limitation, a computer system 900 (particularly a core 940) having an architecture may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such a computer-readable medium may be a medium associated with a user-accessible mass storage as described above and certain memories of the core 940 having non-transitory properties, such as a core internal mass storage 947 or a ROM 945. Software implementing various embodiments of the present disclosure may be stored in such a device and executed by the core 940. Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software may enable the core 940 (particularly the processor therein (including a CPU, GPU, FPGA, etc.)) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in RAM 946 and modifying such a data structure according to a software-defined process. In addition or as an alternative, the computer system may provide functionality as a result of logic (e.g., an accelerator 944) that is hardwired or otherwise contained in a circuit, which may replace the software or operate with the software to perform a specific process or a specific part of a specific process described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits containing logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0098] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. A cross-component prediction method, the method being executed by at least one processor, characterized in that: The method comprises: Obtaining a reconstructed luminance block of an image or video, wherein the reconstructed luminance block is an N×M block; Obtain a first chrominance reference and a second chrominance reference of an adjacent chrominance block associated with the reconstructed luminance block, wherein the first chrominance reference is a block of size N×K, and the second chrominance reference is a block of size K×M; transforming the first chromaticity reference into a first transformed chromaticity reference, and transforming the second chromaticity reference into a second transformed chromaticity reference, wherein the sizes of the first transformed chromaticity reference and the second transformed chromaticity reference are N×M; Inputting the reconstructed luminance block, the first transformed chrominance reference, the second transformed chrominance reference and the auxiliary information into a deep neural network DNN; and The DNN predicts a reconstructed chroma block of the image or the video based on the reconstructed luminance block, the first transformed chroma reference, the second transformed chroma reference and the auxiliary information, wherein the auxiliary information includes at least one of a quantization parameter QP value and block segmentation depth information.

2. The method according to claim 1, characterized in that The input of the DNN network includes the adjacent luminance references of the reconstructed luminance block and the adjacent chrominance references of the reconstructed chrominance block to be predicted, and The predicting the reconstructed chroma block of the image or the video further comprises predicting, by the DNN, the reconstructed chroma block based on the input reconstructed luminance block, the neighboring luminance reference, and the neighboring chroma reference; The adjacent chromaticity references include the first transformed chromaticity reference and the second transformed chromaticity reference.

3. The method according to claim 1, characterized in that It also includes generating a feature map based on the auxiliary information, and connecting the generated feature map with other image-based feature maps for DNN training.

4. The method according to claim 1, characterized in that: Also includes: Generate the input of the DNN, wherein predicting the reconstructed chroma block of the image or the video comprises predicting, by the DNN, the reconstructed chroma block of the image or the video based on the input, Wherein, the input for generating the DNN includes: Reconstructing a luminance block and obtaining adjacent luminance references of the luminance block; downsampling the luma block to obtain the reconstructed luma block as one of the inputs; downsampling the neighboring luma references of the luma block; and transforming the downsampled neighboring luma reference to have the same size as the downsampled luma block, and The input of the DNN includes the downsampled luminance block and the transformed adjacent luminance reference.

5. The method according to claim 4, characterized in that The luminance block is a 2N×2M block, and the adjacent luminance reference includes a 2N×2K first luminance reference block and a 2K×2M second luminance reference block for reference, wherein N, K and M are integers, 2N is the width, 2M is the height, and 2K is the number of rows or columns in the luminance reference.

6. The method according to claim 5, characterized in that The size of the reconstructed luminance block obtained by downsampling the luminance block is N×M, and After down-sampling the adjacent luminance references, the size of the first luminance reference block is N×K, and the size of the second luminance reference block is K×M.

7. A cross-component prediction system, characterized in that: include: at least one memory configured to store computer program code; as well as At least one processor, the processor being configured to access the computer program code and operate according to instructions of the computer program code, wherein the computer program code is used to execute the method described in any one of claims 1 to 6.

8. A non-transitory computer readable medium storing computer code, characterized in that: The computer code is configured to, when executed by at least one processor, cause the at least one processor to perform the method as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Deep intra predictor generating side information

    WO2021069688A1