Video encoding method, apparatus, and medium
By using neural network encoding in the video bitstream and utilizing SEI messages and neural network exchange format, the problem of combining video encoding with neural networks in existing technologies is solved, achieving more efficient video encoding and information transmission.
Patent Information
- Application Number
- CN202180026007.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2021-10-07
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-10-07
AI Technical Summary
Existing video coding technologies are difficult to combine effectively with neural network-based coding, resulting in a lack of compressibility, accuracy, and information loss.
Encoding is performed using a neural network in the video bitstream. The topology information and parameters of the neural network are signaled using supplementary enhancement information (SEI) messages, parameter sets, and metadata container frames. Encoding is performed in conjunction with neural network exchange formats such as NNEF, ONNX, and MPEG NNR.
It improves the compressibility and accuracy of video encoding, ensuring the complete transmission and decoding of neural network information.
Smart Images

Figure CN115398453B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 133,682, filed January 4, 2021, and U.S. Application No. 17 / 476,824, filed September 16, 2021. The entire contents of the prior applications are expressly incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to video coding techniques, for example, video coding techniques involving Supplemental Enhancement Information (SEI) used with neural networks involved in video coding. Background Art
[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / High Efficiency Video Coding (HEVC) standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). In 2015, the two standards organizations formed the Joint Video Exploration Team (JVET) to explore the potential of developing a next-generation video coding standard that would surpass HEVC. In October 2017, they released a joint call for proposals (CfP) for video compression with capabilities exceeding HEVC. As of February 15, 2018, 22 CfP responses had been submitted for Standard Dynamic Range (SDR), 12 for High Dynamic Range (HDR), and 12 for 360 video categories. In April 2018, all received CfP responses were evaluated at the 122MPEG / 10th JVET meeting. As a result of this meeting, JVET officially launched the standardization process for the next generation of video coding, surpassing HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team (JVET). China's Audio Video Coding Standard (AVS) is also under development.
[0005] When neural networks are involved, at least due to the complexity of neural network-based coding, common codecs may not be able to perform the filtering process with them. As a result, there are technical drawbacks, including lack of compressibility, accuracy, and unnecessary discarding of information related to the neural network. Summary of the Invention
[0006] According to an exemplary embodiment, a method and apparatus are included. The apparatus includes a memory for storing computer program code; and one or more processors for accessing the computer program code and operating in accordance with instructions of the computer program code. The computer program code includes: acquisition code for causing the at least one processor to obtain a video bitstream; encoding code for causing the at least one processor to encode the video bitstream at least in part using a neural network; determination code for causing the at least one processor to determine topology information and parameters of the neural network; and signaling code for causing the at least one processor to signal the determined topology information and parameters of the neural network in a plurality of syntax elements associated with the encoded video bitstream.
[0007] According to an exemplary embodiment, the plurality of syntax elements are signaled via one or more of a supplemental enhancement information (SEI) message, a parameter set, and a metadata container box. Further, according to an exemplary embodiment, the neural network includes a plurality of operation nodes, and encoding the video bitstream through the neural network includes: sending input tensor data of the video bitstream to a first operation node of the plurality of operation nodes; processing the input tensor data using any pre-trained constants and variables; and outputting intermediate tensor data, wherein the intermediate tensor data includes a weighted sum of the input tensor data and any trained constants and updated variables. Further, according to an exemplary embodiment, the topology information and parameters are based on encoding the video bitstream by the neural network.
[0008] According to an exemplary embodiment, the signaling of the determined topology information and the parameters includes: providing external link information, storing the determined topology information and the parameters at the external link information. Further, according to an exemplary embodiment, the neural network includes a plurality of operation nodes, wherein encoding the video bitstream by the neural network includes: sending input tensor data of the video bitstream to a first operation node of the plurality of operation nodes; processing the input tensor data using any pre-trained constants and variables; and outputting intermediate tensor data, the intermediate tensor data including a weighted sum of the input tensor data and any trained constants and updated variables. Further, according to an exemplary embodiment, the topology information and parameters are based on encoding the video bitstream by the neural network.
[0009] According to an exemplary embodiment, signaling the determined topology information and the parameters includes explicitly signaling the determined topology information and the parameters via at least one of a Neural Network Exchange Format (NNEF), an Open Neural Network Exchange (ONNX) format, and a Moving Picture Experts Group (MPEG) Neural Network Compression Standard (NNR) format. Furthermore, according to an exemplary embodiment, the neural network includes a plurality of operation nodes, and encoding the video bitstream via the neural network includes sending input tensor data of the video bitstream to a first operation node of the plurality of operation nodes; processing the input tensor data using any pre-trained constants and variables; and outputting intermediate tensor data, wherein the intermediate tensor data includes a weighted sum of the input tensor data and any trained constants and updated variables. Furthermore, according to an exemplary embodiment, the topology information and parameters are based on the encoding of the video bitstream by the neural network. Further, according to an exemplary embodiment, at least one of the Neural Network Exchange Format (NNEF), the Open Neural Network Exchange (ONNX) format, and the Moving Picture Experts Group (MPEG) Neural Network Compression Standard (NNR) format is an MPEG NNR format, and at least one of the parameters is compressed into at least one of a Supplemental Enhancement Information (SEI) message and a data file. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Further features, nature, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.
[0011] Figure 1 is a simplified illustration of a schematic diagram according to an embodiment.
[0012] Figure 2 is a simplified illustration of a schematic diagram according to an embodiment.
[0013] Figure 3is a simplified illustration of a schematic diagram according to an embodiment.
[0014] Figure 4 is a simplified illustration of a schematic diagram according to an embodiment.
[0015] Figure 5 is a simplified illustration of a diagram according to an embodiment.
[0016] Figure 6 is a simplified illustration of a diagram according to an embodiment.
[0017] Figure 7 is a simplified illustration of a diagram according to an embodiment.
[0018] Figure 8 is a simplified illustration of a diagram according to an embodiment.
[0019] Figure 9A is a simplified illustration of a diagram according to an embodiment.
[0020] Figure 9B is a simplified illustration of a diagram according to an embodiment.
[0021] Figure 10 is a simplified illustration of a diagram according to an embodiment.
[0022] Figure 11 is a simplified illustration of a flow chart according to an embodiment.
[0023] Figure 12 is a simplified illustration of a flow chart according to an embodiment.
[0024] Figure 13 is a simplified illustration of a flow chart according to an embodiment.
[0025] Figure 14 is a simplified illustration of a flow chart according to an embodiment.
[0026] Figure 15 is a simplified illustration of a flow chart according to an embodiment.
[0027] Figure 16 is a simplified illustration of a flow chart according to an embodiment.
[0028] Figure 17 is a simplified illustration of a flow chart according to an embodiment.
[0029] Figure 18 is a simplified illustration of an example diagram according to an embodiment. DETAILED DESCRIPTION
[0030] The features of the suggestions discussed below can be used individually or in any combination in any order. In addition, these embodiments can be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.
[0031] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. Communication system 100 may include at least two terminal devices 102 and 103 interconnected via a network 105. For one-way data transmission, a first terminal device 103 may encode video data locally for transmission to another terminal device 102 via network 105. A second terminal device 102 may receive the encoded video data from the other terminal device via network 105, decode the encoded video data, and display the recovered video data. One-way data transmission is common in applications such as media services.
[0032] Figure 1 A second pair of terminal devices 101 and 104 is shown supporting bidirectional transmission of encoded video, such as can occur during a video conference. For bidirectional data transmission, each terminal device 101 and 104 can encode video data captured at a local location for transmission to the other terminal device via network 105. Each terminal device 101 and 104 can also receive encoded video data transmitted by the other terminal device, decode the encoded video data, and display the recovered video data on a local display device.
[0033] exist Figure 1 In the embodiment, terminal devices 101, 102, 103 and 104 may be servers, personal computers and smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network 105 represents any number of networks that transmit encoded video data between terminal devices 101, 102, 103 and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless explained below, the architecture and topology of network 105 may be irrelevant to the operation of the present disclosure.
[0034] As an application example of the disclosed subject matter, Figure 2The placement of the video decoder and encoder in a streaming environment is shown. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0035] A streaming system may include an acquisition subsystem 203, which may include a video source 201, such as a digital camera, that creates, for example, an uncompressed video sample stream 213. Sample stream 213 may be a high-data-volume video sample stream, as compared to an encoded video bitstream, and may be processed by an encoder 202 coupled to camera 201. Encoder 202 may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter, as described in greater detail below. Encoded video bitstream 204 may be a relatively low-data-volume encoded video bitstream, as compared to a sample stream, and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access streaming server 205 to retrieve copies 208 and 206 of encoded video bitstream 204. Client 212 may include a video decoder 211. The video decoder 211 decodes the incoming copy 208 of the encoded video bitstream and produces an output video sample stream 210 that can be presented on a display 209 or another presentation device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video codec / compression standards. Examples of these standards are mentioned above and are further described herein.
[0036] Figure 3 is a functional block diagram of a video decoder 300 according to an embodiment of the present disclosure.
[0037] Receiver 302 may receive one or more encoded video sequences to be decoded by video decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where each encoded video sequence is decoded independently of the other encoded video sequences. The encoded video sequences may be received from channel 301, which may be a hardware / software link to a storage device storing the encoded video data. Receiver 302 may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not shown). Receiver 302 may separate the encoded video sequences from the other data. To mitigate network jitter, a buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter referred to as the "parser"). Buffer memory 303 may not be required, or may be smaller, when receiver 302 receives data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network. Of course, for use on a packet network such as the Internet, a buffer memory 303 may also be required, which may be relatively large and may advantageously have an adaptive size.
[0038] Video decoder 300 may include a parser 304 for reconstructing symbols 313 from an entropy-coded video sequence. These symbols may include information for managing the operation of video decoder 300 and potentially information for controlling a display device, such as display 312, which is not integral to the decoder but may be coupled to the decoder. The control information for the display device may be in the form of a parameter set fragment (not shown) of Supplementary Enhancement Information (SEI) or Video Usability Information (VUI). Parser 304 may parse / entropy decode the received coded video sequence. The coded video sequence may be encoded or decoded according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. Parser 304 may extract, from the coded video sequence, a subgroup parameter set for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0039] Parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer memory 303, thereby creating symbols 313. Parser 304 may receive encoded data and selectively decode specific symbols 313. In addition, parser 304 may determine whether to provide specific symbols 313 to motion compensation prediction unit 306, scaler / inverse transform unit 305, intra prediction unit 307, or loop filter 311.
[0040] Depending on the type of coded video picture or portion thereof (e.g., inter- and intra-pictures, inter- and intra-blocks), as well as other factors, the reconstruction of symbol 313 may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 304. For the sake of brevity, the flow of such subgroup control information between parser 304 and the various units described below is not described.
[0041] In addition to the functional blocks already mentioned, video decoder 300 can be conceptually subdivided into several functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0042] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives quantized transform coefficients as symbols 313 from the parser 304, along with control information including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit 305 may output a block comprising sample values, which may be input to the aggregator 310.
[0043] In some cases, the output samples of the scaler / inverse transform unit 305 may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding reconstructed information extracted from the (partially reconstructed) current picture 309 to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator 310 adds the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 on a per-sample basis.
[0044] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After the extracted samples are motion compensated according to the symbols 313, these samples may be added to the output of the scaler / inverse transform unit (in this case referred to as residual samples or residual signal) by the aggregator 310, thereby generating output sample information. The retrieval of the prediction samples by the motion compensated prediction unit from the address in the reference picture memory may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit 306 in the form of the symbols 313, which may include, for example, X, Y and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (457) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0045] The output samples of aggregator 310 may be employed by various loop filtering techniques in loop filter unit 311. The video compression techniques may include in-loop filter techniques that are controlled by parameters included in the coded video bitstream and available to loop filter unit 311 as symbols 313 from parser 304. However, the video compression techniques may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of a coded video sequence, as well as to previously reconstructed and loop-filtered sample values.
[0046] The output of the loop filter unit 311 may be a sample stream that may be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.
[0047] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 304), current picture 309 can become part of reference picture memory 308 and new current picture memory can be reallocated before starting reconstruction of a subsequent coded picture.
[0048] The video decoder 300 may perform decoding operations according to a predetermined video compression technique, such as that documented in the ITU-T H.265 standard. An encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard as specified in the video compression technology document or standard, particularly in a profile. Compliance also requires that the complexity of the encoded video sequence be within the limits imposed by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, and the like. In some cases, the limits imposed by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata regarding HRD buffer management signaled in the encoded video sequence.
[0049] In an embodiment, receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0050] Figure 4 is a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0051] The video encoder 400 may receive video samples from a video source 401 (not part of the encoder), which may capture video images to be encoded by the video encoder 400 .
[0052] Video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by video encoder 400. The digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, video source 401 may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that, when viewed sequentially, are imparted with motion. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples can be readily understood by those skilled in the art. The following description focuses on samples.
[0053] According to an embodiment, the video encoder 400 may encode and compress pictures of a source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Enforcing an appropriate encoding speed is a function of the controller 402. The controller controls and is functionally coupled to other functional units, such as those described below. For the sake of simplicity, couplings are not shown in the figure. Parameters set by the controller may include rate control-related parameters (e.g., picture skipping, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller 402 may be readily identified by one skilled in the art, as they may pertain to video encoder 400 optimized for a particular system design.
[0054] Some video encoders operate in a manner readily recognizable to those skilled in the art as a "coding loop." For simplicity's sake, the coding loop can include the encoding portion of encoder 402 (hereinafter referred to as the "source encoder," which is responsible for creating symbols based on the input picture to be encoded and the reference pictures), and a (local) decoder 406 embedded within video encoder 400. The "local" decoder 406 reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (since, in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream is input to a reference picture memory 405. Because decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory also correspond bit-accurately between the local and remote encoders. In other words, the reference picture samples "seen" by the encoder's prediction portion are identical to the sample values that the decoder will "see" when using prediction during decoding. This fundamental principle of reference picture synchronization (and the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0055] The operation of the "local" decoder 406 can be combined with the Figure 3 The same is true of the "remote" decoder 300 described in detail. However, additional brief reference is made to Figure 4 , when symbols are available and the entropy encoder 408 and the parser 304 are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder 300 (including the channel 301, the receiver 302, the buffer memory 303 and the parser 304) may not be fully implemented in the local decoder 406.
[0056] At this point, it can be observed that any decoder techniques other than parsing / entropy decoding present in the decoder must also be present in essentially the same functional form in the corresponding encoder. The description of the encoder techniques may be abbreviated, as they are the inverse of the fully described decoder techniques. A more detailed description is required only in certain areas and is provided below.
[0057] As part of its operation, the source encoder 403 may perform motion-compensated predictive coding. This predictive coding encodes an input frame with reference to one or more previously encoded frames in the video sequence, designated as "reference frames." In this manner, the encoding engine 407 encodes the differences between pixel blocks of the input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame.
[0058] The local video decoder 406 can decode the encoded video data of the frame that can be designated as the reference frame based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When the encoded video data is available at the video decoder ( Figure 4 When decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that the video decoder may perform on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 405. In this way, the encoder 400 may locally store a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.
[0059] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can serve as suitable prediction references for the new frame. The predictor 404 may operate on a pixel-by-pixel-block basis based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have prediction references obtained from multiple reference pictures stored in the reference picture memory 405.
[0060] The controller 402 may manage encoding operations of the source encoder 403 , including, for example, setting parameters and subgroup parameters for encoding video data.
[0061] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder may perform lossless compression on the symbols generated by the various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0062] Transmitter 409 can buffer the encoded video sequence created by entropy encoder 408 in preparation for transmission over communication channel 411, which can be a hardware / software link to a storage device where the encoded video data will be stored. Transmitter 409 can combine the encoded video data from source encoder 403 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0063] The controller 402 can manage the operation of the video encoder 400. During encoding, the controller 402 can assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types.
[0064] An intra picture (I picture) is a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow for different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their respective applications and characteristics.
[0065] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0066] Bidirectionally predictive pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0067] A source picture is typically spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictively coded (spatial or intra-predicted) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be predictively coded using spatial prediction with reference to one previously coded reference picture or using temporal prediction. Blocks of a B picture may be predictively coded using spatial prediction with reference to one or two previously coded reference pictures or using temporal prediction.
[0068] Video encoder 400 may perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In operation, video encoder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. Consequently, the encoded video data may conform to the syntax specified by the video coding technique or standard used.
[0069] In an embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source encoder 403 may include such data as, for example, a portion of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant pictures and slices, and other forms of redundant data, SEI messages, VUI parameter set fragments, and the like.
[0070] Figure 5The intra prediction modes used in HEVC and JEM are shown. To capture arbitrary edge directions present in natural video, the number of directional intra modes is extended from 33 as used in HEVC to 65. The additional directional modes in JEM on top of HEVC are shown in Figure 1 (b) is depicted as dashed arrows, and the planar and DC modes remain the same. These more dense directional intra prediction modes are applicable to all block sizes and luma intra prediction and chroma intra prediction. Figure 5 As shown in FIG, the directional intra prediction mode (directional intra prediction mode) identified by the dotted arrow associated with the odd intra prediction mode index is called the odd intra prediction mode. The directional intra prediction mode identified by the solid arrow associated with the even intra prediction mode index is called the even intra prediction mode. In this article, as Figure 5 The directional intra prediction modes indicated by the solid or dashed arrows in are also called angular modes.
[0071] In JEM, a total of 67 intra prediction modes are used for luma intra prediction. To encode an intra mode, a Most Probable Mode (MPM) list of size 6 is built based on the intra modes of neighboring blocks. If the intra mode is not from the MPM list, a flag is signaled to indicate whether the intra mode belongs to the selected mode. In JEM-3.0, there are 16 selected modes, which are uniformly selected for every four angular modes. In JVET-D0114 and JVET-G0060, 16 secondary MPMs are derived instead of a uniformly selected mode.
[0072] Figure 6 N reference layers for intra directional mode are shown. There are block unit 611, segment A 601, segment B 602, segment C 603, segment D 604, segment E 605, segment F 606, first reference layer 610, second reference layer 609, third reference layer 608 and fourth reference layer 607.
[0073] In HEVC and JEM, as well as some other standards such as H.264 / AVC, the reference samples used to predict the current block are limited to the nearest reference line (row or column). In the method of multiple reference line intra prediction, for intra directional mode, the number of candidate reference lines (rows or columns) increases from 1 (i.e., the nearest) to N, where N is an integer greater than or equal to 1. Figure 2The concept of the multiple line intra directional prediction method is illustrated by taking a 4×4 prediction unit (PU) as an example. The intra directional mode can arbitrarily select one of N reference layers to generate a prediction value. In other words, the prediction value p(x, y) is generated from one of the reference samples S1, S2, ... and SN. A signaling flag is sent to indicate which reference layer is selected for the intra directional mode. If N is set to 1, the intra directional prediction method is the same as the conventional method in JEM 2.0. Figure 6 , reference lines 610, 609, 608, and 607 are composed of six segments 601, 602, 603, 604, 605, and 606 and an upper-left reference sample. In this document, a reference layer is also referred to as a reference line. The coordinates of the upper-left pixel in the current block unit are (0, 0), and the coordinates of the upper-left pixel in the first reference line are (-1, -1).
[0074] In JEM, for the luma component, neighboring samples used for intra-prediction sample generation are filtered before the generation process. Filtering is controlled by the given intra-prediction mode and transform block size. If the intra-prediction mode is DC or the transform block size is equal to 4×4, no filtering is performed on the neighboring samples. If the distance between the given intra-prediction mode and the vertical mode (or horizontal mode) is greater than a predetermined threshold, the filtering process is enabled. For neighboring sample filtering, a [1, 2, 1] filter and a bilinear filter are used.
[0075] The Position Dependent intra Prediction Combination (PDPC) method is an intra prediction method that uses a combination of unfiltered boundary reference samples and HEVC-type intra prediction with filtered boundary reference samples. Each prediction sample pred[x][y] at (x, y) is calculated as follows:
[0076] pred[x][y]=(wL*R -1,y +wT*R x,-1 +wTL*R -1,-1 +(64-wL-wT-wTL)*pred[x][y]+32)>>6 (Equation 2-1)
[0077] Among them, R x,-1 、R -1,y denote the unfiltered reference sample located at the upper left corner of the current sample (x, y), and R -1,-1 represents the unfiltered reference sample located at the top left corner of the current block. The weights are calculated as follows,
[0078] wT=32>>((y<<1)>>offset) (Equation 2-2)
[0079] wL=32>>((x<<1)>>offset) (Equation 2-3)
[0080] wTL=-(wL>>4)-(wT>>4) (Equation 2-4)
[0081] Offset=(log2(width)+log2(height)+2)>>2 (Equation 2-5).
[0082] Figure 7 A simplified diagram 700 is shown where DC-mode PDPC weights (wL, wT, wTL) the (0, 0) and (1, 0) positions within a 4x4 block. If PDPC is applied to DC mode, planar mode, horizontal intra mode, and vertical intra mode, no additional boundary filters, such as the HEVC DC-mode boundary filter or horizontal / vertical mode edge filters, are required. Figure 7 shows a reference sample R of PDPC applied to the upper right diagonal pattern x,-1 、R -1,y and R -1,-1 Definition of . The prediction sample pred(x', y') is located at (x', y') in the prediction block. The reference sample R x,-1 The coordinate x of is given by: x=x'+y'+1, and the reference sample R -1,y The coordinate y of is similarly given by: y=x'+y'+1.
[0083] Figure 8 A diagram 800 shows Local Illumination Compensation (LIC), which is based on a linear model of illumination variation using a scaling factor a and an offset b. LIC is adaptively enabled or disabled for each inter-mode coded coding unit (CU).
[0084] When LIC is applied to a CU, a least square error method is used to derive parameters a and b by using the neighboring samples of the current CU and its corresponding reference samples. More specifically, as Figure 8 As shown in , the subsampled (2:1 subsampled) neighboring samples of the CU and the corresponding samples in the reference picture (identified by the motion information of the current CU or sub-CU) are used. IC parameters are derived and applied to each prediction direction separately.
[0085] When the CU is coded in merge mode, the LIC flag is copied from the neighboring block in a manner similar to the motion information copying in merge mode; otherwise, the LIC flag is signaled for the CU to indicate whether LIC is applied.
[0086] Figure 9A The intra prediction modes used in HEVC are shown in FIG900. In HEVC, there are a total of 35 intra prediction modes, of which mode 10 is a horizontal mode, mode 26 is a vertical mode, and modes 2, 18, and 34 are diagonal modes. The intra prediction modes are signaled by three most probable modes (MPMs) and 32 remaining modes.
[0087] Figure 9B In an embodiment of VVC, there are a total of 87 intra prediction modes, of which mode 18 is a horizontal mode, mode 50 is a vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes -1 to -10 and modes 67 to 76 are called wide-angle intra prediction (WAIP) modes.
[0088] According to the PDPC expression, the prediction sample pred(x, y) at position (x, y) is predicted using a linear combination of the intra prediction mode (DC, planar, diagonal) and the reference samples:
[0089] pred(x,y)=(wL×R-1,y+wT×Rx,-1-wTL×R-1,-1+(64-wL-wT+wTL)×pred(x,y)+32)>>6
[0090] Among them, R x,-1 、R -1,y represent the reference sample located at the upper left corner of the current sample (x, y), and R -1,-1 Represents the reference sample located at the upper left corner of the current block.
[0091] For DC mode, for a block with width and height annotations, the weights are calculated as follows:
[0092] wT=32>>((y<<1)>>nScale), wL=32>>((x<<1)>>nScale), wTL=(wL>>4)+(wT>>4),
[0093] Where nScale = (log2(width)-2+log2(height)-2+2)>>2, where wT represents the weighting factor of the reference sample located in the upper reference line with the same horizontal coordinate, wL represents the weighting factor of the reference sample located in the left reference line with the same vertical coordinate, and wTL represents the weighting factor of the upper left reference sample of the current block. nScale specifies the speed at which the weighting factor decreases along the axis (wL decreases from left to right, or wT decreases from top to bottom), that is, the weighting factor decrement rate, and in the current design it is the same along the x-axis (from left to right) and the y-axis (from top to bottom). In addition, 32 represents the initial weighting factor of the neighboring sample, which is also the top (left or upper left) weight assigned to the upper left sample in the current CB. The weighting factor of the neighboring sample during the PDPC process should be equal to or less than this initial weighting factor.
[0094] For planar mode, wTL=0, while for horizontal mode, wTL=wT, and for vertical mode, wTL=wL. PDPC weights can be calculated using only additions and shifts. The value of pred(x, y) can be calculated in a single step using Equation 1.
[0095] The methods proposed herein may be used alone or in combination in any order. Further, each of the methods (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. Depending on the embodiment, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., a CU).
[0096] Figure 10 An exemplary diagram 1000 is shown in which there is an analysis step S1001, which may involve analyzing any one or more of a network structure 1003 and a computational graph, or other representation of some convolutional neural network 1003, and providing an output step S1002, which will be further described below.
[0097] In the context of neural network-based coding, such as in VVC and AVS3, network structure 1003 can involve various neural network-based methods, particularly neural network-based filters. Network structure 1003 represents a neural network-based filter comprising several convolutional layers. As an example, the kernel size is 3*3*M, meaning that for each channel, the convolution kernel size is 3*3, and the number of output layers is M. As in network structure 1003, combining convolutional layers with non-linear activation functions (e.g., ReLU) allows the entire process to be viewed as a nonlinear filter for reconstruction, and after the filtering process, quality can be improved. Depending on the embodiment, the illustrated network structure 1002 may be a simplified block diagram, and given the complexity of neural network-based coding methods, standard codecs may not be able to perform the filtering process. Therefore, according to exemplary embodiments herein, several identifiers in the SEI may be added to indicate whether the current CVS uses neural network-based tools. Network details may also be indicated. Similarly, if the decoder cannot process neural network-based filters, information related to the neural network may be discarded, and the processing may be skipped.
[0098] As discussed further below, exemplary embodiments provide at least two mechanisms for signaling neural network model information. The first mechanism is to explicitly signal one or more topology information and corresponding parameters trained using specific syntax elements defined in the VSEI, and the other mechanism is to provide external linkage information indicating where the corresponding topology information and network parameters exist.
[0099] To signal the network topology and parameters, exemplary embodiments may employ reference to existing formats already developed for network representation, and an example thereof may include the Neural Network Exchange Format (NNEF), a generalized neural network exchange format developed by Khronos. Other possible examples include, for example, embodiments involving Open Neural Network Exchange (ONNX) and MPEG NNR, which is a format for encoded representations of neural networks.
[0100] Ideally, any neural network model can be exported to NNEF and other formats, and network accelerators and libraries can consume data in these formats without compatibility issues with any network architecture. As a practical approach, an embodiment can directly reference external files or code streams with URI information. However, it is also desirable to have a lightweight syntax design to represent video coding specific networks for VVC or HEVC extensions with new neural network-based video coding tools, because a general representation of the network model may be bulky for compressing video formats. Since most of the used network models for video compression are based on Convolutional Neural Networks (CNNs), it is expected that having a compact representation of CNNs in SEI messages will help reduce the total bitrate, as well as enable easy access to network model data according to exemplary embodiments.
[0101] The embodiments herein may represent that a neural network model may be represented by a computational graph, which is a directed graph having a plurality of nodes (e.g., a convolutional neural network 1003 as shown), wherein a node may consist of an operation node and data (e.g., a tensor). According to exemplary embodiments, various network topologies are designed and used; however, for post- / in-loop filtering for video processing, a simple model based on CNN may be employed. In this case, a simple multi-layered feedforward network similar to CNN may be represented by a linear graph starting from the input data, wherein each operation node in the layer produces intermediate processed data. Finally, at S1002, output data may be generated by a plurality of layers, and a representation thereof may be output.
[0102] For example, the convolutional neural network 1003 is an example of a linear computational graph of a CNN, such that when input tensor data is input to an operation node, the operation node processes the input data using pre-trained constants and / or variables and outputs intermediate tensor data. When the operation node is executed, the actual data is consumed by the operation node. Typically for a CNN, a weighted summation of the input data with trained constants and / or updated variables is the output of each operation node.
[0103] According to the specification of this network diagram, once a specific operation node is specified as a single step, the same operation node can be used iteratively. Thus, the embodiments herein advantageously represent the description of the network topology through syntax elements in the SEI message at S1002. In cases where a more complex model design may be required, an external format such as NNEF or ONNX can be used at S1002.
[0104] In order to transmit network parameters, the bit size of the network parameters that are typically trained is too large to be included in the SEI message. To reduce the data size, the MPEG-NNR format can be used to compress the parameters and can be partitioned into multiple data chunks at S1002. Each chunk of compressed parameters can be included in the SEI message or in a separate data file, which may be transmitted in the same bitstream or stored in a remote server. According to an exemplary embodiment, when decoding, all linked chunk data in the SEI message representing the neural network can be spliced and consumed by the neural network library or decoder.
[0105] The following will at least combine Figures 11 to 18 Further description about Figure 10 of these embodiments.
[0106] For example, embodiments herein relate to generating and using SEI messages to carry post-filtering Neural Network (NN) information. Figure 11 An exemplary flow chart 1100 illustrating NN-based post-filtering of messages and their grammatical aspects is shown.
[0107] Such a syntax can be represented in Table 1 below:
[0108] Table 1
[0109]
[0110]
[0111]
[0112]
[0113] nn_partition_flag equal to 0 specifies that all data representing the network topology and trained parameters are included in the SEI message, and nn_partition_flag equal to 1 specifies that the data representing the network topology and trained parameters are partitioned into multiple SEI messages.
[0114] nn_output_pic_format_present_flag equal to 0 specifies that the syntax element indicating the output picture format is not present in the SEI message and the output picture format of the neural network inference process is the same as the output picture format of the decoder, and nn_output_pic_format_present_flag equal to 1 specifies that the syntax element indicating the output picture format is present in the SEI message.
[0115] nn_postfilter_type_idc specifies the post-filter type of the neural network represented by the SEI message, as specified in Table 2 below (NN post-filter type):
[0116] Table 2
[0117] nn_postfilter_type_idc Post filtering type 0 Visual quality improvement of a single input image 1 Improved visual quality for multiple input images 2 Super resolution of a single input image 3 Super-resolution of multiple input images 4..15 Reserve
[0118] num_nn_input_ref_pic specifies the number of input reference pictures. A num_nn_input_ref_pic of 0 specifies that the current output picture of the decoder is the only input data to the neural network, and a num_nn_input_ref_pic greater than 0 specifies that the number of reference pictures used as input data to the neural network is num_nn_input_ref_pic – 1.
[0119] num_partitioned_nn_sei_messages specifies the number of neural network based post-filtering SEI messages to represent the entire neural network topology with corresponding parameters, and when not present, the value of num_partitioned_nn_sei_messages is inferred to be equal to 1.
[0120] nn_sei_message_idx specifies the index of the partial neural network data carried in the SEI message, and when not present, the value of nn_sei_message_idx is inferred to be equal to 0.
[0121] Taking into account the above syntax, flowchart 1100 indicates that at S1101, the post-filter may be initialized, and data may be generated or obtained at S1102, so that at S1103 it can be determined whether an information flag such as network_topology_info_external_present_flag exists, and if so, data including external_nn_topology_info_format_idc may be obtained at S1104, data including num_bytes_external_network_topology_uri_info may be obtained at S1105, data including external_nn_topology_uri_info may be obtained at S1106, or if not, a check may be performed at S1107 to receive network_topology_info(input).
[0122] That is, according to an exemplary embodiment, nn_topology_info_external_present_flag equal to 0 specifies that the data of the neural network topology representation is contained in the SEI message, while nn_topology_info_external_present_flag equal to 1 specifies that the data of the neural network topology representation can exist externally and the SEI message only contains external link information.
[0123] For example, external_nn_topology_info_format_idc at S1104 may specify the external storage format of the neural network topology representation, as specified in the following Table 3 (External NN Topology Information Format Identifier):
[0124] Table 3
[0125] external_nn_topology_info_format_idc Storage format 0 Unrecognized storage format 1 NNEF 2 ONNX 3..15 Reserve
[0126] For example, num_bytes_external_network_topology_uri_info at S1105 specifies the number of bytes of the syntax element external_network_topology_uri_info.
[0127] For example, external_nn_topology_uri_info at S1106 specifies URI information of external neural network topology information. The length of the syntax element may be Ceil(Log2(num_bytes_external_nn_topology_uri_info)) bytes.
[0128] For example, network_topology_info(input) at S1107 may involve the process shown in Table 4 below:
[0129] Table 4
[0130]
[0131]
[0132] For example, see Figure 14 1400 , where, as in S1007 , at S1401 , it is determined that processing according to network_topology_info(input) is present, and the processing may proceed to one or more of the following: generating or obtaining nn_topology_storage_format_idc at S1402 , generating or obtaining nn_topology_compression_format_idc at S1403 , generating or obtaining num_bytes_topology_data at S1404 , and determining whether nn_top_format_idc>0 at S1405 . If the determination at S1405 is “yes,” then at S1406 nn_topology_data_byte[I] at S1406 is obtained.
[0133] For example, nn_topology_storage_format_idc at S1402 specifies the storage format of the neural network topology representation, as specified in the following Table 5 (NN Topology Storage Format Identifier):
[0134] Table 5
[0135] nn_parameter_storage_format storage format 0 Unrecognized storage format 1 NNEF 2 ONNX 3..15 Reserve
[0136] For example, nn_topology_compression_format_idc at S1403 specifies the compression format of the neural network topology, as specified in the following Table 6 (NN Topology Compression Format Identifier):
[0137] Table 6
[0138]
[0139] For example, num_bytes_topology_data at S1404 specifies the number of bytes of the neural network topology payload contained in the SEI message.
[0140] For example, nn_topology_data_byte[I] at S1406 specifies the i-th byte of the neural network topology payload.
[0141] For example, num_variables at S1408 specifies the number of variables that can be used to execute the operation node in the neural network specified by the SEI message.
[0142] For example, num_node_types at S1409 specifies the number of operation node types that can be used to execute the operation node in the neural network specified by the SEI message.
[0143] For example, num_operation_node_executions at S1410 specifies the number of operation node executions having input variables of the neural network specified by the SEI message.
[0144] Return to Figure 11 1110, at S1108, it can be determined whether network_parameter_info_external_present_flag exists, and if so, external_network_parameter_info_format_idc can be generated or obtained at S1109, num_bytes_external_network_parameter_uri_info can be generated or obtained at S1110, and external_nn_parameter_uri_info can be generated or obtained at S1111; otherwise, at S1112, network_parameter_info(input) can be obtained or generated.
[0145] For example, network_parameter_info_external_present_flag equal to 0 at S1108 specifies that data of neural network parameters are contained in the SEI message, and such network_parameter_info_external_present_flag equal to 1 specifies that data of neural network parameters may be present externally and the SEI message contains only external linking information.
[0146] The external_network_parameter_info_format_idc at S1109 specifies the external storage format of the neural network parameters, as specified in the following Table 7 (External Neural Network Parameter Storage Format Identifier):
[0147] Table 7
[0148] external_network_parameter_info_format_idc Storage format 0 Unrecognized storage format 1 NNEF 2 ONNX 3 MPEG-NNR 4..15 Reserve
[0149] num_bytes_external_network_parameter_uri_info at S1110 specifies the number of bytes of the syntax element external_network_parameter_uri_info.
[0150] The external_nn_parameter_uri_info at S1111 specifies the URI information of the external neural network parameter. The length of the syntax element is Ceil(Log2(num_bytes_external_network_parameteruri_info)) bytes.
[0151] For example, network_parameter_info(input) at S1112 represents the following process of Table 8:
[0152] Table 8
[0153]
[0154] As in Figure 17 In the flowchart 1700, as determined at S1701, network_parameter_info (input), it is also indicated that Figure 11 S1112 in the method includes obtaining or otherwise generating nn_parameter_type_idc at S1702, obtaining or otherwise generating nn_parameter_storage_format_idc at S1703, obtaining or otherwise generating nn_parameter_compression_format_idc at S1704, and obtaining or otherwise generating num_bytes_parameter_data at S1705, and obtaining or otherwise generating nn_parameter_data_byte at S1706.
[0155] For example, nn_parameter_type_idc at S1702 specifies the data payload type of the neural network parameters, as specified in Table 9 below (NN parameter payload type):
[0156] Table 9
[0157] nn_parameter_type_idc Parameter type 0 Integer 1 Floating point numbers (Float) 2..15 Reserved
[0158] For example, nn_parameter_storage_format_idc at S1703 specifies the storage format of the neural network parameters, as specified in the following Table 10 (NN parameter storage format identifier):
[0159] Table 10
[0160] nn_parameter_storage_format_idc Storage format 0 Unrecognized storage format 1 NNEF 2 ONNX 3 MPEG-NNR 4..15 Reserve
[0161] For example, nn_parameter_compression_format_idc at S1704 specifies the compression format of the neural network parameters, as specified in the following Table 11 (NN topology compression format identifier):
[0162] Table 11
[0163]
[0164] For example, num_bytes_parameter_data at S1705 specifies the number of bytes of the neural network parameter payload included in the SEI message.
[0165] For example, nn_parameter_data_byte at S1706 specifies the i-th byte of the neural network parameter payload.
[0166] return Figure 11 , at S1113, the process may proceed to, for example Figure 12 In S1201 of the flowchart 1200, it may be determined whether network_input_pic_format_present_flag exists. If so, nn_input_chroma_format_idc may be generated or obtained at S1202, nn_input_bitdepth_minus8 may be generated or obtained at S1203, nn_input_pic_width may be generated or obtained at S1204, nn_input_pic_height may be generated or obtained at S1205, and whether nn_patch_size_present_flag exists may be determined at S1206, and if so, nn_input_patch_width may be obtained or generated at S1207, nn_input_patch_height may be obtained or generated at S1208, and nn_boundary_padding_idc may be obtained or generated at S1209.
[0167] For example, network_input_pic_format_present_flag equal to 0 at S1201 specifies that the syntax element indicating the input picture format is not present in the SEI message, and the input picture format of the neural network inference process is the same as the output picture format of the decoder. nn_input_pic_format_present_flag equal to 1 specifies that the syntax element indicating the input picture format is present in the SEI message.
[0168] nn_input_chroma_format_idc at S1202 may specify chroma samples relative to luma samples according to the following Table 12 (Chroma Format Identifier):
[0169] Table 12
[0170] nn_input_chroma_format_idc Chroma format 0 Monochrome 1 4:2:0 2 4:2:2 3 4:4:4
[0171] For example, nn_input_bitdepth_minus8 (or plus8) at S1203 specifies the bit depth of luma samples and chroma samples in the input picture of the neural network.
[0172] For example, nn_input_pic_width at 1204 specifies the width of the input picture.
[0173] For example, nn_input_pic_height at S1205 specifies the height of the input picture.
[0174] For example, nn_patch_size_present_flag equal to 0 at S1206 specifies that the patch size (patchsize) is equal to the input picture size. nn_patch_size_present_flag equal to 1 specifies that the patch size is explicitly signaled.
[0175] For example, nn_input_patch_width at S1207 specifies the width of the patch used for the neural network inference process.
[0176] For example, nn_input_patch_height at S1208 specifies the height of the patch used for the neural network inference process.
[0177] For example, nn_boundary_padding_idc at S1209 specifies the padding method applied to the boundary of the patch when the patch size is different from the input picture size, for example according to the following Table 13 (Boundary Padding Identifiers):
[0178] Table 13
[0179]
[0180] Return to Figure 11 At S1113, the process can also be Figure 13 Flowchart 1300 and Figure 12 1200 can be performed in parallel or serially, wherein it can be determined at S1301 whether num_network_input_ref_pic>0 exists. If so, num_fwd_ref_pics_as_input can be obtained or generated at S1302, and it can be determined at S1303 whether an indication of NumFwdRefPics>0 exists, and it can be determined at S1306 whether an indication of NumBwdRefPics>0 exists. Further, if such an indication exists at S1303, nearest_fwd_ref_pics_as_input can be determined at S1304, and poc_dist_fwd_ref_pic[i] can be determined at S1305. Further, if such an indication exists at S1303, nearest_bwd_ref_pics_as_input can be determined at S1307, and poc_dist_bwd_ref_pic[i] can be determined at S1308.
[0181] For example, num_fwd_ref_pics_as_input at S1302 specifies the number of forward reference pictures used as input data of the neural network, if (num_nn_input_ref_pic>0), then (NumFwdRefPics=num_fwd_ref_pics_as_input), otherwise (NumFwdRefPics=0).
[0182] For example, nearest_fwd_ref_pics_as_input at S1304 specifies the nearest forward reference picture that has the smallest picture order count distance to the current picture, which is used as input data of the neural network.
[0183] For example, poc_dist_fwd_ref_pic[i] at S1305 specifies the picture order count value of the i-th forward reference picture used as input data of the neural network. The picture order count value of the i-th forward reference picture is equal to the picture order count value of the current picture minus poc_dist_fwd_ref_pic[i].
[0184] For example, nearest_bwd_ref_pics_as_input at S1307 specifies the number of backward reference pictures used as input data of the neural network, so that if (num_nn_input_ref_pic>0), then (NumBwdRefPics=num_bwd_ref_pics_as_input), otherwise (NumBwdRefPics=0).
[0185] For example, poc_dist_bwd_ref_pic[i] at S1308 specifies the picture order count value of the i-th backward reference picture used as input data of the neural network. The picture order count value of the i-th backward reference picture is equal to the picture order count value of the current picture plus poc_dist_bwd_ref_pic[i].
[0186] In addition, as a note, for example, nearest_bwd_ref_pics_used_flag in Table 1 above specifies that the nearest backward reference picture is used as input data for the neural network, where the nearest backward reference picture has the smallest picture order count distance to the current picture.
[0187] You can use, for example, Figure 15 1500 to implement additional operations, wherein, at S1501, there is define_operation_node(i), and the following operations will exist: iteratively define nn_operation_class_idc[i] at S1502, iteratively define nn_operation_function_idc[i] at S1503, iteratively define num_input_variables[i] at S1504, and iteratively define num_output_variables[i] at S1505.
[0188] For example, nn_operation_class_idc[i] at S1502 specifies the class of the i-th operation node, as specified in the following Table 14 (NN operation function):
[0189] Table 14
[0190]
[0191] For example, nn_operation_function_idc[i] at S1503 specifies the function of the i-th operation node, as specified in the following Table 15 (an example Table 15 where nn operation function, nn_operation_class_idc is equal to 7 (activation function)):
[0192] Table 15
[0193]
[0194] For example, num_input_variables[i] at S1504 specifies the number of input variables of the i-th operation node.
[0195] For example, num_output_variables[i] at S1505 specifies the number of output variables of the i-th operation node.
[0196] about Figure 15 The syntax of can be represented by the following Table 16:
[0197] Table 16
[0198]
[0199] You can use, for example, Figure 16 1600 to implement additional operations, wherein, at S1601, there is operation_node_execution(i), and the following operations will exist: iteratively define nn_op_node_idx[i] at S1602, and iteratively define nn_input_variable_idx[i][j] at S1603, and iteratively define nn_output_variable_idx[i][j] at S1604.
[0200] For example, nn_op_node_idx[i] at S1602 specifies the index of the operation node for the i-th operation node execution. The nn_op_node_idx[i]-th operation node is used for this execution.
[0201] For example, nn_input_variable_idx[i][j] at S1603 specifies the variable index of the j-th input variable of the i-th operation node execution.
[0202] For example, nn_output_variable_idx[i][j] at S1604 specifies the variable index of the j-th output variable executed by the i-th operation node.
[0203] about Figure 16 The syntax of can be represented by the following Table 17:
[0204] Table 17
[0205]
[0206] Additional procedures may involve using, for example Figure 15 Iteration defines variables (i), for example according to the following syntax of Table 18:
[0207] Table 18
[0208]
[0209] Looking at Table 16, nn_variable_class_idc[i] specifies the variable class of the i-th variable in the neural network, as specified in Table 19 (NN variable class) below:
[0210] Table 19
[0211]
[0212] According to an exemplary embodiment, it can be determined that when nn_variable_class_idc is equal to 1, the variable is input data of the neural network; when nn_variable_class_idc is equal to 2, the variable is output data of the neural network; when nn_variable_class_idc is equal to 3, the variable is intermediate data between operation nodes; and when nn_variable_class_idc is equal to 4, the variable is pre-trained or predefined constant data.
[0213] Looking further at Table 16, nn_variable_type_idc[I] specifies the variable type of the i-th variable in the neural network, as specified in Table 20 (NN variable type) below:
[0214] Table 20
[0215] nn_parameter_type_idc Parameter type 0 Integer 1 Floating point numbers (Float) 2..15 Reserved
[0216] According to an exemplary embodiment, nn_variable_dimensions[I] of Table 16 specifies the dimension of the i-th variable, and Nn_variable_dimension_size[I][j] further specifies the size of the j-th dimension of the i-th variable. As an illustration according to an exemplary embodiment, when the i-th variable is input data, wherein the number of color components of the input data is 3, the width and height are 1920 and 1080, respectively, nn_variable_class_idc[i] is equal to 1, nn_variable_dimensions[I] is equal to 3, nn_variable_dimension_size[I][0] is equal to 3, nn_variable_dimension_size[I][1] is equal to 1920, and nn_variable_dimension_size[I][2] is equal to 1080.
[0217] The above techniques may be implemented as computer software via computer-readable instructions and physically stored in one or more computer-readable media, or the above techniques may be implemented via one or more specially configured hardware processors. For example, Figure 18 A computer system 1800 is shown that is suitable for implementing certain embodiments of the disclosed subject matter.
[0218] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or similar mechanisms to create a code comprising instructions, which may be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), or the like, or executed through decoding, microcode, or the like.
[0219] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, and the like.
[0220] Figure 18The components shown for computer system 1800 are exemplary in nature and are not intended to limit the scope of use or functionality of computer software implementing embodiments of the present application. Nor should the configuration of components be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of computer system 1800.
[0221] Computer system 1800 may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users via tactile input (e.g., keyboard input, swipe gestures, data glove movements), audio input (e.g., voice, applause), visual input (e.g., gestures), or olfactory input (not shown). The human-computer interface devices may also be used to capture media that is not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0222] The human interface input device may include one or more of the following (only one of which is depicted): keyboard 1801 , mouse 1802 , touchpad 1803 , touch screen 1810 , joystick 1805 , microphone 1806 , scanner 1808 , camera 1807 .
[0223] Computer system 1800 may also include certain human-computer interface output devices. Such human-computer interface output devices can stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback via touch screen 1810 or joystick 1805, although tactile feedback devices that do not function as input devices may also be present), audio output devices (e.g., speakers 1809, headphones (not shown)), visual output devices (e.g., screen 1810 including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities—some of which can output two-dimensional visual output or output in three or more dimensions through means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and cigarette boxes (not shown)), and printers (not shown).
[0224] The computer system 1800 may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable compact discs (CD / DVD ROM / RW) 1820 with CD / DVD 2211 or similar media, a thumb drive 1822, a removable hard drive or solid state drive 1823, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security software dongles (not shown), and the like.
[0225] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0226] Computer system 1800 may also include an interface 1899 to one or more communication networks 1898. For example, network 1898 may be wireless, wired, or optical. Network 1898 may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. Network 1898 also includes local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), wide area digital wired or wireless networks (including cable, satellite, and terrestrial broadcast television), in-vehicle and industrial networks (including CAN Bus), and the like. Some networks 1898 typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1850 and 1851) (e.g., a USB port on computer system 1800); other systems are typically integrated into the core of computer system 1800 by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). By using any of these networks 1898, the computer system 1800 can communicate with other entities. The communication can be one-way, for reception only (e.g., wireless television), one-way, for transmission only (e.g., a CAN bus to certain CAN bus devices), or two-way, such as to other computer systems via a local or wide area digital network. Each of the above networks and network interfaces can use certain protocols and protocol stacks.
[0227] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core 1840 of the computer system 1800 .
[0228] Core 1840 may include one or more central processing units (CPUs) 1841, graphics processing units (GPUs) 1842, graphics adapters 1817, specialized programmable processing units in the form of field programmable gate arrays (FPGAs) 1843, hardware accelerators for specific tasks 1844, and the like. These devices, as well as read-only memory (ROM) 1845, random access memory 1846, and internal mass storage (e.g., an internal non-user-accessible hard drive, solid-state drive, etc.) 1847, may be connected via a system bus 1848. In some computer systems, system bus 1848 may be accessible in the form of one or more physical plugs to allow expansion with additional central processing units, graphics processing units, and the like. Peripheral devices may be attached directly to the core's system bus 1848 or connected via a peripheral bus 1851. Peripheral bus architectures include PCI (Peripheral Controller Interface) and USB (Universal Serial Bus).
[0229] The CPU 1841, GPU 1842, FPGA 1843, and accelerator 1844 can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM 1845 or RAM 1846. Transient data can also be stored in RAM 1846, while permanent data can be stored, for example, in internal mass storage 1847. The use of cache memory allows for rapid storage and retrieval from any memory device, and the cache memory can be closely associated with one or more of the CPU 1841, GPU 1842, mass storage 1847, ROM 1845, RAM 1846, and the like.
[0230] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes of this application, or may be medium and code well known and available to those skilled in the art of computer software.
[0231] As an embodiment and not limitation, a computer system having architecture 1800, in particular core 1840, can provide the function of executing software contained in one or more tangible computer-readable media as a processor (including CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the core 1840 having non-volatility, such as core internal mass storage 1847 or ROM 1845. The software implementing the various embodiments of the present application can be stored in such a device and executed by the core 1840. Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core 1840, in particular the processor therein (including CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in RAM 1846 and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator 1844) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing the executing software, circuitry containing the executing logic, or both. This application includes any suitable combination of hardware and software.
[0232] Although this application has described a number of exemplary embodiments, various modifications, permutations, and equivalent substitutions of the embodiments are within the scope of this application. Therefore, it should be understood that those skilled in the art will be able to design a variety of systems and methods that, although not explicitly shown or described herein, embody the principles of this application and are therefore within the spirit and scope of this application.
Claims
1. A video encoding method, characterized in that: The method comprises: Get video bitstream; encoding the video bitstream at least in part via a neural network; Determining topological information and parameters of the neural network; and, The determined topology information and the parameters of the neural network are signaled in a plurality of syntax elements associated with the coded video bitstream via a supplemental enhancement information (SEI) message; wherein, nn_output_pic_format_present_flag equal to 0 specifies that the syntax element indicating the output picture format is not present in the SEI message and the output picture format of the neural network inference process is the same as the output picture format of the decoder, and nn_output_pic_format_present_flag equal to 1 specifies that the syntax element indicating the output picture format is present in the SEI message; num_nn_input_ref_pic equal to 0 specifies that the current output picture of the decoder is the only input data of the neural network, and num_nn_input_ref_pic greater than 0 specifies that the number of reference pictures used as input data of the neural network is num_nn_input_ref_pic–1; nn_operation_class_idc[i] specifies the class of the i-th operation node; nn_operation_function_idc[i] specifies the function of the i-th operation node; num_input_variables[i] specifies the number of input variables of the i-th operation node; num_output_variables[i] specifies the number of output variables of the i-th operation node; Wherein, the neural network includes a plurality of operation nodes, i is an integer; nn_parameter_type_idc specifies the data payload type of the neural network parameters; num_bytes_parameter_data specifies the number of bytes of the neural network parameter payload contained in the SEI message; nn_parameter_data_byte specifies the i-th byte of the neural network parameter payload.
2. The method according to claim 1, characterized in that Encoding the video bitstream by the neural network includes: Sending input tensor data of the video bitstream to a first operation node among the plurality of operation nodes; Processing the input tensor data using any pre-trained constants and variables; and, Output intermediate tensor data, The intermediate tensor data includes a weighted sum of the input tensor data and any trained constants and updated variables.
3. The method according to claim 2, characterized in that The topology information and parameters are based on encoding of the video bitstream by the neural network.
4. The method according to claim 1, wherein The signaling of the determined topology information and the parameters includes providing external link information, and storing the determined topology information and the parameters at the external link information.
5. The method according to claim 4, characterized in that Encoding the video bitstream by the neural network includes: Sending input tensor data of the video bitstream to a first operation node among the plurality of operation nodes; Processing the input tensor data using any pre-trained constants and variables; and, Output intermediate tensor data, and, The intermediate tensor data includes a weighted sum of the input tensor data and any trained constants and updated variables.
6. The method according to claim 5, characterized in that The topology information and parameters are based on encoding of the video bitstream by the neural network.
7. The method according to claim 1, characterized in that The signaling of the determined topology information and the parameters includes: explicitly signaling the determined topology information and the parameters through at least one of the neural network exchange format NNEF, the open neural network exchange ONNX format and the moving picture experts group MPEG neural network compression standard NNR format.
8. The method according to claim 7, characterized in that Encoding the video bitstream by the neural network includes: Sending input tensor data of the video bitstream to a first operation node among the plurality of operation nodes; Processing the input tensor data using any pre-trained constants and variables; and, Output intermediate tensor data, wherein the intermediate tensor data comprises a weighted sum of the input tensor data and any trained constants and updated variables, and The topology information and parameters are based on the encoding of the video bitstream by the neural network.
9. The method according to any one of claims 1 to 8, characterized in that At least one of the Neural Network Exchange Format (NNEF), the Open Neural Network Exchange (ONNX) format, and the Moving Picture Experts Group (MPEG) Neural Network Compression Standard (NNR) format is an MPEG NNR format, and, At least one of the parameters is compressed into at least one of a supplemental enhancement information (SEI) message and a data file.
10. A video decoding method, characterized in that: The method is applied to a decoder including a local encoder, comprising: Get video bitstream; encoding the video bitstream at least in part via a neural network; Determining topological information and parameters of the neural network; and, The determined topology information and the parameters of the neural network are signaled in a plurality of syntax elements associated with the coded video bitstream via a supplemental enhancement information (SEI) message; wherein, nn_output_pic_format_present_flag equal to 0 specifies that the syntax element indicating the output picture format is not present in the SEI message and the output picture format of the neural network inference process is the same as the output picture format of the decoder, and nn_output_pic_format_present_flag equal to 1 specifies that the syntax element indicating the output picture format is present in the SEI message; num_nn_input_ref_pic equal to 0 specifies that the current output picture of the decoder is the only input data of the neural network, and num_nn_input_ref_pic greater than 0 specifies that the number of reference pictures used as input data of the neural network is num_nn_input_ref_pic–1; nn_operation_class_idc[i] specifies the class of the i-th operation node; nn_operation_function_idc[i] specifies the function of the i-th operation node; num_input_variables[i] specifies the number of input variables of the i-th operation node; num_output_variables[i] specifies the number of output variables of the i-th operation node; Wherein, the neural network includes a plurality of operation nodes, i is an integer; nn_parameter_type_idc specifies the data payload type of the neural network parameters; num_bytes_parameter_data specifies the number of bytes of the neural network parameter payload contained in the SEI message; nn_parameter_data_byte specifies the i-th byte of the neural network parameter payload.
11. A video encoding device, characterized in that: The device comprises: An acquisition module, used to obtain a video bit stream; an encoding module for encoding the video bitstream at least in part via a neural network; a determination module, configured to determine topological information and parameters of the neural network; and A signaling module is configured to signal the determined topology information and the parameters of the neural network in a plurality of syntax elements associated with the coded video bitstream via a supplemental enhancement information (SEI) message; wherein, nn_output_pic_format_present_flag equal to 0 specifies that the syntax element indicating the output picture format is not present in the SEI message and the output picture format of the neural network inference process is the same as the output picture format of the decoder, and nn_output_pic_format_present_flag equal to 1 specifies that the syntax element indicating the output picture format is present in the SEI message; num_nn_input_ref_pic equal to 0 specifies that the current output picture of the decoder is the only input data of the neural network, and num_nn_input_ref_pic greater than 0 specifies that the number of reference pictures used as input data of the neural network is num_nn_input_ref_pic–1; nn_operation_class_idc[i] specifies the class of the i-th operation node; nn_operation_function_idc[i] specifies the function of the i-th operation node; num_input_variables[i] specifies the number of input variables of the i-th operation node; num_output_variables[i] specifies the number of output variables of the i-th operation node; Wherein, the neural network includes a plurality of operation nodes, i is an integer; nn_parameter_type_idc specifies the data payload type of the neural network parameters; num_bytes_parameter_data specifies the number of bytes of the neural network parameter payload contained in the SEI message; nn_parameter_data_byte specifies the i-th byte of the neural network parameter payload.
12. A video encoding device, characterized in that: include: one or more non-transitory computer-readable media for storing computer program code; as well as, One or more computer processors, configured to access the computer program code and operate according to instructions of the computer program code to execute the method according to any one of claims 1 to 9.
13. A non-volatile computer-readable medium storing instructions, characterized in that: When at least one processor executes the instructions, the at least one processor is caused to perform the method according to any one of claims 1 to 9.
14. A method for storing a video stream, characterized in that: Execute the method according to any one of claims 1 to 9 to generate a video stream, and store the video stream.
15. A method for transmitting a video stream, characterized in that: Execute the method according to any one of claims 1 to 9 to generate a video stream, and transmit the video stream.
16. A computer-readable storage medium storing a computer program / instruction and a video stream, wherein: When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented to generate the video code stream.
Citation Information
Patent Citations
METHOD AND APPARATUS FOR VIDEO CODING, computer equipment and storage medium
CN112135134A
Deterministic neural networking interoperability
US20200175396A1
Supplemental enhancement information messages for neural network based video post processing
US20200304836A1