Symbolization and decoding method and apparatus
By dynamically adjusting parameter weights in neural network-based image compression, the method enhances rate-distortion performance and reduces data redundancy, addressing the inefficiencies of existing methods.
Patent Information
- Application Number
- JP2024506874
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-05
- Filing Date
- 2022-08-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-08-01
AI Technical Summary
Existing neural network-based image compression methods face challenges in achieving high rate-distortion performance without requiring online learning, leading to suboptimal compression speed or efficiency.
An encoding and decoding method that dynamically adjusts parameter weights of encoding and decoding networks based on the to-be-encoded data, enhancing the representation ability of the networks and improving rate-distortion performance.
The method improves the rate-distortion performance of image compression by aligning parameter weights with the to-be-encoded data, resulting in more accurate decoding and reduced data redundancy.
Smart Images

Figure 0007704475000003 
Figure 0007704475000004 
Figure 0007704475000005
Abstract
Description
Technical Field
[0001]
[0002] Technical Field This application is related to the field of artificial intelligence, and in particular, to encoding and decoding methods and apparatuses. Technology
Background Art
[0002]
[0003] Background With the development of science and technology, the quantity and resolution of images are increasing. A large number of images not only require a storage medium with a larger capacity, but also a wider transmission frequency band and a longer transmission time, which has become a basic issue in image processing. In order to improve the storage efficiency and transmission efficiency of images, it is necessary to encode large amounts of image data to compress the images.
[0003]
[0004] Image compression based on neural networks can improve image compression efficiency. Existing neural network-based image compression methods are mainly classified into a neural network-based image compression method that requires online learning (briefly referred to as Method 1) and a neural network-based image compression method that does not require online learning (briefly referred to as Method 2). Method 1 has good rate-distortion performance but requires online learning, resulting in a low image compression speed. Method 2 has poor rate-distortion performance but has a high image compression speed.
Summary of the Invention
[0004]
[0005] This application provides an encoding and decoding method and apparatus for improving the rate-distortion performance of data encoding and decoding methods without relying on online training. To achieve the above object, this application uses the following technical solutions.
[0005]
[0006] According to a first aspect, the present application provides an encoding method. The method includes: obtaining to-be-encoded data; then inputting the to-be-encoded data into a first encoding network to obtain target parameters; then constructing a second encoding network based on the target parameters; next, inputting the to-be-encoded data into the second encoding network to obtain a first feature; and finally, encoding the first feature to obtain an encoded bitstream.
[0006]
[0007] In an existing encoding method, an encoding network (i.e., the second encoding network) extracts content features (i.e., the first feature) of to-be-encoded data based on fixed parameter weights, then encodes the content features into a bitstream (i.e., the encoded bitstream), and transmits the bitstream to the decoder side. The decoder side performs decoding and reconstruction on the bitstream to obtain decoded data. It can be seen that in the prior art, the parameter weights of the encoding network are not associated with the to-be-encoded data. However, in the encoding method provided in the present application, first, the to-be-encoded data is input into the first encoding network, the first encoding network generates the parameter weights of the second encoding network based on the to-be-encoded data, the parameter weights of the second encoding network are dynamically adjusted based on the obtained weights, as a result, the parameter weights of the second encoding network are associated with the to-be-encoded data, the representation ability of the second encoding network is enhanced, and the decoded data obtained on the decoder side by decoding and reconstruction with respect to the bitstream obtained by encoding the first feature is closer to the to-be-encoded data. This improves the rate-distortion performance of the encoding and decoding networks.
[0007]
[0008] Optionally, the target parameter is all or part of the parameter weights for the convolution and non-linear activation of the second encoding network.
[0008]
[0009] In a possible implementation, the step of encoding the first feature to obtain an encoded bitstream includes: rounding the first feature to obtain an integer value of the first feature; performing a probability estimation on the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature; and performing entropy encoding on the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to obtain an encoded bitstream.
[0009]
[0010] Entropy encoding is performed on the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to form a bitstream. This can reduce the redundancy of the encoding for outputting the first feature and can further reduce the data transmission volume in the data encoding or decoding (compression) process.
[0010]
[0011] In a possible implementation, the step of performing a probability estimation on the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature includes: performing a probability estimation on the integer value of the first feature based on first information to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0011]
[0012] The probability distribution is estimated based on context information and side information, and as a result, the accuracy of the obtained estimated probability distribution can be improved. This can reduce the bit rate in the entropy encoding process and reduce the overhead of entropy encoding.
[0012]
[0013] According to a second aspect, the present application provides a decoding method. The method includes: obtaining a bitstream to be decoded; then decoding the bitstream to be decoded to obtain an integer value of a first feature and an integer value of a second feature; further, inputting the integer value of the second feature into a first decoding network to obtain the target parameter; next, constructing a second decoding network based on the target parameter; and finally, inputting the integer value of the first feature into the second decoding network to obtain decoded data. The integer value of the first feature is used to obtain the decoded data, and the integer value of the second feature is used to obtain the target parameter.
[0013]
[0014] In an existing decoding method, a decoding network (i.e., the second decoding network) performs decoding and reconstruction on the content value feature (i.e., the integer value of the first feature) of the encoded target data based on fixed parameter weights to obtain the decoded data. It can be seen from the prior art that the parameter weights of the decoding network are not associated with the target data to be decoded. However, in the present application, the content feature and the model feature of the target data to be decoded (i.e., the first feature and the second feature) are encoded in the bitstream to be decoded, and then the decoder side decodes the target bitstream to be decoded to obtain the integer value of the second feature. The integer value of the second feature is input into the first decoding network to obtain the parameter weights of the second decoding network. Then, the parameter weights of the second decoding network are dynamically adjusted based on the parameter weights. As a result, the parameter weights of the second decoding network are associated with the target data to be decoded, the representation ability of the second decoding network is improved, and the decoded data obtained by the second decoding network through decoding and reconstruction is closer to the encoded target data. This improves the rate-distortion performance of the encoding and decoding networks.
[0014]
[0015] Optionally, the target parameter is all or part of the parameter weights for the convolution and non-linear activation of the second encoding network.
[0015]
[0016] Optionally, the bitstream to be decoded includes the first bitstream to be decoded and the second bitstream to be decoded.
[0016]
[0017] In a possible implementation, the step of decoding the bitstream to be decoded to obtain the integer value of the first feature and the integer value of the second feature includes: decoding the first bitstream to be decoded to obtain the integer value of the first feature; and decoding the second bitstream to be decoded to obtain the integer value of the second feature.
[0017]
[0018] In a possible implementation, the step of decoding the first bitstream to be decoded to obtain the integer value of the first feature includes: performing probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain an estimated probability distribution for the integer value of the first feature; and performing entropy decoding on the first bitstream to be decoded based on the estimated probability distribution for the integer value of the first feature to obtain the integer value of the first feature.
[0018]
[0019] In a possible implementation, the step of performing probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain an estimated probability distribution for the integer value of the first feature includes: performing probability estimation on the integer value of the first feature in the first bitstream to be decoded based on the first information to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0019]
[0020] In a possible implementation, the step of decoding a second bitstream to be decoded to obtain an integer value of a second feature includes: performing a probability estimation on the integer value of the second feature in the second bitstream to be decoded to obtain an estimated probability distribution for the integer value of the second feature; and performing entropy decoding on the second bitstream to be decoded based on the estimated probability distribution for the integer value of the second feature to obtain the integer value of the second feature.
[0020]
[0021] In a possible implementation, the step of performing a probability estimation on the integer value of the second feature in the second bitstream to be decoded to obtain an estimated probability distribution for the integer value of the second feature includes: performing a probability estimation on the integer value of the second feature in the second bitstream to be decoded based on first information to obtain an estimated probability distribution for the integer value of the second feature, where the first information includes at least one of context information and side information.
[0021]
[0022] According to a third aspect, the present application provides a decoding method. The method includes: obtaining a bitstream to be decoded; then decoding the bitstream to be decoded to obtain an integer value of a first feature; further inputting the integer value of the first feature into a first decoding network to obtain a target parameter; next, constructing a second decoding network based on the target parameter; and finally, inputting the integer value of the first feature into the second decoding network to obtain decoded data. The integer value of the first feature is used to obtain the decoded data and the target parameter. 1 first decoding network to obtain a target parameter; next, constructing a second decoding network based on the target parameter; and finally, inputting the integer value of the first feature into the second decoding network to obtain decoded data. The integer value of the first feature is used to obtain the decoded data and the target parameter.
[0022]
[0023] In the existing decoding method, the decoding network (i.e., the second decoding network) performs decoding and reconstruction on the content value feature (i.e., the integer value of the first feature) of the encoded target data based on fixed parameter weights to obtain the decoded data. It can be seen from the prior art that the parameter weights of the decoding network are not associated with the target data to be decoded. However, in this application, the decoded target bitstream obtained by encoding the feature (i.e., the first feature) of the target data to be decoded is decoded to obtain the integer value of the first feature. The integer value of the first feature is input into the first decoding network to obtain the parameter weights of the second decoding network. Then, the parameter weights of the second decoding network are dynamically adjusted based on the parameter weights. As a result, the parameter weights of the second decoding network are associated with the target data to be decoded, the representation ability of the second decoding network is improved, and the decoded data obtained by the second decoding network through decoding and reconstruction is closer to the encoded target data. This improves the rate-distortion performance of the encoding and decoding networks.
[0023]
[0024] Optionally, the target parameter is all or part of the parameter weights related to the convolution and non-linear activation of the second encoding network.
[0024]
[0025] In a possible implementation, the step of decoding the target bitstream to be decoded to obtain the integer value of the first feature includes: performing probability estimation on the integer value of the first feature in the target bitstream to be decoded to obtain the estimated probability distribution of the integer value of the first feature; and performing entropy decoding on the target bitstream to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature.
[0025]
[0026] In a possible implementation, the step of performing a probability estimation regarding an integer value of a first feature in a bitstream to be decoded to obtain an estimated probability distribution for the integer value of the first feature includes: performing a probability estimation regarding an integer value of a first feature in a bitstream to be decoded based on first information to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0026]
[0027] According to a fourth aspect, the present application provides an encoding device. The encoding device includes a processing circuit. The processing circuit is configured to: obtain target data to be encoded; input the target data to be encoded into a first encoding network to obtain target parameters; construct a second encoding network based on the target parameters; input the target data to be encoded into the second encoding network to obtain a first feature; and encode the first feature to obtain an encoded bitstream.
[0027]
[0028] Optionally, the target parameters are all or part of the parameter weights regarding the convolution and non-linear activation of the second encoding network.
[0028]
[0029] In a possible implementation, the processing circuit is specifically configured to: round a first feature to obtain an integer value of the first feature; perform a probability estimation regarding the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature; and perform entropy encoding regarding the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to obtain an encoded bitstream.
[0029]
[0030] In a possible implementation, the processing circuit is specifically configured to execute a probability estimation regarding an integer value of a first feature based on first information to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0030]
[0031] According to a fifth aspect, the present application provides a decoding device. The decoding device includes a processing circuit. The processing circuit is specifically configured to: obtain a bitstream to be decoded; decode the bitstream to be decoded to obtain an integer value of a first feature and an integer value of a second feature (the integer value of the first feature is used to obtain decoded data, and the integer value of the second feature is used to obtain a target parameter); input the integer value of the second feature into a first decoding network to obtain a target parameter; construct a second decoding network based on the target parameter; and input the integer value of the first feature into the second decoding network to obtain decoded data.
[0031]
[0032] Optionally, the target parameter is all or part of the parameter weights regarding the convolution and non-linear activation of a second encoding network.
[0032]
[0033] Optionally, the bitstream to be decoded includes a first bitstream to be decoded and a second bitstream to be decoded.
[0033]
[0034] In a possible implementation, the processing circuit is specifically configured to: decode the first bitstream to be decoded to obtain an integer value of a first feature; and decode the second bitstream to be decoded to obtain an integer value of a second feature.
[0034]
[0033] In a possible implementation, the processing circuit: executes a probability estimation regarding an integer value of a first feature in a first bitstream to be decoded, to obtain an estimated probability distribution for the integer value of the first feature; and is specifically configured to execute entropy decoding regarding the first bitstream to be decoded based on the estimated probability distribution for the integer value of the first feature, to obtain the integer value of the first feature.
[0035]
[0036] In a possible implementation, the processing circuit is specifically configured to execute a probability estimation regarding an integer value of a first feature in a first bitstream to be decoded based on first information, to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0036]
[0037]
[0037] In a possible implementation, the processing circuit: executes a probability estimation regarding an integer value of a second feature in a second bitstream to be decoded, to obtain an estimated probability distribution for the integer value of the second feature; and is specifically configured to execute entropy decoding regarding the second bitstream to be decoded based on the estimated probability distribution for the integer value of the second feature, to obtain the integer value of the second feature.
[0037]
[0038] In a possible implementation, the processing circuit is specifically configured to execute a probability estimation regarding an integer value of a second feature in a second bitstream to be decoded based on first information, to obtain an estimated probability distribution for the integer value of the second feature, where the first information includes at least one of context information and side information.
[0038]
[0039] According to the sixth aspect, the present application provides a decoding device. The decoding device includes a processing circuit. The processing circuit is specifically configured to: obtain a bitstream to be decoded; decode the bitstream to be decoded to obtain an integer value of a first feature (where the integer value of the first feature is used to obtain decoded data and target parameters); input the integer value of the first feature into a first decoding network to obtain target parameters; construct a second decoding network based on the target parameters; and input the integer value of the first feature into the second decoding network to obtain decoded data.
[0039]
[0040] Optionally, the target parameters are all or part of the parameter weights related to the convolution and non-linear activation of the second encoding network.
[0040]
[0044] In a possible implementation, the processing circuit is specifically configured to: perform probability estimation on the integer value of the first feature in the bitstream to be decoded to obtain an estimated probability distribution for the integer value of the first feature; and perform entropy decoding on the bitstream to be decoded based on the estimated probability distribution for the integer value of the first feature to obtain the integer value of the first feature.
[0041]
[0042] In a possible implementation, the processing circuit is specifically configured to perform probability estimation on the integer value of the first feature in the bitstream to be decoded based on first information, and obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information.
[0042]
[0043] According to a seventh aspect, an embodiment of the present application further provides an encoder. The encoder includes at least one processor, and when the at least one processor executes program code or instructions, the method in any one of the first aspect or possible implementations of the first aspect is implemented.
[0043]
[0044] Optionally, the encoder may further include at least one memory, and the at least one memory is configured to store program code or instructions.
[0044]
[0045] According to an eighth aspect, an embodiment of the present application further provides a decoder. The decoder includes at least one processor, and when the at least one processor executes program code or instructions, the method in any one of the second aspect or possible implementations of the second aspect is implemented.
[0045]
[0046] Optionally, the decoder may further include at least one memory, and the at least one memory is configured to store program code or instructions.
[0046]
[0047] According to a ninth aspect, an embodiment of the present application further provides a chip including an input interface, an output interface, and at least one processor. Optionally, the chip further includes a memory. The at least one processor is configured to execute code in the memory. When the at least one processor executes the code, the chip implements the method in any one of the first aspect or possible implementations of the first aspect.
[0047]
[0048] Optionally, the chip may be an integrated circuit.
[0048]
[0049] According to a tenth aspect, an embodiment of the present application further provides a terminal. The terminal includes the aforementioned encoding device, decoding device, encoder, decoder, or chip.
[0049]
[0050] According to the eleventh aspect, the present application further provides a computer-readable storage medium configured to store a computer program. The computer program is configured to implement the method in any one of the first aspect or possible implementations of the first aspect.
[0050]
[0051] According to the twelfth aspect, an embodiment of the present application further provides a computer program product including instructions. When the computer program product runs on a computer, the computer implements the method in any one of the first aspect or possible implementations of the first aspect.
[0051]
[0052] The encoding device, decoding device, encoder, decoder, computer storage medium, computer program product, and chip provided in the embodiments are all configured to execute the method provided above. Therefore, for the achievable beneficial effects, please refer to the beneficial effects of the method provided above. Details will not be described again here.
Brief Description of the Drawings
[0052]
[0053] To more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings for explaining the embodiments are briefly described below. The accompanying drawings in the following description only show some embodiments of the present application, and it is obvious that those skilled in the art can further derive other Attached drawings from these accompanying drawings without creative efforts.
Figure 1a
[0054] FIG. 1a is an exemplary block diagram of a coding system according to an embodiment of the present application.
Figure 1b
[0055] FIG. 1b is an exemplary block diagram of a video coding system according to an embodiment of the present application.
Figure 2
[0056] Figure 2 is an exemplary block diagram of a video encoder according to an embodiment of the present application.
Figure 3
[0057] Figure 3 is an exemplary block diagram of a video decoder according to an embodiment of the present application.
Figure 4
[0058] Figure 4 is a schematic diagram of an example of a candidate picture block according to an embodiment of the present application.
Figure 5
[0059] Figure 5 is an exemplary block diagram of a video coding device according to an embodiment of the present application.
Figure 6
[0060] Figure 6 is an exemplary block diagram of a device according to an embodiment of the present application.
Figure 7a
[0061] Figure 7a is a schematic diagram of an application scenario according to an embodiment of the present application.
Figure 7b
[0062] Figure 7b is a schematic diagram of an application scenario according to an embodiment of the present application.
Figure 8
[0063] Figure 8 is a schematic flowchart of an encoding and decoding method according to an embodiment of the present application.
Figure 9
[0064] Figure 9 is a schematic diagram of the structure of an encoding and decoding system according to an embodiment of the present application.
Figure 10
[0065] Figure 10 is a schematic flowchart of another encoding and decoding method according to an embodiment of the present application.
Figure 11
[0066] Figure 11 is a schematic diagram of the structure of another encoding and decoding system according to an embodiment of the present application.
Figure 12
[0067] Figure 12 is a schematic diagram of the structure of yet another encoding and decoding system according to an embodiment of the present application.
Figure 13
[0068] FIG. 13 is a schematic flowchart of yet another encoding and decoding method according to an embodiment of the present application.
Figure 14
[0069] FIG. 14 is a schematic diagram of the structure of yet another encoding and decoding system according to an embodiment of the present application.
Figure 15
[0070] FIG. 15 is a schematic diagram of the performance of the encoding and decoding method according to an embodiment of the present application.
Figure 16
[0071] FIG. 16 is a schematic diagram of an application scenario according to an embodiment of the present application.
Figure 17
[0072] FIG. 17 is a schematic diagram of another application scenario according to an embodiment of the present application.
Figure 18
[0073] FIG. 18 is a schematic diagram of the structure of an encoding and decoding device according to an embodiment of the present application.
Figure 19
[0074] FIG. 19 is a schematic diagram of the structure of another encoding and decoding device according to an embodiment of the present application.
Figure 20
[0075] FIG. 20 is a schematic diagram of the structure of a chip according to an embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0053]
[0076] Hereinafter, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Surely It is clear that the described embodiments are only a part, not all, of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0054]
[0077] The term "and / or" in this specification only describes the association relationship for describing associated objects, indicating that three relationships can exist. For example, A and / or B may represent the following three cases: only A exists, both A and B exist, and only B exists.
[0055]
[0078] In the specification and attached drawings of this application, terms such as "first", "second", etc. are intended to distinguish different objects or different processes of the same object, but do not indicate a specific order of the objects.
[0056]
[0079] Further, the terms "comprising", "having", or other variants thereof in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units is not limited to the listed steps or units, and optionally may further include another unlisted step or unit, or optionally may further include another specific step or unit of the process, method, System product, or device.
[0057]
[0080] In the description of the embodiments of this application, it should be noted that the words "example" or "for example" are used to give an example, illustration, or explanation. Any embodiment or design described as an "example" or "for example" in the embodiments of this application should not be described as being more preferable or having more advantages than another embodiment or design. Strictly speaking, the use of the terms "example", "for example", or the like is intended to present a relative concept in a specific manner.
[0058]
[0081] In the description of this application, unless otherwise specified, "a plurality of" means two or more.
[0059]
[0082] Embodiments of the present application provide AI-based data compression / decompression techniques, and in particular, neural network-based data compression / decompression techniques, and specifically, provide encoding and decoding techniques to improve conventional hybrid data encoding and decoding systems.
[0060]
[0083] Data encoding and decoding includes data encoding and data decoding. Data encoding is performed on the source side (or usually referred to as the encoder side) and typically includes processing raw data (e.g., compressing) to reduce the amount of data required to represent the raw data (for more efficient storage and / or transmission). Data decoding is performed on the destination side (or usually referred to as the decoder side) and typically includes the reverse process to the encoder side to reconstruct the raw data. "Encoding and decoding" of data in the embodiments of the present application should be understood as "encoding" or "decoding" of data. The combination of the encoding unit and the decoding unit is also called a CODEC (encoding and decoding, CODEC).
[0061]
[0084] In the case of lossless data coding, it is possible to reconstruct the raw data. In other words, the reconstructed raw data has the same quality as the raw data (assuming that no transmission loss or other data loss has occurred during storage or transmission). In the case of non-lossless data coding, in order to reduce the amount of data required to represent the raw data, further compression is performed, for example, by quantization, and the raw data cannot be completely reconstructed on the decoder side. In other words, the quality of the reconstructed raw data may be lower or worse than the quality of the raw data.
[0062]
[0085] Embodiments of the present invention may be applicable to video data, picture data, audio data, integer data, and other data having compression / decompression requirements. Hereinafter, embodiments of the present application will be described by using video coding (briefly referred to as video coding). For other types of data (for example, picture data, audio data, integer data, and other data having compression / decompression requirements), refer to the following description. Details are not described in the embodiments of the present application. In the process of coding data such as audio data or integer data as compared with video coding, the data does not need to be divided into blocks and the data may be directly coded.
[0063]
[0086] Video coding generally refers to the processing of a sequence of pictures that form a video or a video sequence. In the field of video coding, the terms "picture", "frame", and "image" may be used synonymously.
[0064]
[0087] Some video coding standards are used for "lossless hybrid video coding" (i.e., spatial and temporal prediction in the pixel domain is combined with 2D transform coding to apply quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in the encoder, the video is typically processed, i.e., encoded, at the block (video block) level. For example, prediction blocks are generated by spatial (intra-picture) prediction and temporal (inter-picture) prediction, and the prediction blocks are subtracted from the current block (the block being processed or scheduled to be processed) to obtain a residual block, which is transformed and quantized in the transform domain to reduce (compress) the amount of data to be transmitted. On the decoder side, an inverse processing unit is applied to the encoded blocks or compressed blocks as compared to the encoder to reconstruct the current block for presentation. Further, the encoder replicates the processing steps of the decoder, and as a result, the encoder and the decoder perform the same prediction (e.g., intra prediction or inter prediction) and / or pixel reconstruction and process, i.e., code, subsequent blocks.
[0065]
[0088] In the following embodiments of the coding system 10, the encoder 20 and the decoder 30 are described with reference to FIGS. 1a to 3.
[0066]
[0089] FIG. 1a is an exemplary block diagram of a coding system 10 according to an embodiment of the present application, for example, a video coding system 10 (abbreviated as the coding system 10) that can utilize the techniques of the present application. The video encoder 20 (or abbreviated as the encoder 20) and the video decoder 30 (or abbreviated as the decoder 30) of the video coding system 10 represent devices that may be configured to execute the techniques according to various examples described herein.
[0067]
[0090] As shown in FIG. 1a, the coding system 10 includes a source device 12 configured to provide encoded picture data 21, such as an encoded picture, to a destination device 14 for decoding the encoded picture data 21.
[0068]
[0091] The source device 12 includes an encoder 20 and additionally, i.e., optionally, may include a picture source 16, a pre-processor (or pre-processing unit) 18, e.g., a picture pre-processor, and a communication interface (or communication unit) 22.
[0069]
[0092] The picture source 16 may include any kind of picture capture device, such as a camera that captures real-world pictures, and / or any kind of picture generation device, such as a computer graphics processor that generates computer animation pictures, or real-world pictures, computer-generated pictures (e.g., screen content or virtual reality (VR) pictures), and / or any other device that acquires and / or provides any combination thereof (e.g., augmented reality (AR) pictures), or it may be or consist of those. The picture source may be any kind of memory or storage that stores any of the above pictures.
[0070]
[0093] To distinguish the processing performed by the pre-processor (or pre-processing unit) 18, the picture (or picture data) 17 may also be referred to as raw picture (or raw picture data) 17.
[0071]
[0094] The pre - processing processor 18 is configured to receive the raw picture data 17, pre - process the raw picture data 17, and obtain the pre - processed picture (or pre - processed picture data) 19. The pre - processing executed by the pre - processing processor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the pre - processing unit 18 may be an optional component.
[0072]
[0095] The video encoder (or encoder) 20 is configured to receive the pre - processed picture data 19 and provide the encoded picture data 21 (further details will be described later, for example, based on FIG. 2).
[0073]
[0096] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or some further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0074]
[0097] The destination device 14 includes a decoder 30 and may further, i.e., optionally, include a communication interface (or communication unit) 28, a post - processing processor (or post - processing unit) 32, and a display device 34.
[0075]
[0098] The communication interface 28 of the destination device 14 is configured to directly receive the encoded picture data 21 (or some further processed version thereof) from the source device 12 or any other source device such as a storage device, and provide the encoded picture data 21 to the decoder 30. For example, the storage device is an encoded picture data storage device.
[0076]
[0099] The communication interfaces 22 and 28 are capable of transmitting or receiving encoded picture data 21 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination thereof.
[0077]
[0100] The communication interface 22 is capable of, for example, packaging the encoded picture data 21 into a suitable format, such as a packet, and / or processing the encoded picture data using any type of transmission encoding or processing for transmission over a communication link or communication network.
[0078]
[0101] The communication interface 28 corresponding to the communication interface 22 is capable of, for example, receiving the transmitted data and processing the transmitted data using any type of corresponding transmission decoding or processing and / or unpackaging to obtain the encoded picture data 21.
[0079]
[0102] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface or a bidirectional communication interface indicated by an arrow regarding the communication channel 13 from the source device 12 to the destination device 14 in FIG. 1a, and can be configured to, for example, send and receive messages, for example, set up a connection, and confirm and exchange any other information related to the communication link and / or data transmission, such as encoded picture data transmission.
[0080]
[0103] Video decoder (or decoder) 30 is configured to receive the encoded picture data 21 and provide the decoded picture data (or decoded picture data) 31. (Further details are described, for example, based on FIG. 3.)
[0081]
[0104] Post-processing processor 32 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), such as the decoded picture 31, to obtain the post-processed picture data 33, such as the post-processed picture. The post-processing executed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, resampling, or any other processing, for example, for preparing the decoded picture data 31 for display by the display device 34.
[0082]
[0105] Display device 34 is configured to receive the post-processed picture data 33 and display the picture, for example, to a user or viewer. The display device 34 may be any type of display that represents the reconstructed picture, such as an integrated or external display or monitor, or may include them. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0083]
[0106] Coding system 10 further includes a training engine 25. The training engine 25 is configured to train the encoder 20 such that the encoder 20 (in particular, the entropy encoding unit 270 of the encoder 20) or the decoder 30 (in particular, the entropy decoding unit 304 of the decoder 30) performs entropy encoding on the target picture block to be encoded based on an estimated probability distribution obtained by estimation. For a detailed description of the training engine 25, refer to the following method embodiments.
[0084]
[0107] Although FIG. 1a depicts the source device 12 and the destination device 14 as separate devices, device embodiments may alternatively include both the source device 12 and the destination device 14, or the functions of both the source device 12 and the destination device 14, i.e., the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented by the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.
[0085]
[0108] As will be apparent to those skilled in the art based on the specification, the functions or the presence and (exact) partitioning of various units in the source device 12 and / or the destination device 14 shown in FIG. 1a may vary depending on the actual device and application.
[0086]
[0109] Figure 1b is an exemplary block diagram of a video coding system 40 according to an embodiment of the present application. The encoder 20 (e.g., video encoder 20), or the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30 can be realized by a processing circuit of a video coding system as shown in Figure 1b, for example, by one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, a processor dedicated to video coding, or any combination thereof. Refer to Figures 2 and 3. Figure 2 is an exemplary block diagram of a video encoder according to an embodiment of the present application, and Figure 3 is an exemplary block diagram of a video decoder according to an embodiment of the present application. The encoder 20 may be implemented by the processing circuit 46 to embody various modules described with respect to the encoder 20 of Figure 2 and / or any other encoder system or subsystem described in this specification. The decoder 30 may be implemented by the processing circuit 46 to embody various modules described with respect to the decoder 30 of Figure 3 and / or any other decoder system or subsystem described in this specification. The processing circuit 46 can be configured to perform various operations as described later. As shown in Figure 5, when the technology is partially realized in software, the device may store software instructions in a suitable non-transitory computer-readable storage medium and may execute the instructions in hardware using one or more processors to execute the technology of the present application. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of an integrated encoder / decoder (CODEC) within a single device as shown in Figure 1b.
[0087]
[0110] The source device 12 and the destination device 14 can be any device within a wide range including any type of handheld or stationary device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video gaming console, video streaming device (such as a content service server or content delivery server, etc.), broadcast receiving device, broadcast transmitting device, monitoring device, etc., and may or may not use any type of operating system. The source device 12 and the destination device 14 can also be devices in a cloud computing scenario, such as virtual machines in a cloud computing scenario. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Therefore, the source device 12 and the destination device 14 can be wireless communication devices.
[0088]
[0111] Virtual scenario applications (applications, APPs) such as virtual reality (VR) applications, augmented reality (AR) applications, or mixed reality (MR) applications may be installed on each of the source device 12 and the destination device 14, and VR applications, AR applications, or MR applications may be user-operated (e.g., tap, touch, slide, ShakeIt may be executed based on (e.g., gesture, or voice control). The source device 12 and the destination device 14 may each capture a picture / video of any object in the environment using a camera and / or a sensor, and may display a virtual object on the display device based on the captured picture / video. The virtual object may be a virtual object in a VR scenario, an AR scenario, or an MR scenario (i.e., an object in a virtual environment).
[0089]
[0112] In this embodiment of the present application, the virtual scenario applications in the source device 12 and the destination device 14 may be built-in applications of the source device 12 and the destination device 14, or may be applications provided by a third-party service provider and installed by the user. This is not particularly limited in this case.
[0090]
[0113] Further, a real-time video transmission application, for example, a live broadcast application, may be installed in each of the source device 12 and the destination device 14. The source device 12 and the destination device 14 may each use a camera to capture a picture / video and can display the captured picture / video on the display device.
[0091]
[0114] In some cases, the video coding system 10 shown in FIG. 1a is merely an example, and the technology of this application is applicable to video coding settings (e.g., video encoding or video decoding) that do not necessarily involve any data communication between the encoding device and the decoding device. In other examples, data is retrieved from local memory and streamed over a network, etc. The video encoding device can encode data and store the encoded data in memory, and / or the video decoding device can retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but encode data to / from memory and / or retrieve data from memory to decode the data.
[0092]
[0115] FIG. 1b is an exemplary block diagram of a video coding system 40 according to this embodiment of the present application. As shown in FIG. 1b, the video coding system 40 may include an imaging device 41, a video encoder 20, and a video decoder 30 (and / or a video encoder / decoder implemented by processing circuit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0093]
[0116] As shown in FIG. 1b, the imaging device 41, the antenna 42, the processing circuit 46, the video encoder 20, the video decoder 30, the processor 43, the memory 44, and / or the display device 45 can communicate with each other. In various examples, the video coding system 40 may include only the video encoder 20 or only the video decoder 30.
[0117] In some examples, the antenna 42 can be configured to transmit or receive an encoded bitstream of video data. Further, in some examples, the display device 45 may be configured to present video data. The processing circuit 46 can include, for example, application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, and the like. The video coding system 40 may also include an optional processor 43. The optional processor 43 can similarly include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, and the like. Further, the memory 44 can be any type of memory, such as, for example, volatile memory (e.g., static random-access memory (SRAM) or dynamic random-access memory (DRAM)) or non-volatile memory (e.g., flash memory). In a non-limiting example, the memory 44 can be implemented by a cache memory. In other examples, the processing circuit 46 can include memory (e.g., a cache) for implementing a picture buffer.
[0094]
[0118] In some examples, a video encoder 20 implemented by a logic circuit may include a picture buffer (implemented by, for example, a processing circuit 46 or a memory 44) and a graphics processing unit (implemented by, for example, a processing circuit 46). The graphics processing unit can be communicatively coupled to the picture buffer. The graphics processing unit is included in the video encoder 20 implemented by the processing circuit 46 and can implement various modules described with reference to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit can be configured to perform various operations described herein.
[0095]
[0119] In some examples, a video decoder 30 may be implemented in a similar manner by a processing circuit 46 to implement various modules described with reference to the video decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, a video decoder 30 implemented by a logic circuit may include a picture buffer (implemented by a processing circuit 46 or a memory 44) and a graphics processing unit (implemented by, for example, a processing circuit 46). The graphics processing unit can be communicatively coupled to the picture buffer. The graphics processing unit is included in the video decoder 30 implemented by the processing circuit 46 and can implement various modules described with reference to FIG. 3 and / or any other decoder system or subsystem described herein.
[0096]
[0120] In one example, antenna 42 can be configured to receive an encoded bitstream of video data. As described, the bitstream to be encoded can include data related to video frame encoding, indicators, index values, mode selection data, etc., described herein, e.g., data related to encoding partitioning (e.g., transform coefficients or quantized transform coefficients, optional indicators (such as those described), and / or data defining the encoding partitioning). The video coding system 40 can further include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present video frames.
[0097]
[0121] In this embodiment of the present application, it should be understood that for the example described with reference to the video encoder 20, the video decoder 30 can be configured to perform reverse processing. Regarding the signaling syntax elements, the video decoder 30 can be configured to receive and analyze such syntax elements and, correspondingly, decode the relevant video data. In some examples, the video encoder 20 can entropy encode the syntax elements into the encoded video bitstream. In such examples, the video decoder 30 can analyze such syntax elements and, correspondingly, decode the relevant video data.
[0098]
[0122] For the sake of simplicity of description, embodiments of the present application are described by referring to the Versatile Video Coding (VVC) reference software developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG), or the High-Efficiency Video Coding (HEVC). Those skilled in the art understand that the embodiments of the present application are not limited to HEVC or VVC.
[0099]
[0123] Encoder and encoding method
[0124] As shown in FIG. 2, the video encoder 20 includes an input end (or input interface) 201, a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output end (or output interface) 272. The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a video encoder or a hybrid video encoder based on a hybrid video codec.
[0100]
[0125] Refer to FIG. 2. The inter prediction unit is a trained target model (also called a neural network). The neural network is configured to process an input picture, a picture area, or a picture block to generate a predictor for the input picture block. For example, a neural network for inter prediction is configured to receive an input picture, a picture area, or a picture block and generate a predictor for the input picture, picture area, or picture block.
[0101]
[0126] The residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, and the mode selection unit 260 form the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 form the backward signal path of the encoder. The backward signal path of the encoder 20 corresponds to the signal path of a decoder (refer to the decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter prediction unit 244, and the intra prediction unit 254 also form the "built-in decoder" of the video encoder 20.
[0102]
[0127] Picture and picture partitioning (picture and block)
[0128] Encoder 20 can be configured to receive a picture (or picture data) 17, for example, a picture within a sequence of pictures forming a video or video sequence, via an input terminal 201. The received picture or picture data may also be a preprocessed picture (or preprocessed picture data) 19. For simplicity, the following description refers to picture 17. Picture 17 may also be referred to as a current picture or a target picture to be encoded (in particular, in video encoding, to distinguish the current picture from other pictures, for example, pictures previously encoded and / or decoded in the same video sequence, i.e., the video sequence including the current picture).
[0103]
[0129] A (digital) picture is or can be considered as a two - dimensional array or matrix of samples having intensity values. The samples in the array may also be referred to as pixels (pix or pels, short for picture elements). The number of samples in the horizontal and vertical directions (or axes) of the array defines the size and / or resolution of the picture. For color representation, usually three color components are used. Specifically, a picture may be represented as or may include three sample arrays. In the RGB format or color space, a picture includes the corresponding red, green, and blue sample arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, which includes a luminance component (sometimes denoted by Y or L) represented by Y and two chrominance components represented by Cb and Cr. The luminance (luma) component Y represents the brightness or gray - level intensity (e.g., they are the same in a gray - scale picture), and the two chrominance (or chroma for short) components Cb and Cr represent the chrominance or color - information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format can be converted or transformed into the YCbCr format and vice versa. The process is also referred to as color conversion or transformation. If a picture is monochrome, it can include only a luminance sample array. Thus, a picture can be, for example, an array of luminance samples in a monochrome format or in 4:2:0, 4:2:2, and 4:4:4 color formats of an array of luminance samples and two corresponding arrays of chrominance samples. Picture in It is possible to convert or transform it into the YCbCr format and vice versa. The process is also referred to as color conversion or transformation. If a picture is monochrome, it can include only a luminance sample array. Thus, a picture can be, for example, an array of luminance samples in a monochrome format or in 4:2:0, 4:2:2, and 4:4:4 color formats of an array of luminance samples and two corresponding arrays of chrominance samples.
[0104]
[0130] In an embodiment, an embodiment of the video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs), as in the H.265 / HEVC and VVC standards. The partitioning unit may be configured to use the same block size and corresponding grid defining the block size for all pictures of the video sequence, or to vary the block size between pictures, subsets, or groups of pictures and partition each picture into corresponding blocks.
[0105]
[0131] In other embodiments, the video encoder may be configured to directly receive blocks 203 of picture 17, such as one, several, or all of the blocks forming picture 17. The picture blocks 203 may also be referred to as current picture blocks or picture blocks to be encoded.
[0106]
[0132] Similar to Picture 17, Block 203 can again be considered as, or be, a two-dimensional array or matrix of samples having intensity values (sample values), although of dimensions smaller than Picture 17. In other words, Block 203 may include one sample array (e.g., the luminance array in the case of a monochrome Picture 17, or the luminance or chrominance arrays in the case of a color picture), or three sample arrays (e.g., one luminance array and two chrominance arrays in the case of a color Picture 17), or any other quantity and / or type of array depending on the color format used. The number of samples in the horizontal and vertical directions (or axes) of Block 203 defines the size of Block 203. Thus, the block may be an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.
[0107]
[0133] In an embodiment, the video encoder 20 shown in FIG. 2 can be configured to encode Picture 17 block by block, for example, encoding and prediction are performed for each Block 203.
[0108]
[0134] In an embodiment, the video encoder 20 shown in FIG. 2 can further be configured to partition and / or encode a picture using slices (also called video slices), where the picture can be partitioned into one or more slices (typically non-overlapping), or encoded using them. Each slice can include one or more blocks (e.g., coding tree units CTUs) or one or more block groups (e.g., tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).
[0109]
[0135] In an embodiment, the video encoder 20 shown in FIG. 2 can be further configured to partition and / or encode a picture using slices / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). A picture can be partitioned into and / or encoded using one or more slices / tile groups (typically non-overlapping), and each slice / tile group can include one or more blocks (e.g., CTUs) or one or more tiles. Each tile can have a rectangular shape and can include one or more blocks (e.g., CTUs), e.g., complete or fragmented blocks.
[0110]
[0136] Residual calculation
[0137] The residual calculation unit 204 can be configured to calculate the residual block 205, for example, by subtracting the sample values of the prediction block 265 (further details regarding the prediction block 265 will be provided later) from the sample values of the picture block 203 (original block) 203 on a sample-by-sample (pixel-by-pixel) basis to obtain the residual block 205 in the sample domain.
[0111]
[0138] Transformation
[0139] The transformation processing unit 206 is configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain the transformation coefficients 207 in the transform domain. The transformation coefficients 207 are also referred to as transform residual coefficients and can also represent the residual block 205 in the transform domain.
[0112]
[0140] The conversion processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the conversion specified in H.265 / HEVC. Compared to the orthogonal DCT transform, such an integer approximation is typically scaled based on a certain factor. To maintain the norm of the residual block processed by using the forward transform and the inverse transform, an additional scale factor is applied as part of the conversion process. The scale factor is typically selected based on several constraints, for example, the scale factor is a power of 2 for shift operations, the bit depth of the conversion coefficients, and the trade-off between accuracy and implementation cost. For example, a specific scale factor is specified for the inverse conversion by the inverse conversion processing unit 212 on the encoder side 20 (and the corresponding inverse conversion by the inverse conversion processing unit 312 on the decoder side 30, for example), and correspondingly, the corresponding scale factor may be specified for the forward conversion by the conversion processing unit 206 on the encoder side 20.
[0113]
[0141] In an embodiment, the video encoder 20 (correspondingly, the conversion processing unit 206) can be configured to output conversion parameters, such as the type of one or more conversions, for example, the type of being after or directly being the encoding or compression executed by the entropy encoding unit 270. As a result, for example, the video decoder 30 can receive and use the conversion parameters for decoding.
[0114]
[0142] Quantization
[0143] The quantization unit 208 may be configured to quantize the conversion coefficients 207 to obtain the quantized conversion coefficients 209, for example, by applying scalar quantization or vector quantization. The quantized conversion coefficients 209 may also be referred to as quantized residual coefficients 209.
[0115]
[0144] The quantization process may reduce the bit depth associated with all or part of the conversion coefficient 207. For example, an n-bit conversion coefficient can be rounded to an m-bit conversion coefficient during quantization, where n is greater than m. The degree of quantization can be changed by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, different scales may be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, and larger quantization steps correspond to coarser quantization. The appropriate quantization step may be specified by the quantization parameter (QP). For example, the quantization parameter may be an index to a predefined set of appropriate quantization steps. For example, a smaller quantization parameter may correspond to finer quantization (smaller quantization step), and a larger quantization parameter may correspond to coarser quantization (larger quantization step), or vice versa. Quantization may include division by the quantization step, and the corresponding and / or inverse quantization may include multiplication by the quantization step, for example, performed by the inverse quantization unit 210. Embodiments according to certain standards such as HEVC may be configured to determine the quantization step using the quantization parameter. Generally, the quantization step may be calculated based on the quantization parameter by using a fixed-point approximation of an expression including division. Additional scale factors are introduced for quantization and dequantization to be able to restore the norm of the residual block, where the norm of the residual block may be modified due to the scale used in the fixed-point approximation of the expression related to the quantization step and the quantization parameter. In one example implementation, the scale of the inverse transform may be combined with the scale of dequantization. Alternatively, a customized quantization table may be used and signaled from the encoder to the decoder, for example, in the bitstream. Quantization is a lossless operation, where larger quantization steps indicate larger losses.
[0116]
[0145] In an embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, after or directly after the encoding or compression performed by the entropy encoding unit 270, such that, for example, the video decoder 30 can receive and use the quantization parameter for decoding.
[0117]
[0146] Inverse quantization
[0147] The inverse quantization unit 210 is configured to apply the inverse of the quantization method applied by the quantization unit 208 to the quantized coefficients, for example, based on the quantization unit 208 or by using the same quantization step as the quantization unit 208, in order to obtain the dequantized coefficients 211. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211 and usually differ from the transform coefficients due to losses caused by quantization, but correspond to the transform coefficients 207.
[0118]
[0148] Inverse transform
[0149] The inverse transform processing unit 212 is configured to apply the inverse of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain the reconstructed residual block 213 (or the corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0119]
[0150] Reconstruction
[0151] The reconstruction unit 214 (e.g., adder 214) is configured to add the conversion block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the sample domain, for example, by adding the sample value of the reconstructed residual block 213 and the sample value of the prediction block 265.
[0120]
[0152] Filtering
[0153] The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter the reconstructed samples to obtain filtered sample values. For example, the loop filter unit is configured to smooth pixel transitions or improve video quality. The loop filter unit 220 may include a deblocking filter, a sample - adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in - loop rescaler) is added. This process is executed before deblocking. In another example, the deblocking filter process may be applied to internal sub - block edges, such as affine sub - block edges, ATMVP sub - block edges, sub - block transform (SBT) edges, and intra sub - partition (ISP) edges. Although the loop filter unit 220 is shown as a loop filter in FIG. 2, in other configurations, the loop filter unit 220 may be implemented as a post - loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0121]
[0154] In an embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be configured to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, after or directly after entropy encoding performed by the entropy encoding unit 270, such that the decoder 30 can receive and use the same or different loop filter parameters for decoding.
[0122]
[0155] Decoded picture buffer
[0156] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use in video data encoding by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, including, for example, a dynamic random-access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or another type of memory device. The decoded picture buffer 230 may be configured to store one or more filtering blocks 221. The decoded picture buffer 230 may be further configured to store other blocks, such as other previously filtered blocks, such as previously reconstructed and filtered blocks 221, of the same current picture or different pictures, such as previously reconstructed pictures, and may also provide, for example, for inter prediction, a fully previously reconstructed, such as decoded, picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer 230 may be further configured to store, for example, one or more unfiltered reconstructed blocks 215 or generally unfiltered reconstructed samples if the reconstructed blocks 215 have not been filtered by the loop filter unit 220, or to store other further processed versions of the reconstructed blocks or samples.
[0123]
[0157] Mode Selection (Partitioning and Prediction)
[0158] The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is also configured to receive or obtain raw picture data, for example, the original block 203 (the current block 203 of the current picture 17), reconstructed picture data, for example, of the same (current) picture, and / or from one or more previously decoded pictures, for example, from the decoded picture buffer 230 or another buffer (for example, a line buffer not shown in FIG. 2), filtered and / or unfiltered reconstructed samples or blocks. The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or predictor 265.
[0124]
[0159] The mode selection unit 260 may be configured to determine or select the partitioning (including non-partitioning) and prediction mode (for example, intra or inter prediction mode) of the current block and generate the corresponding prediction block 265, which is used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.
[0125]
[0160] In an embodiment, the mode selection unit 260 can be configured to select a partitioning and prediction mode (e.g., from those available or supported by the mode selection unit 260). The prediction mode provides an optimal match or a minimum residual (the minimum residual means better compression for transmission or storage), or a minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or considers or balances both. The mode selection unit 260 can be configured to determine the partitioning and prediction mode based on bit rate distortion optimization (RDO), e.g., select a prediction mode that provides minimum bit rate distortion optimization. Terms such as "best," "lowest," "optimal," etc. in this specification do not necessarily mean the general "best," "lowest," "optimal," and may mean a situation where the criteria for termination or selection are met. For example, a value that exceeds or falls below a threshold or other constraint is a "suboptimal choice," but may reduce complexity and processing time.
[0126]
[0161] In other words, the partitioning unit 262 may be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs), and the CTU 203 may further partition into smaller block partitions or sub-blocks (re-forming the blocks), for example, by repeatedly using quad-tree partitioning (QT), binary-tree partitioning (BT), triple-tree partitioning (TT), or any combination thereof, and, for example, may be configured to perform prediction for each of the block partitions or sub-blocks, where the mode selection includes selection of the tree structure of the partitioned block 203 and the prediction mode applied to each of the block partitions or sub-blocks.
[0127]
[0162] Hereinafter, the partitioning (e.g., by the partitioning unit 262) and prediction (e.g., by the inter prediction unit 244 and the intra prediction unit 254) performed by the video encoder 20 will be described in detail.
[0128]
[0163] Partitioning
[0164] The partitioning unit 262 is capable of partitioning (or splitting) a picture block (or CTU) 203 into smaller partitions, for example, smaller square or rectangular blocks. In the case of a picture having three sample arrays, a CTU includes an N×N block of luminance samples together with two corresponding blocks of chrominance samples. The maximum allowable size of the luminance block of a CTU is defined to be 128×128 in the developing Versatile Video Coding (VVC) standard, but may be defined to be a value other than 128×128 in the future, for example, 256×256. The CTUs of a picture may be clustered / grouped as slice / tile groups, tiles, or bricks. A tile covers a rectangular area of a picture, and a tile may be divided into one or more bricks. A brick includes a plurality of CTU rows in a tile. A tile that is not partitioned into a plurality of bricks may be called a brick. However, a brick is a true subset of a tile and is not called a tile. In VVC, two tile group modes, namely, a raster scan slice / tile group mode and a rectangular slice mode, are supported. In the raster scan tile group mode, a slice / tile group includes a sequence of tiles in the tile raster scan of a picture. In the rectangular slice mode, a slice includes a plurality of bricks of a picture that collectively form a rectangular area of the picture. The bricks within a rectangular slice are in the order of the brick raster scan of the slice. These small blocks (which may also be called sub-blocks) may be further partitioned into even smaller partitions. This is also referred to as tree partitioning or hierarchical tree partitioning, where, for example, a root block at root tree level 0 (hierarchical level 0 or depth 0) is recursively partitioned into two or more blocks of nodes at the next lower tree level, for example, tree level 1 (hierarchical level 1 or depth 1).These blocks can again be partitioned into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2 or depth 2), until the partitioning is complete (due to the completion criteria being met, e.g., reaching the maximum tree depth or the minimum block size). Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree using partitioning into two partitions is called a binary-tree (BT), a tree using partitioning into three partitions is called a ternary-tree (TT), and a tree using partitioning into four partitions is called a quad-tree (QT).
[0129]
[0165] For example, a coding tree unit (CTU) may be, or may include, a CTB of luminance samples of a picture having three sample arrays, two corresponding CTBs of chrominance samples, or a CTB of samples of a monochrome picture or a picture (where a picture is coded by using three separate color planes and syntax structures (used to code samples)). Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some N value such that the partitioning of a component into CTBs is a partitioning. A coding unit (CU) may be, or may include, a coding block of luminance samples of a picture having three sample arrays, two corresponding coding blocks of chrominance samples, or a coding block of samples of a monochrome picture or a picture (where a picture is coded by using three separate color planes and syntax structures (used to code samples)). Correspondingly, a coding block (CB) may be an M×N block of samples for some M and N values such that the partitioning of a CTB into coding blocks is a partitioning.
[0130]
[0166] In an embodiment, for example, in accordance with HEVC, by using a quad-tree structure shown as a coding tree, a coding tree unit (CTU) may be divided into a plurality of CUs. A determination of whether to code a picture area using inter (temporal) prediction or intra (spatial) prediction is made at the leaf CU level. Each leaf CU may be further divided into one, two, or four PUs based on the PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder in PU units. After obtaining a residual block by applying a prediction process based on the PU partition type, the leaf CU may be partitioned into transform units (TUs) based on another quad-tree structure similar to the coding tree of the CU.
[0131]
[0167] In an embodiment, for example, according to the latest video coding standard currently under development (referred to as Versatile Video Coding (VVC)), a combined quad-tree nested multi-type tree (e.g., a binary tree and a ternary tree) divides a segmentation structure used to partition coding tree units. In the coding tree structure within a coding tree unit, a CU may have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quad-tree structure. Then, a quad-tree leaf node can be further partitioned by a multi-type tree structure. In a multi-type tree structure, there are four splitting types: vertical binary tree split (SPLIT_BT_VER), horizontal binary tree split (SPLIT_BT_HOR), vertical ternary tree split (SPLIT_TT_VER), and horizontal ternary tree split (SPLIT_TT_HOR). A multi-type tree leaf node is called a coding unit (CU), and as long as the CU is not overly large relative to the maximum transform length, this split is used for prediction and transform processing without further partitioning. In most cases, CUs, PUs, and TUs have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of the CU. VVC has developed a unique signaling mechanism for separating split information in a quadtree using a nested multi-type tree coding structure. In this signaling mechanism, a coding tree unit (CTU) is treated as the root of the quadtree and is first partitioned by the quadtree structure. Each quadtree leaf node (if large enough to allow it) is then further partitioned by a multi-type tree structure.In a multi-type tree structure, a first flag (mtt_split_cu_flag) is signaled to specify whether a node is further partitioned; if the node is further partitioned, a second flag (mtt_split_cu_vertical_flag) is signaled to specify the splitting direction; then, a third flag (mtt_split_cu_binary_flag) is signaled to specify whether the split is a binary tree split or a ternary tree split. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived by the decoder based on predefined rules or tables. In a specific design, for example, in a 64×64 luma block and 32×32 chroma pipeline design in a VVC hardware decoder, if either the width or height of the luminance coding block is greater than 64, TT splitting is prohibited. Also, if either the width or height of the chroma coding block is greater than 32, TT splitting is prohibited. The pipeline design divides a picture into virtual pipeline data units (VPDUs), and a VPDU is defined as a non-overlapping unit within the picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. The size of a VPDU is roughly proportional to the buffer size in most pipeline stages, and thus it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the partitioning of ternary tree (TT) and binary tree (BT) may lead to an increase in the VPDU size.
[0132]
[0168] Furthermore, it should be noted that if a portion of the tree node block extends beyond the bottom or right picture boundary, the tree node block is forcibly split until all samples of all coded CUs are located inside the picture boundary.
[0133]
[0169] For example, the intra sub-partition (ISP) tool can split a luma intra-predicted block vertically or horizontally into two or four sub-partitions according to the block size.
[0134]
[0170] In one example, the mode selection unit 260 of the video encoder 20 can be configured to execute any combination of the above-described partitioning techniques.
[0135]
[0171] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a (pre-determined) set of prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0136]
[0172] Intra prediction
[0173] The intra prediction mode set may include 35 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes like those defined in HEVC, or may include 67 different intra prediction modes, such as non - directional modes like DC (or average) mode and planar mode, or directional modes like those defined in VVC. For example, some conventional angular intra prediction modes are adaptively replaced with wide - angle intra prediction modes for non - square blocks as defined in VVC. As another example, to avoid the division operation for DC prediction, only the longer side is used to calculate the average for non - square blocks. Further, the result of intra prediction in the planar mode may be further modified by using the position dependent intra prediction combination (PDPC) method.
[0137]
[0174] The intra prediction unit 254 is configured to generate an intra prediction block 265 based on an intra prediction mode within the intra prediction mode set, using the reconstructed samples of neighboring blocks of the same current picture.
[0138]
[0175] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output to the entropy encoding unit 270 in the form of a syntax element 266 for including the intra prediction parameter (or generally information indicating the intra prediction mode selected for the block) in the encoded picture data 21, so that as a result, for example, the video decoder 30 can receive and use the prediction parameter for decoding.
[0139]
[0176] The intra prediction modes in HEVC include a DC prediction mode, a planar prediction mode, and 33 angular prediction modes. That is, there are a total of 35 prediction mode candidates. The current block can use the pixels of the reconstructed picture blocks on the left and above as references for performing intra prediction. The picture blocks within the peripheral area of the current block that are used for intra prediction of the current block become reference blocks, and the pixels within the reference blocks are referred to as reference pixels. Among the 35 candidate prediction modes, the DC prediction mode is applicable to an area where the texture of the area is flat in the current block, and all the pixels in the area use the average value of the reference pixels in the reference block as a prediction. The planar prediction mode is applicable to a certain picture block where the texture of the picture block changes smoothly. For the current block that meets this condition, bilinear interpolation is performed by using the reference pixels in the reference block as predictions for all the pixels in the current block. In the angular prediction mode, taking advantage of the feature that the texture of the current block is strongly correlated with the texture of the neighboring reconstructed picture blocks, the values of the reference pixels in the corresponding reference block are copied along a certain angle as predictions for all the pixels of the current block
[0177] The HEVC encoder selects the optimal intra prediction mode from 35 candidate prediction modes for the current block and writes the optimal intra prediction mode into the video bitstream. To improve the coding efficiency of intra prediction, the encoder / decoder derives three most probable modes from the respective optimal intra prediction modes of the reconstructed picture blocks that use intra prediction in the surrounding area. If the optimal intra prediction mode selected for the current block is one of the three most probable modes, the first index is coded to indicate that the selected optimal intra prediction mode is one of the three most probable modes. If the selected optimal intra prediction mode is not one of the three most probable modes, the second index is coded to indicate that the selected optimal intra prediction mode is one of the other 32 modes (modes other than the aforementioned three most probable modes among the 35 prediction mode candidates). The HEVC standard uses a 5-bit fixed-length code as the aforementioned second index.
[0140]
[0178] The method by which the HEVC encoder derives three most probable modes includes: selecting the optimal intra prediction modes of the left neighboring picture block and the upper neighboring picture block of the current block, and putting the optimal intra prediction modes into a set. If the two optimal intra prediction modes are the same, only one intra prediction mode in the set is reserved. If the two optimal intra prediction modes are the same and both are angular prediction modes, two angular prediction modes adjacent to each other in the angular direction are further selected and added to the set. Otherwise, the planar prediction mode, the DC prediction mode, and the vertical prediction mode are sequentially selected and added to the set until the number of modes in the set reaches 3.
[0141]
[0179] After performing entropy decoding on the bitstream, the HEVC decoder obtains the mode information of the current block. The mode information includes an identifier indicating whether the optimal intra prediction mode of the current block is among the three most likely modes, the index of the optimal intra prediction mode of the current block in the three most likely modes, or the index of the optimal intra prediction mode of the current block in another 32 modes.
[0142]
[0180] Inter prediction
[0181] In a possible implementation, the inter prediction mode set depends on available reference pictures (i.e., for example, previously at least partially decoded pictures stored in the DBP 230) and other inter prediction parameters, for example, whether the entire reference picture or only a part of the reference picture (e.g., the search window area around the area of the current block) is used to search for the best matching reference block, and / or for example, whether pixel interpolation, such as 1 / 2 pixel interpolation, 1 / 4 pixel interpolation, and / or 1 / 16 pixel interpolation, is applied.
[0143]
[0182] In addition to the aforementioned prediction modes, the skip mode and / or the direct mode may be further applied.
[0144]
[0183] For example, the merge candidate list in the extended merge prediction mode includes the following five types of candidates in order: spatial MVP from spatially neighboring CUs, temporal MVP from co-located CUs, history-based MVP from the FIFO table, pairwise mean MVP, and zero MV. Bilateral matching-based decoder side motion vector refinement (DMVR) may be used to improve the accuracy of the MV in the merge mode. The merge mode with an MVD (MMVD) is derived from the merge mode using motion vector difference. The MMVD flag is sent immediately after the skip flag and the merge flag are sent, and specifies whether the MMVD mode is used for the CU. An adaptive motion vector resolution (AMVR) method at the CU level may be used. AMVR allows the MVD of the CU to be coded with different precisions. The MVD of the current CU may be adaptively selected based on the prediction mode of the current CU. When the CU is coded in the merge mode, a combined inter / intra prediction (CIIP) mode may be applied to the current CU. To obtain the CIIP prediction, a weighted averaging of the inter and intra prediction signals is performed. In affine motion compensation prediction, the affine motion field of the block is described based on the motion information of a two-control-point (4-parameter) motion vector or a three-control-point (6-parameter) motion vector. Sub ·Subblock-based temporal motion vector prediction (SbTMVP) is similar to temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of sub-CUs within a current CU. Bi-directional optical flow (BDOF), which was previously called BIO, is a simpler version and requires considerably fewer operations, especially in terms of the number of multiplications and the values of the multipliers. In the triangular partition mode, a CU is evenly divided into two triangular parts by diagonal splitting and anti-diagonal splitting. Furthermore, the bi-prediction mode is extended beyond simple averaging to allow for a weighted average of two prediction signals.
[0145]
[0184] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, a picture block 203 (the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one previously reconstructed block, e.g., a reconstructed block of one or more other / different previously decoded pictures 231. For example, a video sequence may include a current picture and a previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of the pictures forming the video sequence or may form the sequence.
[0146]
[0185] For example, the encoder 20 can be configured to select reference blocks from a plurality of reference blocks of the same picture or a plurality of reference blocks of different pictures of other pictures, and provide the reference picture (or reference picture index) and / or the offset (spatial offset) between the position of the reference block (x and y coordinates) and the position of the current block to the motion estimation unit as an inter-prediction parameter. This offset is also called a motion vector (MV).
[0147]
[0186] The motion compensation unit is configured to, for example, acquire inter-prediction parameters, for example, receive them, perform inter-prediction based on or using the inter-prediction parameters, and acquire the inter-prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on the motion / block vector determined through motion estimation, and may further include performing interpolation with sub-pixel accuracy. Interpolation filtering may generate additional pixel samples from known pixel samples, and thus may potentially increase the amount of candidate prediction blocks that can be used for coding a picture block. When receiving the motion vector corresponding to the PU of the current picture block, the motion compensation unit can locate the prediction block indicated by the motion vector in one of the reference picture lists.
[0148]
[0187] When decoding a picture block of a video slice, the motion compensation unit may further generate syntax elements associated with the block and the video slice for use by the video decoder 30. In addition to or instead of the slice and its respective syntax elements, tile groups and / or tiles and their respective syntax elements may be generated or used.
[0149]
[0188] In the process of obtaining the motion vector candidate list in the advanced motion vector prediction (AMVP) mode, the motion vectors (MVs) that can alternatively be added to the motion vector candidate list include the MVs of the spatially neighboring picture blocks of the current block and the MVs of the temporally neighboring picture blocks of the current block. The MVs of the spatially neighboring picture blocks may include the MVs of the left candidate picture block of the current block and the MVs of the upper candidate picture block of the current block. For example, FIG. 4 is a schematic diagram of an example of a candidate picture block according to an embodiment of the present application. As shown in FIG. 4, the set of left candidate picture blocks includes {A0, A1}, the set of upper candidate picture blocks includes {B0, B1, B2}, and the set of temporally neighboring candidate picture blocks includes {C, T}. All three of these sets can alternatively be added to the motion vector candidate list. However, in existing coding standards, the maximum length of the motion vector candidate list for AMVP is 2. Therefore, it is necessary that at most two MVs of picture blocks are determined from the three sets in the specified order and added to the motion vector candidate list. The order may be as follows: the set {A0, A1} of the left candidate picture blocks of the current block is preferentially considered (where A0 is considered first, and if A0 is not available, then A1 is considered); then, the set {B0, B1, B2} of the upper candidate picture blocks of the current block is considered (where B0 is considered first, if B0 is not available, then B1 is considered, and if B1 is not available, then B2 is considered last); and finally, the set {C, T} of the temporally neighboring candidate picture blocks of the current block is considered (where T is considered first, and if T is not available, then C is considered).
[0150]
[0189] After obtaining the motion vector candidate list, based on the rate distortion cost (RD cost), the optimal MV is determined from the motion vector candidate list, and the motion vector candidate with the minimum RD cost is used as the motion vector predictor (MVP) of the current block. The rate distortion cost is calculated according to the following formula: J = SAD + λR
[0190] J represents the RD cost, SAD represents the sum of absolute differences (SAD) obtained by motion estimation based on the motion vector candidate between the pixel values of the predicted block and the pixel values of the current block, R represents the bit rate, and λ represents the Lagrange multiplier.
[0151]
[0191] The encoder side transfers the index of the determined MVP in the motion vector candidate list to the decoder side. Further, in order to obtain the actual motion vector of the current block, motion search is performed in the vicinity area centered on the MVP. The encoder side calculates the motion vector difference (MVD) between the MVP and the actual motion vector, and transfers the MVD to the decoder side. The decoder side analyzes the index, finds the corresponding MVP in the motion vector candidate list based on the index, analyzes the MVD, adds the MVD to the MVP, and obtains the actual motion vector of the current block.
[0152]
[0192] In the process of obtaining the candidate motion information list in the Merge mode, the motion information that can be alternatively added to the candidate motion information list includes the motion information of the picture blocks that are spatially or temporally adjacent to the current block. The spatially adjacent picture blocks and the temporally adjacent picture blocks may be those as shown in FIG. 4. The candidate motion information corresponding to the spatial picture blocks in the candidate motion information list is obtained from five spatially adjacent blocks (A0, A1, B0, B1, B2). If the spatially adjacent blocks are not available, or in the case of the intra-frame prediction mode, the motion information of the spatially adjacent blocks is not added to the candidate motion information list. Based on the picture order count (POC) of the reference frame and the picture order count (POC) of the current frame, after the MV of the block at the corresponding position in the reference frame is scaled, the temporal candidate motion information of the current block is obtained. First, it is determined whether the block at position T in the reference frame is available. If it is not available, the block at position C is selected. After the candidate motion information list is obtained, based on the RD cost, the optimal motion information is determined as the motion information of the current block from the candidate motion information list. The encoder side transmits the index value of the position of the optimal motion information in the candidate motion information list (shown as the merge index) to the decoder side.
[0153]
[0193] Entropy coding
[0194] The entropy coding unit 270 uses an entropy coding algorithm or method (e.g., variable length coding (VLC) method, context-adaptive VLC (CA VLC) a method, an arithmetic coding method, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique) is applied to the quantized residual coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements to obtain encoded picture data 21 that can be output via the output terminal 272, for example, in the form of an encoded bitstream 21, so that parameters can be received and used by a video decoder 30 or the like for decoding. The encoded bitstream 21 may be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.
[0154]
[0195] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal for some blocks or frames without using the transform processing unit 206. In another implementation, the encoder 20 may combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.
[0155]
[0196] Decoder and Decoding Method
[0197] As shown in FIG. 3, for example, the video decoder 30 is configured to receive the encoded picture data 21 (e.g., the encoded bitstream 21) encoded by the encoder 20 and obtain the picture 331 to be decoded. The encoded picture data or bitstream includes information for decoding the encoded picture data, e.g., data representing picture blocks of the encoded video slice (and / or tile group or tile), and related syntax elements.
[0156]
[0198] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 can perform a decoding process that is generally inverse to the encoding process described with reference to the video encoder shown in FIG. 2. 20
[0157]
[0199] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer DPB 230, inter prediction unit 344, and intra prediction unit 354 also form the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 122, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Accordingly, the description made for each unit and function of video encoder 20 is correspondingly applicable to each unit and function of video decoder 30.
[0158]
[0200] Entropy decoding
[0201] Entropy decoding unit 304 analyzes the bitstream 21 (or generally, the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain the quantized coefficients 309 and / or the decoded coding parameters (not shown in FIG. 3), for example, the inter prediction parameters (e.g., reference picture index and motion vector), the intra prediction parameters (e.g., intra prediction mode or index), the transform parameters, the quantization parameters, the loop filter parameters, and / or any or all of the other syntax elements. The entropy decoding unit 304 can be configured to apply a decoding algorithm or method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 can further be configured to provide the inter prediction parameters, the intra prediction parameters, and / or the other syntax elements to the mode application unit 360 and provide the other parameters to the other units of the decoder 30. The video decoder 30 can receive the syntax elements at the video slice level and / or the video block level. In addition to or instead of the slice and its respective syntax elements, the tile group and / or the tile and its respective syntax elements can be received or used.
[0159]
[0202] Inverse quantization
[0203] The inverse quantization unit 310 receives quantization parameters (QPs) (or generally, information related to inverse quantization) and quantized coefficients from the encoded picture data 21 (e.g., by syntax analysis and / or decoding by the entropy decoding unit 304), and is configured to perform inverse quantization on the decoded quantized coefficients 309 based on the quantization parameters to obtain dequantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameters determined by the video encoder 20 for each video block within a video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0160]
[0204] Inverse transform
[0205] The inverse transform processing unit 312 can be configured to receive the dequantized coefficients 311, which may also be referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain the reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may sometimes be referred to as transform block 2 13. The transform may be an inverse transform, such as an inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may be further configured to receive transform parameters or corresponding information from the encoded picture data 21 (e.g., by analysis and / or decoding by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.
[0161]
[0206] Reconstruction
[0207] The reconstruction unit 314 (e.g., adder 314) adds the reconstructed residual block 2 13 to the prediction block 365 to obtain the reconstructed block 315 in the sample domain, e.g., the reconstructed residual block 2It is configured to obtain by adding the sample value of 13 and the sample value of the prediction block 365.
[0162]
[0208] Filtering
[0209] The loop filter unit 320 (either within or after the coding loop) filters the reconstructed block 315 to obtain a filtered block 321, and is configured to smooth pixel transitions or improve video quality. The loop filter unit 320 can include a deblocking filter, a sample - adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one example, the loop filter unit 220 can include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process can be a deblocking filter, an SAO filter, and then an ALF filter. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in - loop rescaler) is added. This process is executed before deblocking. In another example, the deblocking filter process can also be applied to internal sub - block edges, such as affine sub - block edges, ATMVP sub - block edges, sub - block transform (SBT) edges, and intra sub - partition (ISP) edges. The loop filter unit 320 is shown as a loop filter in FIG. 3, but in other configurations, the loop filter unit 320 may be implemented as a post - loop filter.
[0163]
[0210] Decoded picture buffer
[0211] The decoded video block 321 of the picture is then stored in the decoded picture buffer 330, and the decoded picture buffer 330 stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for respective output displays.
[0164]
[0212] Decoder 30 is configured to output the decoded picture 311 for presentation or display to the user, for example, via output terminal 3 3 2.
[0165]
[0213] Prediction
[0214] The inter prediction unit 344 may be functionally identical to the inter prediction unit 244 (in particular, the motion compensation unit), and the intra prediction unit 354 may be functionally identical to the intra Tra prediction unit 254. Based on the partitioning and / or prediction parameters or respective information received from the encoded picture data 21 (e.g., by analysis and / or decoding by the entropy decoding unit 304), it performs partitioning or partitioning determination and prediction. The mode application unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed picture, block, or respective samples (filtered or unfiltered) to obtain the prediction block 365.
[0166]
[0215] When a video slice is coded as an intra coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter coded (e.g., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In the case of inter prediction, the prediction block may be generated from a reference picture within a reference picture list. The video decoder 30 is capable of constructing reference frame lists: list 0 and list 1 by using a default construction technique based on the reference pictures stored in the DPB 330. In embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices), the same or similar processes may be applied. For example, the video may be coded by using I, P, or B tile groups and / or tiles.
[0167]
[0216] The mode application unit 360 is configured to determine prediction information for video blocks of a current video slice by analyzing motion vectors or other syntax elements, and use the prediction information to generate a prediction block for the current video block to be decoded. For example, the mode application unit 360 uses a part of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to individualize video blocks of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information of one or more reference picture lists of the slice, motion vectors of each inter-coded video block of the slice, an inter prediction status of each inter-coded video block of the slice, and other information for decoding video blocks of the current video slice. In embodiments in addition to or instead of a slice (e.g., a video slice), the same or similar processes can be applied to embodiments that use a tile group (e.g., a video tile group) and / or a tile (e.g., a video tile). For example, a video can be coded by using I, P, or B tile groups and / or tiles.
[0168]
[0217] In an embodiment, the video De coder 30 may be further configured to partition and / or decode a picture by using a slice (also referred to as a video slice), where the picture can be partitioned or decoded by using one or more slices (typically non-overlapping). Each slice may include one or more blocks (e.g., CTUs) or a group of one or more blocks (e.g., tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).
[0169]
[0218] In an embodiment, the video decoder 30 shown in FIG. 3 can be further configured to partition and / or decode a picture by using a slice / tile group (also referred to as a video tile group) and / or a tile (also referred to as a video tile). The picture may be partitioned or decoded by using one or more slice / tile groups (typically non-overlapping), and each slice / tile group may include one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., a complete block or a partial block.
[0170]
[0219] For decoding the coded picture data 21, other variations of the video decoder 30 may be used. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-conversion-based decoder 30 can directly inverse quantize the residual signal without using the inverse transform processing unit 312 for some blocks or frames. In another implementation, the video decoder 30 may combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.
[0171]
[0220] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further processing such as clip or shift may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0172]
[0221] It should be noted that further operations may be applied to the derived motion vectors of the current block (including, but not limited to, the control point motion vectors in the affine mode, the sub-block motion vectors in the affine, planar, and ATMVP modes, the temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predefined range based on the representation bits of the motion vector. When the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1 where "^" represents an exponent. For example, when bitDepth is set to 16, the range is -32768 to 32767, or when bitDepth is set to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of a 4×4 sub-block within a single 8×8 block) is restricted such that the maximum difference between the integer parts of the MVs of the 4×4 sub-blocks does not exceed N pixels, for example, does not exceed 1 pixel. In this case, two methods for restricting the motion vector based on bitDepth are provided.
[0173]
[0222] Although the embodiments are mainly described based on video coding, embodiments of the coding system 10, the encoder 20, and the decoder 30, as well as other embodiments described herein, may be configured for still picture processing or coding, i.e., processing or coding of individual pictures independent of any preceding or consecutive pictures in video coding. It should be noted that generally, when picture processing is limited to a single picture 17, the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 may be equally used for still picture processing, for example, residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220 / 320, as well as entropy coding 270 and entropy decoding 304, etc.
[0174]
[0223] FIG. 5 is an exemplary block diagram of a video coding device 500 according to an embodiment of the present application. The video coding device 500 is suitable for implementing the disclosed embodiments as described herein. In an embodiment, the video coding device 500 may be a decoder such as the video decoder 30 of FIG. 1a, or an encoder such as the video encoder 20 of FIG. 1a.
[0175]
[0224] The video coding device 500 includes an inlet port 510 (or input port 510) and a receiver unit (Rx) 520 for receiving data; a processor, logic unit, or central processing unit (CPU) 530 for processing data (for example, in this case, the processor 530 may be a neural network processing unit 530); a transmitter unit (Tx) 540 and an outlet port 550 (or output port 550) for transmitting data; and a memory 560 for storing data. The video coding device 500 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the inlet port 510, the receiver unit 520, the transmitter unit 540, and the outlet port 550 for the inlet or outlet of optical signals or electrical signals.
[0176]
[0225] The processor 530 is implemented by hardware and software. The processor 530 can be implemented as one or more processor chips, cores (for example, multi-core processors), FPGAs, ASICs, and DSPs. The processor 530 communicates with the inlet port 510, the receiver unit 520, the transmitter unit 540, the outlet port 550, and the memory 560. The processor 530 includes a coding module 570 (for example, a neural network-based coding module 570). The coding module 570 implements the embodiments of the disclosure described above. For example, the coding module 570 implements, processes, prepares, or provides various coding operations. Therefore, the coding module 570 brings significant improvements to the functions of the video coding device 500 and affects the switching of the video coding device 500 to different states. Alternatively, the coding module 570 is implemented as instructions stored in the memory 560 and executed by the processor 530.
[0177]
[0226] Memory 560 may include one or more disks, tape drives, and solid state drives, and be used as an overflow data storage device to store a program when such a program is selected for execution, and is also capable of storing instructions and data read during the execution of the program. Memory 560 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0178]
[0227] FIG. 6 is an exemplary block diagram of an apparatus 600 according to an embodiment of the present application. Apparatus 600 can be used as either or both of the source device 12 and the destination device 14 of FIG. 1a.
[0179]
[0228] The processor 602 of apparatus 600 may be a central processing unit. Alternatively, processor 602 may be any other type of device, or a plurality of devices, capable of computing or processing information that currently exists or will be developed in the future. The disclosed implementation can be carried out by using a single processor such as the illustrated processor 602, but by using a plurality of processors, advantages in speed and efficiency can be achieved.
[0180]
[0229] In some implementations, the memory 604 within the device 600 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 604. The memory 604 can include code and data 606 that are accessed by the processor 602 via the bus 612. The memory 604 can further include an operating system 608 and applications 610. The applications 610 include at least one program that enables the processor 602 to execute the methods described herein. For example, the applications 610 can include applications 1 through N, and may further include a video coding application that executes the methods described herein.
[0181]
[0230] The device 600 may further include one or more output devices such as a display 618. In one example, the display 618 may be a touch-sensitive display that combines a touch-sensitive element that is operable to sense touch input with the display. The display 618 may be coupled to the processor 602 via the bus 612.
[0182]
[0231] Although depicted as a single bus in this specification, the bus 612 of the device 600 may include multiple buses. Further, secondary storage may be directly coupled to another component of the device 600 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Accordingly, the device 600 can be implemented in a variety of configurations.
[0183]
[0232] Embodiments of this application are related to the application of neural networks. Hereinafter, for ease of understanding, first, some nouns or terms used in the embodiments of this application will be explained. The nouns or terms are also used as part of the content of the present invention.
[0184]
[0233] (1) Neural network
[0234] A neural network (neural network, NN) is a machine learning model. A neural network may include neurons. A neuron may be an arithmetic unit that uses x s and bias 1 as inputs, where the output of the arithmetic unit may be as follows:
[0185] [Number]
[0235] s = 1, 2,..., or n, where n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron, f is the activation function of the neuron, and is used to introduce non-linear features into the neural network and convert the input signal in the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function may be a sigmoid function. A neural network is a network formed by connecting a plurality of single neurons together. Specifically, the output of one neuron can be the input of another neuron. The input of each neuron is connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field may be a region containing a plurality of neurons.
[0186]
[0236] (2) Deep neural network
[0237] A deep neural network (DNN), also called a multi-layer neural network, may be understood as a neural network having a plurality of hidden layers. In this case, there is no special measurement criterion for "a plurality of". The DNN is divided based on the arrangement of different layers, and the neural network of the DNN may be classified into three types: an input layer, a hidden layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layer is the hidden layer. The layers are fully connected. Specifically, any neuron in the i-th layer is surely connected to some neuron in the (i + 1)-th layer. The DNN may seem complex, but it is not complex from the perspective of the operations in each layer. Briefly speaking, the DNN is represented by the following linear relational expression:
[0187]
Number
[0188]
[0238] (3) Convolutional Neural Network
[0239] A convolutional neural network (CNN) is a deep neural network with a convolutional structure and is a deep learning architecture. In a deep learning architecture, multi-layer learning is performed at different abstraction levels according to machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network. Each neuron in a feed-forward artificial neural network can respond to the image input into the neural network. A convolutional neural network includes a feature extractor that includes convolutional layers and pooling layers. The feature extractor may be considered as a filter. The convolution process may be considered as performing convolution on the input picture or convolutional feature map using a learnable filter.
[0189]
[0240] A convolutional layer is a neuron layer within a convolutional neural network where a convolution process is performed on an input signal. The convolutional layer may include multiple convolution operators. A convolution operator is also called a kernel. In image processing, a convolution operator functions as a filter to extract specific information from an input picture matrix. A convolution operator may essentially be a weight matrix, and the weight matrix is usually predefined. In the process of performing a convolution operation on a picture, the weight matrix is usually used to process pixels at a granularity of 1 pixel (or 2 pixels depending on the value of the stride) horizontally in the input picture to extract specific features from the picture. The size of the weight matrix should be associated with the size of the picture. It should be noted that the depth dimension of the weight matrix is the same as the depth dimension of the input picture. During the convolution operation, the weight matrix is extended across the entire depth of the input picture. Thus, a convolution output of a single depth dimension is generated by convolution with a single weight matrix. However, in many cases, a single weight matrix is not used, and multiple weight matrices of the same size (row × column), i.e., multiple matrices of the same type, are applied. The outputs of the weight matrices are stacked to form the depth dimension of the convolved picture. The dimension in this case may be understood as being determined based on the aforementioned "multiple". Different weight matrices may be used to extract different features from the picture. For example, one weight matrix may be used to extract edge information of the picture, and another weight matrix may be used to extract a specific color of the picture, FurthermoreAnother weight matrix is used to blur out the unwanted noise within the picture. The sizes (rows × columns) of the multiple weight matrices are the same. The sizes of the feature maps extracted from the multiple weight matrices of the same size are also the same, and the multiple extracted feature maps of the same size are combined to form the output of the convolution operation. The weight values of these weight matrices need to be obtained by a large amount of learning in the actual application. Each weight matrix containing the weight values obtained by training can be used to extract information from the input picture so that the convolutional neural network can make correct predictions. When the convolutional neural network has multiple convolutional layers, usually a large number of general features are extracted in the initial convolutional layer. General features may sometimes be called low-level features. As the depth of the convolutional neural network increases, the features extracted in the subsequent convolutional layers become more complex, for example, high-level semantic features. Features with high-level semantic features become more applicable to the problem to be solved.
[0190]
[0241] Since the amount of training parameters usually needs to be reduced, a pooling layer usually needs to be periodically introduced after the convolutional layer. One pooling layer may follow one convolutional layer, or one or more pooling layers may follow multiple convolutional layers. During picture processing, the pooling layer is used only to reduce the spatial size of the picture. The pooling layer includes an average pooling operator and / or a max pooling operator and can perform sampling on the input picture to obtain a picture of a smaller size. The average pooling operator can be used to calculate the pixel values within a specific range in the picture to generate an average value. The average value is used as the result of average pooling. The max pooling operator can be used to select the pixel with the maximum value within a specific range as the max pooling result. Also, similar to the need for the size of the weight matrix of the convolutional layer to be related to the size of the picture, the operator of the pooling layer also needs to be related to the size of the picture. The size of the processed picture output from the pooling layer may be smaller than the size of the picture input to the pooling layer. Each pixel of the picture output from the pooling layer represents the average value or the maximum value of the corresponding sub-area of the picture input to the pooling layer.
[0191]
[0242] Even after the processes executed in the convolutional layer / pooling layer, the convolutional neural network may still not be able to output the necessary output information. As described above, the convolutional layer / pooling layer only extracts features and reduces the parameters brought by the input picture. However, in order to generate the final output information (necessary class information or other related information), the convolutional neural network needs to use a neural network layer to generate the output of one necessary class or the output of a necessary class group. Therefore, the neural network layer may include a plurality of hidden layers. The parameters included in the plurality of hidden layers may be obtained by pre-training based on the relevant training data of a specific task type. For example, the task type may include image recognition, image classification, and super-resolution image reconstruction.
[0192]
[0243] Optionally, in the neural network layer, after a plurality of hidden layers, an output layer of the entire convolutional neural network follows. The output layer has a loss function similar to categorical cross-entropy, and the loss function is specifically used to calculate the prediction error. When the forward propagation of the entire convolutional neural network is completed, backpropagation is started to update the weight values and biases of each layer described above to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network using the output layer and the ideal result.
[0193]
[0244] (4) Recurrent Neural Network
[0245] A recurrent neural network (RNN) is used to process sequence data. Conventional neural network models start from the input layer, go to the hidden layer, and then to the output layer. Each layer is fully connected, and in each layer, nodes are not connected. This ordinary neural network solves many problems, but it is still insufficient for many problems. For example, when the words in a sentence are to be predicted, usually, the previous word needs to be used because adjacent words in the sentence are related. The reason why an RNN is called a recurrent neural network is that the current output of the sequence is also related to the previous output of the sequence. The specific manifestation form is that the network remembers previous information and applies the previous information to the calculation of the current output. Specifically, the nodes in the hidden layer are connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time point. Theoretically, an RNN can process sequence data of any length. The training of an RNN is the same as that of conventional CNNs and DNNs. The error backpropagation algorithm is also used, but there are differences: when an RNN is extended, parameters such as W of the RNN are shared. This point is different from the conventional neural network described in the previous example. Also, when using the gradient descent algorithm, the output at each step depends not only on the network at the current step but also on the network state at several previous steps. This learning algorithm is called the Back propagation Through Time (BPTT) algorithm.
[0194]
[0246] When a convolutional neural network is available, why is a recurrent neural network still necessary? The reason is simple. In a convolutional neural network, there is a premise that the elements are independent of each other, and the input and output are also independent, like cats and dogs. However, in the real world, multiple elements are related to each other. For example, stock prices change over time. As another example, suppose someone says, "I like traveling, and my favorite place is Yunnan Province. In the future, if I have the opportunity, I will go to (_)." Here, people should be able to tell that the person will go to "Yunnan." This is because people make inferences from the context. However, how would a machine do that? That's where RNN comes in. RNN is designed to enable a machine to remember like a human. Therefore, the output of an RNN needs to depend on the current input information and past memory information.
[0195]
[0247] (5) Loss function
[0248] In the process of training a deep neural network, since the output of the deep neural network is expected to be as close as possible to the actually predicted predictor, the current predictor of the network can be compared with the actually predicted target value, and then, based on the difference between the predicted value and the target value, the weight vectors in each layer of the neural network are updated. (Certainly, before the first update, there is usually an initialization process, specifically, parameters are preset for each layer of the deep neural network). For example, when the predictor of the network is large, the weight vector is adjusted so as to reduce the predicted value until the deep neural network can predict the actually predicted target value, or a value extremely close to the actually predicted target value. Therefore, it is necessary to determine in advance how to obtain the difference between the predictor and the target value by comparison. This is the loss function or objective function. The loss function and the objective function are important formulas for measuring the difference between the predictor and the target value. Take the loss function as an example. A larger output value (loss) of the loss function indicates a larger difference. Therefore, the training of a deep neural network is a process of making the loss as small as possible.
[0196]
[0249] (6) Backpropagation algorithm
[0250] The convolutional neural network can correct the parameter values of the initial super-resolution model in the training process according to the error backpropagation (BP) algorithm so that the reconstruction error loss of the super-resolution model becomes smaller. Specifically, the input signal is propagated forward until the error loss occurs at the output, and the parameters of the initial super-resolution model are updated based on the backpropagation error loss information to converge the error loss. The backpropagation algorithm is a backpropagation movement centered on the error loss, which is intended to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0197]
[0251] (7) Adversarial generation network
[0252] A generative adversarial network (GAN) is a deep learning model. The model includes at least two modules: one module is a generative model, and the other module is a discriminative model. The two modules learn through a game with each other and are used to generate better outputs. It is possible that both the generative model and the discriminative model are neural networks, specifically, they may be deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Take GAN for generating pictures as an example. Assume that there are two networks: G (Generator) and D (Discriminator). G is a network that generates pictures. G receives random noise z and generates a picture based on that noise, where the picture is denoted as G(z). D is a network used to determine whether a picture is "real". The input parameter of D is x, where x represents a picture, and the output D(x) represents the probability that x is a real picture. If the value of D(x) is 1, it indicates that the picture is 100% real. If the value of D(x) is 0, it indicates that the picture should not be real. In the process of training a generative adversarial network, the goal of the generative network G is to generate pictures that are as real as possible so as to deceive the discriminative network D, and the goal of the discriminative network D is to accurately distinguish the pictures generated by G from real pictures. In this way, there is a dynamic "gaming" process between G and D, specifically, there is an "adversary" in the "generative adversarial network". The final gaming result is that, ideally, G can generate a picture G(z) that is difficult to distinguish from real pictures, and for D, it becomes difficult to determine whether the picture generated by G is real, specifically, D(G(z)) = 0.5.In this way, an excellent generation model G can be obtained, and pictures can be generated using this model.
[0198]
[0253] Fig. 7a is a schematic diagram of an application scenario according to an embodiment of the present application. As shown in Fig. 7a, in the application scenario, the device acquires data, compresses the acquired data, and then stores the compressed data. The device may integrate the functions of the aforementioned source device and destination device.
[0199]
[0254] 1. The device acquires data.
[0200]
[0255] 2. The device compresses the data to obtain the compressed data.
[0201]
[0256] 3. The device compresses and stores the data.
[0202]
[0257] It should be understood that the device compresses the data to save memory space. Optionally, the device can store the compressed data in an album or a cloud album.
[0203]
[0258] 4. The device decompresses the compressed data to obtain the data.
[0204]
[0259] Fig. 7b is a schematic diagram of an application scenario according to an embodiment of the present application. As shown in Fig. 7b, in the application scenario, the source device acquires data, compresses the acquired data to obtain the compressed data, and then transmits the compressed data to the destination device.
[0205]
[0260] In an embodiment of the present application, the source device can compress the acquired data and then transmit the compressed data to the destination device. This can reduce the transmission bandwidth.
[0206]
[0261] 1. The source device acquires data.
[0207]
[0262] 2. The source device compresses the data to obtain the compressed data.
[0208]
[0263] 3. The source device transmits the compressed data to the destination device.
[0209]
[0264] The source device compresses the data and then transmits the data. This can reduce the transmission bandwidth and improve the transmission efficiency.
[0210]
[0265] 4. The destination device decompresses the compressed data to obtain the data.
[0211]
[0266] FIG. 8 is a flowchart of an encoding and decoding method 800 according to an embodiment of the present application. The encoding and decoding method 800 can be executed by an encoder and a decoder. The encoding and decoding method 800 is described as a series of steps or operations. It should be understood that the encoding and decoding method 800 may be executed in various sequences and / or simultaneously, and is not limited to the execution sequence shown in FIG. 8. As shown in FIG. 8, the encoding and decoding method 800 may include the following steps.
[0212]
[0267] Step 801: The encoder acquires the data to be encoded.
[0213]
[0268] For example, the encoder acquires the data x to be encoded.
[0214]
[0269] Step 802: The encoder inputs the target data to be encoded into the first encoding network to obtain target parameters.
[0215]
[0270] The target parameters may be the parameter weights for all or part of the convolution and non-linear activation of the second encoding network.
[0216]
[0271] Optionally, the first encoding network may include a convolution kernel generator (convolution group or fully connected group). The convolution kernel generator is configured to generate target parameters based on the target data to be encoded.
[0217]
[0272] For example, the encoder inputs the target data x to be encoded into the first encoding network to obtain the target parameter θ g to obtain.
[0218]
[0273] Step 803: The encoder constructs the second encoding network based on the target parameters.
[0219]
[0274] For example, the encoder constructs the second encoding network g g based on the target parameter θ a (x; θ g ).
[0220]
[0275] Step 804: The encoder inputs the target data to be encoded into the second encoding network to obtain the first feature.
[0221]
[0276] The first feature is used to reconstruct the target data to be encoded, and the first feature may also be called the content feature. For example, the first feature may be a three-dimensional feature map of the target data x to be encoded.
[0222]
[0277] For example, the encoder inputs the target data x to be encoded into the second encoding network g a (x; θ g ) to obtain the first feature y. y satisfies y = g a (x; θ g ).
[0223]
[0278] Step 805: The encoder encodes the first feature to obtain an encoded bitstream (i.e., the bitstream to be decoded).
[0224]
[0279] In a possible implementation, for the encoder to encode the first feature to obtain an encoded bitstream may include: the encoder rounds the first feature to obtain an integer value of the first feature, then performs probability estimation on the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature, and then performs entropy encoding on the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to obtain the encoded bitstream. The integer value of the first feature may be referred to as the first value feature or the content rounding feature.
[0225]
[0280] For example, the encoder first rounds the first feature y to obtain an integer value y^ of the first feature. Then, the encoder performs probability estimation on the integer value y^ of the first feature to obtain an estimated probability distribution p(y^) for the integer value of the first feature. Next, the encoder performs entropy encoding on the integer value y^ of the first feature based on the estimated probability distribution p(y^) for the integer value of the first feature to obtain the encoded bitstream.
[0226]
[0281] y ^ satisfies y ^ = round(y), where round is a rounding operation.
[0227]
[0282] p(y ^ ) satisfies p(y ^ ) = h(y ^ ; θ h ), and h(y ^ ; θ h ) is an entropy estimation network.
[0228]
[0283] Optionally, the encoder may perform probability estimation regarding the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature: The encoder may perform probability estimation regarding the integer value of the first feature based on the first information to obtain an estimated probability distribution for the integer value of the first feature. The first information includes at least one of context information and side information.
[0229]
[0284] It should be noted that the probability distribution is estimated based on context information and side information, and as a result, the accuracy of the obtained estimated probability distribution may be improved. This reduces the bit rate in the entropy encoding process and reduces the entropy encoding overhead.
[0230]
[0285] Step 806: The encoder transmits the encoded bitstream to the decoder.
[0231]
[0286] Step 807: The decoder decodes the encoded bitstream to obtain the integer value of the first feature.
[0232]
[0287] In a possible implementation, for the decoder to decode the encoded bitstream and obtain the integer value of the first feature: the decoder may first perform a probability estimation regarding the integer value of the first feature in the bitstream to be decoded to obtain the estimated probability distribution for the integer value of the first feature; and then, based on the estimated probability distribution for the integer value of the first feature, perform entropy decoding on the bitstream to be decoded to obtain the integer value of the first feature.
[0233]
[0288] Step 808: The decoder inputs the integer value of the first feature into the first decoding network to obtain the decoded data.
[0234]
[0289] For example, the decoder inputs the integer value y ^ of the first feature into the decoding network g s (y ^ ; φ) to obtain the decoded data x ^ , where the decoded data x ^ satisfies x ^ = g s (y ^ ; φ), and φ is all or part of the parameter weights for the convolution and non-linear activation of the encoding network.
[0235]
[0290] In the existing encoding method, an encoding network (i.e., the second encoding network) extracts content features (i.e., the first features) of the target data to be encoded based on fixed parameter weights, and then encodes the content features into a bitstream (i.e., the encoded bitstream) and sends the bitstream to the decoder side. The decoder side performs decoding and reconstruction on the bitstream to obtain the decoded data. In the prior art, the parameter weights of the encoding network are not associated with the target data to be encoded. However, in the encoding method provided in the embodiments of this application, first, the target data to be encoded is input into the first encoding network, and the first encoding network generates the parameter weights of the second encoding network based on the target data to be encoded. Therefore, the parameter weights of the second encoding network are dynamically adjusted based on the obtained weights. As a result, the parameter weights of the second encoding network are associated with the target data to be encoded, the representation ability of the second encoding network is enhanced, and the decoded data obtained on the decoder side by decoding and reconstruction with respect to the bitstream obtained by encoding the first features is closer to the target data to be encoded. This improves the rate-distortion performance of the encoding and decoding networks.
[0236]
[0291] The encoding and decoding method 800 provided in this embodiment of this application is applicable to the encoding and decoding system shown in FIG. 9. As shown in FIG. 9, the encoding and decoding system includes a first encoding network 901, a second encoding network 902, a rounding module 903, an entropy estimation network 904, an entropy encoding module 905, an entropy decoding module 906, and a decoding network 907.
[0237]
[0292] As shown in FIG. 9, first, the target data to be encoded is input to the first encoding network 901 to obtain target parameters, and then the parameters of the second encoding network 902 are adjusted based on the target parameters (that is, all or part of the parameter weights for convolution and non-linear activation of the second encoding network 902 are adjusted based on the target parameters).
[0238]
[0293] The target data to be encoded is input to the second encoding network 902 to obtain the first feature.
[0239]
[0294] The rounding module 903 rounds the first feature to obtain an integer value of the first feature.
[0240]
[0295] The entropy estimation network 904 performs probability estimation on the integer value of the first feature to obtain an estimated probability distribution of the integer value of the first feature.
[0241]
[0296] The entropy encoding module 905 performs entropy encoding on the integer value of the first feature based on the estimated probability distribution of the integer value of the first feature to obtain an encoded bit stream.
[0242]
[0297] The entropy decoding module 906 performs entropy decoding on the encoded bit stream based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature.
[0243]
[0298] The integer value of the first feature is input to the decoding network 907 to obtain decoded data.
[0244]
[0299] Figure 10 is a flowchart of an encoding and decoding method 1000 according to an embodiment of the present application. The encoding and decoding method 1000 can be executed by an encoder and a decoder. The encoding and decoding method 1000 is described as a series of steps or operations. It should be understood that the encoding and decoding method 1000 may be executed in various sequences and / or simultaneously, and is not limited to the execution sequence shown in Figure 10. As shown in Figure 10, the encoding and decoding method 1000 may include the following steps.
[0245]
[0300] Step 1001: The encoder obtains target data to be encoded.
[0246]
[0301] Step 1002: The encoder inputs the target data to be encoded into a second encoding network to obtain a first feature.
[0247]
[0302] The first feature is used to reconstruct the target data to be encoded.
[0248]
[0303] Step 1003: The encoder inputs the target data to be encoded into a first encoding network to obtain a second feature.
[0249]
[0304] The second feature is used to reconstruct target parameters. The second feature may be referred to as a model feature. The target parameters are all or part of the parameter weights for convolution and non-linear activation of a second decoding network.
[0250]
[0305] In a possible implementation, the encoder can first split the first feature into two parts (a first sub - feature and a second sub - feature) in the channel dimension. One part is used to reconstruct the target data to be encoded (the first sub - feature), and the other part is used to reconstruct the target parameter (the second sub - feature). Then, the encoder inputs the second sub - feature into the first encoding network to obtain the second feature.
[0251]
[0306] Optionally, in order to enable the second feature to be compressed into a small bit - stream, before the second sub - feature is input into the third encoding network, the second sub - feature may be further transformed through a convolutional network and a fully - connected network. The second sub - feature before transformation may be called the initial model feature, and the second sub - feature obtained by the transformation may be called the model feature.
[0252]
[0307] Step 1004: The encoder encodes the first feature to obtain a first encoded bit - stream.
[0253]
[0308] Step 1005: The encoder encodes the second feature to obtain a second encoded bit - stream.
[0254]
[0309] In a possible implementation, the encoder can encode the first feature and the second feature to obtain a bit - stream to be decoded.
[0255]
[0310] Step 1006: The encoder sends the first bit - stream to be decoded and the second bit - stream to be decoded to the decoder.
[0256]
[0311] Step 1007: The decoder decodes the first bit - stream to be decoded to obtain an integer value of the first feature.
[0257]
[0312] In a possible implementation, for the decoder to decode a first bitstream to be decoded and obtain an integer value of a first feature: the decoder may perform probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain an estimated probability distribution of the integer value of the first feature, and based on the estimated probability distribution of the integer value of the first feature, perform entropy decoding on the bitstream to be decoded to obtain the integer value of the first feature.
[0258]
[0313] In a possible implementation, performing probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain an estimated probability distribution of the integer value of the first feature includes: performing probability estimation on the integer value of the first feature in the first bitstream to be decoded based on first information to obtain an estimated probability distribution of the integer value of the first feature, where the first information includes at least one of context information and side information.
[0259]
[0314] Step 1008: The decoder decodes a second bitstream to be decoded and obtains an integer value of a second feature quantity.
[0260]
[0315] The integer value of the second feature may also be referred to as a model-rounded feature.
[0261]
[0316] In a possible implementation, for the decoder to decode a second bitstream to be decoded and obtain an integer value of a second feature: the decoder performs probability estimation on the integer value of the second feature in the second bitstream to be decoded to obtain an estimated probability distribution of the integer value of the second feature, and based on the estimated probability distribution of the integer value of the second feature, performs entropy decoding on the second bitstream to be decoded to obtain the integer value of the second feature.
[0262]
[0317] In a possible implementation, performing probability estimation on the integer value of the second feature in the second bitstream to be decoded to obtain the estimated probability distribution of the integer value of the second feature includes: performing probability estimation on the integer value of the second feature in the second bitstream to be decoded based on first information to obtain the estimated probability distribution of the integer value of the second feature, where the first information includes at least one of context information and side information.
[0263]
[0319] Step 1009: The decoder inputs the integer value of the second feature into the first decoding network to obtain target parameters.
[0264]
[0319] Step 1010: The decoder constructs the second decoding network based on the target parameters.
[0265]
[0320] Step 1011: The decoder inputs the integer value of the first feature into the second decoding network to obtain the decoded data.
[0266]
[0321] In the existing decoding method, the decoding network (i.e., the second decoding network) performs decoding and reconstruction on the content value feature of the encoded target data (i.e., the integer value of the first feature) based on fixed parameter weights to obtain the decoded data. It can be seen from the prior art that the parameter weights of the decoding network are not associated with the target data to be decoded. However, in this embodiment of the present application, the content feature and the model feature of the target data to be decoded (i.e., the first feature and the second feature) are encoded into the target bitstream to be decoded. Then, the decoder side decodes the target bitstream to be decoded to obtain the integer value of the second feature. The integer value of the second feature is input into the first decoding network to obtain the parameter weights of the second decoding network. Then, the parameter weights of the second decoding network are dynamically adjusted based on the parameter weights. As a result, the parameter weights of the second decoding network are associated with the target data to be decoded, the representation ability of the second decoding network is improved, and the decoded data obtained by the second decoding network through decoding and reconstruction is closer to the encoded target data. This improves the rate-distortion performance of the encoding and decoding networks.
[0267]
[0322] The encoding and decoding method 1000 provided in this embodiment of the present application is applicable to the encoding and decoding system shown in FIG. 11. As shown in FIG. 11, the encoding and decoding system includes a first encoding network 1101, a second encoding network 1102, a first rounding module 1103, a second rounding module 1104, an entropy estimation network 1105, a first entropy encoding module 1106, a second entropy encoding module 1107, a first entropy decoding module 1108, a second entropy decoding module 1109, a first decoding network 1110, and a second decoding network 1111.
[0268]
[0323] As shown in FIG. 11, first, the target data to be encoded is input into the second encoding network 1102 to obtain the first feature, and then the target data to be encoded is input into the first encoding network 1101 to obtain the second feature.
[0269]
[0324] The first rounding module 1103 rounds the first feature to obtain an integer value of the first feature.
[0270]
[0325] The second rounding module 1104 rounds the second feature to obtain an integer value of the second feature.
[0271]
[0326] The entropy estimation network 1105 first performs probability estimation on the integer value of the first feature to obtain an estimated probability distribution of the integer value of the first feature, and then performs probability estimation on the integer value of the second feature to obtain an estimated probability distribution of the integer value of the second feature.
[0272]
[0327] The first entropy encoding module 1106 performs entropy encoding on the integer value of the first feature based on the estimated probability distribution of the integer value of the first feature to obtain a bit stream of the first target to be decoded.
[0273]
[0328] The second entropy encoding module 1107 performs entropy encoding on the integer value of the second feature based on the estimated probability distribution of the integer value of the second feature to obtain a bit stream of the second target to be decoded.
[0274]
[0329] The first entropy decoding module 1108 performs entropy decoding on the bit stream of the first target to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature.
[0275]
[0330] The second entropy decoding module 1109 performs entropy decoding on the bitstream of the second object to be decoded based on the estimated probability distribution of the integer values of the second feature, and obtains the integer value of the second feature.
[0276]
[0331] The integer value of the second feature is first input to the first decoding network 1110 to obtain target parameters, and then the parameters of the second decoding network 1111 are adjusted based on the target parameters (that is, all or part of the parameter weights for convolution and non-linear activation of the second decoding network 1111 are adjusted based on the target parameters).
[0277]
[0332] The integer value of the first feature is input to the second decoding network 1111 to obtain the decoded data.
[0278]
[0333] The encoding and decoding method 1000 provided in this embodiment of the present application is also applicable to the encoding and decoding system shown in FIG. 12. As shown in FIG. 12, the encoding and decoding system includes a first encoding network 1201, a second encoding network 1202, a channel splitting module 1203, a first rounding module 1204, a second rounding module 1205, an entropy estimation network 1206, a first entropy encoding module 1207, a second entropy encoding module 1208, a first entropy decoding module 1209, a second entropy decoding module 1210, a first decoding network 1211, and a second decoding network 1212.
[0279]
[0334] As shown in FIG. 12, the data to be encoded is first input to the second encoding network 1202 to obtain the first feature.
[0280]
[0335] The first feature is input to the channel splitting module 1203 and is split into a first sub-feature and a second sub-feature in the channel dimension.
[0281]
[0336] The second sub - feature is input into the first encoding network 1201 to obtain the second feature.
[0282]
[0337] The first rounding module 1204 rounds the first sub - feature to obtain an integer value of the first feature.
[0283]
[0338] The second rounding module 1205 rounds the second sub - feature to obtain an integer value of the second feature.
[0284]
[0339] The entropy estimation network 1206 first performs probability estimation on the integer value of the first feature to obtain an estimated probability distribution of the integer value of the first feature, and then performs probability estimation on the integer value of the second feature to obtain an estimated probability distribution of the integer value of the second feature.
[0285]
[0340] The first entropy encoding module 1207 performs entropy encoding on the integer value of the first feature based on the estimated probability distribution of the integer value of the first feature to obtain a bitstream of the first object to be decoded.
[0286]
[0341] The second entropy encoding module 1208 performs entropy encoding on the integer value of the second feature based on the estimated probability distribution of the integer value of the second feature to obtain a bitstream of the second object to be decoded.
[0287]
[0342] The first entropy decoding module 1209 performs entropy decoding on the bitstream of the first object to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature.
[0288]
[0343] The second entropy decoding module 1210 performs entropy decoding on the bit stream of the second object to be decoded based on the estimated probability distribution of the integer values of the second feature, and obtains the integer value of the second feature.
[0289]
[0344] The integer value of the second feature is first input to the first decoding network 1211 to obtain target parameters. Then, the parameters of the second decoding network 1212 are adjusted based on the target parameters (i.e., all or part of the parameter weights for convolution and non-linear activation of the second decoding network 1212 are adjusted based on the target parameters).
[0290]
[0345] The integer value of the first feature is input to the second decoding network 1212 to obtain the decoded data.
[0291]
[0346] FIG. 10 is a flowchart of an encoding and decoding method 1300 according to an embodiment of the present application. The encoding and decoding method 1300 can be executed by an encoder and a decoder. The encoding and decoding method 1300 is described as a series of steps or operations. It should be understood that the encoding and decoding method 1300 may be executed in various sequences and / or simultaneously, and is not limited to the execution sequence shown in FIG. 13. As shown in FIG. 13, the encoding and decoding method 1300 may include the following steps.
[0292]
[0347] Step 1301: The encoder obtains the data to be encoded.
[0293]
[0348] Step 1302: The encoder inputs the data to be encoded into the encoding network to obtain the first feature.
[0294]
[0349] Step 1303: The encoder encodes the first feature to obtain the bit stream of the object to be decoded.
[0295]
[0350] Step 1304: The encoder transmits the bitstream to be decoded to the decoder.
[0296]
[0351] Step 1305: The decoder decodes the bitstream to be decoded to obtain the integer value of the first feature.
[0297]
[0352] In a possible implementation, the decoder decodes the bitstream to be decoded to obtain the integer value of the first feature may include: the decoder performs probability estimation on the integer value of the first feature in the bitstream to be decoded to obtain the estimated probability distribution of the integer value of the first feature, and based on the estimated probability distribution of the integer value of the first feature, performs entropy decoding on the bitstream to be decoded to obtain the integer value of the first feature.
[0298]
[0353] In a possible implementation, performing probability estimation on the integer value of the first feature in the bitstream to be decoded to obtain the estimated probability distribution of the integer value of the first feature may include: performing probability estimation on the integer value of the first feature in the bitstream to be decoded based on the first information to obtain the estimated probability distribution of the integer value of the first feature, where the first information includes at least one of context information and side information.
[0299]
[0354] Step 1309: The decoder inputs the integer value of the first feature into the first decoding network to obtain the target parameter.
[0300]
[0355] Optionally, the first decoding network may include a convolutional kernel generator (convolutional group or fully connected group). The convolutional kernel generator is configured to generate the target parameter based on the integer value of the first feature of the data to be encoded.
[0301]
[0356] Step 1307: The decoder constructs a second decoding network based on the target parameter.
[0302]
[0357] Step 1308: The decoder inputs the integer value of the first feature into the second decoding network to obtain the decoded data.
[0303]
[0358] In the existing decoding method, the decoding network (i.e., the second decoding network) performs decoding and reconstruction on the content value feature (i.e., the integer value of the first feature) of the encoded target data based on fixed parameter weights to obtain the decoded data. It can be seen that in the prior art, the parameter weights of the decoding network are not associated with the target data to be decoded. However, in this embodiment of the present application, the decoding target bitstream obtained by encoding the feature of the target data to be decoded is decoded to obtain the integer value of the first feature, and the integer value of the first feature is input into the first decoding network to obtain the parameter weights of the second decoding network. Then, the parameter weights of the second decoding network are dynamically adjusted based on the parameter weights, so that the parameter weights of the second decoding network are associated with the target data to be decoded. As a result, the representation ability of the second decoding network is improved, and the decoded data obtained by the second decoding network through decoding and reconstruction is closer to the encoded target data. This improves the rate-distortion performance of the encoding and decoding networks.
[0304]
[0359] The encoding and decoding method 1300 provided in this embodiment of the present application is applicable to the encoding and decoding system shown in FIG. 14. As shown in FIG. 14, the encoding and decoding system includes an encoding network 1401, a rounding module 1402, an entropy estimation network 1403, an entropy encoding module 1404, an entropy decoding module 1405, a first decoding network 1406, and a second decoding network 1407.
[0305]
[0360] As shown in FIG. 14, first, the target data to be encoded is input into the encoding network 1401 to obtain a first feature.
[0306]
[0361] The rounding module 1402 rounds the first feature to obtain an integer value of the first feature.
[0307]
[0362] The entropy estimation network 1403 performs probability estimation on the integer value of the first feature to obtain an estimated probability distribution of the integer value of the first feature.
[0308]
[0363] The entropy encoding module 1404 performs entropy encoding on the integer value of the first feature based on the estimated probability distribution of the integer value of the first feature to obtain a bit stream of the object to be decoded.
[0309]
[0364] The entropy decoding module 1405 performs entropy decoding on the bit stream of the object to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain an integer value of the first feature.
[0310]
[0365] The integer value of the first feature is first input into the decoding network 1406 to obtain a target parameter, and then the second decoding network 140 7The parameters are adjusted based on the target parameters (i.e., all or part of the parameter weights for convolution and non-linear activation of the second decoding network 1407 are adjusted based on the target parameters).
[0311]
[0366] The integer value of the first feature is input into the second decoding network 1407 to obtain the decoded data.
[0312]
[0367] FIG. 15 is a schematic diagram showing the performance of the encoding and decoding method according to an embodiment of the present application. The coordinate system of FIG. 15 shows the performance of separately encoding and decoding a test set using the embodiment of the present application and the prior art based on the Peak Signal-to-Noise Ratio (PSNR) metric. The test set is the Kodak test set. The Kodak test set contains 24 images in the Portable Network Graphics (PNG) format. The resolution of the 24 pictures may be 768×512 or 512×768. In the coordinate system of FIG. 15, the horizontal axis represents the bits per pixel (BPP), and the vertical axis represents the PSNR. BPP indicates the amount of bits used to save each image. The smaller the value, the smaller the compression bit rate. PSNR is an objective criterion for evaluating an image. The higher the PSNR, the better the image quality.
[0313]
[0368] Line segment A in the coordinate system shown in FIG. 15 represents an embodiment of the present invention, and line segment B represents the prior art. It can be seen from FIG. 15 that when the BPP index is the same, all PSNR indexes in the embodiments of the present application are larger than those in the prior art, and when the image compression quality (i.e., the PSNR index) is the same, all BPP indexes in the embodiments of the present application are smaller than those in the prior art. It can be understood that the rate-distortion performance in the embodiments of the present application is higher than the rate-distortion performance in the prior art, and the rate-distortion performance of the data encoding and decoding method can be improved in the embodiments of the present application.
[0314]
[0369] Scenarios in which the encoding and decoding method provided in the embodiments of the present application can be applied include, but are not limited to, all services related to the capture, storage, and transmission of data such as images, videos, and sounds in electronic devices, cloud services, and video surveillance (for example, shooting of electronic devices and video or audio services, albums, cloud albums, video surveillance, video conferencing, model compression).
[0315]
[0370] Figure 16 is a schematic diagram of an application scenario according to an embodiment of the present application. As shown in Figure 16, in the application scenario, after acquiring picture data (data to be compressed) by shooting, the electronic device inputs the picture data acquired by shooting (for example, picture data in RAW format, YUV format, or RGB format) into the AI encoding unit of the electronic device. The AI encoding unit of the electronic device activates a first encoding network, a second encoding network, and an entropy estimation network to convert the picture data into output features with low redundancy and generate an estimated probability distribution of the output features. The arithmetic encoding unit of the electronic device activates an entropy encoding module to encode the output features into a data file based on the estimated probability distribution of the output features. The file storage unit stores the data file generated by the entropy encoding module in the corresponding storage location of the electronic device. When the electronic device needs to use the data file, the file loading unit loads the data file from the corresponding storage location of the electronic device, and the data file is input into the arithmetic decoding unit. The arithmetic decoding unit activates an entropy decoding module to decode the data file to obtain the output features, and the output features are input into the AI decoding unit. The AI decoding unit activates a first decoding network, a second decoding network, and a third decoding network to perform an inverse transformation on the output features and analyze the output features into image data (for example, RGB image data), that is, compressed data. The AI encoding unit and the AI decoding unit may be deployed in a neural network processing unit (NPU) or a graphics processing unit (GPU) of the electronic device. The arithmetic encoding unit, the file storage unit, the file loading unit, and the arithmetic decoding unit may be deployed in the CPU of the electronic device.
[0316]
[0371] Figure 17 is a schematic diagram of another application scenario according to an embodiment of the present application. As shown in Figure 17, in the application scenario, before uploading to the cloud (e.g., a server), the electronic device encodes the stored picture data (the data to be compressed) (e.g., JPEG encoding) or uploads it directly. The cloud decodes the received picture data (e.g., JPEG decoding) or inputs it directly before inputting it to the AI encoding unit in the cloud. The AI encoding unit in the cloud activates the first encoding network, the second encoding network, and the entropy estimation network to convert the picture data into output features with low redundancy and generate an estimated probability distribution of the output feature amounts. The arithmetic encoding unit in the cloud activates the entropy encoding module to encode the output features into a data file based on the estimated probability distribution of the output features. The file storage unit stores the data file generated by the entropy encoding module in the corresponding storage location in the cloud. When the electronic device needs to use the data file, the electronic device sends a download request to the cloud. After the cloud receives the download request, the file loading unit loads the data file from the corresponding storage location in the cloud, and the data file is input to the arithmetic decoding unit. The arithmetic decoding unit activates the entropy decoding module to decode the data file to obtain the output features, and the output features are input to the AI decoding unit. The AI decoding unit activates the first decoding network, the second decoding network, and the third decoding network to perform an inverse transformation on the output features and analyze the output features into picture data. The cloud encodes the picture data (compressed data) or sends it directly to the electronic device before sending it.
[0317]
[0372] Hereinafter, with reference to FIGS. 18 and 19, an encoding and decoding apparatus configured to execute the foregoing encoding and decoding methods will be described.
[0318]
[0373] To implement the foregoing functions, it is possible to understand that the encoding and decoding apparatus includes corresponding hardware and / or software modules that execute each function. Referring to each example algorithm step described in the embodiments disclosed in this specification, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether the function is executed by hardware or by hardware driven by computer software depends on the design constraints of the technical solution and specific applications. Those skilled in the art may use various methods to implement the described functions for their respective specific applications with reference to the embodiments, but it should not be considered that the implementation goes beyond the scope of this case.
[0319]
[0374] In the embodiments of this application, the encoding and decoding apparatus may be divided into functional modules based on the foregoing method examples. For example, each functional module may be obtained by division based on each corresponding function, or two or more functions may be integrated into one processing module. The foregoing integrated module may be implemented in the form of hardware. In this embodiment, the division into modules is an example and is merely a logical function division, and other divisions may be used in actual implementation.
[0320]
[0375] When each functional module is obtained by a division based on the corresponding function, FIG. 18 is a possible schematic diagram of the configuration of the encoding and decoding apparatus in the foregoing embodiment. As shown in FIG. 18, the apparatus 1800 may include a transceiver unit 1801 and a processing unit 1802. The processing unit 1802 can perform the method executed by the encoding apparatus, the decoding apparatus, the encoder, or the decoder in the foregoing method embodiments, and / or other processing of the technology described in the present specification.
[0321]
[0376] It should be noted that all relevant contents of the steps in the foregoing method embodiments can be cited in the functional descriptions of the corresponding functional modules. Details are not described again here.
[0322]
[0377] When an integrated unit is used, the apparatus 1800 may include a processing unit, a storage unit, and a communication unit. The processing unit may be configured to control and manage the operation of the device 1800. For example, when executing the steps executed by the foregoing units, it may be configured to support the apparatus 1800. The storage unit may be configured to support the apparatus 1800 when storing program code, data, and / or the like. The communication unit may be configured to support communication between the apparatus 1800 and other devices.
[0323]
[0378] The processing unit may be a processor or a controller. The processor can implement or execute various exemplary logic blocks, modules, and circuits described with reference to the content disclosed herein. The processor may alternatively be a combination for implementing arithmetic functions. For example, one or more microprocessors and a digital signal processor (digital signal process or, it may also be a combination including a combination of a DSP and a microprocessor. The storage device may be a memory. The communication unit may specifically be, for example, a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc., and may interact with other electronic devices.
[0324]
[0379] In a possible implementation, the encoding and decoding device in this embodiment of the present application may be a device 1900 having the structure shown in FIG. 19. The device 1900 includes a processor 1901 and a transceiver 1902. The related functions implemented by the transceiver unit 1801 and the processing unit 1802 in FIG. 18 may be implemented by the processor 1901.
[0325]
[0380] Optionally, the device 1900 may further include a memory 1903. The processor 1901 and the memory 1903 communicate with each other via an internal connection path. The related functions implemented by the storage device in FIG. 18 may be implemented by the memory 1903.
[0326]
[0381] The embodiment of the present application further provides a computer storage medium. The computer storage medium stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device can execute the foregoing related method steps to implement the encoding and decoding method in the foregoing embodiment.
[0327]
[0382] The embodiment of the present application further provides a computer program product. When the computer program product is executed on a computer, the computer can execute the foregoing related steps to implement the encoding and decoding method in the foregoing embodiment.
[0328]
[0383] Embodiments of the present application further provide an encoding and decoding apparatus. The apparatus may specifically be a chip, an integrated circuit, a component, or a module. Specifically, the apparatus may include a connected processor and a memory configured to store instructions, or the apparatus may include at least one processor configured to obtain instructions from an external memory. When the apparatus operates, the processor executes the instructions, and as a result, the chip executes the encoding and decoding methods in the embodiments of the foregoing method.
[0329]
[0384] FIG. 20 is a schematic diagram of the structure of chip 2000. Chip 2000 includes one or more processors 2001 and an interface circuit 2002. Optionally, chip 2000 may further include a bus 2003.
[0330]
[0385] Processor 2001 may be an integrated circuit chip and has signal processing capabilities. In the implementation process, each step of the foregoing encoding method can be completed by using the integrated logic circuit of the hardware in processor 2001 or by using instructions in the form of software.
[0331]
[0386] Processor 2001 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another programmable logic device, a discrete gate, or a transistor logic device, or a discrete hardware component. The processor can implement or execute the methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, and the processor may be any conventional processor or the like. or , DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another programmable logic device, a discrete gate, or a transistor logic device, or a discrete hardware component. The processor can implement or execute the methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, and the processor may be any conventional processor or the like.
[0332]
[0387] Interface circuit 2002 can transmit or receive data, instructions, or information. Processor 2001 can process the data, instructions, or other information received via interface circuit 2002 and transmit the information obtained by the processing via interface circuit 2002.
[0333]
[0388] Optionally, the chip further includes a memory. The memory includes a read-only memory and a random access memory and can provide operating instructions and data to the processor. A part of the memory may further include a non-volatile random access memory (NVRAM).
[0334]
[0389] Optionally, the memory stores an executable software module or data structure, and the processor can execute the corresponding operation by starting the operation instructions stored in the memory (the operation instructions can be stored in the operating system).
[0335]
[0390] Optionally, the chip may be used in the electronic device or D S P in the embodiments of this application. Optionally, interface circuit 2002 can be configured to output the execution result of processor 2001. For the encoding method provided in one or more embodiments of this application, refer to the foregoing embodiments. Details are not described again here.
[0336]
[0391] It should be noted that the functions corresponding to each of processor 2001 and interface circuit 2002 can be implemented using hardware design, can be implemented using software design, or can be implemented using a combination of software and hardware. This is not limited in this case.
[0337]
[0392] An electronic device, a computer storage medium, a computer program product, or a chip provided in an embodiment is configured to execute the corresponding method provided above. Therefore, for the achievable beneficial effects, please refer to the beneficial effects of the corresponding method provided above. Details will not be described here again.
[0338]
[0393] It should be understood that the sequence numbers of the foregoing processes do not mean the execution sequence in the embodiments of this application. The execution order of the processes should be determined based on the functions and internal logics of the processes, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0339]
[0394] Those skilled in the art can recognize that in combination with the examples described in the embodiments disclosed in this specification, the units and algorithm steps can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the function is executed by hardware or by software depends on the design constraints of the technical solution and specific applications. Those skilled in the art may use different methods to implement the described functions for each specific application, but it should not be considered that the implementation goes beyond the scope of this case.
[0340]
[0395] It will be clearly understood by those skilled in the art that for the sake of easy and concise description, for the detailed processing steps of the foregoing system, device, and unit, reference may be made to the corresponding processes in the embodiments of the foregoing method. Details will not be described here again.
[0341]
[0396] In some embodiments provided herein, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the described device embodiments are merely illustrative. For example, the division into units is merely a logical functional division, and there may be other divisions in actual implementation. For example, multiple units or components may be combined, integrated into another system, or some functions may be ignored or not executed. Further, the illustrated or described mutual coupling or direct coupling or communication connection may be implemented via some interface. The indirect coupling or communication connection between devices or units may be implemented in electrical, mechanical, or other forms.
[0342]
[0397] The foregoing units described as separate parts may or may not be physically separate, and the parts illustrated as units may or may not be physical units, and may be located in one place, or may be distributed over multiple network units. Some or all of the units can be selected based on the actual requirements for achieving the objectives of the embodiment solution.
[0343]
[0398] Furthermore, the functional units in the embodiments of this application may be integrated into one processing unit, each unit may exist physically alone, or two or more units may be integrated into one unit.
[0344]
[0399] When the function is implemented in the form of a software functional unit and sold or used as an independent product, the function may be stored in a computer-readable storage medium. Based on such an understanding, essentially the technical solution of this case, or the part that contributes to the prior art, or all or part of the technical solution, can be implemented in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for instructing a computer device (which may be a personal computer, a server, a network device, etc.) to execute all or part of the steps of the foregoing method described in the embodiments of this application. The foregoing storage medium includes any medium capable of storing program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, and a compact disk.
[0345]
[0400] The foregoing description is only a specific implementation of this application and is not intended to limit the protection scope of this application. Any modification or substitution that can be easily devised by those skilled in the art within the technical scope disclosed in this case shall be included in the protection scope of this case. Therefore, the protection scope of this application shall follow the protection scope of the claims.
Claims
1. A method of encoding, comprising: obtaining target data to be encoded; inputting the target data to be encoded into a first encoding network to obtain target parameters; constructing a second encoding network based on the target parameters; inputting the target data to be encoded into the second encoding network to obtain a first feature; and encoding the first feature to obtain an encoded bitstream. A method comprising the above steps.
2. The method according to claim 1, wherein one or more of the target parameters are all or part of parameter weights related to convolution and non-linear activation of the second encoding network.
3. In the method according to claim 1 or 2, the step of encoding the first feature to obtain an encoded bitstream comprises: rounding the first feature to obtain an integer value of the first feature; performing probability estimation on the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature; and performing entropy encoding on the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to obtain the encoded bitstream. A method comprising the above steps.
4. In the method according to claim 3, the step of performing probability estimation on the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature comprises: performing probability estimation on the integer value of the first feature based on first information, the first information including at least one of context information and side information. A method comprising the above step.
5. A method of decoding, comprising: obtaining a target bitstream to be decoded; decoding the target bitstream to be decoded to obtain an integer value of a first feature and an integer value of a second feature, the integer value of the first feature being used to obtain decoded data, and the integer value of the second feature being used to obtain target parameters. Inputting the integer value of the second feature into the first decoding network to obtain the target parameter; Constructing a second decoding network based on the target parameter; and Inputting the integer value of the first feature into the second decoding network to obtain the decoded data; A method comprising the above steps.
6. The method according to claim 5, wherein the target parameter is all or part of the parameter weights related to the convolution and non-linear activation of the second decoding network.
7. The method according to claim 5 or 6, wherein the bitstream to be decoded includes a first bitstream to be decoded and a second bitstream to be decoded, and the steps of decoding the bitstream to be decoded to obtain the integer value of the first feature and the integer value of the second feature are: Decoding the first bitstream to be decoded to obtain the integer value of the first feature; and Decoding the second bitstream to be decoded to obtain the integer value of the second feature; A method comprising the above steps.
8. The method according to claim 7, wherein the steps of decoding the first bitstream to be decoded to obtain the integer value of the first feature are: Performing probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain the estimated probability distribution of the integer value of the first feature; and Performing entropy decoding on the first bitstream to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature; A method comprising the above steps.
9. The method according to claim 8, wherein the steps of performing probability estimation on the integer value of the first feature in the first bitstream to be decoded to obtain the estimated probability distribution of the integer value of the first feature are: Performing probability estimation on the integer value of the first feature in the bitstream of the first object to be decoded based on the first information to obtain an estimated probability distribution for the integer value of the first feature, where the first information includes at least one of context information and side information; A method including this.
10. In the method according to claim 7, the step of decoding the bitstream of the second object to be decoded to obtain the integer value of the second feature is: Performing probability estimation on the integer value of the second feature in the bitstream of the second object to be decoded to obtain an estimated probability distribution for the integer value of the second feature; and Based on the estimated probability distribution for the integer value of the second feature, performing entropy decoding on the bitstream of the second object to be decoded to obtain the integer value of the second feature; A method including this.
11. In the method according to claim 10, the step of performing probability estimation on the integer value of the second feature in the bitstream of the second object to be decoded to obtain an estimated probability distribution for the integer value of the second feature is: Performing probability estimation on the integer value of the second feature in the bitstream of the second object to be decoded based on the first information to obtain an estimated probability distribution for the integer value of the second feature, where the first information includes at least one of context information and side information; A method including this.
12. A decoding method comprising: Obtaining a bitstream of an object to be decoded; Decoding the bitstream of the object to be decoded to obtain an integer value of a first feature, where the integer value of the first feature is used to obtain decoded data and target parameters; Inputting the integer value of the first feature into a first decoding network to obtain the target parameters; Constructing a second decoding network based on the target parameters; and Inputting the integer value of the first feature into the second decoding network to obtain decoded data; A method including this.
13. The method according to claim 12, wherein the target parameter is all or part of the parameter weights related to the convolution and non-linear activation of the second decoding network.
14. In the method according to claim 12 or 13, the step of decoding the bit stream to be decoded to obtain an integer value of a first feature is as follows: Performing probability estimation regarding the integer value of the first feature in the bit stream to be decoded to obtain an estimated probability distribution for the integer value of the first feature; and Performing entropy decoding regarding the bit stream to be decoded based on the estimated probability distribution for the integer value of the first feature to obtain the integer value of the first feature; The method comprising the above steps.
15. In the method according to claim 14, the step of performing probability estimation regarding the integer value of the first feature in the bit stream to be decoded to obtain an estimated probability distribution for the integer value of the first feature is as follows: Performing probability estimation regarding the integer value of the first feature in the bit stream to be decoded based on first information to obtain an estimated probability distribution for the integer value of the first feature, wherein the first information includes at least one of context information and side information. The method comprising the above step.
16. An encoding device including a processing circuit, wherein the processing circuit: Obtains target data to be encoded; Inputs the target data to be encoded into a first encoding network to obtain target parameters; Constructs a second encoding network based on the target parameters; Inputs the target data to be encoded into the second encoding network to obtain a first feature; and Encodes the first feature to obtain an encoded bit stream; The device is configured to perform the above steps.
17. The device according to claim 16, wherein the target parameter is all or part of the parameter weights related to the convolution and non-linear activation of the second encoding network.
18. In the device according to claim 16 or 17, the processing circuit: Rounding the first feature to obtain an integer value of the first feature; Performing probability estimation regarding the integer value of the first feature to obtain an estimated probability distribution for the integer value of the first feature; and Performing entropy encoding regarding the integer value of the first feature based on the estimated probability distribution for the integer value of the first feature to obtain the encoded bit stream; An apparatus specifically configured to perform the above.
19. In the apparatus according to claim 18, the processing circuit: Performing probability estimation regarding the integer value of the first feature based on first information to obtain an estimated probability distribution for the integer value of the first feature, wherein the first information includes at least one of context information and side information; An apparatus specifically configured to perform the above.
20. A decoding apparatus including a processing circuit, wherein the processing circuit: Obtaining a bit stream to be decoded; Decoding the bit stream to be decoded to obtain an integer value of a first feature and an integer value of a second feature, wherein the integer value of the first feature is used to obtain decoded data, and the integer value of the second feature is used to obtain a target parameter; Inputting the integer value of the second feature into a first decoding network to obtain the target parameter; Constructing a second decoding network based on the target parameter; and Inputting the integer value of the first feature into the second decoding network to obtain decoded data; An apparatus specifically configured to perform the above.
21. In the apparatus according to claim 20, the target parameter is all or part of the parameter weights regarding convolution and non-linear activation of the second decoding network.
22. In the apparatus according to claim 20 or 21, the bit stream to be decoded includes a first bit stream to be decoded and a second bit stream to be decoded, and the processing circuit: Decoding the first bit stream to be decoded to obtain the integer value of the first feature; and Decoding the second bitstream to be decoded to obtain an integer value of the second feature; An apparatus specifically configured to perform the above.
23. A decoding apparatus including a processing circuit, wherein the processing circuit: Obtaining a bitstream to be decoded; Decoding the bitstream to be decoded to obtain an integer value of a first feature, wherein the integer value of the first feature is used to obtain decoded data and target parameters; Inputting the integer value of the first feature into a first decoding network to obtain the target parameters; Constructing a second decoding network based on the target parameters; and Inputting the integer value of the first feature into the second decoding network to obtain decoded data; An apparatus specifically configured to perform the above.
24. The apparatus according to claim 23, wherein the target parameters are all or part of the parameter weights related to the convolution and non-linear activation of the second decoding network.
25. In the apparatus according to claim 23 or 24, the processing circuit: Performing probability estimation regarding the integer value of the first feature in the bitstream to be decoded to obtain an estimated probability distribution of the integer value of the first feature; and Performing entropy decoding regarding the bitstream to be decoded based on the estimated probability distribution of the integer value of the first feature to obtain the integer value of the first feature; An apparatus specifically configured to perform the above.
26. In the apparatus according to claim 25, the processing circuit: Performing probability estimation regarding the integer value of the first feature in the bitstream to be decoded based on first information, wherein the first information includes at least one of context information and side information; An apparatus specifically configured to perform the above.
27. One or more processors; and A non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor; An encoder including the above, wherein when the program is executed by the processor, the encoder is capable of executing the method according to any one of claims 1 to 2.
28. One or more processors; and A non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor; A decoder including the above, wherein when the program is executed by the processor, the decoder is capable of executing the method according to any one of claims 5 to 6 or any one of claims 12 to 13.
29. A computer-readable storage medium including a computer program, wherein when the computer program is executed on a computer, the computer is capable of executing the method according to any one of claims 1 to 2, claims 5 to 6, or claims 12 to 13.
30. A computer program including computer program code, wherein when the computer program code is executed on a computer, the computer is capable of executing the method according to any one of claims 1 to 2, claims 5 to 6, or claims 12 to 13.
Citation Information
Patent Citations
Learning method and program
JP2017084320A
Image processing device, image processing program, image recognition device, image recognition program, and image recognition system
JP2022078735A
Learning method and recording medium
US20160260014A1
Image compression method and apparatus thereof
US20220329807A1
Image compression method and apparatus
WO2021135715A1