Video image decoding method, video image encoding method, apparatus, and storage medium
By determining quantization parameters based on channel-level and block-level complexity levels and target bit numbers, the method optimizes quantization selection, enhancing image quality and efficiency in video encoding/decoding.
Patent Information
- Application Number
- JP2025504477
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-26
- Filing Date
- 2023-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-06-30
AI Technical Summary
The challenge in video encoding/decoding technology lies in selecting appropriate quantization parameters to balance image quality and efficiency, as inappropriate choices can lead to image distortion.
A method for determining quantization parameters by obtaining channel-level and block-level complexity levels from a code stream, using rate control parameters to optimize the selection of quantization parameters, and adjusting these parameters based on target bit numbers to improve image quality and encoding/decoding efficiency.
This approach enhances image quality and encoding/decoding efficiency by optimizing quantization parameter selection, reducing distortion, and improving the visual experience.
Smart Images

Figure 2025525001000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] The present invention claims priority based on a Chinese patent application with application number 202210887907.9, filed on July 26, 2022. Herein, all of its contents are incorporated into the present invention by reference.
Technical Field
[0002] The present invention relates to the field of video encoding / decoding technology, and particularly to a video image decoding method, a video image encoding method, an apparatus, and a storage medium.
Background Art
[0003] In the field of video processing, video encoding / decoding technology plays an important role. Video encoding / decoding technology is a technology that reduces the amount of video data by encoding and decoding video. Among them, quantization is an important step in the process of video encoding and decoding. Mainly, by using quantization parameters instead of some original data in the code stream, the reduction of the redundancy of the original data in the code stream is realized. The quantization parameter for quantization is written into the code stream during the video encoding process. On the video decoding side, decoding is realized by analyzing the quantization parameter in the code stream. However, quantization is also accompanied by the risk of image distortion. Therefore, selecting an appropriate quantization parameter can improve the quality of the image. Therefore, how to select the quantization parameter is the key to video encoding / decoding technology.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Embodiments of the present invention provide a video image decoding method, a video image encoding method, an apparatus, and a storage medium, which are useful for improving the quality of images by video encoding / decoding and improving the visual experience.
Means for Solving the Problems
[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions.
[0006] As a first aspect, the embodiments of the present invention provide a video image decoding method and a video image encoding method, which are applicable to a chip of a video encoding device, a video decoding device, or a video encoding and decoding device. The method includes obtaining a channel-level complexity level from a code stream, and determining a block-level complexity level of a current block according to at least two channel-level complexity levels, where the code stream is the encoding code stream of the current block, and the channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block; determining a target bit number of the current block based on a rate control parameter, where the rate control parameter includes the block-level complexity level of the current block; determining a quantization parameter of the current block based on the target bit number; and decoding the current block based on the quantization parameter.
[0007] The above quantization parameter plays an important role in the process of video encoding and decoding. By the video image decoding method and the video image encoding method provided by the present invention, a video encoding / decoding device obtains a channel-level complexity level from a code stream, determines a block-level complexity level of a current block according to at least two channel-level complexity levels, determines a target bit number of the current block based on a rate control parameter including the block-level complexity level, and further determines a quantization parameter of the current block based on the target bit number. According to the above process, the video image decoding method and the video image encoding method provided by the present invention can optimize the selection of the quantization parameter, improve the image quality of video encoding / decoding, and improve the visual experience.
[0008] In one possible implementation, the rate control parameter includes the same-level reversible coding average bit number, the reversible coding average bit number, and the code stream buffer fullness. Determining the target bit number of the current block based on the rate control parameter includes determining the same-level reversible coding average bit number and the reversible coding average bit number, determining an initial target bit number based on the same-level reversible coding average bit number and the reversible coding average bit number, and determining the target bit number of the current block based on the code stream buffer fullness, the block-level complexity level of the current block, and the initial target bit number. Here, the same-level reversible coding average bit number is the average value of the predicted values of the bit numbers required to reversibly code the current block and a plurality of decoded image blocks, and the complexity levels of the plurality of decoded image blocks and the current block are the same. The reversible coding average bit number is the average value of the predicted values of the bit numbers required to reversibly code the current block and all decoded image blocks. The code stream buffer fullness is for indicating the fullness of the buffer, and the buffer is for storing the code stream of the image to be processed.
[0009] In such a possible implementation, a method for determining the target bit number of the current block is provided. By introducing the same-level reversible coding average bit number and the reversible coding average bit number to determine the initial target bit number, and determining the target bit number of the current block based on the initial target bit number, the code stream buffer fullness, and the block-level complexity level, the determination of the quantization parameter can be made more accurate. Furthermore, while ensuring the encoding / decoding efficiency, the effect of improving the image quality can be obtained.
[0010] In one possible implementation, determining the same-level reversible coding average number of bits and the reversible coding average number of bits includes determining the reversible coding number of bits of the current block, where the reversible coding number of bits is a predicted value of the number of bits required to reversibly code the current block, updating the same-level reversible coding average number of bits of the current block based on the reversible coding number of bits of the current block and a plurality of past same-level reversible coding average numbers of bits, where the past same-level reversible coding average number of bits is the same-level reversible coding average number of bits of a decoded image block having the same complexity level as the block-level complexity level of the current block, and updating the reversible coding average number of bits of the current block based on the reversible coding number of bits of the current block and all past reversible coding average numbers of bits, where the past reversible coding average number of bits is the reversible coding average number of bits of a decoded image block.
[0011] In such a possible implementation, a method for determining the same-level reversible coding average number of bits and the reversible coding average number of bits based on the reversible coding number of bits of the current block, the past same-level reversible coding average number of bits, and the past reversible coding average number of bits is provided, which helps to improve the feasibility of implementing the technical solution. By corresponding the same-level reversible coding average number of bits to the block-level complexity level, the selection of quantization parameters can be optimized, and further the effect of improving image quality can be obtained.
[0012] In one possible implementation, the current block is the first row block of the image to be processed, the rate control parameter includes a first row quality improvement parameter, and determining the target number of bits of the current block based on the rate control parameter further includes adjusting the target number of bits of the current block based on the first row quality improvement parameter such that the quantization parameter of the current block becomes smaller.
[0013] In such a possible implementation, another method for determining the target number of bits of the current block is provided and applied to the scenario where the current block is the first row block. When the current block is the first row block, prediction becomes more difficult. Since the prediction error is transient, in this implementation, by introducing a first row quality improvement parameter to reduce the quantization parameter of the first row block, the above influence is reduced, and the effect of improving the image quality of video encoding / decoding can be obtained.
[0014] In one possible implementation, the current block is the first column block of the image to be processed, the rate control parameter includes a first column quality improvement parameter, and determining the target number of bits of the current block based on the rate control parameter further includes adjusting the target number of bits of the current block based on the first row quality improvement parameter so that the quantization parameter of the current block becomes smaller. In such a possible implementation, another method for determining the target number of bits of the current block is provided and applied to the scenario where the current block is the first column block. When the current block is the first column block, prediction becomes more difficult. Since the prediction error is transient, in this implementation, by introducing a first column quality improvement parameter to reduce the quantization parameter of the first column block, the above influence is reduced, and the effect of improving the image quality of video encoding / decoding can be obtained.
[0015] In one possible implementation, obtaining the channel-level complexity level from the code stream is to obtain the complexity information bits of the current block from the code stream, where the complexity information bits are for indicating the channel-level complexity level of the current block, and determining the channel-level complexity level based on the complexity information bits.
[0016] In such a possible implementation, on the decoding side, a method for obtaining the channel-level complexity level is provided. The channel-level complexity level is obtained by analyzing information bits for indicating the channel-level complexity level of the current block in the encoded code stream. In one possible scenario, the complexity information bits may be 1 bit or 3 bits, and the most significant bit in the complexity information bits is for indicating whether the current channel-level complexity level is the same as the complexity level of the same channel of the previous image block of the current block, and the change value between the two. Taking a YUV image as an example, when the current channel-level complexity level is the complexity level of the U channel of the current block, the complexity level of the same channel of the previous image block indicates the complexity level of the U channel of the image block decoded before the current block. When the result of the determination by the most significant bit is the same, the complexity information bits are 1 bit; when the result of the determination is not the same, the complexity information bits are 3 bits, and the lower 2 bits indicate the change value between the channel-level complexity level of the current block and the complexity level of the same channel of the previous image block of the current block. Based on the change value and the channel complexity level of the same channel of the previous image block, the currently required channel-level complexity level can be determined. Also, the above scenario is only an exemplary description, and the protection scope of this possible implementation is not limited thereto. It can be understood that on the decoding side, by providing a specific method for obtaining the channel-level complexity level, the feasibility of the technical solution is improved.
[0017] As a second aspect, an embodiment of the present invention provides a video image encoding method, which is applied to a chip of a video encoding device. The method includes obtaining channel-level texture information of a current block, determining a channel-level complexity level of the current block based on the channel-level texture information, and determining a block-level complexity level of the current block according to at least two channel-level complexity levels, where the channel-level complexity level is used to indicate the degree of complexity of the channel-level texture of the current block; determining a target bit number of the current block based on a rate control parameter, where the rate control parameter includes the block-level complexity level of the current block; determining a quantization parameter of the current block based on the target bit number; and encoding the current block based on the quantization parameter.
[0018] In such a possible implementation, a method for obtaining the channel-level complexity level of the current block is provided. On the encoding side, the channel-level complexity level of the current block is determined based on the channel-level texture information of the current block. On the decoding side, the channel-level complexity level of the current block is obtained from the received encoded code stream. By providing a method for obtaining the channel-level complexity level of the current block in the encoding and decoding processes, the feasibility of the technical solution can be improved, and it becomes easy to determine other rate control parameters and the target bit number based on the channel-level complexity level obtained later, thereby optimizing the quantization parameter.
[0019] In one possible implementation, obtaining the channel-level texture information of the current block and determining the channel-level complexity level of the current block based on the channel-level texture information includes using, as a processing unit, the image block of at least one channel in the current block, dividing the processing unit into at least two sub-units, determining the texture information of each sub-unit, and in the processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit.
[0020] In such a possible implementation, a method for determining the channel-level complexity level of the current block applied to the encoding side is provided, improving the feasibility of the technical solution. Among them, using, as a processing unit, the image block of at least one channel in the current block and further dividing each processing unit into at least two sub-units helps to improve the accuracy of the complexity information.
[0021] In one possible implementation, determining the texture information of each sub-unit includes obtaining the original pixel values of the sub-unit, the original pixel values or reconstructed values of the adjacent column on the left side of the sub-unit, and the reconstructed values of the adjacent row above the sub-unit, calculating the horizontal texture information and vertical texture information of the corresponding sub-unit, and selecting the minimum value from the horizontal texture information and vertical texture information as the texture information of the corresponding sub-unit.
[0022] In such a possible implementation, an implementation for determining the texture information of the sub-unit is provided, improving the feasibility of the technical solution.
[0023] In one possible implementation, in the processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit includes dividing the texture information of each sub-unit into the complexity levels of the corresponding sub-units based on a plurality of thresholds, where the plurality of thresholds are preset, and determining the block-level complexity level of the current block based on the complexity levels of each sub-unit.
[0024] In such a possible implementation, by setting a plurality of thresholds, it provides a way to divide the texture information of each sub-unit into the complexity levels of the corresponding sub-units and determine the block-level complexity level of the current block based on the complexity levels of each sub-unit, which helps to improve the feasibility of the technical solution.
[0025] In one possible implementation, determining the block-level complexity level of the current block based on the complexity levels of each sub-unit includes mapping the complexity levels of each sub-unit to the corresponding channel-level complexity levels based on preset rules, and determining the block-level complexity level of the current block based on the channel-level complexity levels of each channel.
[0026] In such a possible implementation, it provides a method of mapping the complexity levels of each sub-unit to the corresponding channel-level complexity levels and further determining the block-level complexity level of the current block according to the channel-level complexity level of the current block, which helps to improve the feasibility of the technical solution.
[0027] In one possible implementation, mapping the complexity levels of each sub-unit to the corresponding channel-level complexity levels based on preset rules includes determining the channel-level complexity level based on a plurality of thresholds and the sum of the complexity levels of each sub-unit, where the plurality of thresholds are preset.
[0028] In such a possible implementation, a method for determining the channel-level complexity level based on a plurality of thresholds and the sum of the complexity levels of each subunit is provided, which helps to improve the feasibility of the technical solution.
[0029] In one possible implementation, mapping the complexity level of each subunit to the corresponding channel-level complexity level based on preset rules includes determining the level configuration of the complexity level of the subunit and determining the corresponding channel-level complexity level based on the level configuration.
[0030] In such a possible implementation, a method for determining the corresponding channel-level complexity level based on the level configuration of the complexity level of the subunit is provided, which helps to improve the feasibility of the technical solution.
[0031] In one possible implementation, determining the block-level complexity level of the current block based on each channel-level complexity level is obtaining the maximum value, minimum value or weighted value of each channel-level complexity level as the block-level complexity level of the current block, or determining the block-level complexity level of the current block based on a plurality of thresholds and the sum of each channel-level complexity level, wherein the plurality of thresholds are preset.
[0032] In such a possible implementation, two methods for determining the block-level complexity level of the current block based on each channel-level complexity level are provided, which helps to improve the feasibility of the technical solution.
[0033] In one possible implementation, the current block is a multi-channel image block, and each channel component of the multi-channel image block jointly or independently determines the same-level reversible coding bit number and the target bit number.
[0034] In such a possible implementation, each channel component of the multi-channel image block jointly or independently determines the same-level reversible coding bit number and the target bit number, which helps to improve the feasibility of implementing the technical solution.
[0035] In one possible implementation, the current block is a multi-channel image block, and each channel component of the multi-channel image block jointly or independently determines the channel-level complexity level.
[0036] In such a possible implementation, each channel component of the multi-channel image block jointly or independently determines the channel-level complexity level, which helps to improve the feasibility of implementing the technical solution.
[0037] As a second aspect, an embodiment of the present invention provides a video encoding / decoding device, and the device has a function of implementing the video image decoding method or the video image encoding method according to any one of the first aspects. This function can be realized by hardware or by the hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the function.
[0038] As a third aspect, a video encoder is provided. The video encoder includes a processor and a memory. The memory is for storing computer execution instructions. When the video encoder operates, the processor executes the computer execution instructions stored in the memory so as to cause the video encoder to execute the video image decoding method or the video image encoding method according to any one of the first aspects.
[0039] As a fourth aspect, a video decoder is provided. The video decoder includes a processor and a memory. The memory is for storing computer-executable instructions. When the video decoder operates, the processor executes the computer-executable instructions stored in the memory so as to cause the video decoder to execute the video image decoding method or the video image encoding method according to any one of the first aspects.
[0040] As a fifth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer is caused to execute the video image decoding method or the video image encoding method according to any one of the first aspects.
[0041] As a sixth aspect, a computer program product including instructions is provided. When the computer program product is executed on a computer, the computer is caused to execute the video image decoding method according to any one of the first aspects.
[0042] As a seventh aspect, an electronic device is provided. The electronic device includes a video encoding / decoding device. The processing circuit is configured to execute the video image decoding method or the video image encoding method according to any one of the first aspects.
[0043] As an eighth aspect, a chip is provided. The chip includes a processor. The processor is coupled to a memory. Program instructions are stored in the memory. When the program instructions stored in the memory are executed by the processor, the video image decoding method or the video image encoding method according to any one of the first aspects is executed.
[0044] As a ninth aspect, a video encoding and decoding system is provided. The system includes a video encoder and a video decoder. The video encoder is configured to execute the video image decoding method or the video image encoding method according to any one of the first aspects. The video decoder is configured to execute the video image decoding method or the video image encoding method according to any one of the first aspects.
[0045] For any one of the implementation forms in the second aspect to the ninth aspect, the technical effect can be referred to the corresponding implementation form in the first aspect. Here, the description thereof is omitted.
Brief Description of the Drawings
[0046]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Best Mode for Carrying Out the Invention
[0047] In the description of the present invention, unless otherwise specified, " / " means "or". For example, A / B represents A or B. "And / or" in this specification only describes the relationship of related objects and indicates that three types of relationships may exist. For example, A and / or B may indicate that A exists alone, A and B exist simultaneously, and B exists alone. Also, "at least one" means one or more, and "a plurality" means two or more. Characters such as "first" and "second" do not limit the quantity or execution order. And characters such as "first" and "second" do not necessarily limit that they are different. In the present invention, terms such as "exemplary" or "for example" are used as an example, illustration, or explanation. In the present invention, any embodiment or design described as "exemplary" or "for example" should not be construed as being more preferable or advantageous than other embodiments or designs. To be precise, using terms such as "exemplary" or "for example" is intended to indicate related concepts by specific forms.
[0048] First, technical terms related to the embodiments of the present invention will be introduced.
[0049] 1. Video encoding / decoding technology
[0050] Video encoding / decoding technology includes video encoding technology and video decoding technology, and may also be collectively referred to as video encoding and decoding technology.
[0051] Among them, the video sequence has a series of redundant information such as spatial redundancy, temporal redundancy, visual redundancy, information entropy redundancy, structural redundancy, knowledge redundancy, and importance redundancy. In order to remove as much redundant information as possible in the video sequence and reduce the amount of data representing the video, a video coding technology is provided to achieve the effects of reducing storage space and saving transmission bandwidth. The video coding technology is also called video compression technology.
[0052] In order to obtain the data stored or transmitted based on the video compression technology, it is necessary to be realized by the corresponding video decoding technology.
[0053] Within the internationally applicable range, the video compression coding standard is for standardizing video coding and decoding methods. For example, Advanced Video Coding (AVC) in Part 10 of the MPEG-2 and MPEG-4 standards formulated by the Motion Picture Experts Group (MPEG), and H.263, H.264, and H.265 (also known as High Efficiency Video Coding standard, HEVC) formulated by the International Telecommunication Union - Telecommunication Standardization Sector (ITU-T).
[0054] In addition, in the coding algorithm based on the architecture of hybrid coding, the compression coding methods can be combined and used.
[0055] The basic processing unit in the video encoding and decoding process is an image block, which is obtained by dividing an image of one frame per sheet on the encoding side. Usually, the divided image blocks are processed one by one row by row. Here, the currently processed image block is called the current block, and the processed image block is called the post-encoded image block, or the post-decoded image block, or the post-encoded / decoded image block. Taking HEVC as an example, in HEVC, a Coding Tree Unit (CTU), a Coding Unit (CU), a Prediction Unit (PU), and a Transform Unit (TU) are defined. The CTU, CU, PU, and TU can all be image blocks obtained by division. Among them, both the PU and TU are divided based on the CU.
[0056] 2. Video Sampling
[0057] Since a pixel is the smallest complete sampling of a video or an image, data processing for an image block is performed in pixel units. Here, each pixel records color information. One sampling method is to represent colors by RGB, which includes three image channels, where R represents red, G represents green, and B represents blue. Another sampling method is to represent colors by YUV, which also includes three image channels, where Y represents luminance, U represents the first chrominance Cb, and V represents the second chrominance Cr. Since humans are more sensitive to luminance than to chrominance, storage space can be reduced by storing more luminance and less chrominance. Specifically, during video encoding and decoding, usually, YUV formats including 420 sampling format, 422 sampling format, etc. are used for video sampling. This sampling format determines the number of samples of the two chrominances based on the number of luminance samples. For example, if a CU has 4×2 pixels, the format is as follows.
[0058] TIFF2025525001000002.tif12170
[0059] TIFF2025525001000003.tif111156
[0060] TIFF2025525001000004.tif38163
[0061] TIFF2025525001000005.tif93155
[0062] TIFF2025525001000006.tif33163
[0063] The sampled luminance encoding unit, the first chrominance encoding unit, and the second chrominance encoding unit described above are used as data units for each channel for subsequent processing on the current block.
[0064] The encoding / decoding method provided by the present invention is applied to a video encoding and decoding system. The video encoding and decoding system is also referred to as a video encoding / decoding system. FIG. 1 shows the configuration of the video encoding and decoding system.
[0065] As shown in FIG. 1, the video encoding and decoding system includes a source device 10 and a destination device 11. The source device 10 generates encoded video data. The source device 10 may be referred to as a video encoding device or a video encoding apparatus. The destination device 11 is capable of decoding the encoded video data generated by the source device 10. The destination device 11 may be referred to as a video decoding device or a video decoding apparatus. The source device 10 and / or the destination device 11 may include at least one processor and a memory coupled to the at least one processor. The memory may include, but is not limited to, a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, or any other arbitrary medium for storing desired program code in the form of instructions or data structures accessible by a computer.
[0066] The source device 10 and the destination device 11 may include various devices. For example, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, mobile phones such as so-called "smartphones", televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or other similar electronic devices such as those.
[0067] The destination device 11 receives the encoded video data from the source device 10 via the link 12. The link 12 may include one or more media and / or devices capable of transmitting the encoded video data from the source device 10 to the destination device 11. In one example, the link 12 may include one or more communication media that enable the source device 10 to directly transmit the encoded video data to the destination device 11 in real time. In this example, the source device 10 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to the destination device 11. The one or more communication media may include wireless and / or wired communication media such as, for example, the radio frequency (RF) spectrum, one or more physical transmission lines, etc. The one or more communication media may form part of a packet-based network such as, for example, a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices that enable communication from the source device 10 to the destination device 11.
[0068] In other examples, the encoded video data may be output from the output interface 103 to the storage device 13. Similarly, the input interface 113 may access the encoded video data from the storage device 13. The storage device 13 may include various locally accessible data storage media. For example, it may include Blu-ray discs, digital video discs (DVDs), compact disc read-only memories (CD-ROMs), flash memories, or other suitable digital storage media for storing the encoded video data.
[0069] In other examples, the storage device 13 may correspond to another intermediate storage device for storing the encoded video data generated by the file server or the source device 10. In this example, the destination device 11 may obtain the video data stored in the storage device 13 from the storage device 13 by streaming or downloading. The file server may be any type of server that can store the encoded video data and transmit the encoded video data to the destination device 11. For example, the file server may include a World Wide Web (Web) server (e.g., for websites), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, and a local disk drive.
[0070] The destination device 11 may access the encoded video data via a data connection of any standard (e.g., an Internet connection). Examples of the type of data connection may include a wireless communication channel suitable for accessing the encoded video data stored in the file server, a wired connection (e.g., a cable modem, etc.), or a combination of both. The method by which the encoded video data is transmitted from the file server may be streaming, downloading, or a combination of both.
[0071] The encoding / decoding method according to the present invention is not limited to the scenario of wireless applications. Exemplarily, the encoding / decoding method according to the present invention is applied to video encoding and decoding that supports the following various multimedia applications. For example, wireless television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding of video data stored in a data storage medium, decoding of video data stored in a data storage medium, or other applications. In some examples, the video encoding and decoding system is configured to support unidirectional or bidirectional video transmission, so as to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0072] Note that the video encoding and decoding system shown in FIG. 1 is only an example of a video encoding and decoding system and does not limit the video encoding and decoding system in the present invention. The encoding / decoding method provided by the present invention can also be applied to a scenario where there is no data communication between the encoding device and the decoding device. In other examples, the video data to be encoded or the encoded video data may be retrieved from local memory or streamed over a network. The video encoding device may encode the video data to be encoded and store the encoded video data in memory. The video decoding device may obtain the encoded video data from memory and decode the encoded video data.
[0073] In FIG. 1, the source device 10 includes a video source 101, a video encoder 102, and an output interface 103. In some examples, the output interface 103 may include a modulator / demodulator (modem) and / or a transmitter. The video source 101 includes a video capture device (e.g., a video camera), a video archive containing previously captured video data, a video input interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these video data sources.
[0074] The video encoder 102 may encode the video data from the video source 101. In some examples, the source device 10 may directly transmit the encoded video data to the destination device 11 via the output interface 103. In other examples, the encoded video data may be stored in the storage device 13 for later access by the destination device 11 for decoding and / or playback.
[0075] In the example of FIG. 1, the destination device 11 includes a display device 111, a video decoder 112, and an input interface 113. In some examples, the input interface 113 includes a receiver and / or a modem. The input interface 113 may receive the encoded video data via the link 12 and / or from the storage device 13. The display device 111 may be integrated with the destination device 11 or may be external to the destination device 11. Generally, the display device 111 displays the decoded video data. The display device 111 may include various display devices such as, for example, a liquid crystal display, a plasma display, an organic light emitting diode display, or other types of display devices.
[0076] Optionally, the video encoder 102 and the video decoder 112 may each be integrated with an audio encoder and decoder, and include a suitable multiplexer / demultiplexer unit or other hardware and software to process the encoding of both audio and video in a common data stream or individual data streams.
[0077] The video encoder 102 and the video decoder 112 may include at least one microprocessor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the encoding / decoding method provided by the present invention is implemented by software, the instructions used in the software may be stored in a suitable non-volatile computer-readable storage medium, and the present invention may be implemented by executing the instructions by at least one processor.
[0078] The video encoder 102 and the video decoder 112 in the present invention may operate according to a video compression standard (e.g., HEVC) or may operate according to other industry standards. The present invention is not particularly limited.
[0079] FIG. 2 is a schematic block diagram of a video encoder 102 according to an embodiment of the present invention. In the video encoder 102, a prediction module 21, a conversion module 22, a quantization module 23, and an entropy encoding module 24 may perform processes of prediction, conversion, quantization, and entropy encoding, respectively. The video encoder 102 further includes a preprocessing module 20 and an adder 202. Among them, the preprocessing module 20 may include a splitting module and a coding rate control module. For the reconstruction of a video block, the video encoder 102 further includes an inverse quantization module 25, an inverse conversion module 26, an adder 201, and a reference image memory 27.
[0080] As shown in FIG. 2, the video encoder 102 receives video data. The preprocessing module 20 obtains input parameters of the video data. Here, the input parameters include information such as the resolution of the image in the video data, the sampling format of the image, the pixel color depth (bits per pixel, bpp), and the bit width. Here, bpp refers to the number of bits occupied by one pixel component in a unit pixel. The bit width refers to the number of bits occupied by a unit pixel. For example, if one pixel is represented by the values of three pixel components of RGB and each pixel component occupies 8 bits (bits), the pixel color depth of the pixel is 8, and the bit width of the pixel is 3×8 = 24 bits.
[0081] The splitting module in the preprocessing module 20 splits the image into original blocks. This splitting may include splitting into slices, image blocks, or other relatively large units, and (for example) video block splitting based on the four-tree structure of the largest coding unit (LCU) and CU. Exemplarily, the video encoder 102 is a component that encodes video blocks within a video slice to be encoded. Generally, a slice may be split into a plurality of original blocks (and may be split into a set of original blocks called image blocks). Usually, the splitting module determines the sizes of the CU, PU, and TU. Also, the splitting module is further used to determine the size of the coding rate control unit. The coding rate control unit refers to the basic processing unit in the coding rate control module. For example, in the coding rate control module, the coding rate control unit calculates complexity information for the current block and further calculates the quantization parameter of the current block based on the complexity information. Here, the splitting policy of the splitting module may be preset or may be continuously adjusted based on the image during the encoding process. When the splitting policy is a preset policy, correspondingly, the same splitting policy is preset on the decoding side to obtain the same image processing unit. The image processing unit is any one of the above image blocks and corresponds one-to-one with the encoding side. When the splitting policy is continuously adjusted based on the image during the encoding process, the splitting policy can be directly or indirectly incorporated into the code stream. Correspondingly, the decoding side obtains the corresponding parameters from the code stream, obtains the same splitting policy, and obtains the same image processing unit.
[0082] The coding rate control module in the preprocessing module 20 is used to generate quantization parameters so that the quantization module 23 and the inverse quantization module 25 perform correlated calculations. Here, in the process of calculating the quantization parameters, the coding rate control module may, for example, obtain and calculate the image information of the current block such as the above input information, or obtain and calculate the reconstructed value reconstructed by the adder 201, but the present invention is not limited thereto.
[0083] The prediction module 21 may provide a prediction block to the adder 202 to generate a residual block, and provide the prediction block to the adder 201 to obtain a reconstructed block by reconstruction. The reconstructed block is used as a reference pixel for subsequent prediction. Here, the video coder 102 generates a pixel difference value by subtracting the pixel value of the prediction block from the pixel value of the original block. The pixel difference value is the residual block, and the data in the residual block may include a luminance difference and a chrominance difference. The adder 201 represents one or more components that perform this subtraction. The prediction module 21 may further send related syntax elements to the entropy coding module 24 to merge into the code stream.
[0084] The conversion module 22 may divide and convert the residual block into one or more TUs. The conversion module 22 may convert the residual block from a pixel domain to a conversion domain (for example, a frequency domain). For example, the discrete cosine transform (DCT) or the discrete sine transform (DST) is used to convert the residual block to obtain conversion coefficients. The conversion module 32 may send the obtained conversion coefficients to the quantization module 23.
[0085] The quantization module 23 may be quantized by a quantization unit. Here, the quantization unit may be the same as the CU, TU, and PU, and may be further divided in the splitting module. The quantization module 23 quantizes the transform coefficients so as to further reduce the code rate and obtain quantization coefficients. Here, the quantization process can reduce the bit depth associated with some or all of the coefficients. By adjusting the quantization parameters, the degree of quantization can be changed. In some possible embodiments, the quantization module 23 may subsequently perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding module 24 may perform the scan.
[0086] After quantization, the entropy coding module 24 may entropy code the quantization coefficients. For example, the entropy coding module 24 may perform context-adaptive variable-length coding (CAVLC), context-based adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding, or other entropy coding methods or techniques. After entropy coding by the entropy coding module 24, a code stream is obtained, and the code stream is transmitted to the video decoder 112 or stored for subsequent transmission or search by the video decoder 112.
[0087] The inverse quantization module 25 and the inverse transform module 26 apply inverse quantization and inverse transform respectively. The adder 201 adds the inverse-transformed residual block and the predicted residual block to generate a reconstructed block, and the reconstructed block is later used as a reference pixel for predicting the original block. The reconstructed block is stored in the reference image memory 27.
[0088] Figure 3 is a schematic configuration diagram of the video decoder 112 in an embodiment of the present invention. As shown in Figure 3, the video decoder 112 includes an entropy decoding module 30, a prediction module 31, an inverse quantization module 32, an inverse transform module 33, an adder 301, and a reference image memory 34. Here, the entropy decoding module 30 includes an analysis module and a code rate control module. In some possible embodiments, the video decoder 112 may execute a decoding flow that is exemplarily inverse to the encoding flow described for the video encoder 102 shown in Figure 2.
[0089] In the decoding process, the video decoder 112 receives the encoded video code stream from the video encoder 102. The analysis module in the entropy decoding module 30 of the video decoder 112 generates quantization coefficients and syntax elements by entropy decoding the code stream. The entropy decoding module 30 transmits the syntax elements to the prediction module 31. The video decoder 112 may receive the syntax elements at the video slice level and / or the video block level.
[0090] The code rate control module in the entropy decoding module 30 generates quantization parameters based on the information of the image to be decoded obtained by the analysis module so that the inverse quantization module 32 performs a correlation calculation. The code rate control module may calculate the quantization parameters based on the reconstructed block reconstructed by the adder 301.
[0091] The inverse quantization module 32 performs inverse quantization (e.g., dequantization) on the quantization coefficients provided from the code stream and decoded by the entropy decoding module 30 and the generated quantization parameters. The inverse quantization process may include determining the degree of quantization using the quantization parameters calculated for each video block in the video slice by the video encoder 102, and similarly determining the degree of application of inverse quantization. The inverse transform module 33 applies an inverse transform (e.g., a transform method such as DCT, DST) to the inverse quantized transform coefficients, and generates an inverse transformed residual block in the image area for each inverse transform unit from the inverse quantized transform coefficients. Here, the size of the inverse transform unit is the same as the size of the TU. The inverse transform method and the transform method use the corresponding forward transform and inverse transform in the same transform method. For example, the inverse transform of DCT, DST is an inverse DCT, inverse DST or a conceptually similar inverse transform process.
[0092] After the prediction module 31 generates the prediction block, the video decoder 112 adds the inverse transformed residual block from the inverse transform module 33 and the prediction block to generate a decoded video block. The adder 301 indicates one or more components that perform this addition operation. Optionally, the decoded block may be filtered using a deblocking filter to remove block artifacts. The decoded image blocks in a given frame or image are stored in the reference image memory 34 as reference pixels for later prediction.
[0093] The present invention provides one possible implementation of video encoding / decoding. As shown in FIG. 4, FIG. 4 is a schematic diagram of the video encoding / decoding flow provided by the present invention. The video encoding / decoding implementation includes processes 1 to 5. Processes 1 to 5 may be executed by any one or more of the above-described source device 10, video encoder 102, destination device 11, or video decoder 112.
[0094] Process 1: Divide the image of one frame into one or more parallel encoding units that do not overlap with each other. There is no dependency between the one or more parallel encoding units, and they can be encoded and decoded completely in parallel / independently, like the parallel encoding unit 1 and the parallel encoding unit 2 shown in FIG. 4.
[0095] Process 2: Each parallel encoding unit may be further divided into one or more independent encoding units that do not overlap with each other. Although the independent encoding units do not depend on each other, they may share the header information of some parallel encoding units.
[0096] The independent encoding unit may include three components, namely luminance Y, first chrominance Cb, and second chrominance Cr, or three components of RGB, or may include only one of these components. When the independent encoding unit includes three components, the sizes of these three components may be exactly the same or different, specifically related to the input format of the image. This independent encoding unit may be understood as one or more processing units composed of N channels included in each parallel encoding unit. For example, the three components of Y, Cb, and Cr are three channels that make up this parallel encoding unit, and each may be an independent encoding unit. Or, when Cb and Cr are collectively referred to as chrominance channels, this parallel encoding unit includes an independent encoding unit composed of a luminance channel and an independent encoding unit composed of chrominance channels.
[0097] Process 3: Each independent encoding unit may be further divided into one or more encoding units that do not overlap with each other. Each encoding unit within the independent encoding unit may depend on each other. For example, multiple encoding units may perform preliminary encoding and preliminary decoding with reference to each other.
[0098] When the sizes of the symbolization unit and the independent symbolization unit are the same (i.e., the independent symbolization unit is divided into only one symbolization unit), the size may be any of the sizes described in Process 2.
[0099] The symbolization unit may include three components of luminance Y, first chrominance Cb, and second chrominance Cr (or three components of RGB), or may include only one of these components. When including three components, the sizes of these components may be exactly the same or different. Specifically, it is related to the input format of the image.
[0100] Note that Process 3 is one selectable step in the video encoding and decoding method, and the video encoder / decoder may perform encoding / decoding on the residual coefficients (or residual values) of the independent symbolization unit obtained by Process 2.
[0101] Process 4: The symbolization unit may be further divided into one or more non-overlapping prediction groups (PG). PG may be abbreviated as Group. Each PG is encoded and decoded according to the selected prediction mode, obtains the prediction value of the PG, and constitutes the prediction value of the entire symbolization unit. Based on the prediction value and the original value of the symbolization unit, the residual value of the symbolization unit is obtained.
[0102] Process 5: Based on the residual value of the symbolization unit, the symbolization units are grouped to obtain one or more non-overlapping residual blocks (RB). The residual coefficients of each RB are encoded and decoded according to the selected mode to form a residual coefficient stream. Specifically, it may be divided into two types: the case of converting the residual coefficients and the case of not converting the residual coefficients.
[0103] Here, the selection modes of the encoding and decoding methods for the residual coefficients in Process 5 include, but are not limited to, semi-fixed length encoding method, exponential Golomb encoding method, Golomb-Rice encoding method, truncated unary encoding method, run-length encoding method, method of directly encoding the original residual value, etc.
[0104] For example, the video encoder may directly encode the coefficients in the RB.
[0105] Also for example, the video encoder may perform a transformation such as DCT, DST, Hadamard transform, etc. on the residual block and then encode the transformed coefficients.
[0106] As a possible example, when the RB is small, the video encoder may directly uniformly quantize each coefficient in the RB and then perform binary encoding. When the RB is large, it may be further divided into a plurality of coefficient groups (CG), and then each CG may be uniformly quantized and further binary encoded. In some embodiments of the present invention, the coefficient group (CG) and the quantization group (QG) may be the same.
[0107] The following is an exemplary description of the part for encoding the residual coefficients in the semi-fixed length encoding method. First, define the maximum value of the absolute values of the residuals within one RB block as the modified maximum (mm). Next, determine the number of encoding bits for the residual coefficients within this RB block (the number of encoding bits for the residual coefficients within the same RB block is the same). For example, if the critical limit (CL) of the current RB block is 2 and the current residual coefficient is 1, 2 bits are required to encode the residual coefficient 1, which is represented as 01. When the CL of the current RB block is 7, it indicates encoding an 8-bit residual coefficient and a 1-bit sign bit. Determining the CL is to find the minimum value of M that satisfies the condition that all residuals in the current sub-block are within the range of [−2^(M−1), 2^(M−1)]. When both of the two boundary values of −2^(M−1) and 2^(M−1) exist simultaneously, M should only be incremented by 1, that is, M + 1 bits are required to encode all residuals in the current RB block. When only one of the two boundary values of −2^(M−1) and 2^(M−1) exists, it is necessary to encode one Trailing bit to determine whether the boundary value is −2^(M−1) or 2^(M−1). When neither −2^(M−1) nor 2^(M−1) exists for all residuals, there is no need to encode this Trailing bit.
[0108] Also, in a specific case, the video encoder may directly encode the original value of the image instead of the residual value.
[0109] The video encoder 102 and the video decoder 112 can be implemented by another implementation form. For example, it can be implemented using a general-purpose digital processor system such as the encoding and decoding device 50 shown in FIG. 5. The encoding and decoding device 50 may be a part of the devices in the video encoder 102 or a part of the devices in the video decoder 112.
[0110] The encoding and decoding device 50 may be applied to the encoding side or the decoding side. The encoding and decoding device 50 includes a processor 501 and a memory 502. The processor 501 is connected to the memory 502 (for example, connected to each other via a bus 504). Optionally, the encoding and decoding device 50 further includes a communication interface 503, and the communication interface 503 connects the processor 501 and the memory 502 and is used for transmitting and receiving data.
[0111] The memory 502 may be a random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory 502 is for storing relevant program codes and video data.
[0112] The processor 501 may be one or more central processing units (CPUs) such as CPU 0 and CPU 1 shown in FIG. 5 for example. When the processor 501 is one CPU, this CPU may be a single-core CPU or a multi-core CPU.
[0113] The processor 501 is for reading the program codes stored in the memory 502 and executing the operations of any one of the embodiments corresponding to FIG. 6 and its various executable embodiments.
[0114] Hereinafter, by combining the video encoding and decoding system shown in FIG. 1, the video encoder 102 shown in FIG. 2, and the video decoder 112 shown in FIG. 3, the encoding / decoding method provided by the present invention will be described in detail.
[0115] FIG. 6 is a flowchart of a video image decoding method and a video image encoding method provided by the present invention. The method includes the following steps.
[0116] In S601, at least two channel-level complexity levels of the current block in the image to be processed are obtained, and the block-level complexity level of the current block is determined according to at least two channel-level complexity levels. The channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block.
[0117] The video image decoding method and the video image encoding method provided by the present invention are applied to an encoding and decoding scenario of video data. The video data consists of a plurality of frame images, one frame is a still image, and a dynamic video is generated by synthesizing a temporally continuous frame sequence. The image to be processed is an image to be encoded / decoded in the video data. When encoding / decoding a video, the image to be processed is divided into a plurality of image blocks, and each image block is processed one by one row by row with the image block as the basic processing unit. Here, the image block being processed is called the current block.
[0118] In this exemplary embodiment, the image to be processed may be a multi-channel image. At least two channel-level complexity levels of the current block in the image to be processed are obtained, and the block-level complexity level of the current block is determined according to at least two channel-level complexity levels. That is, at least two channel complexity levels of the current multi-channel image block are obtained, and based on this, the block-level complexity level of the current block is determined. For example, the image to be processed may be an image in YUV format. In this case, the complexity levels of the Y channel and the U channel are obtained, and based on this, the block-level complexity level of the current block is determined. It can be understood that obtaining the complexity levels of YV or UV or YUV can also determine the block-level complexity level of the current block.
[0119] The multi-channel according to the present invention is not limited to the above three YUV channels, and may have more channels. For example, when the image sensor is a four-channel sensor, the corresponding image to be processed includes four-channel image information, and when the image sensor is a five-channel sensor, the corresponding image to be processed includes five-channel image information.
[0120] The plurality of channels in the present invention may include at least one or more of a Y channel, a U channel, a V channel, a Co channel, a Cg channel, an R channel, a G channel, a B channel, an alpha channel, an IR channel, a D channel, and a W channel. For example, the plurality of channels may include a Y channel, a U channel, and a V channel, or the plurality of channels may include an R channel, a G channel, and a B channel, or the plurality of channels may include an R channel, a G channel, a B channel, and an alpha channel, or the plurality of channels may include an R channel, a G channel, a B channel, and an IR channel, or the plurality of channels may include an R channel, a G channel, a B channel, and a W channel, or the plurality of channels may include an R channel, a G channel, a B channel, an IR channel, and a W channel, or the plurality of channels may include an R channel, a G channel, a B channel, and a D channel, or the plurality of channels may include an R channel, a G channel, a B channel, a D channel, and a W channel. Here, in addition to the RGB color light-sensitive channels, there are also an IR channel (an infrared or near-infrared light-sensitive channel), a D channel (a dark channel mainly by infrared light or near-infrared light), and a W channel (a panchromatic light-sensitive channel). Different sensors have different channels. For example, the sensor type may be an RGB sensor, an RGBIR sensor, an RGBW sensor, an RGBIRW sensor, an RGBD sensor, an RGBDW sensor, etc.
[0121] Texture is a visual feature that reflects a uniform phenomenon in an image and is used to represent the attribute of the organized arrangement of a surface structure having a gentle change or a periodic change on the object surface. The channel-level complexity level indicates the degree of complexity of the channel-level texture of the current block. The more complex the channel-level texture information of the current block is, the higher the channel-level complexity level becomes. Similarly, the simpler the channel-level texture information is, the lower the channel-level complexity level becomes.
[0122] Obtaining at least two channel-level complexity levels of the current block in the image to be processed and determining the block-level complexity level of the current block according to at least two channel-level complexity levels have different implementation forms on the encoding side and the decoding side of video encoding / decoding. In one embodiment, On the encoding side, it is possible to obtain the channel-level texture information of the current block and determine the block-level complexity level of the current block based on the channel-level texture information. Specifically, the process can be realized by using at least one channel image block of the current block as a processing unit, dividing each processing unit into at least two sub-units, determining the texture information of each sub-unit, and in each processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit.
[0123] Among them, at least one channel image block in the current block is used as a processing unit, and each processing unit is divided into at least two sub-units. Taking the case where the image to be processed is a YUV image as an example, in one possible scenario, the process can be realized by further dividing the Y channel of the current block as one processing unit into four sub-units. In another possible scenario, the process can also be realized by taking two channels, namely the Y channel and the U channel, as one processing unit and further dividing the processing unit into two sub-units. It should be understood that the above scenarios are only illustrative explanations, and the protection scope of this exemplary embodiment is not limited thereto. For example, the U channel, the V channel, the YU channel, the YV channel, the UV channel, or the YUV channel may be used as one processing unit, and the number of the sub-units may be any integer greater than or equal to 2.
[0124] In one possible implementation, determining the texture information of each sub-unit is achieved by obtaining the original pixel values of the sub-unit, the original pixel values or reconstructed values of the adjacent column on the left side of the sub-unit, and the reconstructed values of the adjacent row above the sub-unit, and accordingly calculating the horizontal texture information and vertical texture information of the sub-unit, and selecting the minimum value from the horizontal texture information and vertical texture information as the texture information of the corresponding sub-unit.
[0125] Specifically, the texture information may be pixel point information. Taking the case where the image to be processed is a YUV image and each of the Y, U, and V channels is regarded as one processing unit as an example, after dividing the processing unit into at least two sub-units, determining the texture information of the sub-units may be to calculate the horizontal complexity and vertical complexity of the sub-units based on the original pixel values of the sub-units, the original pixel values or reconstructed values of the adjacent column on the left side of the sub-units, and the reconstructed values of the adjacent row above the sub-units, and select the minimum value from the horizontal complexity and vertical complexity as the texture information of the corresponding sub-units. The horizontal complexity and vertical complexity may be calculated based on the degree of difference of pixel points in the horizontal and vertical directions of the sub-units. The scenario is only an exemplary explanation, and it can be understood that the texture information can also be obtained by other methods. For example, it may be determined by image information other than pixel points, or the weighted values of the horizontal complexity and vertical complexity may be selected as the texture information of the corresponding sub-units, or other methods that can achieve the same effect may also be used. In this embodiment, it is not particularly limited.
[0126] In one possible implementation, in each of the processing units, determining the block-level complexity level of the current block based on the texture information of each sub-unit may be to divide the texture information of each sub-unit into the complexity levels of the corresponding sub-units based on a plurality of thresholds in the processing unit, where the plurality of thresholds are preset, and determine the block-level complexity level of the current block based on the complexity levels of each sub-unit.
[0127] Specifically, the process may set two thresholds, i.e., threshold 1 and threshold 2, and divide the complexity levels of the sub-units into three levels of 0 to 2 based on the set thresholds. If it is determined that the complexity of the obtained texture information of the sub-unit is less than or equal to threshold 1, the texture information of the sub-unit is divided into level 0. If it is determined that the complexity of the obtained texture information of the sub-unit is greater than threshold 1 and less than threshold 2, the texture information of the sub-unit is divided into level 1. If it is determined that the complexity of the obtained texture information of the sub-unit is greater than or equal to threshold 2, the texture information of the sub-unit is divided into level 2. It should be understood that the above scenario is only an exemplary scenario, and the protection scope of the present invention is not limited thereto.
[0128] In one possible implementation, determining the block-level complexity level of the current block based on the complexity levels of the respective sub-units is achieved by mapping the complexity levels of the respective sub-units to the corresponding channel-level complexity levels based on preset rules, and determining the block-level complexity level of the current block based on the channel-level complexity levels.
[0129] Mapping the complexity levels of the respective sub-units to the corresponding channel-level complexity levels based on the preset rules is achieved as follows.
[0130] Implementation 1: Based on a plurality of thresholds and the sum of the complexity levels of the respective sub-units, determine the channel-level complexity level, where the plurality of thresholds are preset.
[0131] In one embodiment, the process is implemented by determining a plurality of channel-level complexity levels based on the plurality of thresholds, adding the complexity levels of each subunit, and dividing the obtained sum into corresponding channel-level complexity levels. Specifically, the thresholds may be three thresholds of 2, 4, and 7, and these three thresholds may divide the channel-level complexity levels into five levels from level 0 to level 4. The calculated complexity levels of the subunits are three levels from level 0 to level 2. Add the complexity levels of each subunit. If the sum of the complexity levels of the subunits is less than 2, the corresponding channel-level complexity level is 0. If the sum of the complexity levels of the subunits is 2 or more and less than 4, the corresponding channel level is 1. If the sum of the complexity levels of the subunits is 4, the corresponding channel-level complexity level is 2. If the sum of the subunit complexity levels is greater than 4 and less than 7, the corresponding channel-level complexity level is 3. If the sum of the complexity levels of the subunits is 7 or more, the corresponding channel-level complexity level is 4. It is understood that the scenario is only an exemplary scenario and the protection scope of the present invention is not limited thereto.
[0132] Implementation form 2: Determine the level configuration of the complexity level of the subunit, and determine the corresponding channel-level complexity level based on the level configuration.
[0133] In one embodiment, in one specific scenario, the process is implemented by, for example, determining the corresponding channel-level complexity level based on the complexity levels of the subunits being a plurality of levels and the levels and arrangement methods to which each subunit belongs. For example, when the processing unit is divided into four subunits and the complexity levels of the four subunits are 1, 2, 2, and 2 respectively, if the preset rule includes a determination method that there are three 2s in the complexity levels of each subunit, the channel-level complexity level is 2, and the corresponding channel-level complexity level is 2. It is understood that the scenario is only an exemplary explanation, and other subunit division methods and determination methods also belong to the protection scope of the present invention.
[0134] In one possible implementation, there are multiple forms of determining the block-level complexity level of the current block based on the complexity levels of each channel level. Exemplarily, the following forms are included.
[0135] Implementation 1: Obtain the maximum value, minimum value, or weighted value of the complexity level of each channel level as the block-level complexity level of the current block. Implementation 2: Determine the block-level complexity level of the current block based on a plurality of thresholds and the sum of the complexity levels of each channel level. The plurality of thresholds are preset.
[0136] Also, in the present invention, each channel component of the multi-channel image block, that is, the current block, determines the channel-level complexity level jointly or independently. For example, when the image to be processed is a YUV image, the U channel and the V channel may share one channel-level complexity level, or may determine their respective channel-level complexity levels individually.
[0137] On the decoding side, it is implemented to obtain the channel-level complexity level of the current block from the code stream. Here, the code stream is the encoded code stream of the current block. Specifically, the decoding side receives the encoded code stream of the current block transmitted from the encoding side. There is complexity information bit in the encoded code stream for indicating the channel-level complexity level. Based on this information bit, the decoding side obtains the channel-level complexity level determined by the channel-level texture information on the encoding side, and determines the block-level complexity level based on the channel-level complexity level. Here, since the implementation method of determining the block-level complexity level by at least two channel-level complexity levels is the same as that on the encoding side, the description is omitted here.
[0138] The implementation method for determining the channel-level complexity level by the complexity information bits is as follows. Obtain the complexity information bits of the current block from the encoded code stream, and determine the channel-level complexity level based on the complexity information bits. Here, the complexity information bits may be 1 bit or 3 bits, and the most significant bit in the complexity information bits is for indicating whether the current channel-level complexity level is the same as the complexity level of the same channel of the previous image block of the current block, and the change value between the two. Taking a YUV image as an example, when the current channel-level complexity level is the complexity level of the U channel of the current block, the complexity of the same channel of the previous image block indicates the complexity level of the U channel of the decoded image block before the current block. If it is determined to be the same by the most significant bit, the complexity information bits are 1 bit. If it is not the same, the complexity information bits are 3 bits, and the lower 2 bits indicate the change value between the channel-level complexity of the current block and the complexity level of the same channel of the previous image block of the current block. Based on the change value and the channel complexity level of the same channel of the previous image block, the currently required channel-level complexity level can be determined. It should be understood that the above scenario is only an exemplary scenario, and the protection scope of the present invention is not limited thereto. For example, the complexity information bits may indicate whether the complexity of the U channel of the current block is the same as that of the Y channel of the current block, and the change value in the case of difference, but the present invention is not particularly limited thereto.
[0139] In S602, determine the target number of bits of the current block according to the rate control parameter. The rate control parameter includes the block-level complexity level of the current block.
[0140] In one possible implementation, the rate control parameter includes the block-level complexity level of the current block calculated by the step S601 and is used to determine the target number of bits of the current block. The target number of bits is the number of bits required to predict the encoding of the current block. The rate control parameter includes at least one of the image bit width bpc, the target number of bits per pixel bpp, the image format, the average number of bits of the same-level reversible encoding, the average number of bits of reversible encoding, the code stream buffer fullness, the leading row quality improvement parameter, and the leading column quality improvement parameter. The target number of bits of the current block is determined by one or more of the above-mentioned rate control parameters.
[0141] Here, the average number of bits of the same-level reversible encoding is the average value of the predicted values of the number of bits required to reversibly encode the current block and a plurality of coded image blocks, and the complexity levels of the plurality of decoded image blocks and the current block are the same. The average number of bits of reversible encoding is the average value of the predicted values of the number of bits required to reversibly encode the current block and all decoded image blocks. The code stream buffer fullness is for indicating the fullness of the buffer, and the buffer is for storing the code stream of the image to be processed. The leading row quality improvement parameter is for reducing the influence caused by the difficulty of predicting the leading row block and the transitional prediction error when the current block is the leading row block in the image to be processed, so that the quantization parameter of the leading row block becomes smaller.
[0142] In one possible implementation, the rate control parameter includes the same-level reversible coding average bit number, the reversible coding average bit number, and the code stream buffer fullness. Determining the target bit number of the current block according to the rate control parameter includes determining the same-level reversible coding average bit number and the reversible coding average bit number, determining an initial target bit number based on the same-level reversible coding average bit number and the reversible coding average bit number, and determining the target bit number of the current block based on the code stream buffer fullness, the block-level complexity level of the current block, and the initial target bit number. Here, the calculation of the code stream buffer fullness is affected by the initial transmission delay mechanism. The initial transmission delay mechanism refers to the influence of some invalid bits existing in the buffer on the buffer fullness before the coded code stream of the current block is stored in the buffer.
[0143] In one embodiment, determining the same-level reversible coding average bit number and the average coding bit number includes determining the reversible coding bit number of the current block, where the reversible coding bit number is a predicted value of the number of bits required to reversibly code the current block, updating the same-level reversible coding average bit number of the current block based on the reversible coding bit number of the current block and a plurality of past same-level reversible coding average bit numbers, and updating the reversible coding average bit number of the current block based on the reversible coding bit number of the current block and all past reversible coding average bit numbers. Here, the past same-level reversible coding average bit number is the same-level reversible coding average bit number of the decoded image block having the same complexity level as the block complexity level of the current block, and the past reversible coding average bit number is the reversible coding average bit number of the decoded image block.
[0144] In one possible implementation, when the current block is the top row block in the image to be processed, the rate control parameter further includes a top row quality improvement parameter. Determining the target number of bits for the current block according to the rate control parameter includes determining an initial target number of bits based on the same-level reversible coding average number of bits and the reversible coding average number of bits, and determining the target number of bits for the current block based on the code stream buffer fullness, the block-level complexity level of the current block, the top row quality improvement parameter, and the initial target number of bits. Here, the top row quality improvement parameter mainly improves the image quality of the current block by reducing the quantization parameter of the current block.
[0145] In one possible implementation, when the current block is the leftmost column block in the image to be processed, the rate control parameter further includes a leftmost column quality improvement parameter. Determining the target number of bits for the current block according to the rate control parameter includes determining an initial target number of bits based on the same-level reversible coding average number of bits and the reversible coding average number of bits, and determining the target number of bits for the current block based on the code stream buffer fullness, the block-level complexity level of the current block, the leftmost column quality improvement parameter, and the initial target number of bits. Here, the leftmost column quality improvement parameter mainly improves the image quality of the current block by reducing the quantization parameter of the current block.
[0146] Also, in the process, the channel components of the multi-channel image block, that is, the current block, jointly or independently determine the same-level reversible coding number of bits and the target number of bits.
[0147] In S603, determine the quantization parameter of the current block based on the target number of bits.
[0148] In one possible implementation, determining the quantization parameter of the current block based on the target number of bits involves calculating and obtaining the reference quantization parameter of the current block based on the same-level reversible coding average number of bits, the target number of bits, and the sampling rate corresponding to the image format of the image to be processed. Further, based on the reference quantization parameter, each component quantization parameter corresponding to each channel of the current block is calculated.
[0149] In S604, the current block is encoded / decoded based on the quantization parameter.
[0150] In this step, the video encoding / decoding device encodes / decodes the current block based on the quantization parameter of the current block. Optionally, during encoding, the video encoder incorporates the channel-level complexity level of the current block into the code stream, i.e., the complexity information bits. Or, the quantization parameter of the current block is incorporated into the code stream. Correspondingly, on the decoding side, the complexity information bits in the code stream are obtained, the quantization parameter is calculated, and decoding is performed. Or, the decoding side obtains the quantization parameter in the code stream and performs decoding. Of course, the video encoder may incorporate the above two pieces of information into the code stream.
[0151] Next, taking the case where the image to be processed is a YUV image as an application scenario, by combining the flow shown in FIG. 7, a series of specific embodiments are used to explain in detail the video image decoding method and the video image encoding method.
[0152] In step S701, the block-level complexity level of the current block is determined. This step is for determining the complexity level of the current block, and the implementation processes on the video encoding side and the video decoding side are different.
[0153] On the video encoding side, this process is mainly realized by the flow shown in FIG. 8.
[0154] In S801, the complexity of the texture information of the current block is determined. This step is realized by using, as a processing unit, the image block of at least one channel of the current block, dividing each processing unit into at least two sub-units, and determining the complexity of the texture information of each sub-unit.
[0155] Specifically, taking the YUV444 format as an example, as shown in FIG. 9, each 16x2 channel may be divided into four 4x2 sub-blocks. The texture information of the current block is the pixel information of the current block. To calculate the complexity of the current block, it is necessary to use the pixel points of the following three parts. As shown in FIG. 9, (1) the original pixel value of the current block, (2) the leftmost column of the current block, that is, the original pixel value of the adjacent column on the left side of sub-block 1 (it should be noted that when the original pixel value cannot be obtained, the reconstructed value may be used); (3) the reconstructed value of the previous row adjacent to the current block, that is, the gray grid area in FIG. 9, are included.
[0156] Next, taking sub-block 1 in FIG. 9 as an example, the process of calculating the complexity of the sub-block, that is, the process of determining the texture information of each sub-unit, will be described. Usually, it can be obtained by calculating the horizontal complexity and the vertical complexity of each sub-block. It is realized as follows. The calculation method of the horizontal complexity is the sum of the absolute values of the pixel values of the current column and the adjacent column on its left side. The calculation method of the vertical complexity is the sum of the absolute values of the pixel values of the current row and the adjacent row above it.
[0157] In one embodiment, when the current block is the left boundary of the current video slice, the calculation of the horizontal complexity uses the filling value of the left boundary, and the filling value is the pixel value of the current column. Correspondingly, when the current block is the upper boundary of the current video slice, the calculation of the vertical complexity uses the filling value of the upper boundary, and the filling value is the pixel value of the current row.
[0158] Specifically, the process is realized as follows.
[0159] First, calculate the horizontal complexity sub_comp_hor of sub-block 1. Specifically, it is calculated by the set ori_pix[i][j] consisting of the pixel values constituting sub-block 1, the pixel values or reconstructed values of the left adjacent column of sub-block 1, and the pixel values of the upper adjacent row. Here, i and j represent the row and column where the pixel value is located. The pixel value at the first row and first column of sub-block 1 is represented as ori_pix[0][0], and other pixel values are estimated and shown in this way. The horizontal complexity sub_comp_hor of sub-block 1 refers to the degree of difference between pixel points in the horizontal direction of sub-block 1. Specifically, it is calculated as follows.
[0160]
Equation
[0161] Obtain the absolute value of the above formula as the horizontal complexity between the pixel value at the first row and first column in sub-block 1 and the adjacent pixel value on its left.
[0162]
Equation
[0163] Obtain the absolute value of the above formula as the horizontal complexity between each pixel value in the first row in sub-block 1.
[0164]
Equation
[0165] Obtain the absolute value of the above formula as the horizontal complexity between the pixel value at the second row and first column in sub-block 1 and the adjacent pixel value on its left.
[0166]
Equation
[0167] The absolute value of the above formula is obtained as the horizontal complexity between each pixel value in the second row in sub-block 1.
[0168] The vertical complexity sub_comp_ver of sub-block 1 refers to the degree of difference between pixel points in the vertical direction of sub-block 1. Similarly, the vertical complexity sub_comp_ver is calculated as follows.
[0169]
Equation
[0170] The absolute value of the above formula is obtained as the vertical complexity between the pixel values in the first row and the adjacent upper row in sub-block 1.
[0171]
Equation
[0172] The absolute value of the above formula is obtained as the vertical complexity between the pixel values in the second row and the first row in sub-block 1.
[0173] After obtaining the multiple horizontal complexities and multiple vertical complexities of sub-block 1, the minimum value among them is defined as the complexity comp of the texture information of sub-block 1. That is, the complexity of sub-block 1 is
Equation
[0174] The calculation method of the complexity of the texture information of the said sub-blocks 2, 3, and 4 is the same as that of sub-block 1, and the method of calculating the complexity of the texture information of the sub-blocks by dividing the textures of the U channel and the V channel into sub-blocks is the same as that of the Y channel, so the description is omitted here.
[0175] Also, when it is necessary to merge a plurality of channels to jointly calculate complexity, for example, when it is necessary to merge the U channel and the V channel to jointly calculate complexity, it is realized by the following formula.
[0176]
Number
[0177] In step S802, based on the complexity of the texture information of the sub-blocks of each channel, determine the block-level complexity level of the current block.
[0178] In one embodiment, this step is realized by the following three steps. In S8021, determine the complexity level of the sub-blocks of each channel according to the complexity of the texture information of the sub-blocks of each channel.
[0179] This process is realized by setting a plurality of thresholds. In one embodiment, specifically, it is realized as follows.
[0180] Implementation form 1: Set two thresholds of thres1 and thres2.
[0181]
Number
[0182] When bpc is less than 8, the defaults of the two thresholds are 2 and 6 respectively.
[0183] The complexity sub_comp of the texture information of each sub-block is divided into three complexity levels of level 0, level 1, and level 2 by the two thresholds. Specifically, it is divided as follows. When sub_comp <= thres1, sub_comp_level = 0. When thres1 < sub_comp < thres2, sub_comp_level = 1. When sub_comp >= thres2, sub_comp_level = 2.
[0184] Embodiment 2: Set four thresholds thres1, thres2, thres3, and thres4. They are set as thres1 = 2 * (1 << (bpc - 8)), thres2 = 4 * (1 << (bpc - 8)), thres3 = 6 * (1 << (bpc - 8)), thres4 = 8 * (1 << (bpc - 8)), where bpc >= 8.
[0185] Based on the four thresholds, divide the complexity sub_comp_level of the texture information of each sub-block into levels 0, 1, 2, 3, and 4. Specifically, divide as follows. When sub_comp <= thres1, sub_comp_level = 0. When thres1 < sub_comp < thres2, sub_comp_level = 1. When thres2 < sub_comp < thres3, sub_comp_level = 2. When thres3 < sub_comp < thres4, sub_comp_level = 3. When sub_comp >= thres4, sub_comp_level = 4.
[0186] In S8022, map the complexity level of the sub-blocks of each channel to the channel-level complexity level.
[0187] Mapping the complexity level of the sub - blocks of each channel to the channel - level complexity level is achieved by presetting a plurality of thresholds and based on threshold mapping or a preset mapping policy. In one embodiment, there are three implementation forms as follows.
[0188] Implementation form 1: Set a plurality of thresholds, calculate the sum sub_comp_level of the thresholds of each sub - block within the corresponding channel, and map the complexity of the texture information of each sub - block to the channel - level complexity level of the corresponding channel according to the plurality of thresholds. Taking the channel sub - block shown in FIG. 9 above as an example, this process is specifically as follows. Set three thresholds 2, 4, and 7, calculate the sum sum_sub_comp_level of the complexity levels of sub - blocks 1 to 4, and map the complexity sub_comp_level(0, 1, 2) of the texture information of each of the four 3 - level sub - blocks to one 5 - level channel - level complexity level comp_level(0, 1, 2, 3, 4) based on the three thresholds.
[0189] The process of mapping the complexity of the texture information of each sub - block to the obtained channel - level complexity level comp_level(0, 1, 2, 3, 4) according to the thresholds is as follows. When sum_sub_comp_level < 2, comp_level = 0; when 2 <= sum_sub_comp_level < 4, comp_level = 1; when sum_sub_comp_level == 4, comp_level = 2; when 4 < sum_sub_comp_level < 7, comp_level = 3; when 7 <= sum_sub_comp_level, comp_level = 4.
[0190] Embodiment 2: Preset four thresholds 5, 7, 10, and 12, add all sub_comp_levels, and obtain comp_level according to the four thresholds. Here, the possible values of comp_level are (0, 1, 2, 3, 4). In this embodiment, different from Embodiment 1, four 5-level sub_comp_levels may be mapped to one 5-level comp_level.
[0191] Embodiment 3: Obtain comp_level according to a preset logical rule. In one example, taking the luminance channel of the current block as an example, the determination method of this logical rule is as follows.
[0192] Regarding the level composition of the complexity level composed of the sub_comp_levels of the four determined sub-blocks, When the level composition of the complexity level includes two 0s, or one 0 and three 1s, If there are two consecutive 0s and the number of 2s is less than 2, then comp_level = 0; otherwise, comp_level = 1. When the level composition of the complexity level includes three 2s, or there are two consecutive 2s, If there are three 2s, then comp_level = 4; otherwise, comp_level = 3. In other cases, comp_level = 2.
[0193] Taking the 16x2 chrominance channel of the current block as an example, the determination method of the logical rule is as follows.
[0194] When the level composition of the complexity level includes two 0s and the number of 2s is less than 2, If there are three 0s, or two consecutive 0s and the number of 2s is 0, then comp_level = 0; otherwise, comp_level = 1.
[0195] If the level composition of the complexity level includes two 2s, or there is one 2 and three 1s, If there are three 2s, or two consecutive 2s, and the number of 0s is 0, then comp_level = 4; otherwise, comp_level = 3. In other cases, comp_level = 2.
[0196] Also, in another possible implementation, taking the 8x2 or 8x1 chrominance channel of the current block as an example, the determination method of the above logical rules may be as follows.
Number
[0197] In S8023, determine the block-level complexity level of the current block based on the complexity level of each channel level.
[0198] Taking the case where the image to be processed is a YUV image as an example, the channel-level complexity levels of the three channels of Y, U, and V may be determined by the steps S8021 and S8022. In one embodiment, determining the block-level complexity level of the current block based on the complexity level of each channel level may be realized as follows.
[0199] Implementation form 1: Determine the block-level complexity level blk_comp_level of the current block based on the sum of the complexity levels of each channel level.
[0200]
Number
[0201]
Table 1
[0202] In Table 1 above, sample_rate is the sampling rate of the image to be processed, and format_bias is the bias amount set when calculating the quantization parameter.
[0203] In steps S8021 to S8023, the channel-level complexity level may be shared among multiple channels or may be independent. For example, in the case of an image in YUV or YCoCg format, the luminance uses one complexity, the chrominance shares one complexity, or the three channels determine the channel complexity level individually.
[0204] Here, when the first chrominance and the second chrominance share one complexity level, the calculation method of the complexity level is as follows.
[0205] Implementation form 1: Obtain the minimum value, maximum value, or weighted value of the complexity levels of two chrominance channel levels.
[0206] Implementation form 2: Obtain the minimum value, maximum value, or weighted value of the complexity of the texture information of two chrominance channels as the chrominance texture complexity, and obtain the chrominance channel-level complexity level according to the method of steps S8021 to S8023.
[0207] On the video decoding side, obtaining the complexity level of each channel from the code stream encoded by the encoding side is achieved by obtaining the complexity information bits of the current block from the encoded code stream and determining the complexity level of each channel based on the complexity information bits. Here, the complexity information bits may be 1 bit or 3 bits, and the most significant bit in the complexity information bits indicates whether the complexity level of the current channel is the same as the complexity level of the same channel of the previous image block of the current block, and the change value between the two.
[0208] Specifically, taking the case where the image to be processed is a YUV image as an example, for the decoding side to obtain the complexity level of each channel, when the complexity level of the current channel is the complexity level of the U channel of the current block, the complexity of the same channel of the previous image block is the complexity level of the U channel of the decoded image block before the current block. When it is determined to be the same by the most significant bit, the complexity information bits are 1 bit; when they are not the same, the complexity information bits are 3 bits, and the lower 2 bits indicate the change value between the complexity of the current block at the channel level and the complexity level of the same channel of the previous image block of the current block. Based on the change value and the complexity level of the same channel of the previous image block, the currently required complexity level of each channel can be determined. It should be understood that the above scenario is only an exemplary scenario and the protection scope of the present invention is not limited thereto. For example, the complexity information bits may indicate whether the complexity of the U channel of the current block is the same as that of the Y channel of the current block and the change value in the case of difference.
[0209] In step S702, determine the reversible coding average bit number of the current block at the same level and the reversible coding average bit number.
[0210] When the first chrominance and the second chrominance of the YUV image share a complexity level of one channel, determining the reversible coding average bit number and the reversible coding average bit number of the current block at the same level is realized as follows.
[0211] In S7021, determine the reversible coding bit number pred_lossless_bits.
[0212]
Number
[0213] However, cu_bits indicates the actual coding bit number of the current block and is determined based on the coding bit number of the coded / decoded image. width and height indicate the width and height of the coding block respectively. luma_qp indicates the quantization parameter of the luminance channel of the coded / decoded image. chroam_qp indicates the quantization parameter of the chrominance channel of the coded / decoded image. a and b are weight values, and their settings are related to the prediction mode, and the default values are both 8. In the case of the IBC mode, a, b >= 8. In the case of the point prediction mode, a, b >= 8. In the case of the palette mode, the original value mode, and the residual skip mode, since the quantization parameter becomes invalid, a, b == 0.
[0214] Also, when the buffer buffer is full and it is in the residual skip mode,
Number
[0215] In S7022, determine the reversible coding average bit number lossless_bits[blk_comp_level] at the same level.
[0216] The lossless_bits[blk_comp_level] of the same-level reversible coding average number of bits corresponds to the block-level complexity level of the current block. When the block-level complexity level of the current block is the same as that of the decoded image block, it is updated by the following formula.
[0217]
Number
[0218] In one embodiment, a specific setting method for the update rate d is that for the first four image blocks of any complexity, the update rates d are set to 3 / 4, 5 / 8, 1 / 2, and 3 / 8 respectively, and in other cases, they are all set to 1 / 4.
[0219] In S7023, determine the reversible coding average number of bits avg_lossless_bits.
[0220] Different from the same-level reversible coding average number of bits, the reversible coding average number of bits avg_lossless_bits is updated for each block. The specific update method is as follows.
[0221]
Number
[0222] In step S703, determine the target number of bits of the current block.
[0223] In one embodiment, determining the target number of bits of the current block is realized by the following steps.
[0224] In S7031, determine the initial target number of bits.
[0225] The initial target number of bits is the target number of bits calculated without considering the buffer fullness, and is calculated by the following method.
[0226] (1) Calculate the quality ratio quality_ratio.
[0227] [Number] However, comp_offset is a preset value. One setting method is to refer to Table 2 below.
[0228] [Table 2]
[0229] The calculation method of the bpp is as follows.
[0230] [Number] However, end_target_fullness is a preset value. In one embodiment, one specific setting value of end_target_fullness may be (delay_bits - 1533) * 3 / 4. However, delay_bits is a preset value and is related to the initial transmission delay mechanism. delay_bits is the number of delay bits.
[0231] The initial transmission delay mechanism includes the following features: a) When the video slice starts transmission, it is transmitted after delaying delay_blks image blocks, and these image blocks do not perform underflow processing. b) The buffer state of the end buffer of the video slice is fixed at the delay_bits number of delay bits (padded with zeros if insufficient). As shown in Figure 10, Figure 10 is a schematic diagram of the initial transmission delay mechanism. One or more image blocks in the video slice are between the initial position and the second position, and the maximum value of the corresponding buffer increases based on the position between the initial position and the second position. And the image blocks located between the threshold position and the final position of the slice have the maximum value of the corresponding buffer decreasing based on the position between the threshold position and the final position. Between the second position and the threshold position, the buffer size corresponding to the image blocks does not change.
[0232] The delay_bits is determined by the following calculation.
[0233]
Number
[0234] Also, when the image to be processed is a YUV444 or RGB image, the quality_ratio needs to be limited to a value range of 0 to 0.6.
[0235] In one embodiment, after the quality ratio is determined, the quality ratio can also be updated by calculating the average complexity level ave_comp_level of all previous image blocks. Specifically, it is realized as follows.
[0236]
Number
[0237] (2) Determine the initial target number of bits. The initial target number of bits pre_target_bits is determined by the following formula.
[0238]
Number
[0239] In S7032, according to the buffer state and the block-level complexity level of the current block, limit the initial target number of bits and determine the target number of bits of the final current block. Specifically, it is realized as follows.
[0240] (1) The buffer state can be represented by the buffer fullness. The buffer fullness fullness is determined by the following formula.
[0241]
Number
[0242] Also, in the above process, when determining the available_buffer_size, as shown in FIG. 10, considering the influence of the initial transmission delay function, the available_buffer_size varies according to the position of the current block in the video slice.
[0243] For the first delay_blks blocks of the video slice, that is, from the initial position to the second position, the available_buffer_size linearly increases from delay_bits to max_buffer_size. The increasing step size start_step is start_step = (max_buffer_size ‐ delay_bits) / delay_blks; where max_buffer_size represents the maximum available buffer size and is a preset fixed value.
[0244] The available_buffer_size remains unchanged from the second position to the threshold position and is always equal to max_buffer_size.
[0245] From the threshold position to the final position, the available_buffer_size linearly decreases from max_buffer_size to delay_bits. The decreasing step size end_step is end_step = ‐(max_buffer_size ‐ delay_bits) / (end_blks ‐ thres_blks); where end_blks represents the number of blocks at the final position and thres_blks represents the number of blocks at the threshold position.
[0246] Here, since the calculation of the delay_blks was described in detail in the part where the quality ratio is calculated, the description is omitted here.
[0247] (2) After determining the buffer fullness, further determine the upper and lower limits for restricting the target number of bits. Specifically, it is realized as follows. When calculating the lower limit min_bits,
Number
[0248] When calculating the upper limit max_bits,
Number
[0249] Considering the initial transmission delay mechanism, the process of determining the above upper and lower limits needs to also consider the influence of the initial transmission delay mechanism on the buffer fullness fullness.
[0250] (3) Based on the determined upper and lower limits, restrict the initial target number of bits and obtain the target number of bits for the current block. Specifically,
Number
[0251] Also, when the current block is the first row block of the image to be processed, it is difficult to predict the parameters of the first row block, and the prediction error is transitional. Therefore, when the current block is the first row block, the quality of the current block can be improved by introducing the first row quality improvement parameter. This process is mainly realized by the quantization parameter of the first row block becoming smaller.
[0252] Specifically, in the process of determining the above target number of bits, it is realized as follows.
[0253] If the current block is the first-line block, increase bpp by 2.
[0254] For all first-line blocks in the image to be processed, the adjustment of the bpp parameter is realized by setting the increase amount of the bpp of the image block in the first line to bpp_delta_row, and from the first block to the last block in the first line, bpp_delta_row gradually decreases from 2.5 to 0.5.
[0255] If the current block is the first-line block, in the process of determining the target number of bits by restricting the initial target number of bits according to the buffer state and the block-level complexity level, improve the image quality of the current block by the following method. After restricting the target number of bits according to the buffer state and complexity, if the current block is the first-line block of the slice and target_bits < 7, where 7 is a preset empirical threshold, target_bits increases, and the increased target_bits must be within a predetermined range. Specifically,
Number
[0256] If the current block is the first-line block, further improve the quality of the current block by determining the upper limit according to the following formula.
[0257]
Number
[0258] When the current block is the leading row block, improving the quality of the current block by the leading row quality improvement parameter is understood to be executed only when certain conditions are met. For example, the leading row quality improvement is performed only when the complexity level of the current block is low.
[0259] Also, when the current block is the leading column block of the image to be processed, the quality of the current block can be improved by introducing the leading column quality improvement parameter. This process is mainly realized by the quantization parameter of the leading column block becoming smaller.
[0260] When the current block is the leading column block, increase bpp by 2.
[0261] For all leading column blocks in the image to be processed, the adjustment of the bpp parameter is realized by setting the increase amount of the bpp of the image block in the leading column to bpp_delta_col, and from the first block to the last block in the leading row, bpp_delta_col gradually decreases from 2.5 to 0.5.
[0262] When the current block is the leading column block, in the process of determining the target bit number by restricting the initial target bit number according to the buffer state and the block-level complexity level, the image quality of the current block can be improved by the following method. After restricting the target bit number according to the buffer state and complexity, when the current block is the leading column block of the slice and target_bits < 7, where 7 is a preset empirical threshold, target_bits increases, and the increased target_bits must be within a predetermined range. Specifically,
Number
[0263] If the current block is the leading column block, the quality of the current block can be further improved by determining the upper limit according to the following formula.
[0264]
Number
[0265] It should be understood that when the current block is the leading column block, improving the quality of the current block by the leading column quality improvement parameter is only executed when certain conditions are met. For example, the leading row quality improvement is only performed when the complexity level of the current block is high.
[0266] In step S704, determine the quantization parameter of the current block. In this step, determining the quantization parameter of the current block is realized as follows.
[0267] (1) When calculating the reference quantization parameter ref_qp,
Number
[0268] (2) Calculate the quantization parameter of each component. Taking the case where the image to be processed is a YUV image as an example, this process is to calculate the quantization parameter of each channel of YUV. Specifically, it is realized as follows.
[0269] When calculating the offset amount, bias = bias_init * format_bias; However, bias_init and format_bias are preset values. Here, bias_init refers to Table 3 below, and format_bias refers to Table 1 above. When calculating the quantization parameter of the luminance channel, luma_qp = Clip3(0, luma_max_qp, ref_qp ‐ sample_rate * bias) When calculating the chrominance channel quantization parameter, chroma_qp = Clip3(0, chroma_max_qp, ref_qp + bias) However, refer to Table 3 for bias_init and Table 2 for format_bias.
[0270]
Table 3
[0271] In Table 3 described above, comp_level[0] represents the luminance component and comp_level[1] represents the chrominance component.
[0272] Also, for a YUV420 format image, two YUV420s may be combined into one YUV444 image for processing. In this case, the complexity level of the luminance component is determined by the weighted values of both.
[0273]
Equation
[0274] The complexity level of the luminance component may take the maximum or minimum value of the two. The present invention is not particularly limited thereto.
[0275] In step S705, video encoding / decoding is performed on the current block based on the quantization parameter.
[0276] In one embodiment, the video encoding / decoding device encodes / decodes a current block based on the quantization parameter of the current block. In the encoding process, it is understood that the video encoder incorporates the channel-level complexity level of the current block into the code stream, or incorporates the quantization parameter of the current block into the code stream. Accordingly, on the decoding side, the channel-level complexity level in the code stream is obtained to calculate the quantization parameter for decoding. Alternatively, the decoding side obtains the quantization parameter in the code stream and performs decoding. Of course, the video encoder may incorporate both types of information into the code stream.
[0277] Also, in the process of determining the quantization parameter, each channel component of the current block may individually determine parameters such as the same-level reversible encoding average bit number, reversible encoding average bit number, and target bit number of the current block. The process is realized as follows.
[0278] Determining the same-level reversible encoding average bit number and reversible encoding average bit number of each channel component of the current block is specifically as follows.
[0279] (1) Determine the reversible encoding bit number pred_lossless_bits[i] of each channel component of the current block.
[0280]
Number
[0281] (2) Determine the same-level reversible coding average number of bits lossless_bits[i][comp_level[i]] for each channel component.
Number
[0282] (3) Determine the reversible coding average number of bits avg_lossless_bits.
Number
[0283] Determining the target number of bits for each channel component of the current block is realized by the following method.
[0284] (1) Determine the quality ratio quality_ratio.
[0285] Embodiment 1: Determine the quality ratio quality_ratio[i] for each channel component of the current block.
Number
[0286] Embodiment 2: When calculating quality_ratio, still adopt the cu (Coding Unit) level, and the calculation process is the same as that of step S7031, but it is necessary to merge the ave_lossless_bits[i] into the variable at the cu level.
[0287] (2) Determine the target number of bits target_bits.
[0288] Embodiment 1:
Number
[0289] Embodiment 2: Corresponding to Embodiment 2 in step (1). In this case, the calculation method of the target bit number target_bits is the same as that in step S703 above, and the obtained target_bits is already a variable at the cu level. Similarly, for the subsequent step (3), it is necessary to merge the lossless_bits[i] at the cb level to the cu level.
[0290] (3) Limit the target bit number according to the buffer state and complexity. Since this step is the same as step S703 above, the description is omitted here. Determining the quantization parameter of each channel component of the current block is realized as follows.
[0291] Embodiment 1: Determine the quantization parameter of each channel component. [Number]
[0292] When step (3) becomes effective, that is, when the upper limit and the lower limit play a restrictive role and the value of target_bits is changed to the upper limit or the lower limit, based on the ratio of target_bits[i] obtained in steps (1) to (2), redistribute target_bits to obtain new target_bits[i]. When target_bits is not changed in step (3), the value of target_bits[i] does not change.
[0293] Embodiment 2: When target_bits was not separated before, separate target_bits according to the complexity level at this time.
[0294] Embodiment 3: All variables before this step are at the cu level. At this time, the reference quantization parameter ref_qp is separated according to the complexity, and the obtained qp after separation is used as the final luminance and chrominance qp.
[0295] TIFF2025525001000046.tif24148
[0296] TIFF2025525001000047.tif27162
[0297] In addition, it is necessary to merge the cb-level variables into the cu-level variables. Assuming the cu variable to be merged is temp[i], the merging process is specifically as follows.
Number
[0298] In another specific embodiment of the present invention, furthermore, the rate control parameters of the video image decoding method and the video image encoding method can also be made into fixed-point decimals. Specifically, it is realized by the following process.
[0299] S1: Encoding control initialization
[0300] TIFF2025525001000049.tif123163
[0301] All of the above parameters are intermediate parameters of the encoding control initialization process and are used to convert the rate control parameters into fixed-point numbers. Here, WarmUp[i] (0 <= i <= 4) represents the update rate parameters for the first few blocks and is used to update the same-level reversible encoding bit count AdjComplexity. ComplexityShift represents the number of shift bits due to fixed-point operations related to complexity. InfoRatioShift represents the number of shift bits due to fixed-point operations related to the quality ratio. BppShift represents the number of shift bits due to fixed-point operations related to bpp. FullnessShift represents the number of shift bits due to fixed-point operations related to fullness. AvgComplexityShift represents the number of shift bits due to fixed-point operations related to the average number of reversible encoding bits. ChromaSampleRateShift represents the number of shift bits due to fixed-point operations related to the sampling rate. K1Shift represents the number of shift bits due to fixed-point operations related to k1 (see Table 4). K2Shift represents the number of shift bits due to fixed-point operations related to k2. K3Shift represents the number of shift bits due to fixed-point operations related to k3. K4Shift represents the number of shift bits due to fixed-point operations related to k4. BiasShift represents the number of shift bits due to fixed-point operations related to Bias. K2, K3, and K4 are fixed empirical values used in the encoding control algorithm. DelayBits represents the number of delay bits. TargetBpp represents the target bpp and is externally configured. TransmissionDelayCu represents the number of CUs for the initial transmission delay. EndDecreaseBits represents the number of bits that need to be decreased at the end of the slice due to the initial delay function. RcBufferSize represents the buffer size considered by the encoding control. MuxWordSize represents the number of bits occupied by the header information required for the sub-stream parallel function. EndControlBlocks represents the number of blocks that need to operate on the buffer maximum value at the end of the slice due to the initial transmission delay.DecreaseStepLog2 represents the logarithmic value of the step size for defining the decrease of MaxBufferSize used by the coding rate control module in each control block at the end of the slice, and this value is in the header of the code stream. EndControlBegin represents the index of the block where control starts from the end of the slice. SliceWidthInCu represents how many widths of CUs are in the slice width. SliceHeightInCu represents how many heights of CUs are in the slice height. EndTargetFullness represents the target fullness at the end of the slice. RemainBlksLog2 represents the binary highest bit number of the total quantity of coded units in one slice. MaxBufferSize represents the maximum value of the buffer.
[0302] According to BitDepth[0] (representing the bpc of the Y channel), the ImageFormat searches from Table 1 to obtain the initialization values of AdjComplexity (the number of bits for the same-level reversible coding), AvgComplexity (the average number of bits for reversible coding), ComplexityOffset (the bias value for complexity calculation), MaxComp (the maximum number of bits for reversible coding), and K1 (an empirical value).
[0303]
Table 4
[0304] The above Table 4 shows the correspondence between AdjComplexity, AvgComplexity, ComplexityOffset, MaxComp, K1 and BitDepth[0], ImageFormat.
[0305] According to ImageFormat (image format), the initialization values of ChromaSampleRate (sampling rate), InvElem (multiplier required to remove division related to the sampling rate), InvElemShift (shift value required to remove division related to the sampling rate), and FormatBias (bias value for different image formats with respect to qp) can be obtained from Table 5 below.
[0306]
Table 5
[0307] The above Table 5 shows the correspondence between ChromaSampleRate, InvElem, InvElemShift, FormatBias, and ImageFormat.
[0308] S2: Determine the quantization parameter. This step is realized as follows.
[0309] S21: Calculate the quantization parameter MasterQp of the coding unit according to the luminance complexity level ComplexityLevel[0] and the chrominance complexity level ComplexityLevel[1] of the current coding unit.
[0310] S22: Calculate the quantization parameters Qp[0] and Qp[1] of the luminance coding block and the chrominance coding block of the current coding unit according to MasterQp.
[0311] Here, calculating the quantization parameter MasterQp of the coding unit according to the luminance complexity level ComplexityLevel[0] and the chrominance complexity level ComplexityLevel[1] of the current coding unit is realized as follows. TIFF2025525001000052.tif236131
[0312] However, bppAdj represents the adjustment value of bpp. BitsRecord represents the total number of bits that have been currently encoded / decoded. CurrBlocks represents the number of blocks that have been currently encoded / decoded. maxComp represents the conversion value required for MaxComp to make the encoding control a fixed decimal point number. complexityOffset represents the conversion value required for ComplexityOffset to make the encoding control a fixed decimal point number. RcBufferSizeMaxBit represents the buffer size of the code stream. The value of RcBufferSize is equal to the value of rc_buffer_size, and the value of RcBufferSizeMaxBit represents the binary highest bit number of RcBufferSize. shiftCur represents the current shift value. tmp represents the intermediate variable generated in the process of making the encoding control a fixed decimal point number. fullness represents the fullness. infoRatio represents the quality ratio. relativeComplexity represents the relative reversible encoding bit number. minRate1, minRate2, and minRate13 represent the intermediate variables for calculating minRate. minRate represents the lower limit of targetRate. targetRate represents the target bit number. bppOffset1, bppOffset2, and bppOffset3 represent the intermediate variables for calculating bppOffset. bppOffset represents the bias value of bpp. maxRate represents the upper limit of targetRate. InverseTable is a preset table, and InverseTable = { 1024, 512, 341, 256, 205, 171, 146, 128, 114, 102, 93, 85, 79, 73, 68, 64, 60, 57, 54, 51, 49, 47, 45, 43, 41, 39, 38, 37, 35, 34, 33, 32} is defined.
[0313] Calculating the quantization parameters Qp[0] and Qp[1] of the luminance encoding block and the chrominance encoding block of the current encoding unit according to the MasterQp is realized as follows.
[0314] Obtain BiasInit from Table 6 below according to the luminance complexity level ComplexityLevel[0] and chrominance complexity level ComplexityLevel[1] of the current encoding unit.
[0315] Table 6: Definition of BiasInit
[0316] [Table 6]
[0317] TIFF2025525001000054.tif48162
[0318] S3: Update the rate control parameter
[0319] TIFF2025525001000055.tif70164
[0320] TIFF2025525001000056.tif49163
[0321] TIFF2025525001000057.tif39156
[0322] In addition, any cases not specifically described in the above technical solution can be performed on the decoding side or the encoding side.
[0323] In addition, if there is no contradiction, some or all of the above-mentioned multiple embodiments can be combined to form a new embodiment.
[0324] The embodiments of the present invention provide a video encoding / decoding device, which may be a video encoding / decoding device or a video encoder or a video decoder. Specifically, the video encoding / decoding device is for executing the steps executed by the video encoding / decoding device in the above-mentioned video image decoding method and video image encoding method. The video encoding / decoding device provided by the embodiments of the present invention may include a module corresponding to the corresponding step.
[0325] Embodiments of the present invention can perform classification of functional modules on a video encoding / decoding device according to the examples of the above-described method. For example, each functional module may be separated according to each function, or two or more functions may be integrated into one processing module. The integrated module may be implemented in the form of hardware or in the form of a software functional module. The classification of the modules in the embodiments of the present invention is general and is only a classification of logical functions, and there may be other classification methods when actually implemented.
[0326] When separating each functional module corresponding to each function, FIG. 11 shows one possible configuration schematic diagram of a video encoding / decoding device according to the above-described embodiment. As shown in FIG. 11, the video encoding / decoding device 1100 includes a complexity level determination module 1101, a rate control parameter determination module 1102, a quantization parameter determination module 1103, and an encoding / decoding module 1104.
[0327] The complexity level determination module 1101 is for obtaining at least two channel-level complexity levels of the current block in the image to be processed and determining the block-level complexity level of the current block according to the at least two channel-level complexity levels, and the channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block.
[0328] The rate control parameter determination module 1102 is for determining the target number of bits of the current block according to the rate control parameter, and the rate control parameter includes the block-level complexity level of the current block.
[0329] The quantization parameter determination module 1103 is for determining the quantization parameter of the current block based on the target number of bits.
[0330] The symbolization / decoding module 1104 is for symbolizing / decoding the current block based on quantization parameters.
[0331] In one example, the rate control parameter includes the same-level reversible coding average number of bits, the reversible coding average number of bits, and the code stream buffer fullness, and the rate control parameter determination module is specifically used for determining the same-level reversible coding average number of bits and the reversible coding average number of bits, determining an initial target number of bits based on the same-level reversible coding average number of bits and the reversible coding average number of bits, and determining the target number of bits of the current block based on the code stream buffer fullness, the block-level complexity level of the current block, and the initial target number of bits. Here, the same-level reversible coding average number of bits is the average value of the predicted values of the number of bits required for reversibly coding the current block and a plurality of decoded image blocks, and the complexity levels of the plurality of decoded image blocks and the current block are the same. The reversible coding average number of bits is the average value of the predicted values of the number of bits required for reversibly coding the current block and all decoded image blocks. The code stream buffer fullness is for indicating the fullness of the buffer, and the buffer is for storing the code stream of the image to be processed.
[0332] In one example, the rate control parameter determination module specifically determines the same-level reversible coding average bit number and the reversible coding average bit number, and determines the reversible coding bit number of the current block, where the reversible coding bit number is a predicted value of the number of bits required to reversibly code the current block. Based on the reversible coding bit number of the current block and multiple past same-level reversible coding average bit numbers, it updates the same-level reversible coding average bit number of the current block, where the past same-level reversible coding average bit number is the same-level reversible coding average bit number of the decoded image block having the same complexity level as the block complexity level of the current block. Based on the reversible coding bit number of the current block and all past reversible coding average bit numbers, it updates the reversible coding average bit number of the current block, where the past reversible coding average bit number is the reversible coding average bit number of the decoded image block. It is used for the above.
[0333] In one possible implementation, the current block is the first row block of the image to be processed, the rate control parameter includes the first row quality improvement parameter, and the fact that the rate control parameter determination module is specifically used to determine the target bit number of the current block based on the rate control parameter further includes adjusting the target bit number of the current block based on the first row quality improvement parameter so that the quantization parameter of the current block becomes smaller.
[0334] In one possible implementation, the current block is the first column block of the image to be processed, the rate control parameter includes the first column quality improvement parameter, and the fact that the rate control parameter determination module is specifically used to determine the target bit number of the current block based on the rate control parameter further includes adjusting the target bit number of the current block based on the first row quality improvement parameter so that the quantization parameter of the current block becomes smaller.
[0335] In one example, the complexity level determination module is specifically used to obtain at least two channel-level complexity levels of the current block in the image to be processed. On the encoding side, the channel-level texture information of the current block is obtained, and the channel-level complexity level of the current block is determined based on the channel-level texture information. Alternatively, on the decoding side, the channel-level complexity level is obtained from the code stream, and the code stream is the encoded code stream of the current block.
[0336] In one example, the complexity level determination module is specifically used to obtain the channel-level complexity level from the code stream. Obtaining the complexity information bits of the current block from the code stream, where the complexity information bits are for indicating the channel-level complexity level of the current block, and determining the channel-level complexity level based on the complexity information bits.
[0337] In one example, the complexity level determination module is specifically used to obtain the channel-level texture information of the current block and determine the channel-level complexity level of the current block based on the channel-level texture information. Using at least one channel's image block of the current block as a processing unit, dividing the processing unit into at least two sub-units, determining the texture information of each sub-unit, and in the processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit.
[0338] In one example, the complexity level determination module is specifically used to determine the texture information of each sub-unit. It includes obtaining the original pixel value of the sub-unit, the original pixel value or reconstructed value of the adjacent column on the left side of the sub-unit, and the reconstructed value of the adjacent row above the sub-unit, and calculating the horizontal texture information and vertical texture information of the sub-unit accordingly, and selecting the minimum value from the horizontal texture information and vertical texture information as the texture information of the corresponding sub-unit.
[0339] In one example, the complexity level determination module is specifically used to determine the block-level complexity level of the current block based on the texture information of each sub-unit in the processing unit. In the processing unit, it includes dividing the texture information of each sub-unit into corresponding sub-unit complexity levels based on a plurality of thresholds, where the plurality of thresholds are preset, and determining the block complexity level of the current block based on the complexity level of each sub-unit.
[0340] In one example, the complexity level determination module is specifically used to determine the block complexity level of the current block based on the complexity level of each sub-unit. It includes mapping the complexity level of each sub-unit to the corresponding channel-level complexity level based on a preset rule, and determining the block-level complexity level of the current block based on the complexity level of each channel.
[0341] In one example, the complexity level determination module is specifically used to map the complexity level of each sub-unit to the corresponding channel-level complexity level based on a preset rule. It determines the channel-level complexity level based on a plurality of thresholds and the sum of the complexity levels of each sub-unit, where the plurality of thresholds are preset.
[0342] In one example, the complexity level determination module is specifically used to map the complexity level of each sub-unit to the complexity level of the corresponding channel level based on preset rules. Determine the level composition of the sub-unit complexity level, and determine the corresponding channel-level complexity level based on the level composition.
[0343] In one example, the complexity level determination module is specifically used to determine the block-level complexity level of the current block based on the complexity level of each channel level. Obtaining the maximum value, minimum value or weighted value of the complexity level of each channel level as the block-level complexity level of the current block, or determining the block-level complexity level of the current block based on a plurality of thresholds and the sum of the complexity levels of each channel level, the plurality of thresholds are preset. All relevant contents of each step according to the embodiments of the method can be incorporated into the description of the functions of the corresponding functional modules, so the description is omitted here.
[0344] Of course, the video encoding / decoding device provided by the embodiments of the present invention includes the above-mentioned modules, but is not limited thereto. For example, the video encoding / decoding device can also include a storage module.
[0345] The storage module can be used to store the program code and data of the video encoding / decoding device.
[0346] The embodiments of the present invention further provide an electronic device. The electronic device includes the video encoding / decoding device 1100, and the video encoding / decoding device 1100 executes the method executed by any one of the video decoders according to the above description.
[0347] Embodiments of the present invention further provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program is executed on a computer, the computer is caused to execute the method executed by any one of the video decoders described above.
[0348] For the interpretation and beneficial effects description of the related content in any one of the above computer-readable storage media, reference may be made to the corresponding embodiments described above, and thus the description is omitted here.
[0349] Embodiments of the present invention further provide a chip. A control circuit for realizing the functions of the video encoding / decoding device 100 and one or more ports are integrated in the chip. Optionally, the functions supported by the chip can be referred to the above description, and thus the description is omitted here. Those skilled in the art will understand that all or some of the steps for realizing the above embodiments can be completed by instructing the related hardware by a program. The program can be stored in a computer-readable storage medium. The above storage medium may be a read-only memory, a random access memory, etc. The above processing unit or processor may be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0350] Embodiments of the present invention further provide a computer program product containing instructions. When executed by a computer, these instructions cause the computer to perform any one of the methods in the above embodiments. The computer program product contains one or more computer instructions. When the computer loads and executes the computer program instructions, the flow or functions of the embodiments of the present invention are wholly or partially generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, radio, microwave, etc.). The computer-readable storage medium may be any available medium accessible by the computer or a data storage device such as a server or data center integrated with one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD), etc.
[0351] Note that the devices for storing the above-mentioned computer instructions or computer programs provided by the embodiments of the present invention include, for example, but are not limited to, the memory, computer-readable storage medium, and communication chip, etc., all of which have non-transitory properties.
[0352] The above-described embodiments can be implemented in whole or in part by software, hardware, firmware, or a combination thereof. When implemented by a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the flow or function of the embodiments of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, radio, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center integrated with one or more available media. The available media may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0353] In this specification, the present invention has been described with reference to each embodiment. However, when implementing the process of the present invention protected by the present invention, those skilled in the art can understand and realize other changes to the disclosed embodiments by referring to the drawings, the disclosed content, and the appended claims. In the claims, the term "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. A single processor or other unit can implement several functions listed in the claims. Although several measures are described in different dependent claims, these measures may not necessarily be unable to bring good effects when combined.
[0354] Although the present invention has been described with reference to specific features and embodiments, it is obvious that various modifications and combinations are possible without departing from the spirit and scope of the present invention. Accordingly, this specification and the drawings are merely exemplary descriptions of the present invention defined by the appended claims, and any and all modifications, changes, combinations, or equivalents within the scope of this application are considered to be covered. Obviously, those skilled in the art can make various changes and deformations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and deformations of the present invention fall within the scope of the claims of the present invention and the scope of its equivalent technologies, the present invention is intended to include these changes and deformations.
Claims
1. A video image decoding method, comprising: obtaining a channel-level complexity level from a code stream, and determining a block-level complexity level of a current block according to at least two channel-level complexity levels, wherein the code stream is an encoded code stream of the current block, and the channel-level complexity level is for indicating a degree of complexity of a channel-level texture of the current block; determining a target bit number of the current block based on a rate control parameter, wherein the rate control parameter includes the block-level complexity level of the current block; determining a quantization parameter of the current block based on the target bit number; decoding the current block based on the quantization parameter. A video image decoding method comprising the above steps.
2. The rate control parameter includes a same-level reversible coding average bit number, a reversible coding average bit number, and a code stream buffer fullness degree, and determining the target bit number of the current block based on the rate control parameter includes: determining the same-level reversible coding average bit number and the reversible coding average bit number; determining an initial target bit number based on the same-level reversible coding average bit number and the reversible coding average bit number; determining the target bit number of the current block based on the code stream buffer fullness degree, the block-level complexity level of the current block, and the initial target bit number, wherein the same-level reversible coding average bit number is an average value of predicted values of bit numbers required for reversibly coding the current block and a plurality of decoded image blocks, and complexity levels of the plurality of decoded image blocks are the same as a complexity level of the current block, the reversible coding average bit number is an average value of predicted values of bit numbers required for reversibly coding the current block and all decoded image blocks, the code stream buffer fullness degree is for indicating a fullness degree of a buffer, and the buffer is for storing a code stream of an image to be processed. The video image decoding method according to claim 1.
3. Determining the same-level reversible coding average bit number and the reversible coding average bit number includes: Determining the number of reversible coding bits of the current block, where the number of reversible coding bits is a predicted value of the number of bits required to reversibly code the current block, Updating the average number of reversible coding bits at the same level of the current block based on the number of reversible coding bits of the current block and a plurality of past average reversible coding bits at the same level, where the past average reversible coding bits at the same level are the average number of reversible coding bits at the same level of decoded image blocks having the same complexity level as the complexity level of the current block at the block level, Updating the average number of reversible coding bits of the current block based on the number of reversible coding bits of the current block and all past average reversible coding bits, where the past average reversible coding bits are the average number of reversible coding bits of decoded image blocks, The video image decoding method according to claim 2.
4. The current block is the first row block of the image to be processed, and the rate control parameter includes a first row quality improvement parameter, When determining the target number of bits of the current block based on the rate control parameter, the video image decoding method further includes Adjusting the target number of bits of the current block based on the first row quality improvement parameter so that the quantization parameter of the current block becomes smaller. The video image decoding method according to any one of claims 1 to 3.
5. The current block is the first column block of the image to be processed, and the rate control parameter includes a first column quality improvement parameter, When determining the target number of bits of the current block based on the rate control parameter, the video image decoding method further includes Adjusting the target number of bits of the current block based on the first column quality improvement parameter so that the quantization parameter of the current block becomes smaller. The video image decoding method according to any one of claims 1 to 3.
6. Obtaining the channel-level complexity level from the code stream is Obtaining the complexity information bits of the current block from the code stream, where the complexity information bits are for indicating the channel-level complexity level of the current block, determining the channel-level complexity level based on the complexity information bits; The video image decoding method according to any one of claims 1 to 5.
7. The most significant bit in the complexity information bits indicates whether the current channel-level complexity level is the same as the complexity level of the same channel of the previous image block of the current block, and the change value between the two. When they are the same, the complexity information bit is 1 bit. When they are not the same, the complexity information bit is 3 bits. The video image decoding method according to claim 6.
8. A video image encoding method, comprising: obtaining the channel-level texture information of the current block, determining the channel-level complexity level of the current block based on the channel-level texture information, and determining the block-level complexity level of the current block according to at least two channel-level complexity levels, wherein the channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block; determining the target number of bits of the current block based on a rate control parameter, wherein the rate control parameter includes the block-level complexity level of the current block; determining the quantization parameter of the current block based on the target number of bits; encoding the current block based on the quantization parameter; A video image encoding method comprising the above steps.
9. Obtaining the channel-level texture information of the current block and determining the channel-level complexity level of the current block based on the channel-level texture information includes: using at least one channel's image block of the current block as a processing unit, dividing the processing unit into at least two sub-units, and determining the texture information of each sub-unit; In the processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit. The video image encoding method according to claim 8.
10. Determining the texture information of each sub-unit includes: Obtain the original pixel value of the sub-unit, the original pixel value or the reconstructed value of the adjacent column on the left side of the sub-unit, and the reconstructed value of the adjacent row above the sub-unit, and calculate the horizontal texture information and the vertical texture information of the corresponding sub-unit; Select the minimum value from the horizontal texture information and the vertical texture information as the texture information of the corresponding sub-unit; The video image encoding method according to claim 9.
11. In the processing unit, determining the block-level complexity level of the current block based on the texture information of each sub-unit includes: In the processing unit, dividing the texture information of each sub-unit into the complexity levels of the corresponding sub-units based on a plurality of thresholds, where the plurality of thresholds are preset; Determining the block-level complexity level of the current block based on the complexity levels of each sub-unit; The video image encoding method according to claim 9.
12. Determining the block-level complexity level of the current block based on the complexity levels of each sub-unit includes: Mapping the complexity level of each sub-unit to the corresponding channel-level complexity level based on a preset rule; Determining the block-level complexity level of the current block based on the complexity levels of each channel; The video image encoding method according to claim 11.
13. Mapping the complexity level of each sub-unit to the corresponding channel-level complexity level based on the preset rule includes: Determining the channel-level complexity level based on a plurality of thresholds and the sum of the complexity levels of each sub-unit; The plurality of thresholds are preset. The video image encoding method according to claim 12.
14. Mapping the complexity level of each sub-unit to the corresponding channel-level complexity level based on the preset rule includes: Determining the level configuration of the complexity level of the sub-unit and determining the corresponding channel-level complexity level based on the level configuration; The video image encoding method according to claim 12.
15. The rate control parameter includes the same-level reversible coding average bit number, the reversible coding average bit number, and the code stream buffer fullness degree. Determining the target bit number of the current block based on the rate control parameter includes: Determining the same-level reversible coding average bit number and the reversible coding average bit number; Determining an initial target bit number based on the same-level reversible coding average bit number and the reversible coding average bit number; Determining the target bit number of the current block based on the code stream buffer fullness degree, the complexity level of the block level of the current block, and the initial target bit number. Here, the same-level reversible coding average bit number is the average value of the predicted values of the bit numbers required to reversibly code the current block and a plurality of decoded image blocks, and the complexity levels of the plurality of decoded image blocks and the current block are the same. The reversible coding average bit number is the average value of the predicted values of the bit numbers required to reversibly code the current block and all decoded image blocks. The code stream buffer fullness degree is for indicating the fullness degree of the buffer, and the buffer is for storing the code stream of the image to be processed. The video image coding method according to claim 8.
16. Determining the same-level reversible coding average bit number and the reversible coding average bit number includes: Determining the reversible coding bit number of the current block, where the reversible coding bit number is the predicted value of the bit number required to reversibly code the current block; Updating the same-level reversible coding average bit number of the current block based on the reversible coding bit number of the current block and a plurality of past same-level reversible coding average bit numbers, where the past same-level reversible coding average bit number is the same-level reversible coding average bit number of the decoded image blocks having the same complexity level as the complexity level of the block level of the current block; Updating the reversible coding average bit number of the current block based on the reversible coding bit number of the current block and all past reversible coding average bit numbers, where the past reversible coding average bit number is the reversible coding average bit number of the decoded image blocks. The video image encoding method according to claim 15.
17. wherein the current block is a first row block of the image to be processed, and the rate control parameter includes a first row quality improvement parameter, when determining the target number of bits of the current block based on the rate control parameter, the video image encoding method further adjusts the target number of bits of the current block based on the first row quality improvement parameter so that the quantization parameter of the current block becomes smaller. The video image encoding method according to claim 8.
18. wherein the current block is a first column block of the image to be processed, and the rate control parameter includes a first column quality improvement parameter, when determining the target number of bits of the current block based on the rate control parameter, the video image encoding method further adjusts the target number of bits of the current block based on the first column quality improvement parameter so that the quantization parameter of the current block becomes smaller. The video image encoding method according to claim 8.
19. A video image decoding apparatus, a complexity level determination module for obtaining a channel-level complexity level from a code stream and determining a block-level complexity level of a current block according to at least two channel-level complexity levels, wherein the code stream is an encoded code stream of the current block, and the channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block; a rate control parameter determination module for determining a target number of bits of the current block based on a rate control parameter, wherein the rate control parameter includes the block-level complexity level of the current block; a quantization parameter determination module for determining a quantization parameter of the current block based on the target number of bits; and an encoding / decoding module for decoding the current block based on the quantization parameter. A video image decoding apparatus.
20. A video image encoding apparatus, A complexity level determination module for obtaining channel-level texture information of a current block, determining a channel-level complexity level of the current block based on the channel-level texture information, and determining a block-level complexity level of the current block according to at least two channel-level complexity levels, wherein the channel-level complexity level is for indicating the degree of complexity of the channel-level texture of the current block. A rate control parameter determination module for determining a target bit number of the current block based on a rate control parameter, wherein the rate control parameter includes the block-level complexity level of the current block. A quantization parameter determination module for determining a quantization parameter of the current block based on the target bit number. An encoding / decoding module for encoding the current block based on the quantization parameter. A video image encoding device including the above.
21. A video decoder for executing the video image decoding method according to any one of Claims 1 to 7.
22. A video encoder for executing the video image encoding method according to any one of Claims 8 to 18.
23. A video encoding and decoding system, comprising a video encoder and / or a video decoder, wherein the video encoder is for executing the video image encoding method according to any one of Claims 8 to 18, and the video decoder is for executing the video image decoding method according to any one of Claims 1 to 7. A video encoding and decoding system.
24. A computer-readable storage medium, wherein a program is stored in the computer-readable storage medium, and when the program is executed on the computer, the computer is caused to execute the method according to any one of Claims 1 to 18. A computer-readable storage medium.
Citation Information
Patent Citations
Image coder, image coding method, image transmitter, image transmission method and recording medium
JP1998224786A
Moving image encoding device, moving image encoding method, and moving image encoding program
JP2011259408A
Image encoding device, image encoding method, and program
JP2015115735A