Video image reconstruction method, device, computer equipment and storage medium
By combining residual signals and prediction signals in video encoding for image reconstruction, the problem of information loss under string prediction mode is solved, and the encoding and codec performance is improved.
Patent Information
- Application Number
- CN202010844965.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-08-20
AI Technical Summary
In video encoding, under the string prediction mode, the prediction signal is directly composed of reference pixels, resulting in information loss and affecting the encoding and decoding performance.
By combining the residual signal and prediction signal of the coded encoding unit for image reconstruction, the encoding and codec performance under the string prediction mode is improved.
The signal quality of the encoding unit reconstruction under the string prediction method is improved and the encoding and codec performance is improved.
Smart Images

Figure CN114079782B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of video coding and decoding technology, and in particular to an image reconstruction method, apparatus, computer equipment and storage medium. Background Art
[0002] In current video compression technologies, such as VVC (Versatile Video Coding) and AVS3 (Audio Video Coding Standard 3), a string prediction encoding and decoding method is introduced.
[0003] In the related technology, for the string predicted coding unit, its prediction signal is directly composed of the searched reference pixels, which is equivalent to using the prediction signal directly as the reconstructed signal. In the process of encoding a coding unit, there is usually a certain amount of information loss, resulting in a certain difference between the prediction signal and the original signal, thereby affecting the performance of the codec. Summary of the invention
[0004] The embodiments of the present application provide a video image reconstruction method, apparatus, computer equipment and storage medium, which can combine the residual signal and prediction signal of the encoded coding unit to reconstruct the image when encoding and decoding by string prediction, thereby improving the encoding and decoding performance under the string prediction method. The technical solution is as follows:
[0005] According to one aspect of an embodiment of the present application, a video image reconstruction method is provided, the method comprising:
[0006] In response to encoding and decoding in a string prediction manner, a residual code of an encoded target coding unit is obtained; the residual code is obtained by encoding a residual coefficient of the target coding block;
[0007] Decoding the residual coding of the target coding unit to obtain a residual coefficient of the target coding unit;
[0008] Acquire a residual signal of the target coding unit based on the residual coefficient of the target coding unit;
[0009] A reconstructed signal of the target coding unit is obtained based on the residual signal of the target coding unit and the prediction signal of the target coding unit.
[0010] According to one aspect of an embodiment of the present application, a video image encoding method is provided, the method comprising:
[0011] Obtaining an original signal of an unencoded target coding unit;
[0012] Based on the reference signal, the original signal of the target coding unit is predicted by a string prediction method to obtain a prediction signal of the target coding unit and a residual signal of the target coding unit;
[0013] Acquire residual coding of the target coding unit based on the residual signal of the target coding unit;
[0014] The motion information corresponding to the prediction signal of the target coding unit and the residual coding of the target coding unit are added to the encoded video code stream.
[0015] According to one aspect of an embodiment of the present application, a video image reconstruction device is provided, the device comprising:
[0016] A residual coding acquisition module, configured to acquire residual coding of an encoded target coding unit in response to encoding and decoding in a string prediction manner; the residual coding is obtained by encoding residual coefficients of the target coding block;
[0017] A coefficient decoding module, used for decoding the residual coding of the target coding unit to obtain the residual coefficient of the target coding unit;
[0018] A residual signal acquisition module, configured to acquire a residual signal of the target coding unit based on a residual coefficient of the target coding unit;
[0019] A signal reconstruction module is used to obtain a reconstructed signal of the target coding unit based on a residual signal of the target coding unit and a prediction signal of the target coding unit.
[0020] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the above-mentioned video image reconstruction method.
[0021] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned video image reconstruction method.
[0022] In another aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned video image reconstruction method.
[0023] The technical solution provided by the embodiments of the present application may have the following beneficial effects:
[0024] When the encoder / decoder performs encoding and decoding through the string prediction method, the residual information is introduced into the string prediction process, and the string prediction is completely processed with residual signals, which improves the reconstructed signal quality of the coding unit based on the string prediction method, thereby improving the encoding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a basic flow chart of a video encoding process exemplarily shown in this application;
[0026] Figure 2 is a schematic diagram of an inter-frame prediction mode provided by an embodiment of the present application;
[0027] Figure 3 is a schematic diagram of candidate motion vectors provided by an embodiment of the present application;
[0028] Figure 4 is a schematic diagram of an intra-frame block copy mode provided by an embodiment of the present application;
[0029] Figure 5 is a schematic diagram of an intra-frame string replication mode provided by an embodiment of the present application;
[0030] Figure 6 is a simplified block diagram of a communication system provided by one embodiment of the present application;
[0031] Figure 7 is a schematic diagram of the placement of a video encoder and a video decoder in a streaming transmission environment exemplarily shown in the present application;
[0032] Figure 8 is a flow chart of a video image reconstruction method provided by an embodiment of the present application;
[0033] Fig. 9 yes Figure 8 A schematic diagram of coefficient decoding involved in the illustrated embodiment;
[0034] Fig.10 This is a schematic diagram of an image reconstruction framework in a coding and decoding process provided by an embodiment of the present application;
[0035] Fig.11 is a block diagram of a video image reconstruction device provided by an embodiment of the present application;
[0036] Fig.12 It is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0038] Before introducing the embodiments of the present application, Figure 1 A brief introduction to video encoding technology. Figure 1 A basic flow chart of a video encoding process is exemplarily shown.
[0039] Video signal refers to an image sequence consisting of multiple frames. A frame is a representation of the spatial information of a video signal. Taking the YUV mode as an example, a frame includes a brightness sample matrix (Y) and two chrominance sample matrices (Cb and Cr). From the perspective of how the video signal is obtained, it can be divided into two methods: captured by a camera and generated by a computer. Due to different statistical characteristics, the corresponding compression encoding methods may also be different.
[0040] In some mainstream video coding technologies, such as H.265 / HEVC (High Efficient Video Coding), H.266 / VVC (Versatile Video Coding), AVS (Audio Video Coding Standard) (such as AVS3), a hybrid coding framework is used to perform the following operations and processing on the input original video signal:
[0041] 1. Block Partition Structure: The input image is divided into several non-overlapping processing units, each of which will perform similar compression operations. This processing unit is called CTU (Coding Tree Unit) or LCU (Large Coding Unit). Below CTU, more detailed divisions can be made to obtain one or more basic coding units, called CU (Coding Unit). Each CU is the most basic element in the encoding process. The following describes various encoding methods that may be used for each CU.
[0042] 2. Predictive Coding: It includes intra-frame prediction and inter-frame prediction. The original video signal is predicted by the selected reconstructed video signal to obtain the residual video signal. The encoder needs to choose the most suitable one among many possible predictive coding modes for the current CU and inform the decoder. Intra-frame prediction means that the predicted signal comes from the area that has been encoded and reconstructed in the same image. Inter-frame prediction means that the predicted signal comes from other images that have been encoded and are different from the current image (called reference images).
[0043] 3. Transform coding and quantization (Transform & Quantization): The residual video signal is transformed into the transform domain through transformation operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform), which are called transform coefficients. The signal in the transform domain is further subjected to lossy quantization operations, which loses certain information, making the quantized signal conducive to compression expression. In some video coding standards, there may be more than one transform mode to choose from, so the encoder also needs to select one of the transforms for the current CU and inform the decoder. The degree of quantization is usually determined by the quantization parameter. A larger QP (Quantization Parameter) value means that coefficients with a larger value range will be quantized to the same output, which usually results in greater distortion and lower bit rate; on the contrary, a smaller QP value means that coefficients with a smaller value range will be quantized to the same output, which usually results in less distortion and a higher bit rate.
[0044] 4. Entropy Coding or Statistical Coding: The quantized transform domain signal will be statistically compressed and encoded according to the frequency of occurrence of each value, and finally output as a binary (0 or 1) compressed bit stream. At the same time, the encoding generates other information, such as the selected mode, motion vector, etc., which also needs to be entropy encoded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC) or context-based binary arithmetic coding (CABAC).
[0045] 5. Loop Filtering: The encoded image can be reconstructed into a decoded image after undergoing inverse quantization, inverse transformation and prediction compensation (the reverse operation of 2 to 4 above). Compared with the original image, the reconstructed image has some information that is different from the original image due to the influence of quantization, resulting in distortion. Filtering the reconstructed image, such as deblocking, SAO (Sample Adaptive Offset) or ALF (Adaptive Lattice Filter), can effectively reduce the distortion caused by quantization. Since these filtered reconstructed images will be used as a reference for subsequent encoded images and used to predict future signals, the above filtering operation is also called loop filtering, and the filtering operation in the encoding loop.
[0046] According to the above encoding process, it can be seen that at the decoding end, for each CU, after the decoder obtains the compressed bitstream, it first performs entropy decoding to obtain various mode information and quantized transform coefficients. Each coefficient is dequantized and inversely transformed to obtain a residual signal. On the other hand, based on the known encoding mode information, the prediction signal corresponding to the CU can be obtained. After adding the two, the reconstructed signal can be obtained. Finally, the reconstructed value of the decoded image needs to undergo a loop filtering operation to generate the final output signal.
[0047] Some mainstream video coding standards, such as HEVC, VVC, AVS3, etc., all adopt a block-based hybrid coding framework. They divide the original video data into a series of coding blocks, and combine video coding methods such as prediction, transform and entropy coding to achieve video data compression. Among them, motion compensation is a commonly used prediction method for video coding. Motion compensation is based on the redundant characteristics of video content in the time domain or spatial domain, and derives the prediction value of the current coding block from the encoded area. This type of prediction method includes: inter-frame prediction, intra-frame block copy prediction, intra-frame string copy prediction, etc. In specific coding implementations, these prediction methods may be used alone or in combination. For coding blocks using these prediction methods, it is usually necessary to explicitly or implicitly encode one or more two-dimensional displacement vectors in the bitstream to indicate the displacement of the current block (or the same-position block of the current block) relative to one or more reference blocks.
[0048] It should be noted that in different prediction modes and different implementations, the displacement vector may have different names. This article uniformly describes them in the following way: 1) The displacement vector in the inter-frame prediction mode is called the motion vector (MV); 2) The displacement vector in the IBC (Intra Block Copy) prediction mode is called the block vector (BV); 3) The displacement vector in the ISC (Intra String Copy) prediction mode is called the string vector (SV). Intra-frame string copy is also called "string prediction" or "string matching".
[0049] MV refers to the displacement vector used in the inter-frame prediction mode, which points from the current image to the reference image, and its value is the coordinate offset between the current block and the reference block, where the current block and the reference block are in two different images. In the inter-frame prediction mode, motion vector prediction can be introduced. By predicting the motion vector of the current block, the predicted motion vector corresponding to the current block is obtained, and the difference between the predicted motion vector corresponding to the current block and the actual motion vector is encoded and transmitted. Compared with directly encoding and transmitting the actual motion vector corresponding to the current block, it is beneficial to save bit overhead. In the embodiment of the present application, the predicted motion vector refers to the predicted value of the motion vector of the current block obtained by the motion vector prediction technology.
[0050] BV refers to the displacement vector used in the IBC prediction mode, and its value is the coordinate offset between the current block and the reference block, where both the current block and the reference block are in the current image. In the IBC prediction mode, block vector prediction can be introduced. By predicting the block vector of the current block, the predicted block vector corresponding to the current block is obtained, and the difference between the predicted block vector corresponding to the current block and the actual block vector is encoded and transmitted. Compared with directly encoding and transmitting the actual block vector corresponding to the current block, it is beneficial to save bit overhead. In an embodiment of the present application, the predicted block vector refers to the predicted value of the block vector of the current block obtained by the block vector prediction technology.
[0051] SV refers to the displacement vector used in the ISC prediction mode, and its value is the coordinate offset between the current string and the reference string, where both the current string and the reference string are in the current image. In the ISC prediction mode, string vector prediction can be introduced. By predicting the string vector of the current string, a predicted string vector corresponding to the current string is obtained, and the difference between the predicted string vector corresponding to the current string and the actual string vector is encoded and transmitted. Compared with directly encoding and transmitting the actual string vector corresponding to the current string, it is beneficial to save bit overhead. In the embodiment of the present application, the predicted string vector refers to the predicted value of the string vector of the current string obtained by the string vector prediction technology.
[0052] Several different prediction modes are introduced below:
[0053] 1. Inter-frame prediction mode
[0054] like Figure 2 As shown in the figure, inter-frame prediction uses the correlation in the video time domain and uses the pixels of the adjacent encoded images to predict the pixels of the current image, so as to effectively remove the redundancy in the video time domain and effectively save the bits of the encoded residual data. Where P is the current frame, Pr is the reference frame, B is the current block to be encoded, and Br is the reference block of B. B' and B have the same coordinate position in the image, Br coordinates are (xr, yr), and B' coordinates are (x, y). The displacement between the current block to be encoded and its reference block is called the motion vector (MV), that is:
[0055] MV=(xr-x,yr-y).
[0056] Considering that adjacent blocks in the temporal or spatial domain have a strong correlation, MV prediction technology can be used to further reduce the bits required for encoding MV. In H.265 / HEVC, inter-frame prediction includes two MV prediction technologies: Merge and AMVP (Advanced Motion Vector Prediction).
[0057] The Merge mode will establish an MV candidate list for the current PU (Prediction Unit), which contains 5 candidate MVs (and their corresponding reference images). It will traverse these 5 candidate MVs and select the one with the lowest rate-distortion cost as the optimal MV. If the codec establishes the candidate list in the same way, the encoder only needs to transmit the index of the optimal MV in the candidate list. It should be noted that HEVC's MV prediction technology also has a skip mode, which is a special case of the Merge mode. After the Merge mode finds the optimal MV, if the current block is basically the same as the reference block, then there is no need to transmit the residual data, only the MV index and a skip flag.
[0058] The MV candidate list established in Merge mode includes two cases: spatial domain and temporal domain. For B Slice (B frame image), it also includes a combined list. Among them, the spatial domain provides up to 4 candidate MVs, and its establishment is as follows: Figure 3 The spatial list is established in the order of A1→B1→B0→A0→B2, where B2 is a substitute, that is, when one or more of A1, B1, B0, A0 does not exist, the motion information of B2 needs to be used; the temporal domain only provides one candidate MV at most, and its establishment is as follows Figure 3 As shown in part (b) of , the MV of the same-position PU is scaled as follows:
[0059] curMV=td*colMV / tb;
[0060] Among them, curMV represents the MV of the current PU, colMV represents the MV of the co-located PU, td represents the distance between the current image and the reference image, and tb represents the distance between the co-located image and the reference image. If the PU at position D0 on the co-located block is not available, it is replaced by the co-located PU at position D1. For the PU in B Slice, since there are two MVs, its MV candidate list also needs to provide two MVPs (Motion Vector Predictor). HEVC generates a combination list for B Slice by combining the first 4 candidate MVs in the MV candidate list in pairs.
[0061] Similarly, the AMVP mode uses the MV correlation of neighboring blocks in the spatial and temporal domains to establish a MV candidate list for the current PU. Unlike the Merge mode, the AMVP mode selects the optimal predicted MV from the MV candidate list, and performs differential encoding with the optimal MV obtained through motion search for the current block to be encoded, that is, encoding MVD = MV-MVP, where MVD is the motion vector residual (Motion Vector Difference); by establishing the same list, the decoding end only needs the sequence number of MVD and MVP in the list to calculate the MV of the current decoding block. The MV candidate list of the AMVP mode also includes both spatial and temporal scenarios. The difference is that the length of the MV candidate list of the AMVP mode is only 2.
[0062] As mentioned above, in the AMVP mode of HEVC, MVD needs to be encoded. In HEVC, the resolution of MVD is controlled by the use_integer_mv_flag in the slice_header. When the value of the flag is 0, MVD is encoded at 1 / 4 (brightness) pixel resolution; when the value of the flag is 1, MVD is encoded at an integer (brightness) pixel resolution. An adaptive motion vector resolution (AMVR) method is used in VVC. This method allows each CU to adaptively select the resolution of the encoded MV. In the normal AMVP mode, the optional resolutions include 1 / 4, 1 / 2, 1 and 4 pixel resolutions. For a CU with at least one non-zero MVD component, a flag is first encoded to indicate whether the quarter brightness sampling MVD precision is used for the CU. If the flag is 0, the MVD of the current CU is encoded at 1 / 4 pixel resolution. Otherwise, a second flag needs to be encoded to indicate that the CU uses 1 / 2 pixel resolution or other MVD resolution. Otherwise, a third flag is encoded to indicate whether 1 pixel resolution or 4 pixel resolution is used for the CU.
[0063] 2. IBC Prediction Model
[0064] IBC is an intra-frame coding tool adopted in the HEVC Screen Content Coding (Screen Content Coding, referred to as SCC) extension, which significantly improves the coding efficiency of screen content. In AVS3 and VVC, IBC technology is also adopted to improve the performance of screen content coding. IBC uses the spatial correlation of screen content video and uses the coded image pixels on the current image to predict the pixels of the current block to be coded, which can effectively save the bits required for coding pixels. Figure 4 As shown in Figure 1, the displacement between the current block and its reference block in IBC is called BV (block vector). H.266 / VVC uses a BV prediction technique similar to inter-frame prediction to further save the bits required to encode BV. VVC uses a mode similar to the AMVP mode in inter-frame prediction to predict BV and allows BVD to be encoded with 1 or 4 pixel resolution.
[0065] 3. ISC Forecast Model
[0066] ISC technology divides a coding block into a series of pixel strings or unmatched pixels according to a certain scanning order (such as raster scanning, round-trip scanning and Zig-Zag scanning, etc.). Similar to IBC, each string searches for a reference string of the same shape in the encoded area of the current image, derives the predicted value of the current string, and effectively saves bits by encoding the residual between the current string pixel value and the predicted value instead of directly encoding the pixel value. Figure 5The schematic diagram of intra-frame string replication is given. The dark gray area is the encoded area, the 28 white pixels are string 1, the 35 light gray pixels are string 2, and the 1 black pixel represents the unmatched pixel. The displacement between string 1 and its reference string is Figure 5 The string vector 1 in ; the displacement between string 2 and its reference string is Figure 5 The string vector 2 in .
[0067] The intra-frame string replication technology needs to encode the SV corresponding to each string in the current coding block, the string length, and the flag of whether there is a matching string. Among them, SV represents the displacement of the string to be encoded to its reference string. The string length represents the number of pixels contained in the string. In different implementations, there are many ways to encode the string length. The following are several examples (some examples may be used in combination):
[0068] 1) Encode the length of the string directly in the bitstream;
[0069] 2) The number of pixels to be processed in the subsequent sequence is encoded in the bitstream, and the decoding end calculates the length of the current sequence according to the size N of the current block, the number of pixels N1 that have been processed, and the number of pixels to be processed N2 obtained by decoding, L = N-N1-N2;
[0070] 3) A flag is encoded in the bitstream to indicate whether the string is the last string. If it is the last string, the length of the current string L = N-N1 is calculated based on the size of the current block N and the number of processed pixels N1. If a pixel does not find a corresponding reference in the referenceable area, the pixel value of the unmatched pixel will be directly encoded.
[0071] A decoding process of string prediction is shown in Table 1 below:
[0072] Table 1
[0073]
[0074]
[0075] The relevant semantics are described as follows:
[0076] The match type flag isc_match_type_flag[i] for string prediction intra prediction is a binary variable. A value of '1' indicates that the i-th part of the current coding unit is a string; a value of '0' indicates that the i-th part of the current coding unit is an unmatched pixel. See 8.3 for the parsing process. IscMatchTypeFlag[i] is equal to the value of isc_match_type_flag[i]. If isc_match_type_flag[i] does not exist in the bitstream, the value of IscMatchTypeFlag[i] is 0.
[0077] The last flag of string prediction intra prediction isc_last_flag[i] is a binary variable. The value of '1' indicates that the i-th part of the current coding unit is the last part of the current coding unit, and the length of this part StrLen[i] is equal to NumTotalPixel-NumCodedPixel; the value of '0' indicates that the value of '1' indicates that the i-th part of the current coding unit is not the last part of the current coding unit, and the length of this part StrLen[i] is equal to NumTotalPixel-NumCodedPixel-NumRemainingPixelMinus1[i]-1. See 8.3 for the parsing process. IscLastFlag[i] is equal to the value of isc_last_flag[i].
[0078] The value of the next remaining pixel number next_remaining_pixel_in_cu[i] represents the number of pixels remaining in the current coding unit that have not yet been decoded after the decoding of the i-th part of the current coding unit is completed. The value of NextRemainingPixelInCu[i] is equal to the value of next_remaining_pixel_in_cu[i].
[0079] String prediction intra prediction unmatched pixel Y component value isc_unmatched_pixel_y[i]
[0080] The U component value of the unmatched pixel in the intra prediction isc_unmatched_pixel_u[i]
[0081] String prediction intra prediction unmatched pixel V component value isc_unmatched_pixel_v[i]
[0082] The above three values are 10-bit unsigned integers representing the values of the Y, Cb or Cr components of the unmatched pixels in the i-th part of the current coding unit. IscUnmatchedPixelY[i], IscUnmatchedPixelU[i] and IscUnmatchedPixelV[i] are equal to the values of isc_unmatched_pixel_y[i], isc_unmatched_pixel_u[i] and isc_unmatched_pixel_v[i] respectively.
[0083] like Figure 6, which shows a simplified block diagram of a communication system provided by an embodiment of the present application. The communication system 200 includes a plurality of devices, which can communicate with each other through, for example, a network 250. For example, the communication system 200 includes a first device 210 and a second device 220 interconnected through the network 250. Figure 6 In an embodiment, the first device 210 and the second device 220 perform unidirectional data transmission. For example, the first device 210 may encode video data, such as a video picture stream captured by the first device 210, for transmission to the second device 220 via the network 250. The encoded video data is transmitted in the form of one or more encoded video code streams. The second device 220 may receive the encoded video data from the network 250, decode the encoded video data to recover the video data, and display the video picture based on the recovered video data. Unidirectional data transmission is more common in applications such as media services.
[0084] In another embodiment, the communication system 200 includes a third device 230 and a fourth device 240 that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, each of the third device 230 and the fourth device 240 may encode video data (e.g., a video picture stream collected by the device) for transmission to the other of the third device 230 and the fourth device 240 via the network 250. Each of the third device 230 and the fourth device 240 may also receive encoded video data transmitted by the other of the third device 230 and the fourth device 240, and may decode the encoded video data to recover the video data, and may display the video picture on an accessible display device based on the recovered video data.
[0085] exist Figure 6 In the embodiment of the present invention, the first device 210, the second device 220, the third device 230 and the fourth device 240 may be computer devices such as servers, personal computers and smart phones, but the principles disclosed in the present application may not be limited thereto. The present application embodiment is applicable to PC (Personal Computer), mobile phones, tablet computers, media players and / or dedicated video conferencing equipment. The network 250 represents any number of networks that transmit encoded video data between the first device 210, the second device 220, the third device 230 and the fourth device 240, including, for example, wired and / or wireless communication networks. The communication network 250 can exchange data in circuit switching and / or packet switching channels. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purpose of the present application, unless explained below, the architecture and topology of the network 250 may be irrelevant to the operation disclosed in the present application.
[0086] As an example, Figure 7 The video encoder and the video decoder are shown in a streaming environment. The subject matter disclosed in the present application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CD (Compact Disc), DVD (Digital Versatile Disc), memory stick, etc.
[0087] The streaming system may include an acquisition subsystem 313, which may include a video source 301 such as a digital camera, which creates an uncompressed video picture stream 302. In an embodiment, the video picture stream 302 includes samples taken by a digital camera. The video picture stream 302 is depicted as a thick line to emphasize the high data volume of the video picture stream compared to the encoded video data 304 (or the encoded video code stream), and the video picture stream 302 can be processed by an electronic device 320, which includes a video encoder 303 coupled to the video source 301. The video encoder 303 may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream 302, the encoded video data 304 (or the encoded video code stream 304) is depicted as a thin line to emphasize the lower data volume of the encoded video data 304 (or the encoded video code stream 304), which can be stored on the streaming server 305 for future use. One or more streaming client subsystems, such as Figure 7 , can access the streaming server 305 to retrieve the copies 307 and 309 of the encoded video data 304. The client subsystem 306 can include, for example, a video decoder 310 in an electronic device 330. The video decoder 310 decodes the incoming copy 307 of the encoded video data and produces an output video picture stream 311 that can be presented on a display 312 (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data 304, video data 307, and video data 309 (e.g., video bitstreams) can be encoded according to certain video encoding / compression standards.
[0088] It should be noted that the electronic device 320 and the electronic device 330 may include other components (not shown). For example, the electronic device 320 may include a video decoder (not shown), and the electronic device 330 may also include a video encoder (not shown). The video decoder is used to decode the received encoded video data; the video encoder is used to encode the video data.
[0089] It should be noted that the technical solution provided in the embodiments of the present application can be applied to the H.266 / VVC standard, the H.265 / HEVC standard, AVS (such as AVS3) or the next-generation video coding and decoding standard, and the embodiments of the present application are not limited to this.
[0090] In the current AVS3 standard, for the string prediction coding unit, its prediction signal is directly composed of the searched reference pixels, skipping the transformation, quantization and coefficient encoding processes, which is equivalent to using the prediction signal directly as the reconstruction signal without the participation of the residual signal. The complete use of the prediction signal without residual information will bring a large difference from the original signal, affecting the coding performance.
[0091] The present application proposes an image reconstruction method that introduces residual information into string prediction. When a CU is encoded based on string prediction, the residual signal of the CU is encoded. When the signal of the CU encoded by string prediction is reconstructed, the string prediction is completely processed for the residual signal, that is, the signal is reconstructed in combination with the residual signal and the prediction signal of the CU, thereby improving the reconstructed signal quality of the coding unit based on the string prediction method, thereby improving the encoding and decoding performance.
[0092] In the method provided in the embodiment of the present application, the execution subject of each step can be a decoding end device or an encoding end device. In the process of video decoding and video encoding, the technical solution provided in the embodiment of the present application can be used to perform image reconstruction. Both the decoding end device and the encoding end device can be computer devices, which refer to electronic devices with data calculation, processing and storage capabilities, such as PCs, mobile phones, tablet computers, media players, dedicated video conferencing equipment, servers, etc.
[0093] In addition, the method provided in the present application can be used alone or in combination with other methods in any order. The encoder and decoder based on the method provided in the present application can be implemented by one or more processors or one or more integrated circuits.
[0094] Please refer to Figure 8 , which shows a flow chart of a video image reconstruction method provided by an embodiment of the present application. For the sake of convenience, only the computer device is used as the subject of each step. The method may include the following steps:
[0095] Step 801, in response to encoding and decoding in a string prediction manner, obtaining a residual code of an encoded target coding unit; the residual code is obtained by encoding a residual coefficient of the target coding block.
[0096] In an embodiment of the present application, when the encoder performs string prediction-based encoding on a target coding unit, the original signal of the unencoded target coding unit can be obtained; based on a reference signal, the original signal of the target coding unit is predicted by string prediction to obtain a prediction signal of the target coding unit and a residual signal of the target coding unit; based on the residual signal of the target coding unit, the residual code of the target coding unit is obtained; and the motion information corresponding to the prediction signal of the target coding unit and the residual code of the target coding unit are added to the encoded video code stream.
[0097] In a possible implementation, when encoding and decoding is performed in a string prediction manner, the encoder / decoder directly obtains the residual coding of the target coding unit, so as to decode the residual coefficients of the encoded target coding unit.
[0098] The above residual coding is obtained by transforming, quantizing, entropy coding / statistical coding the residual signal of the target coding unit during the process of encoding the target coding unit by the encoder.
[0099] In a possible implementation, when the encoder / decoder reconstructs an image of an encoded target coding unit, it first determines whether residual coding decoding of the target encoded image is required.
[0100] For example, the encoder / decoder obtains residual decoding indication information of the target coding unit in response to encoding and decoding by string prediction; the residual decoding indication information indicates whether to decode the residual of the target coding block;
[0101] The encoder / decoder performs residual decoding on the target coding block in response to the residual indication information, and obtains the residual coding of the target coding unit.
[0102] In a possible implementation manner, the residual decoding indication information includes at least one of the following information:
[0103] 1) A first index obtained by decoding in a sequence header corresponding to the target coding unit, where the first index is used to indicate a residual coefficient of a coding unit decoded in a corresponding sequence by a string prediction decoding method.
[0104] In an embodiment of the present application, when the encoder performs video encoding by string prediction, an index may be added to the sequence header, where the index indicates the residual coefficients of the coding units in all CUs of the sequence that need to be decoded using the string prediction technology.
[0105] For example, when the encoder / decoder decodes the first index in the sequence header and determines that there are residual coefficients of the target coding unit, it is determined that the residual coefficients of the target coding unit need to be decoded.
[0106] 2) A second index obtained by decoding in an image header corresponding to the target coding unit, where the second index is used to indicate a residual coefficient of a coding unit decoded in a corresponding image by a string prediction decoding method.
[0107] In an embodiment of the present application, when the encoder performs video encoding by string prediction, an index may be added to the image header, where the index indicates the residual coefficients of the coding units in all CUs of the image that need to be decoded using the string prediction technology.
[0108] For example, when the encoder / decoder decodes the second index in the picture header and determines that there are residual coefficients of the target coding unit, it is determined that the residual coefficients of the target coding unit need to be decoded.
[0109] 3) A third index obtained by decoding in a slice header of the target coding unit, where the third index is used to indicate a residual coefficient of the coding unit decoded by a string prediction decoding method in a corresponding slice.
[0110] In an embodiment of the present application, when the encoder performs video encoding by string prediction, an index may be added to the slice (slice / patch) header, where the index indicates the residual coefficients of the coding units that need to be decoded using the string prediction technology in all CUs of the slice.
[0111] For example, when the encoder / decoder decodes the third index in the slice header and determines that there are residual coefficients of the target coding unit, it is determined that the residual coefficients of the target coding unit need to be decoded.
[0112] 4) a fourth index obtained by decoding in the largest coding unit LCU of the target coding unit, where the fourth index is used to indicate a residual coefficient of a coding unit decoded by a string prediction decoding method in the corresponding LCU.
[0113] In an embodiment of the present application, when the encoder performs video encoding by string prediction, an index may be added to the LCU, where the index indicates the residual coefficients of the coding units that need to be decoded using the string prediction technology in all CUs of the LCU.
[0114] For example, when the encoder / decoder decodes the fourth index in the LCU and determines that there is a residual coefficient of the target coding unit, it is determined that the residual coefficient of the target coding unit needs to be decoded.
[0115] 5) a fifth index obtained by decoding in the target coding unit, where the fifth index is used to indicate a residual coefficient of the target coding unit.
[0116] In an embodiment of the present application, when the encoder performs video encoding by string prediction, an index may be added to the target coding unit, where the index indicates the current residual coefficient.
[0117] For example, when the encoder / decoder decodes the fifth index in the target coding unit and determines that there is a residual coefficient of the target coding unit, it is determined that the residual coefficient of the target coding unit needs to be decoded.
[0118] 6) A component type of the target coding unit, where the component type includes a luminance component or a chrominance component.
[0119] For example, if the CU is currently a luminance component, it indicates that the CU needs to decode the residual coefficients of the string prediction technology;
[0120] Alternatively, if the CU is currently a chroma component, it indicates that the CU needs to decode the residual coefficients of the string prediction technology.
[0121] That is, if the target coding unit is a luminance component or a chrominance component, the encoder / decoder determines that the residual coefficient of the target coding unit needs to be decoded.
[0122] 7) A coding block identifier of a color component of each pixel in the target coding unit, wherein the coding block identifier is used to indicate whether the color component of the corresponding pixel is non-zero.
[0123] In an embodiment of the present application, for the current CU, the encoder / decoder decodes the cbf (indicating whether the component has a non-zero coefficient) of each color component from the bitstream to indicate that the CU needs to decode the coefficients of the string prediction technology; for example, when the cbf indicates the presence of a non-zero coefficient, the encoder / decoder determines that the residual coefficients of the target coding unit need to be decoded; or, when the cbf indicates that there are no non-zero coefficients, the encoder / decoder determines that the residual coefficients of the target coding unit need to be decoded.
[0124] 8) The size of the target coding unit.
[0125] In an embodiment of the present application, the encoder / decoder may indicate the coefficients of the decoding string prediction technology required for the CU based on the size of the decoding block (which can be measured by length*width, long side size, short side size, etc.); for example, when the size of the CU exceeds a threshold, the encoder / decoder determines that the residual coefficients of the decoding string prediction technology need to be applied to the CU.
[0126] 9) Coefficient indication information in the target coding unit, where the coefficient indication information is used to indicate whether all coefficients of the target coding unit are zero.
[0127] In the embodiment of the present application, the encoder / decoder may indicate that the coefficients of the CU need to be decoded using the string prediction technology, based on whether all coefficients in the CU are zero.
[0128] For example, if the coefficients in the target coding unit are not all zero, it is determined that the residual coefficients of the decoding string prediction technology need to be applied to the target decoding unit.
[0129] Alternatively, if all coefficients in the target coding unit are 0, it is determined that residual coefficients of the target decoding unit need to be decoded using a string prediction technique.
[0130] 10) Whether the target coding unit contains isolated points.
[0131] In an embodiment of the present application, the encoder / decoder may indicate that the CU needs to decode the coefficients of the string prediction technology according to whether the string prediction content contains an isolated point. For example, when an isolated point is contained, it is determined that the CU needs to decode the residual coefficients of the string prediction technology, or when an isolated point is not contained, it is determined that the CU needs to decode the residual coefficients of the string prediction technology.
[0132] Step 802: decode the residual coding of the target coding unit to obtain the residual coefficient of the target coding unit.
[0133] In a possible implementation manner, the encoder / decoder may decode the residual coding by using Scan Region based Coefficient Coding (SRCC) to obtain the residual coefficients.
[0134] like Fig. 9 As shown, it shows a schematic diagram of coefficient decoding involved in an embodiment of the present application. Fig. 9 As shown in the figure, SRCC technology uses (SRx, SRy) to determine the quantized coefficient area to be scanned in a transform unit (TU), where SRx is the horizontal coordinate of the rightmost non-zero coefficient in the coefficient matrix, and SRy is the vertical coordinate of the bottom non-zero coefficient in the coefficient matrix. Only the coefficients in the scan area determined by (SRx, SRy) need to be encoded. The scanning order of encoding is a reverse Z-shaped scan from the lower right corner to the upper left corner, and each coefficient is encoded in turn.
[0135] In a possible implementation, the residual coding of the target coding unit includes at least two sub-coding blocks;
[0136] When the encoder / decoder decodes the residual coding of the target coding unit and obtains the residual coefficient of the target coding unit, it can skip the scanning of the all-zero sub-coding block in the at least two sub-coding blocks, and perform coefficient decoding based on the scanning area on the non-all-zero sub-coding blocks in the at least two sub-coding blocks to obtain the residual coefficient of the target coding unit.
[0137] In one possible implementation, the encoder / decoder skips scanning of the all-zero sub-coding block among the at least two sub-coding blocks, and takes the width of the target coding block as a basic unit, performs coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the horizontal direction to obtain the residual coefficient of the target coding unit.
[0138] In one possible implementation, the encoder / decoder skips scanning of the all-zero sub-coding block among the at least two sub-coding blocks, and takes the height of the target coding block as a basic unit, performs coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the vertical direction, and obtains the residual coefficient of the target coding unit.
[0139] In an embodiment of the present application, the encoder / decoder can skip unnecessary scanning areas by marking non-zero sub-blocks. For example, when decoding coefficients, the encoder / decoder first decodes whether each NxN sub-block coefficient in the CU is 0. If so, the scanning of this all-zero area is skipped. Wherein, N can be one or more preset values, such as 1, 2, 4, ..., w*h / 2, the width (w) of the current CU, or the height (h).
[0140] In an embodiment of the present application, when decoding the residual coding, the encoder / decoder can also scan by changing the scanning order. For example, when using a scanning method along the horizontal direction (including raster-scan, traverse-scan, etc.), the CU block width is used as the basic unit; when using a scanning method along the vertical direction (including raster-scan, traverse-scan, etc.), the CU block height is used as the basic unit.
[0141] In another possible implementation, the encoder / decoder uses other coefficient decoding methods other than SRCC, such as the coefficient decoding method of the first stage of AVS3.
[0142] Step 803: Obtain a residual signal of the target coding unit based on the residual coefficient of the target coding unit.
[0143] In a possible implementation manner, obtaining a residual signal of the target coding unit based on a residual coefficient of the target coding unit includes:
[0144] Dequantizing the residual coefficients of the target coding unit to obtain transform coefficients of the target coding unit;
[0145] An inverse transform skipping process is performed on the transform coefficients of the target coding unit, or an inverse transform process is performed on the transform coefficients of the target coding unit to obtain a residual signal of the target coding unit.
[0146] Transform skip (TS) technology refers to the use of the existing spatial residual signal as the transform coefficient directly without converting the residual video signal into the transform domain, without going through DFT and DCT. Subsequently, similar to other residual blocks based on transform methods, further lossy quantization operations are performed to lose certain information, making the quantized signal conducive to compression expression. Subsequent operations may include encoding whether each position of the residual block is 0, as well as the size and sign of the non-zero coefficient.
[0147] In some video coding standards, such as HEVC and VVC, if the encoder needs to select a transform skip mode for the current coding CU, it will explicitly inform the decoder. However, the AVS video coding standard only uses the implicit selection of transform skip mode (ISTS) for screen content, as shown in Table 2 below:
[0148] Table 2
[0149]
[0150] Based on Table 2 above, the encoder needs to select the (DCT2, DCT2) mode or the transform skip (TS) mode by cost comparison. In order to hide the ISTS flag, it is assumed that when the encoder selects the (DCT2, DCT2) mode, the number of non-zero coefficients of the current block is an even number. If the actual number of non-zero coefficients is an odd number, the encoder will adjust the transform coefficients to make the number of non-zero coefficients an even number. Similarly, when the encoder selects the transform skip mode, the number of non-zero coefficients obtained is an odd number. If the actual number of non-zero coefficients is an even number, the encoder will adjust the transform coefficients to make the number of non-zero coefficients an odd number.
[0151] In a possible implementation manner, performing inverse transform skip processing on the transform coefficients of the target coding unit includes:
[0152] In response to the non-zero number of transform coefficients of the target coding unit satisfying a first condition, performing inverse transform skip processing on the transform coefficients of the target coding unit;
[0153] or,
[0154] In response to obtaining a first flag bit of the target coding unit by decoding from a bitstream, and the first flag bit indicating inverse transform skipping, performing inverse transform skipping processing on transform coefficients of the target coding unit;
[0155] or,
[0156] In response to the existence of an isolated point in the target coding unit, performing an inverse transform skipping process on the transform coefficients of the target coding unit;
[0157] or,
[0158] In response to the absence of an isolated point in the target coding unit, an inverse transform skipping process is performed on the transform coefficients of the target coding unit.
[0159] In the embodiment of the present application, before the encoder / decoder performs inverse transform skipping, it can be determined whether inverse transform skipping is to be performed. The determination method may include the following:
[0160] 1) Determine whether to perform inverse transform skipping by implicitly expressing transform skipping mode;
[0161] 2) Decode the bit stream to get the flag bit and determine whether to skip the inverse transformation;
[0162] 3) directly apply the transform skip mode to perform inverse transform skip without selection;
[0163] 4) Predict whether there are isolated points based on the current CU string and determine whether to skip the inverse transformation.
[0164] In a possible implementation manner, the inverse transformation process is performed on the transformation coefficient of the target coding unit, including:
[0165] In response to the non-zero number of transform coefficients of the target coding unit satisfying a second condition, performing inverse transform processing on the transform coefficients of the target coding unit using a transform kernel indicated by the second condition;
[0166] or,
[0167] In response to a second flag bit of the target coding unit obtained by decoding from a bitstream, performing an inverse transform skipping process on a transform coefficient of the target coding unit through a transform core indicated by the second flag bit;
[0168] or,
[0169] The transform coefficients of the target coding unit are subjected to inverse transform skipping processing through the specified transform kernel.
[0170] In the embodiment of the present application, before the encoder / decoder performs the inverse transformation, it can be determined whether the inverse transformation is required, and if the inverse transformation is required, the transformation kernel to be used is determined; the inverse transformation methods may include the following:
[0171] 1) Perform inverse transformation using the implicitly expressed transformation kernel;
[0172] 2) Decode from the bitstream whether to use DCT2 or other transform kernels for inverse transform;
[0173] 3) Use DCT2 or some other transform kernel directly without selection.
[0174] In a possible implementation manner, the dequantizing the residual coefficients of the target coding unit to obtain the transform coefficients of the target coding unit includes:
[0175] Adjusting the quantization parameter according to a specified quantization parameter adjustment method, where the quantization parameter adjustment method includes increasing or decreasing;
[0176] Based on the adjusted quantization parameter, the residual coefficient of the target coding unit is dequantized to obtain the transform coefficient of the target coding unit.
[0177] In an embodiment of the present application, when the above method is executed by an encoder, before the encoder performs inverse quantization, the QP can be amplified based on the QP used in the quantization step of the encoder in the encoding process of the target coding unit according to business needs or scenario requirements to allow a larger residual signal to appear, thereby expanding the application of coefficient coding in string replication technology, thereby improving the scope of application of the coding, or the QP can be reduced to improve the coding accuracy.
[0178] Step 804: Obtain a reconstructed signal of the target coding unit based on the residual signal of the target coding unit and the prediction signal of the target coding unit.
[0179] In a possible implementation manner, the encoder / decoder superimposes the residual signal and the prediction signal of the target coding unit to obtain a reconstructed signal of the target coding unit.
[0180] In an embodiment of the present application, when the above method is executed by an encoder, before obtaining the reconstructed signal obtained by decoding the target coding unit based on the residual signal of the target coding unit and the prediction signal of the target coding unit, the encoder also obtains an adjusted matching threshold based on the adjusted quantization parameter. The matching threshold is a threshold used to indicate whether pixels match during the string prediction process; based on the adjusted matching threshold, the prediction signal of the target coding unit is obtained through string prediction.
[0181] In another exemplary scheme of the embodiments of the present application, the encoder uses an unadjusted matching threshold to obtain a prediction signal of the target coding unit through a string prediction method, combined with motion information (such as the SV corresponding to each string in the above-mentioned target coding unit, the string length, and a flag as to whether there is a matching string, etc.) and a reconstructed reference signal.
[0182] That is to say, in the scheme shown in the embodiment of the present application, the encoder can directly obtain the prediction signal based on the string prediction method through motion information, and derive its reconstructed signal by superimposing it with the residual signal obtained by decoding; or, based on the above-mentioned adjustment of QP, it can also adjust the threshold of whether the pixels in the string prediction match, that is, the adjusted matching threshold corresponds to the adjusted QP, and then obtain the prediction signal based on the string prediction method through motion information, and derive its reconstructed signal by superimposing it with the residual signal obtained by decoding.
[0183] To summarize, the scheme shown in the embodiment of the present application introduces residual information into the string prediction process when the encoder / decoder performs encoding and decoding through the string prediction method, performs complete residual signal processing on the string prediction, improves the reconstructed signal quality of the coding unit based on the string prediction method, and thus improves the encoding and decoding performance.
[0184] In addition, the scheme shown in the embodiment of the present application, during the image reconstruction process on the encoder side, the QP used for inverse quantization is enlarged or reduced based on the QP used for encoding, and accordingly, when obtaining the prediction signal, the matching threshold is also adjusted, so that the accuracy of image reconstruction can be flexibly adjusted according to business or scenario requirements, that is, the flexibility of image reconstruction accuracy control is improved.
[0185] In addition, the scheme shown in the embodiment of the present application, during the image reconstruction process, when decoding the residual coefficients, skips the search for the all-zero sub-region in the TU, and searches and decodes the non-all-zero sub-region in the TU, thereby improving the decoding efficiency of the residual coefficients and further improving the efficiency of image reconstruction.
[0186] The image reconstruction method proposed in this application can be applied in an encoder or a decoder. In the encoder, the reconstructed signal of the encoded CU is used as a reference signal for subsequent CU encoding, and in the decoder, the reconstructed signal of the encoded CU is used for subsequent CU decoding and video playback.
[0187] Please refer to Fig.10 , which shows a schematic diagram of an image reconstruction framework in the encoding and decoding process provided by an embodiment of the present application. Fig.10As shown, in a possible application example of the above scheme of the present application, the encoder 101 divides the image to obtain the original signal 1011 of the coding unit CU1, and predicts and encodes the original signal 1011 of the coding unit CU1 in a string prediction manner through the reconstructed signal in the reconstructed signal buffer 1012 to obtain a prediction signal 1013, and a residual signal 1014 between the original signal 1011 and the prediction signal 1013; the encoder 101 changes and quantizes the residual signal 1014 to obtain a residual coefficient 1015, and performs entropy encoding or statistical encoding on the residual coefficient 1015 to obtain a residual code 1016; wherein the residual code 1016 and the motion information 1017 corresponding to the above prediction signal 1013 (the SV corresponding to each string, the string length, and the mark of whether there is a matching string, etc.) are added to the encoded bit stream as the encoding result of CU1 and transmitted to the decoder 102.
[0188] After the above-mentioned CU1 is encoded in the encoder 101, the encoder 101 also predicts the prediction signal 1013 through the motion information 1017, and performs coefficient decoding, inverse quantization & inverse transformation / inverse transformation skipping processing on the residual coding 1016 to obtain the residual coefficient 1015 and the residual signal 1014 in turn, and superimposes and fuses the residual signal 1014 and the prediction signal 1013 to obtain the reconstructed signal of CU1, and adds the reconstructed signal of CU1 to the reconstructed signal buffer 1012 to be used as a reference signal for encoding subsequent CUs.
[0189] In the decoder 102, when the decoder 102 decodes CU1 by string prediction, it is determined that the residual coefficients of CU1 need to be decoded, and the motion information 1017 and residual code 1016 of CU1 are obtained from the bitstream, and the reconstructed signal in the reconstructed signal buffer 1021 is used as a reference signal, and the prediction signal 1013 is obtained through the motion information 1017, and the residual code 1016 is subjected to coefficient decoding, inverse quantization & inverse transformation / inverse transformation skipping processing to obtain the residual coefficient 1015 and the residual signal 1014 in turn, and the residual signal 1014 is fused with the prediction signal 1013 to obtain the reconstructed signal of CU1. The reconstructed signal of CU1 is used for video playback on the one hand, and is also used as a reference signal for decoding subsequent CUs on the other hand.
[0190] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0191] Please refer to Fig.11, which shows a block diagram of a video image reconstruction device provided by an embodiment of the present application. The device has the function of implementing the above method example, and the function can be implemented by hardware, or by hardware executing corresponding software. The device can be the computer device introduced above, or it can be set on a computer device. The device 1100 may include:
[0192] The residual coding acquisition module 1101 is used to obtain the residual coding of the encoded target coding unit in response to encoding and decoding in a string prediction manner; the residual coding is obtained by encoding the residual coefficients of the target coding block;
[0193] A coefficient decoding module 1102, configured to decode the residual coding of the target coding unit to obtain the residual coefficient of the target coding unit;
[0194] A residual signal acquisition module 1103, configured to acquire a residual signal of the target coding unit based on a residual coefficient of the target coding unit;
[0195] The signal reconstruction module 1104 is used to obtain a reconstructed signal of the target coding unit based on the residual signal of the target coding unit and the prediction signal of the target coding unit.
[0196] In a possible implementation, the residual coding acquisition module 1101 is used to:
[0197] In response to encoding and decoding by string prediction, obtaining residual decoding indication information of the target coding unit; the residual decoding indication information indicates whether to decode the residual of the target coding block;
[0198] In response to the residual indication information indicating that residual decoding is performed on the target coding block, residual coding of the target coding unit is obtained.
[0199] In a possible implementation manner, the residual decoding indication information includes at least one of the following information:
[0200] A first index obtained by decoding in a sequence header corresponding to the target coding unit, the first index being used to indicate a residual coefficient of a coding unit decoded in a corresponding sequence by a decoding method using string prediction;
[0201] A second index obtained by decoding in a picture header corresponding to the target coding unit, the second index being used to indicate a residual coefficient of a coding unit decoded in a corresponding picture by a string prediction decoding method;
[0202] A third index obtained by decoding in a slice header of the target coding unit, the third index being used to indicate a residual coefficient of a coding unit decoded by a string prediction decoding manner in a corresponding slice;
[0203] A fourth index obtained by decoding in the largest coding unit LCU of the target coding unit, the fourth index being used to indicate a residual coefficient of a coding unit decoded by a string prediction decoding manner in the corresponding LCU;
[0204] a fifth index obtained by decoding in the target coding unit, the fifth index being used to indicate a residual coefficient of the target coding unit;
[0205] A component type of the target coding unit, wherein the component type includes a luminance component or a chrominance component;
[0206] A coding block identifier of a color component of each pixel in the target coding unit, wherein the coding block identifier is used to indicate whether the color component of the corresponding pixel is non-zero;
[0207] The size of the target coding unit;
[0208] Coefficient indication information in the target coding unit, the coefficient indication information is used to indicate whether all coefficients of the target coding unit are zero;
[0209] And, whether the target coding unit contains isolated points.
[0210] In a possible implementation, the residual signal acquisition module 1103 is used to:
[0211] Dequantizing the residual coefficients of the target coding unit to obtain transform coefficients of the target coding unit;
[0212] An inverse transform skipping process is performed on the transform coefficients of the target coding unit, or an inverse transform process is performed on the transform coefficients of the target coding unit to obtain a residual signal of the target coding unit.
[0213] In a possible implementation, the residual signal acquisition module 1103 is used to:
[0214] In response to the non-zero number of transform coefficients of the target coding unit satisfying a first condition, performing inverse transform skip processing on the transform coefficients of the target coding unit;
[0215] or,
[0216] In response to obtaining a first flag bit of the target coding unit by decoding from a bitstream, and the first flag bit indicates to perform inverse transform skipping, performing inverse transform skipping processing on transform coefficients of the target coding unit;
[0217] or,
[0218] In response to the existence of an isolated point in the target coding unit, performing an inverse transform skipping process on the transform coefficients of the target coding unit;
[0219] or,
[0220] In response to the absence of an isolated point in the target coding unit, an inverse transform skipping process is performed on the transform coefficients of the target coding unit.
[0221] In a possible implementation, the residual signal acquisition module 1103 is used to:
[0222] In response to the non-zero number of transform coefficients of the target coding unit satisfying a second condition, performing inverse transform processing on the transform coefficients of the target coding unit using a transform kernel indicated by the second condition;
[0223] or,
[0224] In response to a second flag bit of the target coding unit obtained by decoding from a bitstream, performing an inverse transform skipping process on a transform coefficient of the target coding unit through a transform core indicated by the second flag bit;
[0225] or,
[0226] The inverse transform skipping process is performed on the transform coefficients of the target coding unit through the specified transform kernel.
[0227] In a possible implementation, the residual signal acquisition module 1103 is used to:
[0228] Adjusting the quantization parameter according to a specified quantization parameter adjustment method, wherein the quantization parameter adjustment method includes increasing or decreasing;
[0229] Based on the adjusted quantization parameter, the residual coefficient of the target coding unit is dequantized to obtain the transform coefficient of the target coding unit.
[0230] In a possible implementation manner, the device further includes:
[0231] a matching threshold adjustment module, configured to obtain an adjusted matching threshold based on the adjusted quantization parameter before the signal reconstruction module 1104 obtains a reconstructed signal obtained by decoding the target coding unit based on the residual signal of the target coding unit and the prediction signal of the target coding unit, wherein the matching threshold is a threshold used to indicate whether pixels are matched in a string prediction process;
[0232] The prediction signal acquisition module is used to acquire the prediction signal of the target coding unit through string prediction based on the adjusted matching threshold.
[0233] In a possible implementation, the residual coding of the target coding unit includes at least two sub-coding blocks;
[0234] The coefficient decoding module 1102 is used to skip scanning of the all-zero sub-coding blocks in the at least two sub-coding blocks, perform coefficient decoding based on the scanning area on the non-all-zero sub-coding blocks in the at least two sub-coding blocks, and obtain the residual coefficients of the target coding unit.
[0235] In a possible implementation, the coefficient decoding module 1102 is used to skip scanning of the all-zero sub-coding block in the at least two sub-coding blocks, and take the width of the target coding block as a basic unit, perform coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the horizontal direction, and obtain the residual coefficient of the target coding unit.
[0236] In a possible implementation, the coefficient decoding module 1102 is used to skip scanning of the all-zero sub-coding block in the at least two sub-coding blocks, and take the height of the target coding block as a basic unit, perform coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the vertical direction, and obtain the residual coefficient of the target coding unit.
[0237] To summarize, the scheme shown in the embodiment of the present application introduces residual information into the string prediction process when the encoder / decoder performs encoding and decoding through the string prediction method, performs complete residual signal processing on the string prediction, improves the reconstructed signal quality of the coding unit based on the string prediction method, and thus improves the encoding and decoding performance.
[0238] In addition, the scheme shown in the embodiment of the present application, during the image reconstruction process on the encoder side, the QP used for inverse quantization is enlarged or reduced based on the QP used for encoding, and accordingly, when obtaining the prediction signal, the matching threshold is also adjusted, so that the accuracy of image reconstruction can be flexibly adjusted according to business or scenario requirements, that is, the flexibility of image reconstruction accuracy control is improved.
[0239] In addition, the scheme shown in the embodiment of the present application, during the image reconstruction process, when decoding the residual coefficients, skips the search for the all-zero sub-region in the TU, and searches and decodes the non-all-zero sub-region in the TU, thereby improving the decoding efficiency of the residual coefficients and further improving the efficiency of image reconstruction.
[0240] Please refer to Fig.12 , which shows a block diagram of a computer device provided by an embodiment of the present application. The computer device may be the encoding end device described above, or the decoding end device described above. The computer device 150 may include: a processor 151, a memory 152, a communication interface 153, an encoder / decoder 154, and a bus 155.
[0241] The processor 151 includes one or more processing cores. The processor 151 executes various functional applications and information processing by running software programs and modules.
[0242] The memory 152 may be used to store a computer program, and the processor 151 may be used to execute the computer program to implement the above-mentioned video image reconstruction method.
[0243] The communication interface 153 may be used to communicate with other devices, such as receiving and transmitting audio and video data.
[0244] The encoder / decoder 154 may be used to implement encoding and decoding functions, such as encoding and decoding audio and video data.
[0245] The memory 152 is connected to the processor 151 via a bus 155 .
[0246] In addition, the memory 152 can be implemented by any type of volatile or non-volatile storage device or a combination thereof. The volatile or non-volatile storage device includes but is not limited to: a magnetic disk or an optical disk, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM (Erasable Programmable Read-Only Memory), an SRAM (Static Random-Access Memory), a ROM (Read-Only Memory), a magnetic storage device, a flash memory, and a PROM (Programmable read-only memory).
[0247] Those skilled in the art will understand that Fig.12 The structure shown in the figure does not constitute a limitation on the computer device 150, and the computer device 150 may include more or less components than shown in the figure, or combine some components, or adopt a different arrangement of components.
[0248] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set implements the above-mentioned video image reconstruction method when executed by a processor.
[0249] In an exemplary embodiment, a computer program product or a computer program is also provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned video image reconstruction method.
[0250] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0251] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A video image reconstruction method, It is characterized in that The method comprises: In response to encoding and decoding in a string prediction manner, obtaining residual decoding indication information of an encoded target coding unit; the residual decoding indication information indicates whether to decode the residual of the target coding block; In response to the residual decoding indication information indicating that the target coding block is residually decoded, a residual code of the target coding unit is obtained; the residual code is obtained by encoding the residual coefficient of the target coding block; Decoding the residual coding of the target coding unit to obtain a residual coefficient of the target coding unit; Acquire a residual signal of the target coding unit based on the residual coefficient of the target coding unit; A reconstructed signal of the target coding unit is obtained based on the residual signal of the target coding unit and the prediction signal of the target coding unit.
2. The method according to claim 1, It is characterized in that The residual decoding indication information includes at least one of the following information: A first index obtained by decoding in a sequence header corresponding to the target coding unit, the first index being used to indicate a residual coefficient of a coding unit decoded in a corresponding sequence by a decoding method using string prediction; A second index obtained by decoding in a picture header corresponding to the target coding unit, the second index being used to indicate a residual coefficient of a coding unit decoded in a corresponding picture by a string prediction decoding method; A third index obtained by decoding in a slice header of the target coding unit, the third index being used to indicate a residual coefficient of a coding unit decoded by a string prediction decoding manner in a corresponding slice; A fourth index obtained by decoding in the largest coding unit LCU of the target coding unit, the fourth index being used to indicate a residual coefficient of a coding unit decoded by a string prediction decoding manner in the corresponding LCU; a fifth index obtained by decoding in the target coding unit, the fifth index being used to indicate a residual coefficient of the target coding unit; A component type of the target coding unit, wherein the component type includes a luminance component or a chrominance component; A coding block identifier of a color component of each pixel in the target coding unit, wherein the coding block identifier is used to indicate whether the color component of the corresponding pixel is non-zero; The size of the target coding unit; Coefficient indication information in the target coding unit, the coefficient indication information is used to indicate whether all coefficients of the target coding unit are zero; And, whether the target coding unit contains isolated points.
3. The method according to claim 1, It is characterized in that The obtaining the residual signal of the target coding unit based on the residual coefficient of the target coding unit includes: Dequantizing the residual coefficients of the target coding unit to obtain transform coefficients of the target coding unit; An inverse transform skipping process is performed on the transform coefficients of the target coding unit, or an inverse transform process is performed on the transform coefficients of the target coding unit to obtain a residual signal of the target coding unit.
4. The method according to claim 3, It is characterized in that The performing inverse transform skipping processing on the transform coefficients of the target coding unit comprises: In response to the non-zero number of transform coefficients of the target coding unit satisfying a first condition, performing inverse transform skip processing on the transform coefficients of the target coding unit; or, In response to obtaining a first flag bit of the target coding unit by decoding from a bitstream, and the first flag bit indicates to perform inverse transform skipping, performing inverse transform skipping processing on transform coefficients of the target coding unit; or, In response to the existence of an isolated point in the target coding unit, performing an inverse transform skipping process on the transform coefficients of the target coding unit; or, In response to the absence of an isolated point in the target coding unit, an inverse transform skipping process is performed on the transform coefficients of the target coding unit.
5. The method according to claim 3, It is characterized in that The inverse transformation process of the transformation coefficient of the target coding unit includes: In response to the non-zero number of transform coefficients of the target coding unit satisfying a second condition, performing inverse transform processing on the transform coefficients of the target coding unit using a transform kernel indicated by the second condition; or, In response to a second flag bit of the target coding unit obtained by decoding from a bitstream, performing an inverse transform skipping process on a transform coefficient of the target coding unit through a transform core indicated by the second flag bit; or, The inverse transform skipping process is performed on the transform coefficients of the target coding unit through the specified transform kernel.
6. The method according to claim 3, It is characterized in that The dequantizing the residual coefficient of the target coding unit to obtain the transform coefficient of the target coding unit includes: Adjusting the quantization parameter according to a specified quantization parameter adjustment method, wherein the quantization parameter adjustment method includes increasing or decreasing; Based on the adjusted quantization parameter, the residual coefficient of the target coding unit is dequantized to obtain the transform coefficient of the target coding unit.
7. The method according to claim 6, It is characterized in that Before obtaining a reconstructed signal obtained by decoding the target coding unit based on the residual signal of the target coding unit and the prediction signal of the target coding unit, the method further includes: Based on the adjusted quantization parameter, obtaining an adjusted matching threshold, wherein the matching threshold is a threshold used to indicate whether pixels match during string prediction; Based on the adjusted matching threshold, a prediction signal of the target coding unit is obtained through string prediction.
8. The method according to any one of claims 1 to 7, It is characterized in that The residual coding of the target coding unit includes at least two sub-coding blocks; The decoding of the residual coding of the target coding unit to obtain the residual coefficient of the target coding unit includes: Skip scanning of an all-zero sub-coding block among the at least two sub-coding blocks, perform coefficient decoding based on a scanning area on a non-all-zero sub-coding block among the at least two sub-coding blocks, and obtain residual coefficients of the target coding unit.
9. The method according to claim 8, It is characterized in that The step of skipping scanning of an all-zero sub-coding block in the at least two sub-coding blocks, performing coefficient decoding based on a scanning area on a non-all-zero sub-coding block in the at least two sub-coding blocks, and obtaining a residual coefficient of the target coding unit includes: Skip scanning of the all-zero sub-coding block in the at least two sub-coding blocks, take the width of the target coding block as a basic unit, perform coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the horizontal direction, and obtain the residual coefficient of the target coding unit.
10. The method according to claim 8, It is characterized in that The step of skipping scanning of an all-zero sub-coding block in the at least two sub-coding blocks, performing coefficient decoding based on a scanning area on a non-all-zero sub-coding block in the at least two sub-coding blocks, and obtaining a residual coefficient of the target coding unit includes: Skip scanning of the all-zero sub-coding block in the at least two sub-coding blocks, take the height of the target coding block as a basic unit, perform coefficient decoding based on the scanning area on the non-all-zero sub-coding block by scanning in the vertical direction, and obtain the residual coefficient of the target coding unit.
11. A video image encoding method, It is characterized in that The method comprises: Obtaining an original signal of an unencoded target coding unit; Based on the reference signal, the original signal of the target coding unit is predicted by a string prediction method to obtain a prediction signal of the target coding unit and a residual signal of the target coding unit; Acquire residual coding of the target coding unit based on the residual signal of the target coding unit; The motion information corresponding to the prediction signal of the target coding unit and the residual coding of the target coding unit are added to the encoded video code stream, so that during decoding, the residual decoding indication information of the target coding unit is obtained, and the residual decoding of the target coding block is performed in response to the residual decoding indication information indicating the residual decoding of the target coding block, the residual coding of the target coding unit is obtained, the residual coding of the target coding unit is decoded to obtain the residual coefficient of the target coding unit, the residual signal of the target coding unit is obtained based on the residual coefficient of the target coding unit, and the reconstructed signal of the target coding unit is obtained based on the residual signal of the target coding unit and the prediction signal of the target coding unit, and the residual decoding indication information indicates whether to decode the residual of the target coding block.
12. A video image reconstruction device, It is characterized in that The device comprises: A residual coding acquisition module, configured to obtain residual decoding indication information of an encoded target coding unit in response to encoding and decoding in a string prediction manner; the residual decoding indication information indicates whether to perform residual decoding on a target coding block; in response to the residual decoding indication information indicating to perform residual decoding on the target coding block, obtain residual coding of the target coding unit; the residual coding is obtained by encoding residual coefficients of the target coding block; A coefficient decoding module, used for decoding the residual coding of the target coding unit to obtain the residual coefficient of the target coding unit; A residual signal acquisition module, configured to acquire a residual signal of the target coding unit based on a residual coefficient of the target coding unit; A signal reconstruction module is used to obtain a reconstructed signal of the target coding unit based on a residual signal of the target coding unit and a prediction signal of the target coding unit.
13. A computer device, It is characterized in that The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Decoding method, coding method, decoding device and coding device
CN106998470A