Decoding method, encoding method, decoding apparatus, encoding apparatus, device, medium, and program product
By using a dual-stream encoding and decoding method, additional encoding and detail enhancement processing are performed on specified areas of the video frame at the decoding end, which solves the problem of decoding quality degradation caused by quantization in video encoding and improves the reconstruction quality and encoding and decoding efficiency of the video frame.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-09-15
- Publication Date
- 2026-05-07
AI Technical Summary
The poor video decoding quality caused by quantization during video encoding is mainly due to the large difference between the reconstructed pixels and the original pixels, which reduces the quality of the video frames.
A dual-stream encoding and decoding method is adopted. The encoding end performs additional encoding on a specified area of the video frame to generate a detail stream, and the decoding end performs image reconstruction and detail enhancement processing to restore the lost details in the video frame.
It improves the reconstruction quality of pixels within a specified area in a video frame, reduces the accuracy loss caused by prediction residuals and quantization during video encoding, and improves encoding and decoding efficiency.
Smart Images

Figure CN2025121293_07052026_PF_FP_ABST
Abstract
Description
Decoding methods, encoding methods, devices, equipment, media and program products
[0001] This application claims priority to Chinese Patent Application No. 2024115530184, filed on October 31, 2024, entitled "Decoding Method, Encoding Method, Apparatus, Device, Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of audio and video technology, and more particularly to the field of video encoding and decoding, specifically to a decoding method, an encoding method, a decoding device, an encoding device, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0003] Video encoding and decoding refers to the process of encoding and decoding data streams; encoding converts the data stream into a compressed format for storage and transmission, while decoding restores the compressed bitstream to the original data format for playback and processing.
[0004] Currently, video encoding processes involve quantizing the residual information of video frames. This quantization operation is lossy, primarily sacrificing some information to make the quantized signal more suitable for compression. However, due to the effects of quantization, the residual signal reconstructed from the compressed bitstream during video decoding differs from the difference between the original pixels and the predicted pixels of the video frame. This difference degrades the quality of the reconstructed pixels during video decoding, resulting in poor video decoding quality. Summary of the Invention
[0005] This application provides a decoding method, encoding method, apparatus, device, medium, and program product that can enhance details in video frames, improve prediction accuracy, and thus improve encoding and decoding efficiency.
[0006] On one hand, embodiments of this application provide a decoding method, which includes:
[0007] The video frame bitstream data is obtained, which includes the original bitstream and detail bitstream of the video frame. The original bitstream is obtained by encoding the video frame. The detail bitstream is obtained by encoding a specified region in the video frame. The specified region refers to the region in the video frame that needs to be enhanced in terms of detail.
[0008] Image reconstruction processing is performed on the detail stream to obtain the reconstructed detail image of the video frame; the reconstructed detail image contains the reconstructed pixel information of pixels in a specified region;
[0009] The original bitstream is decoded to obtain the reconstructed original image of the video frame;
[0010] Based on the reconstructed detail image, a specified region in the original reconstructed image is subjected to detail enhancement processing to obtain the reconstructed frame image of the video frame.
[0011] On the other hand, embodiments of this application provide a decoding device, which includes:
[0012] The acquisition unit is used to acquire the bitstream data of video frames. The bitstream data includes the original bitstream and the detail bitstream of the video frame. The original bitstream is obtained by encoding the video frame. The detail bitstream is obtained by encoding a specified region in the video frame. The specified region refers to the region in the video frame that needs to be enhanced in terms of detail.
[0013] The processing unit is used to perform image reconstruction processing on the detail stream to obtain the reconstructed detail image of the video frame; the reconstructed detail image contains the reconstructed pixel information of pixels in a specified region;
[0014] The processing unit is also used to decode the original bitstream to obtain the reconstructed original image of the video frame;
[0015] The processing unit is also used to perform detail enhancement processing on a specified region in the reconstructed original image based on the reconstructed detail image, so as to obtain the reconstructed frame image of the video frame.
[0016] In this embodiment, the bitstream data includes the original bitstream and detail bitstream of the video frame. The original bitstream is obtained by encoding the entire video frame at the encoding end, while the detail bitstream is obtained by encoding a specified region within the video frame that requires detail enhancement. In other words, during encoding, additional encoding is performed on specified regions within the video frame that may suffer from encoding loss, and this is transmitted to the decoding end to avoid information loss caused by encoding only the video frame. Thus, the decoding end performs image reconstruction processing using the detail bitstream of the video frame, obtaining a reconstructed detail image of the video frame. This reconstructed detail image contains reconstructed pixel information of the pixels within the specified region requiring detail enhancement. That is, by reconstructing the detail image from the encoding process, the detailed content of the regions (i.e., the specified regions) lost during the encoding process of the video frame can be restored. Simultaneously, the decoding end also reconstructs the reconstructed original image of the video frame using the original bitstream, thus achieving the reconstruction of the entire video frame. Thus, the decoding end can perform detail enhancement processing on a specified region in the reconstructed original image based on the reconstructed detail image to obtain the reconstructed frame image of the video frame. Since the reconstructed frame image is obtained by optimizing the pixels in the specified region of the reconstructed original image using the reconstructed detail image, the reconstructed pixel information of the pixels contained in the reconstructed frame image can better restore the pixel information of the corresponding pixels in the original video frame, improving the reconstruction quality of the pixels in the specified region. Therefore, this embodiment of the application, through a dual-stream encoding and decoding method, uses the detail stream to perform detail enhancement on the specified region to be enhanced in the reconstructed video frame during the process of restoring the entire video frame from the original bitstream at the decoding end. This effectively improves the reconstruction quality of the pixels in the specified region that are lost during the encoding of the video frame, thereby reducing the accuracy loss caused by prediction residuals and quantization during video encoding and improving encoding and decoding efficiency.
[0017] In another aspect, embodiments of this application provide an encoding method, which includes:
[0018] Acquire the video frames to be encoded, and encode the video frames to obtain the original bitstream of the video frames; and,
[0019] By identifying a specified region within a video frame, a detailed image of the video frame can be obtained; the specified region refers to the area in the video frame where the detail needs to be enhanced.
[0020] The detail images are encoded to obtain the detail bitstream of the video frames;
[0021] The original bitstream and detail bitstream of the video frame are sent to the decoding end for joint image decoding.
[0022] Furthermore, embodiments of this application provide an encoding device, which includes:
[0023] An acquisition unit is used to acquire video frames to be encoded and to encode the video frames to obtain the original bitstream of the video frames; and,
[0024] The processing unit is used to determine a specified region in a video frame and obtain a detailed image of the video frame; the specified region refers to the area in the video frame that needs to be enhanced in terms of detail.
[0025] The processing unit is also used to encode the detail images to obtain the detail bitstream of the video frames;
[0026] The processing unit is also used to send the original bitstream and detail bitstream of the video frame to the decoding end for joint image decoding.
[0027] In this embodiment, the encoding end supports multiple filtering strategies to select target pixels with residual information greater than a residual threshold from the video frame, thereby determining the specified regions in the video frame that require detail enhancement. These multiple filtering strategies better meet the encoding requirements of users, are applicable to different encoding scenarios, and improve the encoding experience. Furthermore, by performing additional / secondary encoding only on pixels within the specified regions requiring detail enhancement at the encoding end, not only is significant resource waste avoided, but detail enhancement is also achieved, significantly improving encoding quality and efficiency. Correspondingly, at the decoding end, during the process of reconstructing the entire video frame from the original bitstream, the detail bitstream is used to enhance the specified regions in the reconstructed video frame, effectively improving the reconstruction quality of pixels within the specified regions that suffer losses during the encoding process. This reduces the accuracy loss caused by prediction residuals and quantization during video encoding, thereby improving encoding and decoding efficiency.
[0028] On the other hand, embodiments of this application provide a computer device, which includes:
[0029] A processor is used to load and execute computer programs;
[0030] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described encoding and decoding methods.
[0031] On the other hand, this application provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed by the above-described encoding and decoding methods.
[0032] On the other hand, this application provides a computer program product comprising a computer program that, when executed by a processor, causes a computer device to perform the aforementioned encoding and decoding methods. Attached Figure Description
[0033] Figure 1 is a schematic diagram of the encoding process of an encoder encoding an image frame;
[0034] Figure 2 is a schematic diagram of a complete video encoding and decoding process;
[0035] Figure 3 is a schematic diagram of the orientation of the reference pixel of the current coding block during intra-frame prediction;
[0036] Figure 4 is a flowchart illustrating a partial detail enhancement scheme for a frame image provided in an exemplary embodiment of this application;
[0037] Figure 5 is a schematic diagram of the architecture of a video encoding and decoding system provided in an exemplary embodiment of this application;
[0038] Figure 6 is a flowchart illustrating a decoding method provided in an exemplary embodiment of this application;
[0039] Figure 7 is a schematic diagram of a detail enhancement process for reconstructing an original image based on a reconstructed detail image provided in an exemplary embodiment of this application;
[0040] Figure 8 is a flowchart illustrating an encoding method provided in an exemplary embodiment of this application;
[0041] Figure 9 is a schematic diagram of determining a specified region from a video frame based on the residual information of pixels, provided by an exemplary embodiment of this application;
[0042] Figure 10a is a schematic diagram of filtering target pixels from a video frame based on residual comparison results, provided by an exemplary embodiment of this application;
[0043] Figure 10b is a schematic diagram of a process for jointly selecting target pixels based on brightness value and residual information, provided in an exemplary embodiment of this application.
[0044] Figure 10c is a schematic diagram of a process for filtering target pixels based on the presence of non-zero residual information around the current pixel, provided by an exemplary embodiment of this application.
[0045] Figure 10d is a schematic diagram of determining target pixel points based on a difference image according to an exemplary embodiment of this application;
[0046] Figure 11a is a schematic diagram of identifying a specified region from a video frame according to an exemplary embodiment of this application;
[0047] Figure 11b is a schematic diagram of another method for identifying a specified region from a video frame, provided by an exemplary embodiment of this application.
[0048] Figure 12 is a schematic diagram of a decoding device provided in an exemplary embodiment of this application;
[0049] Figure 13 is a schematic diagram of an encoding device provided in an exemplary embodiment of this application;
[0050] Figure 14 is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0051] This application provides a video encoding and decoding scheme based on video encoding and decoding technology, specifically including an encoding scheme and a decoding scheme for video frames. To more clearly understand the technical solution provided by this application, the key terms involved in this application are first introduced below:
[0052] 1. Video.
[0053] A video is a file composed of at least two video frames (or image frames) linked together in sequence; that is, a video frame is the smallest or most basic unit of a video; in other words, a video is a dynamic picture composed of a series of consecutive video frames, and each video frame is a still image that makes up the video.
[0054] When playing video, multiple video frames are output continuously in the order of their playback time. When the number of consecutive video frames exceeds 24 frames per second, the human eye perceives a smooth and continuous visual effect based on the principle of visual persistence. Video is represented as a video signal, typically an electrical signal. Transmitting video signals enables the transmission and storage of video over a network. Based on the acquisition method, video signals can be acquired through two methods: those captured by cameras and those generated by computer devices. Due to the different statistical characteristics of different video signals, their corresponding compression encoding methods may also differ.
[0055] II. Video encoding and decoding technology.
[0056] Video encoding technology refers to the encoding method that uses compression technology to convert a video file in one format into another. Specifically, video encoding and decoding technology is based on two processes: encoding and decoding. Decoding is the reverse of encoding. Encoding supports the conversion of video frames into a data format that is easier to transmit, aiming to reduce file size and facilitate storage and transmission. Decoding, being the reverse of encoding, allows the compressed data to be restored to its original format for playback on a display.
[0057] The following is an introduction to the current mainstream video coding technologies:
[0058] Modern mainstream video coding technologies, taking international video coding standards such as HEVC (High Efficiency Video Coding) (e.g., HEVC / H.265), VVC (Versatile Video Coding) (e.g., VVC / H.266), and AVS (Audio Video Coding Standard) as examples, employ a hybrid coding framework to perform the following series of operations and processing on the input raw video signal:
[0059] 1) Block partition structure: Based on the size of the input image (such as video frames to be compressed, encoded, or decoded in a video), the input image is divided into several non-overlapping processing units. During encoding and decoding, similar compression operations can be performed on each processing unit, avoiding the difficulties of directly encoding and decoding a single frame. These partitioned processing units can be called CTUs (Coding Tree Units) or LCUs (Largest Coding Units). CTUs can be further subdivided into one or more basic coding units, called CUs (Coding Units or Coding Blocks). Each CU is the most basic element in an encoding / decoding process; subsequent embodiments of this application will use each CU as an example to illustrate the encoding and decoding process.
[0060] 2) Predictive Coding: Predictive coding is based on the correlation between discrete signals (such as the spatial correlation between pixels in different parts of the same video frame, or the temporal correlation between pixels in different video frames). It uses one or more previous signals to predict the value of the current signal, thereby encoding the residual (or prediction error) between the actual value and the predicted value of the current signal. This avoids the high computational complexity and waste of compression resources caused by directly compressing all video frames.
[0061] Predictive coding mainly includes intra-frame prediction and inter-frame prediction. Specifically: a) Intra-frame prediction: predicts that the current coding unit's prediction signal comes from a previously encoded and reconstructed region within the same image; b) Inter-frame prediction: predicts that the current coding unit's prediction signal comes from another image that has already been encoded and is different from the image to which the current coding unit belongs (this other image can be called a reference image). During video encoding and decoding, when encoding a unit to be encoded (such as the aforementioned CU) in the original video signal (such as a video frame) at the encoding end, if any predictive coding method (such as intra-frame prediction or inter-frame prediction) is used, the reconstructed video signal from the original video signal (e.g., when the predictive coding method is intra-frame prediction, the reconstructed video signal belongs to the current image; when the predictive coding method is inter-frame prediction, the reconstructed video signal comes from an image previously reconstructed by the current image) needs to be used to predict the unit to be encoded, obtaining the residual video signal of the current unit to be encoded (such as the aforementioned residual). Then, the residual video signal is compressed and encoded to generate a bitstream, which is then transmitted to the decoding end. Correspondingly, the encoding end also needs to inform the decoding end of any predictive coding method used in the encoding process, so that after receiving the encoded bitstream (i.e., the bitstream mentioned above, or image bitstream, video bitstream, compressed bitstream, etc.), the decoding end can use the same predictive coding method as the encoding process to reconstruct the image during the decoding process of the encoded bitstream.
[0062] 3) Transform Coding and Quantization: The residual video signal undergoes transform operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform) to be converted into the transform domain, where these coefficients are called transform coefficients. This allows for further lossy quantization of the signal in the transform domain, losing some redundant information and making the quantized signal more suitable for compression. Some video coding standards may use one or more transform methods; therefore, during video encoding and decoding, the encoder needs to select a transform method for the current encoding CU and inform the decoder of this method, enabling the decoder to perform the inverse transform using the corresponding transform method during decoding. It is worth noting that the fineness of the quantization operations mentioned above is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a wider range of values will be quantized into the same output, which usually leads to greater distortion and a lower bit rate (i.e., the number of bits of data transmitted per unit time). Conversely, a smaller QP value means that coefficients with a smaller range of values will be quantized into the same output, which usually leads to less distortion and a higher bit rate.
[0063] 4) Entropy Coding or Statistical Coding: The quantized transform domain signal will undergo statistical compression coding (i.e., statistical compression coding based on the frequency of each value) to finally output a binary (0 or 1) compressed bitstream. Simultaneously, other information generated during encoding, such as the selected mode (e.g., prediction mode) and motion vectors, also requires entropy coding to reduce the bit rate. The aforementioned statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Statistical coding can include, but is not limited to, Variable Length Coding (VLC) or Content Adaptive Binary Arithmetic Coding (CABAC).
[0064] 5) Loop Filtering: Based on the image encoded in the preceding steps, a series of operations such as inverse quantization, inverse transform, and prediction compensation (i.e., the reverse operations of steps 2) to 4) are performed to obtain the reconstructed decoded image. Compared with the original image, the reconstructed image differs in some information due to the influence of quantization, resulting in distortion. Therefore, filtering the reconstructed image can effectively reduce the distortion caused by quantization. Filters can include, but are not limited to, deblocking filters (DF), SAO (Sampled Adaptive Offset), or ALF (Adaptive Loop Filter), etc. These filtered reconstructed images can serve as reference information for subsequent encoded images, used to predict future signals. Therefore, the above filtering operations are also called loop filtering, or filtering operations within the encoding loop.
[0065] The following section, using the video encoder shown in Figure 1, describes the basic video encoding process (steps 1 through 5) described above. In Figure 1, the current coding block to be encoded is the k-th CU in the current image frame (as shown in Figure 1). k Taking [x, y] as an example, k is a positive integer, and k is less than or equal to the total number of CUs contained in the current image frame. k [x, y] represents the pixel (or simply pixel) with coordinates [x, y] in the k-th CU, where x represents the x-coordinate of the pixel and y represents the y-coordinate of the pixel; s k After processing such as motion compensation or intra-frame prediction, the predicted signal can be obtained from [x,y]. Predicted signal and the original signal s k Perform a difference operation on [x,y] to obtain the residual video signal u. k [x,y]; then, for the residual video signal u k After transformation and quantization of [x,y], the quantized data is obtained. The quantization output data has two data flow directions:
[0066] Data Flow 1: The encoder sends the quantized output data to the entropy encoder for entropy encoding, obtaining the encoded bitstream. This bitstream is then stored in a buffer, awaiting transmission to the decoder. Upon receiving the bitstream, the decoder performs entropy decoding on each CU unit to obtain various mode information and quantized transform coefficients. These transform coefficients are then dequantized and inverse transformed to obtain the residual signal. Simultaneously, the decoder can obtain the prediction signal corresponding to the current CU unit based on the known mode information from the encoder. Adding the residual signal and the prediction signal yields the reconstructed signal. Finally, the reconstructed value (or reconstructed signal) of the decoded image is filtered through a loop filter to generate the final output signal.
[0067] Data flow direction 2: The encoding end can perform inverse quantization and inverse transform on the quantized output data to obtain the inverse transformed residual video signal u′. k [x,y]; then, the inverse-transformed residual video signal u′ k [x,y] and the predicted signal Adding them together yields a new prediction signal. And new prediction signals The new prediction signal is saved in the buffer of the current image. After intra-frame prediction processing, the following is obtained: And the new prediction signal The reconstructed signal s′ can be obtained after loop filtering. k [x,y], and reconstruct the signal s ′ k [x,y] is stored in the decoded image buffer for use in generating the reconstructed video. Reconstructed signal s′ k [x,y] is obtained after motion compensation prediction processing. in It can represent a reference block, m x and m y These represent the horizontal and vertical components of the motion vector of the reference block, respectively.
[0068] As mentioned earlier, the decoding process is the reverse of the encoding process. Therefore, the decoding process of the compressed bitstream obtained by the encoding end based on the aforementioned steps can be roughly described as follows: On the one hand, for each coded block (such as those divided according to the encoding end's block partitioning method), after obtaining the compressed bitstream, the decoding end first performs entropy decoding to obtain various mode information generated during the encoding process and quantized transform coefficients, etc. The decoding end then performs inverse quantization and inverse transform on the data obtained from the entropy decoding based on the decoded mode information and transform coefficients, to obtain the residual information (or residual signal) of the current coded block. On the other hand, the decoding end predicts the current coded block based on the known encoding mode information (such as the prediction mode used by the encoding end when encoding the current coded block) to obtain the prediction information corresponding to the current coded block. In this way, the decoding end can add the residual information and prediction information of the current coded block to obtain the reconstructed signal of the current coded block; the reconstructed signals of all coded blocks corresponding to the video frame constitute the decoded image of the video frame, which includes the reconstructed pixel information (or reconstructed value) of each pixel. Then, the reconstructed values of pixels in the decoded image need to undergo a loop filter operation to reduce the distortion caused by quantization, thereby obtaining the final restored output signal (i.e., the reconstructed image).
[0069] The complete flowchart of the video encoding and decoding processes described above can be found in Figure 2. As shown in Figure 2, after the encoding end acquires the original image to be encoded (such as a video frame in a video), it performs predictive encoding on the pixels in the original image to obtain the predicted information of the pixels in the original image. The predicted information is then subtracted from the actual pixel information of the pixel to obtain the residual information of the pixel. The encoding end performs transformation and other processing on the residual information, converting it to the transform domain to obtain the signal in the transform domain (which can be called the transform coefficients). The transform coefficients are then quantized, losing some information to obtain quantization coefficients that are beneficial for compression. The quantization coefficients are then entropy encoded to obtain a binary (0 or 1) compressed bitstream. The encoding end transmits the compressed bitstream to the decoding end, which performs entropy encoding on the compressed bitstream to convert the binary compressed bitstream into quantization coefficients. The decoding end then performs inverse quantization on the quantization coefficients to obtain transform coefficients, and then performs inverse transform on the transform coefficients to obtain the residual information of the current encoded block. Simultaneously, the decoding end acquires the encoding mode information and predicts the prediction information of the current encoding block based on the encoding mode information. In this way, the decoding end predicts the information based on the residual information of the current encoding block, and thus obtains the reconstruction information of the current encoding block. The above decoding process is performed on each encoding block in the video frame to obtain the reconstructed image corresponding to the video frame.
[0070] Based on the encoding and decoding process shown in Figure 2, it is easy to see that due to the effects of predictive coding and quantization in the encoding process, there will be a difference between the reconstructed image and the original image, which greatly reduces the image quality of the frame image. For example, as shown in Figure 3, the prediction mode used in the encoding process is intra-frame prediction. When predictive coding is performed on the current block (such as coding block CU or LCU), the reference pixels of the current coding block come from the left region 301 and the upper region 302 of the current coding block in the video frame. For a pixel in the left and top regions of the current block (i.e., a pixel in the current block that is close to the left region 301 and the top region 302), such as pixel 303, since pixel 303 is close to the reference pixel (i.e., a pixel located in the left region 301 and the top region 302), the two are statistically strongly correlated. That is, the pixel information of the reference pixel is more reliable for the pixel information of pixel 303. Therefore, the prediction information of pixel 303 obtained based on the pixel information of the pixels in the left region 301 and the top region 302 is more accurate, and the residual information generated by the prediction of pixel 303 is relatively small. The residual information can be called the absolute value of the residual, or the absolute value of the residual, which is equal to |true pixel information - predicted information|.
[0071] Conversely, for pixels located in the right and bottom regions of the current block, such as pixel 304, since pixel 304 is far from the reference pixel, the statistical correlation between the two is weak. This means the pixel information from the reference pixel is more reliable for pixel 304. Therefore, the prediction information for pixel 304 obtained based on the pixel information in edge region 301 and upper region 302 is inaccurate, resulting in a relatively large residual information for pixel 304. Furthermore, considering that when quantizing the residual information of a pixel, the larger the residual information (e.g., the larger the absolute value of the residual), the more information is lost during quantization, resulting in a greater loss of accuracy. Therefore, for the current block using intra-frame prediction mode, when the quantization step size is large (corresponding to a large quantization parameter), the loss caused by quantization is not evenly distributed, but has a more significant impact on the right and bottom regions of the current block. Thus, when reconstructing the current block at the decoding end, the degree of reconstruction of pixels in the left and top regions of the current block is greater than that of pixels in the right and bottom regions of the current block. In other words, the details in the right and bottom regions of the current block need to be enhanced.
[0072] To reduce the accuracy loss caused by predictive coding and quantization in certain areas of video frames, the video encoding and decoding scheme provided in this application is specifically a frame image detail enhancement scheme. This scheme can correct problems such as image quality degradation caused by detail loss after video frame compression, thereby improving the user's subjective perception of video decoding quality. For example, this scheme supports adjusting and optimizing the prediction residual of a certain area (such as the right and bottom areas of the current block) by identifying or deriving the prediction residual adjustment value of the pixels to be enhanced in the current block using intra-frame prediction mode (pixels in the right or bottom area of the current block as shown in Figure 3), thereby reducing the overall block prediction residual, greatly improving the prediction accuracy of intra-frame prediction, and thus improving coding efficiency.
[0073] For example, a schematic diagram of the general process of the frame image detail enhancement scheme provided in this application embodiment can be seen in Figure 4. As shown in Figure 4:
[0074] At the encoding end, the encoder acquires the video frame to be encoded. On one hand, it performs encoding processing on the video frame, specifically including block segmentation, predictive coding, transform, quantization, and entropy coding, to obtain the original bitstream of the video frame. On the other hand, during the encoding process, the encoder determines whether pixels in the video frame require detail enhancement; it then identifies the designated regions containing the pixels requiring detail enhancement, obtaining the detail image of the video frame. The encoder then encodes this detail image (using the same process as described above), obtaining the detail bitstream of the video frame. Finally, the encoder packages the original bitstream and the detail bitstream into a single bitstream data stream and sends it to the decoder for joint image decoding.
[0075] At the decoding end, the decoder acquires the bitstream data of the video frame to be decoded. This bitstream data includes the original bitstream and the detail bitstream of the video frame. On one hand, the decoder performs image reconstruction processing (i.e., decoding) on the detail bitstream of the video frame to obtain a reconstructed detail image of the video frame. This reconstructed detail image contains the reconstructed pixel information of pixels in the specified region of the video frame that needs detail enhancement. On the other hand, the decoder decodes the original bitstream of the video frame to obtain a reconstructed original image of the video frame. Decoding is the reverse process of encoding. In this way, the decoder can perform detail enhancement processing on the specified region in the reconstructed original image based on the decoded reconstructed detail image to obtain the reconstructed frame image of the video frame.
[0076] Therefore, in the process of encoding video frames, this embodiment of the application, in addition to performing traditional encoding processing on the video frames to obtain the original bitstream, also identifies the regions (i.e., designated regions) of pixels in the video frames where significant information loss occurred during the encoding process, and obtains detail images. These detail images are then also encoded to obtain detail bitstreams. Thus, at the decoding end, the reconstructed detail image based on the detail bitstream can be used to enhance the details of pixels within the designated region of the reconstructed original image based on the original bitstream. This detail enhancement aims to adjust or optimize the reconstructed pixel information of corresponding pixels within the designated region of the reconstructed original image using the reconstructed pixel information of pixels within the designated region of the reconstructed detail image. This ensures that the reconstructed pixel information of pixels within the adjusted or optimized region better matches the actual pixel information of the corresponding pixels in the original video frame, thereby improving the quality and efficiency of video encoding and decoding.
[0077] The frame image detail enhancement scheme provided in this application embodiment can be applied to any product with relevant video encoding / decoding or video compression functions; the product here may include an application program or a computer device.
[0078] Optionally, if the application has video encoding and decoding capabilities, then when the application deploys the frame image partial detail enhancement scheme provided in this application embodiment, all video frames that need to be encoded and decoded by the application can achieve partial detail enhancement effects. Here, an application can refer to a computer program that performs one or more specific tasks; according to the application's operating method, applications can include: clients installed on a terminal, small programs that can be used without downloading and installation (as subroutines of the client), web (World Wide Web) applications opened through a browser, etc.; according to the application's functional type, applications can include, but are not limited to: IM (Instant Messaging) applications, content interaction applications, etc. Instant messaging applications refer to applications that facilitate instant messaging and social interaction based on the Internet, and can include, but are not limited to: social applications with communication functions, map applications with social interaction functions, game applications, etc. Content interaction applications refer to applications capable of content interaction, such as online banking, sharing platforms, personal spaces, news applications, etc.
[0079] Optionally, a computer device can refer to a physical device with video encoding and decoding capabilities. When such a computer device is equipped with the frame image detail enhancement scheme provided in this application embodiment, all video frames encoded and decoded by the computer device can achieve partial detail enhancement effects. The type of computer device can include a terminal or a server. A terminal can include, but is not limited to, smartphones (such as smartphones running the Android system or smartphones running the Internetworking Operating System (IOS), tablets, portable personal computers, mobile internet devices (MIDs), in-vehicle devices, head-mounted devices, etc. This application embodiment does not limit the type of terminal device, which is stated here. A server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0080] It should be understood that the above is merely a brief introduction to the possible product forms to which the frame image partial detail enhancement scheme provided in the embodiments of this application may be applied; in practical applications, the embodiments of this application do not limit the products to which the frame image partial detail enhancement scheme may be applied. For example, the frame image partial detail enhancement scheme provided in the embodiments of this application can also be deployed in the form of a plug-in in an application or a computer device. For ease of explanation, the following description will take the deployment of the frame image partial detail enhancement scheme in an application running on a computer device as an example, and this is hereby explained.
[0081] A schematic diagram of the architecture of a video encoding and decoding system based on a frame image partial detail enhancement scheme can be seen in Figure 5. As shown in Figure 5, it is assumed that the frame image partial detail enhancement scheme is deployed in a social application, and the social application runs on a computer device—a terminal. The terminals in the video encoding and decoding system may include terminal 501 and terminal 502, and the video encoding and decoding system also includes server 503. This application embodiment does not limit the number and type of terminals and servers in the video encoding and decoding system. Among them, terminal 501 is the terminal device held by user 1, terminal 502 is the terminal device held by user 2, and user 1 and user 2 are two users who have established a communication session in the social application. Server 503 is the backend device corresponding to terminal 501 and terminal 502, mainly providing backend technology and services for the social application in terminal 501 and terminal 502. Among them, the terminals (terminal 501 and terminal 502) and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.
[0082] In a specific implementation, in the video encoding and decoding system shown in Figure 2, if user 1 wants to share a video with user 2, user 1 can send the video through the social conversation page displayed on the screen of terminal 501. After receiving the video, terminal 501 encodes the video using the frame image partial detail enhancement scheme provided in this application embodiment to obtain a compressed bitstream. The compressed bitstream includes the bitstream data of each video frame in the video. Terminal 501 sends the compressed bitstream to server 503, which forwards it to terminal 502. After receiving the compressed bitstream, terminal 502 decodes the bitstream data of each video frame in the compressed bitstream, specifically by using the frame image partial detail enhancement scheme provided in this application embodiment to decode the bitstream data, obtaining a reconstructed frame image of the video frame. Therefore, by using the dual bitstream method provided in this application embodiment during video encoding and decoding, additional encoding of the details to be enhanced in the video frame can improve the degree to which some details in the video frame are restored to their original image effect during decoding, ensuring the quality and efficiency of video encoding.
[0083] Based on the above-described details of the frame image enhancement scheme provided in the embodiments of this application, the following points should also be noted:
[0084] (1) The block partitioning information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering (e.g., model parameters of cross-component prediction models, or template selection methods), are carried in the encoded bitstream when necessary. In this way, the decoder parses the encoded bitstream and analyzes existing information to determine the same block partitioning information, prediction, transform, quantization, entropy coding, and loop filtering mode information or parameter information as the encoder, thus ensuring that the decoded image obtained by the encoder is the same as the decoded image obtained by the decoder. It is worth noting that, depending on the compression process of the mode information or parameter information during encoding, the decoder can use, but is not limited to, two parsing methods when parsing the mode information or parameter information based on the encoded bitstream: Optionally, the mode information or parameter information can be directly obtained by parsing the bits in the encoded bitstream; for example, in the encoded bitstream, a parameter value of 1 is used to select template region 1, and a value of 0 is used to select template region 2, etc. Optionally, mode information or parameter information can be implicitly derived from the encoded bitstream. This implicit deriving can be roughly understood as parsing intermediate parameters from the encoded bitstream, performing calculations on these intermediate parameters, and deriving mode information or parameter information based on the calculation results. The decoding end can parse any mode information or parameter information from the encoded bitstream sent by the encoding end using either of the above two methods; there is no limitation on this approach.
[0085] (2) Figure 5 above is merely a schematic diagram of the architecture of an exemplary video encoding and decoding system provided in this application embodiment. In practical applications, this architecture can be adapted. For example, the server in the video encoding and decoding system is a distributed server, that is, server 503 is not a single device, but multiple servers distributed in different locations. Alternatively, server 503 may not exist in the video encoding and decoding system; instead, terminals 501 and 502 can communicate directly.
[0086] (3) The data collection and processing in this application embodiment should strictly comply with the requirements of relevant laws and regulations. The acquisition of personal information must be based on the knowledge or consent of the individual (or have a legal basis for information acquisition), and subsequent data use and processing should be carried out within the scope of laws and regulations and the authorization of the personal information subject. For example, when this application embodiment is applied to specific products or technologies, such as when the terminal sends a video, it is necessary to obtain the permission or consent of the uploader or creator of the video, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions.
[0087] Based on the above introduction to the frame image partial detail enhancement scheme and the product or scenario architecture in which it is applied, the following describes a more detailed frame image partial detail enhancement method proposed in the embodiments of this application, with reference to the accompanying drawings. Specifically, the frame image partial detail enhancement method includes an encoding method and a decoding method. The encoding method mainly describes the specific implementation process of the encoder encoding the video frame to obtain the original bitstream and detail bitstream of the video frame. The decoding method mainly describes the specific implementation process of the decoder, after receiving the original bitstream and detail bitstream of the video frame, combining the original bitstream and detail bitstream to perform joint image decoding and obtain the reconstructed frame image of the video frame. For ease of explanation, the encoding method and decoding method will be described in detail with different embodiments below, which are hereby explained.
[0088] Please refer to Figure 6, which is a flowchart illustrating a decoding method provided in an exemplary embodiment of this application; the flowchart shown in Figure 6 is a flowchart of the decoding end, which can be executed by a computer device held by the decoding end; the method may include, but is not limited to, steps S601-S604:
[0089] S601: Obtain the bitstream data of video frames.
[0090] A video frame refers to the current video frame to be decoded among multiple video frames. After receiving the compressed video bitstream from the encoder, the decoder decodes each video frame sequentially according to the arrangement order (or playback order) of the video frames in the video. When the decoder needs to decode any video frame, it obtains the bitstream data of that video frame from the compressed video bitstream.
[0091] The video frame bitstream data includes the original bitstream and the detail bitstream. The original bitstream of the video frame can be understood as the bitstream obtained by the encoder after encoding the entire video frame. The detail bitstream of the video frame can be understood as the bitstream obtained by the encoder after encoding a specified area in the video frame. The specified area refers to the area in the video frame that needs to be enhanced in detail, that is, the image area where information / precision is greatly lost during the encoding process of the video frame.
[0092] S602: Perform image reconstruction processing on the detail bitstream to obtain the reconstructed detail image of the video frame.
[0093] After acquiring the detail bitstream of a video frame, the decoder first decodes the detail bitstream to obtain the residual information of pixels in a specified region of the reconstructed detail image. This residual information is obtained by the encoder subtracting the actual pixel information and the predicted information of each pixel, and then encoding it into the detail bitstream. Specifically, the decoder's decoding process for the detail bitstream of the video frame includes: entropy decoding to obtain the encoding mode information used by the encoder when encoding pixels in the specified region of the video frame, as well as the quantized coefficients; entropy decoding is the inverse process of entropy encoding mentioned above, aiming to restore the binarized detail bitstream to the quantized coefficients. The decoder then performs inverse quantization on the quantized coefficients to obtain transform coefficients; inverse quantization is the inverse process of quantization mentioned above, aiming to convert the quantized coefficients into transform coefficients in the transform domain. The decoder then performs inverse transform on the transform coefficients to obtain the residual information of each pixel in the specified region of the video frame; inverse transform is the inverse process of transform mentioned above, aiming to convert the transform coefficients into residual signals (i.e., residual information).
[0094] As mentioned earlier, the designated region refers to the area in the video frame to be enhanced in terms of detail, not the entire video frame. Therefore, to facilitate the decoding end's understanding of the location of the designated region within the video frame, this embodiment supports the encoding end encoding the location of the designated region in the video frame as an identifier into the detail stream. Thus, the decoding end can also obtain the identifier by decoding the detail stream of the video frame. This identifier indicates the location of the designated region in the video frame. For example, the identifier can be expressed as coordinates of the two vertices of a diagonal line in the video frame. The decoding end can determine the location of the designated region from the video frame based on the location indicated by the identifier, and then decode the pixels within that region to obtain the residual information of the pixels within the designated region in the reconstructed detail image. Therefore, by adding an identifier to the detail stream, the decoding end can quickly understand the location of the designated region in the video frame, enabling secure and rapid transmission of the location of the designated region between the encoding and decoding ends. It should be noted that, in addition to adding an identifier to the detail bitstream, the position of a specified region in a video frame can also be transmitted by offline communication between the decoder and encoder to transmit the position information of the specified region (i.e., the position of the specified region in the video frame); the embodiments of this application do not limit the method of transmitting the position information of the specified region between the encoder and decoder.
[0095] Then, the decoder obtains the encoding mode information from decoding the detail stream. This encoding mode information mainly includes the prediction mode (such as intra-frame prediction mode or inter-frame prediction mode) used by the encoder when predictively encoding pixels in a specified region of the video frame. In this way, the decoder can use the same predictive coding method as the encoder to perform prediction processing on the pixels in the specified region according to this encoding mode information, and obtain the prediction information of the pixels in the specified region in the reconstructed detail image to be reconstructed.
[0096] Finally, after obtaining the residual information and prediction information of the pixels in the specified region in the reconstructed detail image based on the above steps, the decoder can obtain the reconstructed detail image of the video frame based on the residual information and prediction information of the pixels in the specified region. Specifically, the residual information and prediction information of the pixels in the specified region are added together to obtain the reconstructed pixel information of the corresponding pixels, thereby obtaining the reconstructed detail image of the video frame.
[0097] Therefore, this embodiment of the application identifies a designated area in a video frame where pixels with significant information loss during encoding processing are located, obtaining a detail image. This detail image is then subjected to additional encoding processing to obtain a detail bitstream. In this way, the detail bitstream can be reconstructed at the decoding end, thereby restoring the reconstructed detail image of the designated area in the video frame. This allows for the supplementation of pixels in the designated area of the video frame where information loss due to encoding occurs, thus improving the quality and efficiency of video encoding and decoding.
[0098] S603: Decode the original bitstream to obtain the reconstructed original image of the video frame.
[0099] Specifically, the decoder's decoding process for the original bitstream of the video frame includes: First, entropy decoding of the original bitstream to obtain the encoding mode information used by the encoder when encoding each pixel in the video frame, as well as the quantized coefficients; then, inverse quantization of the quantized coefficients to obtain transform coefficients; then, inverse transform of the transform coefficients to obtain the residual information of each pixel in the video frame. Second, the decoder acquires the encoding mode information obtained from decoding the original bitstream. This encoding mode information mainly includes the prediction mode used by the encoder when predictively encoding each pixel in the video frame; the decoder uses the same predictive coding method as the encoder to predict each pixel in the video frame based on this encoding mode information to obtain the prediction information of each pixel in the reconstructed original image. Finally, based on the residual information and the prediction information of each pixel in the video frame, the decoder obtains the reconstructed original image of the video frame; specifically, the residual information and the prediction information of each pixel in the video frame are added together to obtain the reconstructed pixel information of the corresponding pixel, thus obtaining the reconstructed original image of the video frame.
[0100] It should be noted that the specific implementation process of the decoder in decoding the original bitstream of the video frame is similar to the specific implementation process of the decoder in reconstructing the image from the detail bitstream of the video frame. The above only provides a brief introduction to the decoding process of the original bitstream; for details, please refer to the aforementioned description of the decoding process for the detail bitstream, which will not be repeated here. Furthermore, the decoder can perform decoding on the original bitstream first, or it can perform image reconstruction on the detail bitstream first. That is, the execution order of steps S602 and S603 by the decoder in this embodiment is not limited.
[0101] S604: Perform detail enhancement processing on a specified region in the original reconstructed image based on the reconstructed detail image to obtain the reconstructed frame image of the video frame.
[0102] After obtaining the reconstructed detail image and the reconstructed original image of the video frame based on the aforementioned steps, considering that the reconstructed detail image includes the reconstructed pixel information of pixels within the area to be enhanced (i.e., the specified area) in the reconstructed original image, the decoder can utilize the reconstructed pixel information of pixels within the specified area in the reconstructed detail image to optimize the reconstructed pixel information of corresponding pixels within the specified area in the reconstructed original image. Optimization refers to adjusting the reconstructed pixel information of corresponding pixels within the specified area in the reconstructed original image according to the direction in which the reconstructed pixel information of corresponding pixels within the specified area in the reconstructed original image approaches the true pixel information of corresponding pixels within the specified area in the original video frame. Therefore, even if the encoder loses some information due to quantization during the encoding of the video frame, this embodiment, by performing additional encoding on the specified area in the video frame where information loss is significant, ensures that the decoder can still obtain most of the information that may have been lost during the encoding process in the specified area. This further improves the prediction accuracy of the reconstructed pixel information of pixels within the specified area in the video frame while ensuring relatively accurate reconstructed pixel information for pixels outside the specified area, thereby improving the overall reconstruction quality of the reconstructed frame image.
[0103] It is important to note that the image size of the reconstructed detail image is the same as that of the original reconstructed image, and the location of the specified region in the reconstructed detail image is the same as that in the original reconstructed image. There is a one-to-one correspondence between the pixels in the reconstructed detail image and the pixels in the original reconstructed image. This one-to-one correspondence means that a pixel within a specified region in the reconstructed detail image and a pixel within a specified region in the original reconstructed image both correspond to the corresponding pixel (i.e., at the same position) within the specified region of the video frame. Based on this, the decoding end performs detail enhancement / optimization processing on the specified region in the original reconstructed image according to the reconstructed detail image. Specifically, it overlays the reconstructed pixel information of the pixels within the specified region in the reconstructed detail image onto the reconstructed pixel information of the corresponding pixels within the specified region in the original reconstructed image; simultaneously, it retains the reconstructed pixel information of the pixels in the original reconstructed image that were not overlaid, thus obtaining the reconstructed frame image corresponding to the video frame.
[0104] As shown in Figure 7, the specified region 701 in the video frame is in the same position in the reconstructed original image 702 and the reconstructed detail image 703. When the decoding end performs detail enhancement processing on the specified region in the reconstructed original image 702 based on the reconstructed detail image 703, it specifically superimposes the reconstructed pixel information of the pixel point (such as pixel point 704) in the specified region in the reconstructed detail image 703 with the reconstructed pixel information of the corresponding pixel point (such as pixel point 705) in the specified region in the reconstructed original image 702 to obtain the new reconstructed pixel information of pixel point 705. Meanwhile, the reconstructed original image 702 also includes pixels outside the specified area (such as pixel 706), and retains the reconstructed pixel information of pixel 706 to obtain the reconstructed frame image corresponding to the video frame; assuming that the reconstructed pixel information of pixel 704 is the first reconstructed pixel information and the reconstructed pixel information of pixel 705 is the second reconstructed pixel information, then the reconstructed pixel information of pixel 705 in the specified area of the reconstructed frame image = the first reconstructed pixel information + the second reconstructed pixel information, and the reconstructed pixel information of pixel 706 outside the specified area of the reconstructed frame image is the reconstructed pixel information of pixel 706 in the reconstructed original image.
[0105] It's worth noting that the decoding end performs detail enhancement processing on a specified region in the reconstructed original image based on the reconstructed detail image. Besides the direct overlay method mentioned above, other methods can also be used. For example, a weighted overlay method can be used to optimize the reconstructed pixel information of pixels within a specified region in the reconstructed original image. That is, the reconstructed pixel information of pixels within a specified region in the reconstructed detail image is overlaid onto the corresponding reconstructed pixel information within the specified region in the reconstructed original image at a certain ratio. This ratio can be compressed into the detail bitstream at the encoding end; thus, the decoding end can directly obtain this ratio by decoding the detail bitstream. When optimizing the reconstructed pixel information of pixels within a specified region in the reconstructed original image using a weighted overlay method, the difference between the reconstructed pixel information of the pixels within the specified region in the reconstructed original image and the actual pixel information of the pixel in the video frame can be minimized. This avoids errors caused by excessive overlay of reconstructed pixel information and improves the quality of the reconstructed pixel information.
[0106] In this embodiment, the decoding end performs image reconstruction processing on the detail bitstream of the video frame to obtain a reconstructed detail image of the video frame. This reconstructed detail image contains reconstructed pixel information of pixels within a specified region to be enhanced. That is, by reconstructing the detail image from the encoding process, the detailed content of the portion of the video frame lost during encoding (i.e., the specified region) can be restored. Simultaneously, the decoding end also reconstructs the original reconstructed image of the video frame from the original bitstream. Thus, the decoding end can perform detail enhancement processing on the specified region in the reconstructed original image based on the reconstructed detail image to obtain a reconstructed frame image of the video frame. Since the reconstructed frame image is obtained by optimizing the pixels within the specified region of the reconstructed original image using the reconstructed detail image, the reconstructed pixel information of the pixels in the reconstructed frame image has a better degree of restoration than the pixel information of the corresponding pixels in the original video frame; that is, the reconstructed pixel information of the pixels in the reconstructed frame image is the same as or similar to the pixel information of the corresponding pixels in the original video frame. Therefore, this application embodiment uses a dual-stream encoding and decoding method. During the process of restoring the entire video frame using the original bitstream at the decoding end, the detail bitstream is used to enhance the details of the specified areas in the restored video frame that need to be enhanced. This effectively improves the reconstruction quality of pixels in the specified areas that are lost during the encoding process of the video frame, thereby reducing the accuracy loss caused by prediction residuals and quantization during video encoding and improving encoding and decoding efficiency.
[0107] The above-described embodiment in Figure 6 mainly focuses on the decoding end and describes the specific implementation process of the decoding method provided in this application. The following describes the specific implementation process of the encoding method executed by the encoding end with reference to Figure 8. Please refer to Figure 8, which is a flowchart illustrating an encoding method provided in an exemplary embodiment of this application. The flowchart shown in Figure 8 can be executed by a computer device held by the encoding end. The method may include, but is not limited to, steps S801-S804:
[0108] S801: Acquire the video frame to be encoded, and encode the video frame to obtain the original bitstream of the video frame.
[0109] The video frame to be encoded is any frame in the video to be encoded, and the encoder performs the same encoding process on each frame in the video to be encoded to obtain a compressed bitstream; in this embodiment of the application, the method of enhancing the details of a frame image is introduced by taking the encoding process of a single video frame as an example, and is hereby described.
[0110] The decoding end acquires the video frame to be encoded and encodes it into a raw bitstream. This encoding process is performed on the entire video frame. The specific implementation process can be found in the aforementioned introduction to video encoding and decoding technologies, including but not limited to: first, dividing the video frame into blocks; performing predictive coding (e.g., using intra-frame prediction mode) on the current block (i.e., the block to be encoded) to obtain the predicted information of the pixels in the current block; then subtracting the predicted information from the actual pixel information of the pixels in the current block to obtain the residual information of the pixels in the current block; then, sequentially transforming and quantizing the residual information of the pixels in the current block to obtain quantization coefficients; performing processing on each pixel in the video frame to obtain the quantization coefficients of each pixel; and finally performing entropy coding on these quantization coefficients to obtain a binarized raw bitstream.
[0111] S802: Determine a specified region in a video frame to obtain a detailed image of the video frame.
[0112] As mentioned above, considering the information loss caused by predictive coding and quantization during the video frame encoding process shown in step S801, it is supported to perform additional encoding on the pixels in the specified area of the video frame where the information loss is large, so that the information lost at the encoding end can still be obtained at the decoding end, thereby ensuring the video frame reconstruction quality at the decoding end.
[0113] In practical applications, when the encoding end performs quantization operations on residual information, the larger the residual information, the greater the quantization loss. Therefore, this application embodiment supports selecting the region containing pixels with large residual information from the video frame based on the residual information of the pixels as the designated region to be enhanced in detail for additional encoding. For example, the specific implementation process of determining the designated region from the video frame based on the residual information of the pixels can be seen in Figure 9, including but not limited to steps s11-s13:
[0114] s11: Obtain the residual information of each pixel in the video frame.
[0115] During the encoding process of the video frame shown in step S801, the encoder predicts the prediction information of each pixel in the video frame through predictive coding, and then subtracts the corresponding prediction information from the actual pixel information of each pixel to obtain the residual information of each pixel. Thus, when determining a specified region from the video frame, it is only necessary to directly obtain the residual information of each pixel generated during the encoding process of the video frame.
[0116] s12: Determine the target pixel points in the video frame whose residual information meets the residual conditions.
[0117] The residual condition is used to determine whether residual information needs additional encoding. Residual information is the difference between the actual pixel information and the predicted information of a pixel. In this embodiment, considering that larger residual information results in greater information loss during quantization, the residual condition aims to filter out larger residual information from video frames to facilitate additional encoding of this larger residual information into a detail stream. This embodiment provides multiple filtering strategies to determine whether a pixel is a target pixel whose residual information meets the residual condition; these strategies are described in detail below:
[0118] (1) Each video frame corresponds to a residual threshold T, where T is a natural number greater than or equal to zero. In this case, the filtering strategy is to filter pixels in the video frame whose residual information is greater than the residual threshold T, and use these pixels as target pixels whose residual information meets the residual condition; at this time, the residual condition is that the residual information is greater than the residual threshold T. This method of directly comparing the residual information with the residual threshold T to filter target pixels can directly filter out larger residual information from the video frame, which has the advantages of simplicity and convenience.
[0119] In the specific implementation, the encoder compares the residual information of each pixel in the video frame with a residual threshold T to obtain the residual comparison result for each pixel. The residual comparison result of any pixel indicates the relationship between the residual information of that pixel and the residual threshold T: the residual information of that pixel is greater than the residual threshold T, the residual information of that pixel is equal to the residual threshold T, or the residual information of that pixel is greater than the residual threshold T. Then, based on the residual comparison result of each pixel, the pixels in the video frame whose residual information is greater than the residual threshold are identified as target pixels.
[0120] A flowchart illustrating the process of selecting target pixels from a video frame based on residual comparison results is shown in Figure 10a. As shown in Figure 10a, assuming the residual threshold T for the video frame is 2, and the residual information of pixel 1, pixel 2, pixel 3, and pixel 4 in the video frame is 1, etc., then the residual information of pixels 1, 2, 3, and 4 are compared with the residual threshold T respectively. The residual comparison result for pixel 1 indicates that the residual information of pixel 1 is less than the residual threshold T, the residual information of pixel 2 is less than the residual threshold T, the residual information of pixel 3 is greater than the residual threshold T, and the residual information of pixel 4 is greater than the residual threshold T. Thus, pixels 3 and 4, whose residual comparison results indicate that their residual information is greater than the residual threshold T, are selected as target pixels, while pixels 1 and 2, whose residual comparison results indicate that their residual information is less than or equal to the residual threshold T, are assumed to have zero residual.
[0121] (2) Divide the video frame into N brightness intervals based on the brightness values of the pixels, and set a matching residual threshold for each brightness interval, where N is a positive integer. In this case, the filtering strategy is as follows: determine the target brightness interval to which the brightness value of the pixel or the average brightness value of the adjacent area of the pixel belongs, and take the pixels whose brightness value or the average brightness value of the adjacent area of the pixel is less than the residual threshold matching the target brightness interval as target pixels whose residual information meets the residual condition; at this time, the residual condition is: the brightness value of the pixel or the average brightness value of the adjacent area of the pixel belongs to the target brightness interval, and the brightness value of the pixel or the average brightness value of the adjacent area of the pixel is less than the residual threshold matching the target brightness interval.
[0122] In the specific implementation, the detailed process of selecting target pixels based on brightness values and residual information can be found in the flowchart shown in Figure 10b; as shown in Figure 10b:
[0123] 1) Obtain N brightness intervals arranged in ascending order of brightness values, and a corresponding residual threshold for each brightness interval. The sources of the brightness values obtained here can include: the original brightness values of each pixel in the video frame, or the reconstructed brightness values of each pixel obtained by reconstructing the original bitstream of the video frame. As shown in Figure 10b, assuming the minimum brightness value of a pixel in the video frame is 0 and the maximum brightness value is 255, then divide the video frame into N brightness intervals in order from darkest to brightest, i.e., brightness values from 0 to 255. The N brightness intervals are ordered as L1→L2→L3→……→LN, such as brightness interval L1=[0,16), brightness interval L2=[16,31), brightness interval L3=[31,47), etc. For each brightness interval Li, i = 1, 2, 3, ..., N, a residual threshold Ti is set (e.g., T1 = 2, T2 = 3, T3 = 4, T4 = 2). Considering that the smaller the maximum brightness value within a brightness interval, the darker the pixels falling into that interval, the more information is likely to be lost during prediction encoding and quantization. Therefore, a smaller residual threshold can be set for that brightness interval to ensure that more pixels with brightness values falling into that interval can be used as target pixels for additional encoding. Thus, this method of dividing brightness values into brightness intervals and setting matching residual thresholds for different brightness intervals can better adapt to the differences in prediction residual information of pixels with different brightness values during video frame encoding and decoding.
[0124] 2) Obtain the brightness value of each pixel in the video frame and determine the target brightness range to which the brightness value of each pixel belongs.
[0125] Optionally, it supports directly comparing the brightness value of each pixel with N brightness intervals to determine the target brightness interval to which each pixel's brightness value belongs. In short, the brightness value of each pixel in the video frame is compared with the brightness values included in the N brightness intervals to determine the target brightness interval to which each pixel's brightness value belongs. As shown in Figure 10b, assuming the brightness value of pixel 1 is 8, the brightness value of pixel 2 is 19, the brightness value of pixel 3 is 42, and the brightness value of pixel 4 is 43; then, the target brightness interval to which the brightness value of pixel 1 belongs is determined to be brightness interval L1, the target brightness interval to which the brightness value of pixel 2 belongs is determined to be brightness interval L2, and the target brightness interval to which the brightness values of pixels 3 and 4 belong is determined to be brightness interval L3.
[0126] Optionally, considering the small difference between the brightness value of a pixel and its surrounding pixels in a video frame, it is also possible to compare the average brightness value of all pixels in the adjacent region of a pixel with N brightness intervals to determine the target brightness interval to which each pixel's brightness value belongs. Specifically, the average brightness value of the adjacent region of each pixel in the video frame is obtained. Here, the adjacent region of a pixel can refer to the area in the video frame where the pixel is located, which contains the pixel, such as an area of size 8×8. The average brightness value of the adjacent region is obtained by averaging the brightness values of all pixels in the adjacent region of the pixel. As shown in Figure 10b, assuming the size of the adjacent region of pixel 1 is 2×2, and the four brightness values in the adjacent region of pixel 1 are brightness value 2 for pixel 1, brightness value 8 for pixel 2, brightness value 10 for pixel 5, and brightness value 12 for pixel 6, then the average brightness value of the adjacent region of pixel 1 is calculated to be 2+8+10+12 / 4=8. The average brightness value 8 falls into the brightness interval L1, so the brightness interval L1 is taken as the target brightness interval to which the brightness value of pixel 1 belongs.
[0127] 3) Compare the residual information of each pixel in the video frame with the residual threshold that matches the target brightness range, and determine the pixels with residual information greater than the corresponding residual threshold as target pixels; and determine the pixels with residual information less than or equal to the corresponding residual threshold as zero residual.
[0128] As shown in Figure 10b, when directly comparing the brightness values of pixels and N brightness intervals, the residual information of pixel 1 is compared with the residual threshold T1 that matches brightness interval L1; the residual information of pixel 2 is compared with the residual threshold T2 that matches brightness interval L2; the residual threshold of pixel 3 is compared with the residual threshold T3 that matches brightness interval L3; and the residual threshold of pixel 4 is compared with the residual threshold T3 that matches brightness interval L3. If the residual information of pixel 1 is 1, the residual information of pixel 2 is 4, the residual information of pixel 3 is 7, and the residual information of pixel 4 is 5, and T1 = 2, T2 = 3, and T3 = 4, then pixels 2, 3, and 5 are determined to be target pixels with residual information greater than the corresponding residual threshold.
[0129] As shown in Figure 10b, when the target pixel is determined based on both the brightness value and the residual information, if the average brightness value 8 of the adjacent area of pixel 1 falls within the brightness interval L1, then the residual information of pixel 1 is compared with the residual threshold T1 that matches the brightness interval L1. Under the assumption that T1 = 2 and the residual information of pixel 1 is 1, pixel 1 is determined not to be a target pixel with a residual information greater than the residual threshold. Similarly, pixels 2, 3, and 4 are judged in the same way as pixel 1 to obtain the results of whether pixels 2, 3, and 4 are target pixels.
[0130] In summary, the embodiments of this application can effectively improve the accuracy of target pixel selection by simultaneously combining brightness values and residual information to filter target pixels that require additional encoding.
[0131] (3) Dynamically set the residual threshold based on the presence of non-zero residual information around the current pixel. The filtering strategy is as follows: set M value intervals based on the residual information of multiple pixels adjacent to the current pixel in the video frame, where M is a positive integer, and set a matching residual threshold for each value interval, where N is a positive integer; assign values to the residual information of multiple pixels adjacent to any pixel, and add the assigned values of the multiple pixels; determine the target value interval to which any pixel belongs based on the addition result; when the residual information of any pixel is greater than the residual threshold matching the target value interval, any pixel is regarded as the target pixel whose residual information meets the residual condition.
[0132] In the specific implementation, the detailed process of filtering target pixels based on the presence of non-zero residual information around the current pixel can be seen in the flowchart shown in Figure 10c; as shown in Figure 10c:
[0133] Preset operation: Assume that the residual information of pixels with non-zero residual information in the video frame is assigned a first preset value (e.g., the first preset value is 1), and the residual information of pixels with zero residual information in the video frame is assigned a second preset value (e.g., the second preset value is 0); then, for Q adjacent pixels of a pixel (Q is a positive integer, e.g., Q = 8), if the range of the sum of the residual information of the 8 pixels surrounding the pixel (e.g., the 8 pixels located above, below, left, right, upper left, lower left, upper right, and lower right of the pixel) after being assigned values is [0, 8], the maximum value in this range is the sum of Q first preset values (i.e., 8), and the minimum value is the sum of Q second preset values (i.e., 0). Then, M value intervals are set according to the range [0,8]. For example, if M=3, then value intervals S1=[0,4), S2=[4,7), and S3=[7,8] can be set. The maximum value in these M=3 value intervals is the sum of Q first preset values, and the minimum value is the sum of Q second preset values. Furthermore, a matching residual threshold is set for each value interval. Considering that if the number of residual information of pixels around the current pixel that are assigned the second preset value is large, that is, the smaller the sum of the values assigned to multiple pixels around the current pixel, it can to a certain extent indicate that the pixel has high accuracy in prediction encoding, that is, the pixel is likely to be a pixel that does not need additional encoding, then a higher residual threshold can be set for value intervals with smaller values, and a lower residual threshold can be set for value intervals with larger values, so as to achieve additional encoding of the residual information of pixels in areas with poor accuracy in prediction encoding. For example, the residual threshold T1 is set to 7 for the value interval S1, the residual threshold T2 is set to 4 for the value interval S2, and the residual threshold T3 is set to 2 for the value interval S3.
[0134] Real-time filtering operation: When it is necessary to determine whether a pixel in a video frame is a target pixel, M value intervals are obtained in ascending order of value, along with a corresponding residual threshold for each interval; the value here is the sum of the residual information of the Q surrounding pixels after assignment. Then, the residual information of each of the Q pixels adjacent to the pixel to be judged is obtained, and the residual information with non-zero values among the Q pixels is marked as a first preset value. The first predicted values marked in the Q pixels are then added together to obtain the preset value sum; based on this sum, the target value interval to which the pixel belongs is determined. Finally, the residual information of the pixel is compared with the residual threshold matching the target value interval; if the residual information of the pixel is greater than the residual threshold matching the target value interval, the pixel is determined as the target pixel. As shown in Figure 10c, assuming the residual information of the eight pixels surrounding pixel 1001 is as follows: residual information of pixel 1 is 0, residual information of pixel 2 is 1, residual information of pixel 3 is 5, residual information of pixel 4 is 0, residual information of pixel 5 is 3, residual information of pixel 6 is 7, residual information of pixel 7 is 1, and residual information of pixel 8 is 0, then the residual information of pixels 2, 3, 5, 6, and 7 is assigned a first preset value of 1, and the sum of the assigned preset values is calculated to be 5. If the sum of the preset values 5 falls within the value range S2 = [4, 7), then the residual information of pixel 1001 is compared with the residual threshold T2 that matches the value range S2; if the sum of the preset values 5 is greater than the residual threshold T2, then pixel 1001 is determined as the target pixel. It is worth noting that the process of determining whether other pixels in the video frame, excluding pixel 1001, are target pixels is the same as the process of determining whether pixel 1001 is a target pixel. Please refer to the above explanation, which will not be repeated here.
[0135] It should be noted that: 1) Figure 10c is illustrated using the example of a first preset value of 1 and a second preset value of 0; in other implementations, the first preset value can be 0 and the second preset value can be 1. In this case, when setting the residual threshold for M value intervals, the principle to be followed is to set a lower residual threshold for value intervals with smaller values and a higher residual threshold for value intervals with larger values; or, in other implementations, the first and second preset values can be set to other numbers. This application embodiment does not limit the numerical value of the first and second preset values, only the residual threshold of the corresponding value interval needs to be set. 2) Figure 10c is illustrated using Q=8 as an example. In other implementations, Q can also be other numbers, such as Q=4, and only the pixels above, below, to the left and to the right of the current pixel are selected for judgment. 3) For pixels located at the edge in a video frame, when the number of pixels P around it is less than Q, P is a positive integer, and QP pixels can be selected around its adjacent pixels for judgment; or, P pixels can be used directly for judgment.
[0136] In summary, the embodiments of this application support using the residual information of each pixel among multiple adjacent pixels in a video frame to help determine the degree of information loss when encoding a pixel. If there are many pixels with large residual information among multiple adjacent pixels, it means that the degree of information loss when encoding the pixel may be large, thereby improving the accuracy of target pixel selection.
[0137] (4) Target pixel selection is performed based on the enhancement algorithm. The selection strategy is as follows: the original bitstream of the video frame is reconstructed at the encoding end to obtain the reconstructed image of the video frame; and the enhancement algorithm is applied to the reconstructed image to obtain the enhanced reconstructed image. In this case, the pixels in the difference image between the enhanced reconstructed image and the unenhanced reconstructed image are taken as the target pixels whose residual information meets the residual conditions.
[0138] In other words, this application embodiment supports the encoding end to perform image reconstruction based on the original bitstream of the video frame encoded on this side, obtain the reconstructed image corresponding to the video frame, and then perform image enhancement processing on the reconstructed image to obtain the enhanced reconstructed image. The image enhancement processing here aims to use enhancement algorithms to enhance areas with poor quality (such as dark brightness or unclear) in the reconstructed image, so as to improve the image quality of the enhanced reconstructed image. This application embodiment does not limit the type of enhancement algorithm, such as the edge enhancement algorithm (used to enhance the details of the edges of the reconstructed image). Thus, the enhanced reconstructed image and the reconstructed image are subjected to a difference operation to obtain a difference image. The difference image includes the enhanced pixel information of the pixels in the enhanced reconstructed image that are enhanced relative to the reconstructed image. This enhanced pixel information can be considered as information that may be lost during the encoding process. At this time, the target pixel point containing residual information that meets the residual condition is determined. The target pixel point is the pixel point in the difference image whose enhanced pixel information is a non-zero value. Therefore, the embodiments of this application support image enhancement processing on the reconstructed image of the original bitstream at the decoding end, thereby analyzing the enhanced reconstructed image in the video frame that can be enhanced in detail. In this way, the target pixels in the video frame that have suffered severe information loss during the encoding process can be deduced, thereby improving the accuracy of target pixel selection.
[0139] As shown in Figure 10d, the encoder performs image reconstruction on the original bitstream of the video frame to obtain a reconstructed image 1002. Image enhancement processing is then performed on this reconstructed image 1002 to obtain an enhanced reconstructed image 1003. Next, a difference image 1004 is calculated between the enhanced reconstructed image 1003 and the reconstructed image 1002. Pixels in the difference image 1004 whose enhanced pixel information is non-zero are considered target pixels. For example, if pixels 1, 2, and 3 have the same pixel information in the enhanced reconstructed image 1003 as they do in the reconstructed image 1002, then the enhanced pixel information of pixels 1, 2, and 3 in the difference image is zero, and therefore pixels 1, 2, and 3 are not target pixels. Conversely, if pixels 4 and 5 have different pixel information in the enhanced reconstructed image 1003 as they do in the reconstructed image 1002, then the enhanced pixel information of pixels 4 and 5 in the difference image is non-zero, and therefore pixels 4 and 5 are target pixels.
[0140] It should be noted that the four filtering strategies provided in this application are examples, intended to filter pixels that need to be optimized or enhanced from the pixels of video frames; in practical applications, the filtering strategies can also be other strategies, and this application does not limit them.
[0141] s13: Based on the position of the target pixel in the video frame, mark the specified area in the video frame to obtain the detailed image of the video frame.
[0142] After identifying the target pixels requiring additional encoding from the video frame based on the aforementioned steps, it is necessary to identify the region where the target pixels are located within the video frame. This region is designated as the area in the video frame to be enhanced for detail, and can be referred to as the specified region in the video frame, thereby obtaining a detail image of the video frame. This detail image includes the pixel information of the pixels within the specified region of the video frame. By dividing the video frame into regions and gradually narrowing down the scope to determine the specified region, the accuracy of determining the specified region is effectively improved.
[0143] In specific implementation, the video frame can be divided into regions to obtain at least two first regions corresponding to the video frame; for example, the video frame can be divided according to a fixed region size (e.g., 8×8) to obtain at least two first regions, where the region size of each first region is smaller than the size of the video frame. It is determined whether any of the at least two first regions contains a target pixel with residual information greater than a residual threshold (pixels filtered according to any of the filtering strategies provided in step s12 above); if any of the at least two first regions contains a target pixel, then that first region is identified as a designated region in the video frame. It is worth noting that the embodiments of this application do not limit the size of the fixed region.
[0144] As shown in Figure 11a, assuming the video frame size is 16×16 and the fixed region size is 8×8, dividing the video frame according to this fixed region size yields four first regions: first region 1101, first region 1102, first region 1103, and first region 1104. Each of the four first regions is then checked to determine if it contains target pixels. If first region 1101 contains target pixels 1, 2, and 3, and first region 1104 contains target pixels 4 and 5, while first regions 1102 and 1103 do not contain target pixels, then first regions 1101 and 1104 can be designated as specific regions within the video frame, thus obtaining a detail image of the video frame. This detail image retains the pixel information of the pixels within first regions 1101 and 1104 (including target pixels and non-target pixels with residual information less than or equal to the residual threshold).
[0145] Furthermore, considering that a large number of non-target pixels with residual information less than or equal to the residual threshold may exist within the first region determined by the fixed-size region, this application embodiment supports further dividing the first region into second regions with smaller sizes, and only encodes the second regions containing target pixels, greatly saving encoding resources and improving encoding efficiency. In specific implementation, assuming that any first region contains target pixels, it supports dividing any first region containing target pixels into regions to obtain at least two second regions corresponding to any first region. The area size of the second region is smaller than that of the first region; for example, dividing the 8×8 first region into 4×4 regions yields four second regions. Then, it is determined whether any of the at least two second regions contains target pixels with residual information greater than the residual threshold; if any of the at least two second regions contains target pixels, then that second region is identified as a designated region in the video frame.
[0146] As shown in Figure 11b, taking the first region 1101 as an example, the size of the first region 1101 is 8×8. Dividing the first region 1101 into two regions of 4×4, we obtain the second region 11011, the second region 11012, the second region 11013, and the second region 11014. Target pixels 1, 2, and 3 are located in the second region 11011, while the second regions 11012, 11013, and 11014 do not contain any target pixels. Therefore, the second regions 11011 and 11014 are designated regions in the video frame. Similarly, the first region 1104 is further divided to obtain the second region 11041, which contains target pixels 4 and 5. In summary, the designated regions in the video frame are determined to be the second region 11011 and the second region 11041.
[0147] It should be understood that if the second region (such as the second region 11011 and the second region 11041) contains a large number of non-target pixels with residual information less than or equal to the residual threshold, in order to improve coding efficiency, it is also supported to continue to divide this part of the second region. The embodiments of this application do not limit the number of times the region is divided, which is hereby stated.
[0148] In summary, steps s11-s13 described above provide multiple strategies for determining target pixels that need to be enhanced from video frames on the encoding side. Compared with selecting target pixels from a single dimension, this greatly improves the accuracy of selecting target pixels that need to be enhanced in video frames, thereby effectively improving the quality of video encoding and decoding.
[0149] S803: Encodes the detail image to obtain the detail bitstream of the video frame.
[0150] The detail image is an image with the same dimensions as the video frame, containing one or more designated regions to be enhanced. After obtaining the detail image corresponding to the video frame based on the aforementioned steps, the encoder encodes the pixels within the designated regions of the detail image to obtain the detail bitstream of the video frame. This encoding process specifically includes operations such as transform, quantization, and entropy coding; the specific implementation methods for transform, quantization, and entropy coding operations can be found in the aforementioned descriptions and will not be repeated here.
[0151] S804: Sends the raw bitstream and detail bitstream of the video frame to the decoding end for joint image decoding.
[0152] The encoder can package the original bitstream and detail bitstream of the same video frame into a single bitstream data, and compress the bitstream data of each video frame into a compressed bitstream before sending it to the decoder. Upon receiving the compressed bitstream, the decoder can reconstruct the image of each video frame based on the bitstream data of each frame in the compressed bitstream, obtaining the reconstructed frame image of each video frame. This enables video playback and sharing on the decoder side. It should be noted that the specific implementation process of the decoder performing joint image decoding based on the original bitstream and detail bitstream of the video frame after receiving the compressed bitstream can be found in the description of the specific implementation process of the embodiment shown in Figure 6 above, and will not be repeated here.
[0153] In summary, at the encoding end, this application embodiment supports the use of multiple filtering strategies to select target pixels with residual information greater than the residual threshold from the video frame to determine the specified region in the video frame. These multiple filtering strategies can better meet the encoding requirements of users, are applicable to different encoding scenarios, and improve the encoding experience. Furthermore, at the encoding end, additional encoding / secondary encoding is performed only on pixels within the specified region requiring detail enhancement, without significant resource waste, achieving detail enhancement and significantly improving encoding quality and efficiency. Correspondingly, at the decoding end, during the process of restoring the entire video frame from the original bitstream, the detail bitstream is used to enhance the specified region requiring detail enhancement in the restored video frame. This effectively improves the reconstruction quality of pixels within the specified region that suffered loss during the video frame encoding process, reducing the accuracy loss caused by prediction residuals and quantization during video encoding and improving encoding / decoding efficiency.
[0154] Figure 12 shows a schematic diagram of a decoding device provided in an exemplary embodiment of this application; the decoding device can be used to perform some or all of the steps in the method embodiment shown in Figure 6. Referring to Figure 12, the device includes the following units:
[0155] The acquisition unit 1201 is used to acquire the bitstream data of the video frame. The bitstream data includes the original bitstream and the detail bitstream of the video frame. The original bitstream is obtained by encoding the video frame. The detail bitstream is obtained by encoding a specified region in the video frame. The specified region refers to the region in the video frame that needs to be enhanced in terms of detail.
[0156] The processing unit 1202 is used to perform image reconstruction processing on the detail bitstream to obtain the reconstructed detail image of the video frame; the reconstructed detail image contains the reconstructed pixel information of pixels in a specified region;
[0157] The processing unit 1202 is also used to decode the original bitstream to obtain the reconstructed original image of the video frame;
[0158] The processing unit 1202 is also used to perform detail enhancement processing on a specified region in the reconstructed original image based on the reconstructed detail image, so as to obtain the reconstructed frame image of the video frame.
[0159] In one implementation, the processing unit 1202, when performing image reconstruction processing on the detail bitstream to obtain the reconstructed detail image of the video frame, specifically performs the following:
[0160] The detail stream is decoded to obtain the residual information of pixels in the specified region of the reconstructed detail image;
[0161] Obtain encoding mode information;
[0162] Based on the encoding mode information, the pixels in the specified area are predicted to obtain the predicted information of the pixels in the specified area in the reconstructed detail image.
[0163] Based on the residual and prediction information of pixels within a specified region, a reconstructed detail image of the video frame is obtained.
[0164] In one implementation, the detail bitstream includes an identifier, which indicates the position of a specified region in a video frame; the processing unit 1202, when decoding the detail bitstream to obtain residual information of pixels within the specified region in the reconstructed detail image, specifically performs the following:
[0165] The identifier is obtained by decoding the detailed bitstream;
[0166] Based on the location indicated by the identifier, the pixels within the specified area are decoded to obtain the residual information of the pixels within the specified area in the reconstructed detail image.
[0167] In one implementation, the pixels in the reconstructed detail image correspond one-to-one with the pixels in the reconstructed original image; the processing unit 1202, when performing detail enhancement processing on a specified region in the reconstructed original image based on the reconstructed detail image to obtain the reconstructed frame image of the video frame, specifically performs the following:
[0168] The reconstructed pixel information of pixels within a specified region in the reconstructed detail image is superimposed onto the reconstructed pixel information of the corresponding pixels within the specified region in the original reconstructed image; and...
[0169] Preserve the reconstructed pixel information of pixels that were not superimposed in the original image.
[0170] According to one embodiment of this application, the units in the decoding device shown in FIG12 can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This can achieve the same operation without affecting the technical effect of the embodiments of this application. The above units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the decoding device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together. According to another embodiment of this application, the decoding device shown in FIG12 and the decoding method of the embodiments of this application can be implemented by running a computer program (including program code) capable of executing the steps involved in the corresponding method shown in FIG6 on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM). Computer programs can be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and run therein.
[0171] Based on the same inventive concept, the decoding device provided in the embodiments of this application solves the problem in a similar principle and with similar beneficial effects as the decoding processing method in the embodiments of this application. For details, please refer to the implementation principle and beneficial effects of the method. For the sake of brevity, these will not be repeated here.
[0172] Figure 13 shows a schematic diagram of an encoding device provided in an exemplary embodiment of this application; the decoding device can be used to perform some or all of the steps in the method embodiment shown in Figure 8. Referring to Figure 13, the device includes the following units:
[0173] Acquisition unit 1301 is used to acquire video frames to be encoded, and to encode the video frames to obtain the original bitstream of the video frames; and,
[0174] Processing unit 1302 is used to determine a specified region in a video frame and obtain a detail image of the video frame; the specified region refers to the region in the video frame that needs to be enhanced in terms of detail.
[0175] The processing unit 1302 is also used to encode the detail image to obtain the detail bitstream of the video frame;
[0176] The processing unit 1302 is also used to send the original bitstream and detail bitstream of the video frame to the decoding end for joint image decoding.
[0177] In one implementation, the processing unit 1302, when determining a specified region in a video frame to obtain a detailed image of the video frame, specifically performs the following:
[0178] Obtain residual information for each pixel in a video frame; the residual information of a pixel is obtained by predictive encoding of the pixel and is used to characterize the difference between the actual pixel information and the predicted information of the pixel.
[0179] Identify target pixels in video frames whose residual information meets the residual conditions;
[0180] Based on the position of the target pixel in the video frame, a specified area is marked in the video frame to obtain a detailed image of the video frame.
[0181] In one implementation, the video frame corresponds to a residual threshold; the processing unit 1302, when determining the target pixel point whose residual information meets the residual condition from the video frame, is specifically used for:
[0182] The residual information of each pixel in the video frame is compared with the residual threshold to obtain the residual comparison result of each pixel; the residual comparison result is used to indicate the magnitude relationship between the residual information of the pixel and the residual threshold;
[0183] Based on the residual comparison results of each pixel, pixels with residual information greater than the residual threshold are identified from the video frame; these pixels are designated as target pixels.
[0184] In one implementation, when processing unit 1302 determines target pixels whose residual information meets the residual conditions from a video frame, it specifically performs the following:
[0185] Obtain N brightness intervals arranged in ascending order of brightness values, and the corresponding residual threshold for each brightness interval; N is a positive integer.
[0186] Obtain the brightness value of each pixel in the video frame and determine the target brightness range to which the brightness value of each pixel belongs;
[0187] The residual information of each pixel in the video frame is compared with the residual threshold that matches the target brightness range, and the pixels with residual information greater than the corresponding residual threshold are identified as target pixels.
[0188] In one implementation, the processing unit 1302, when determining the target brightness range to which the brightness value of each pixel belongs, specifically performs the following:
[0189] The brightness value of each pixel in the video frame is compared with the brightness values included in N brightness intervals to determine the target brightness interval to which the brightness value of each pixel belongs.
[0190] In one implementation, the processing unit 1302, when determining the target brightness range to which the brightness value of each pixel belongs, specifically performs the following:
[0191] Obtain the average brightness value of the adjacent region of each pixel in the video frame; the average brightness value is obtained by averaging the brightness values of all pixels in the adjacent region of the pixel.
[0192] The average brightness value of the adjacent region of each pixel in the video frame is compared with the brightness values included in N brightness intervals to determine the target brightness interval to which the average brightness value of the adjacent region of each pixel belongs; the target brightness interval to which the average brightness value of the adjacent region of each pixel belongs is used as the target brightness interval to which the brightness value of the pixel belongs.
[0193] In one implementation, when processing unit 1302 determines target pixels whose residual information meets the residual conditions from a video frame, it specifically performs the following:
[0194] Obtain M value intervals arranged in ascending order of value, and a residual threshold matching each value interval; the maximum value in the M value intervals is the sum of Q first preset values, and the minimum value is the sum of Q second preset values; M and Q are positive integers;
[0195] Obtain the residual information of each pixel in the Q pixels adjacent to any pixel in the video frame, and mark the residual information of the Q pixels with non-zero residual information as the first preset value, and mark the residual information of the Q pixels with zero residual information as the second preset value.
[0196] The first preset value marked in Q pixels is added together to obtain the sum of the preset values;
[0197] Based on the sum of preset values, determine the target value range to which any pixel belongs;
[0198] If the residual information of any pixel is greater than the residual threshold that matches the target value range, then that pixel is determined as the target pixel.
[0199] In one implementation, when processing unit 1302 determines target pixels whose residual information meets the residual conditions from a video frame, it specifically performs the following:
[0200] Image reconstruction is performed on the original bitstream to obtain the reconstructed images corresponding to the video frames;
[0201] The reconstructed image is then enhanced to obtain the enhanced reconstructed image.
[0202] The enhanced and reconstructed images are interpolated to obtain a difference image; the difference image contains target pixels that meet the residual conditions.
[0203] In one implementation, the processing unit 1302, when identifying a specified region in the video frame based on the position of the target pixel in the video frame to obtain a detailed image of the video frame, specifically performs the following:
[0204] The video frame is divided into regions to obtain at least two first regions corresponding to the video frame;
[0205] If any of the at least two first regions contains the target pixel, then that first region is identified as the designated region in the video frame.
[0206] In one implementation, when the processing unit 1302 identifies any first region as a designated region in a video frame, it specifically performs the following functions:
[0207] Divide any first region including the target pixel into regions to obtain at least two second regions corresponding to any first region; the area size of the second region is smaller than the area size of the first region.
[0208] If any of the at least two second regions contains the target pixel, then that second region is identified as the designated region in the video frame.
[0209] According to one embodiment of this application, the units in the encoding device shown in FIG13 can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This can achieve the same operation without affecting the technical effect of the embodiment of this application. The above units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the encoding device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together. According to another embodiment of this application, the encoding device shown in FIG13 and the encoding method of the embodiment of this application can be implemented by running a computer program (including program code) capable of performing the steps involved in the corresponding method shown in FIG8 on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM). Computer programs can be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and run therein.
[0210] Based on the same inventive concept, the principle and beneficial effects of the encoding device provided in the embodiments of this application in solving the problem are similar to the principle and beneficial effects of the encoding processing method in the embodiments of this application in solving the problem. For the sake of brevity, the principle and beneficial effects of the method implementation can be referred to.
[0211] Figure 14 shows a schematic diagram of a computer device provided in an exemplary embodiment of this application. Referring to Figure 14, the computer device includes a processor 1401, a communication interface 1402, and a computer-readable storage medium 1403. The processor 1401, communication interface 1402, and computer-readable storage medium 1403 can be connected via a bus or other means. The communication interface 1402 is used to receive and send data. The computer-readable storage medium 1403 can be stored in the memory of the computer device. The computer-readable storage medium 1403 is used to store computer programs, including program instructions. The processor 1401 is used to execute the program instructions stored in the computer-readable storage medium 1403. The processor 1401 (or CPU (Central Processing Unit)) is the computing and control core of the computer device, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve corresponding method flows or corresponding functions.
[0212] This application embodiment also provides a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the processing system of the computer device. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by the processor 1401, which may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here may be high-speed RAM memory or non-volatile memory, such as at least one disk storage device; optionally, it may also be at least one computer-readable storage medium located remotely from the aforementioned processor.
[0213] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor 1401 loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-described decoding method embodiment; in a specific implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1401 and executed to implement the decoding method and encoding method described above.
[0214] Based on the same inventive concept, the principle and beneficial effects of the computer device provided in the embodiments of this application in solving the problem are similar to the principle and beneficial effects of the decoding method and encoding method provided in the embodiments of this application in solving the problem. Please refer to the principle and beneficial effects of the implementation of the method. For the sake of brevity, they will not be repeated here.
[0215] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a blockchain node device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned decoding and encoding methods.
[0216] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0217] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0218] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A decoding method, characterized in that, The method is performed by a computer device, and the method includes: The video frame bitstream data is obtained, including the original bitstream and detail bitstream of the video frame; the original bitstream is obtained by encoding the video frame; the detail bitstream is obtained by encoding a specified region in the video frame, where the specified region refers to the region in the video frame to be enhanced in detail. The detailed bitstream is subjected to image reconstruction processing to obtain the reconstructed detailed image of the video frame; the reconstructed detailed image contains the reconstructed pixel information of the pixels in the specified region; The original bitstream is decoded to obtain the reconstructed original image of the video frame; Based on the reconstructed detail image, the specified region in the reconstructed original image is subjected to detail enhancement processing to obtain the reconstructed frame image of the video frame.
2. The method as described in claim 1, characterized in that, The step of performing image reconstruction processing on the detailed bitstream to obtain the reconstructed detailed image of the video frame includes: The detail stream is decoded to obtain the residual information of pixels in the specified region of the reconstructed detail image; Obtain encoding mode information; Based on the encoding mode information, the pixels in the specified region are predicted to obtain the predicted information of the pixels in the specified region in the reconstructed detail image; Based on the residual information and the prediction information of the pixels within the specified region, a reconstructed detail image of the video frame is obtained.
3. The method as described in claim 1 or 2, characterized in that, The detail stream includes an identifier, which indicates the position of the specified region in the video frame; the decoding process of the detail stream to obtain the residual information of the pixels within the specified region in the reconstructed detail image includes: The identifier is obtained by decoding the detailed bitstream; Based on the location indicated by the identifier, the pixels within the specified area are decoded to obtain the residual information of the pixels within the specified area in the reconstructed detail image.
4. The method according to any one of claims 1-3, characterized in that, The pixels in the reconstructed detail image correspond one-to-one with the pixels in the reconstructed original image; the step of performing detail enhancement processing on the specified region in the reconstructed original image based on the reconstructed detail image to obtain the reconstructed frame image of the video frame includes: The reconstructed pixel information of pixels within the specified region in the reconstructed detail image is superimposed onto the reconstructed pixel information of corresponding pixels within the specified region in the original reconstructed image; and... The reconstructed pixel information of the un-overlaid pixels in the reconstructed original image is retained.
5. An encoding method, characterized in that, The method is performed by a computer device, and the method includes: Acquire the video frame to be encoded, and encode the video frame to obtain the original bitstream of the video frame; and, A designated region is determined within the video frame to obtain a detailed image of the video frame; the designated region refers to the area in the video frame where the detail needs to be enhanced. The detailed image is encoded to obtain the detailed bitstream of the video frame; The original bitstream and the detail bitstream of the video frame are sent to the decoding end for joint image decoding.
6. The method as described in claim 5, characterized in that, The step of determining a specified region in the video frame to obtain a detailed image of the video frame includes: Obtain residual information for each pixel in the video frame; the residual information of the pixel is obtained by predictive encoding of the pixel, and the residual information of the pixel is used to characterize the difference between the actual pixel information and the predicted information of the pixel. Determine the target pixels whose residual information meets the residual conditions from the video frames; Based on the position of the target pixel in the video frame, a designated area is identified in the video frame to obtain a detailed image of the video frame.
7. The method as described in claim 5 or 6, characterized in that, The video frame corresponds to a residual threshold; determining the target pixel points whose residual information meets the residual conditions from the video frame includes: The residual information of each pixel in the video frame is compared with the residual threshold to obtain the residual comparison result of each pixel; the residual comparison result is used to indicate the magnitude relationship between the residual information of the pixel and the residual threshold; Based on the residual comparison results of each pixel, pixels in the video frame whose residual information is greater than the residual threshold are determined; the pixels whose residual information is greater than the residual threshold are designated as target pixels.
8. The method according to any one of claims 5-7, characterized in that, Determining the target pixel points in the video frame whose residual information meets the residual conditions includes: Obtain N brightness intervals arranged in ascending order of brightness values, and a residual threshold matching each brightness interval; N is a positive integer; Obtain the brightness value of each pixel in the video frame, and determine the target brightness range to which the brightness value of each pixel belongs; The residual information of each pixel in the video frame is compared with the residual threshold that matches the target brightness range, and the pixels whose residual information is greater than the corresponding residual threshold are determined as target pixels.
9. The method according to any one of claims 5-8, characterized in that, Determining the target brightness range to which the brightness value of each pixel belongs includes: The brightness value of each pixel in the video frame is compared with the brightness values included in the N brightness intervals to determine the target brightness interval to which the brightness value of each pixel belongs.
10. The method according to any one of claims 5-9, characterized in that, Determining the target brightness range to which the brightness value of each pixel belongs includes: The average brightness value of the adjacent region of each pixel in the video frame is obtained; the average brightness value is obtained by averaging the brightness values of all pixels in the adjacent region of the pixel. The average brightness value of the adjacent region of each pixel in the video frame is compared with the brightness values included in the N brightness intervals to determine the target brightness interval to which the average brightness value of the adjacent region of each pixel belongs; the target brightness interval to which the average brightness value of the adjacent region of the pixel belongs is taken as the target brightness interval to which the brightness value of the pixel belongs.
11. The method according to any one of claims 5-10, characterized in that, Determining the target pixel points in the video frame whose residual information meets the residual conditions includes: Obtain M value intervals arranged in ascending order of value, and a residual threshold matching each value interval; the maximum value in the M value intervals is the sum of Q first preset values, and the minimum value is the sum of Q second preset values; M and Q are positive integers; Obtain the residual information of each of the Q pixels adjacent to any given pixel in the video frame, and mark the residual information of the Q pixels with non-zero residual information as a first preset value, and mark the residual information of the Q pixels with zero residual information as a second preset value; The first preset values marked in the Q pixels are added together to obtain the sum of the preset values; Based on the sum of the preset values, the target value range to which any pixel belongs is determined; If the residual information of any pixel is greater than the residual threshold that matches the target value range, then any pixel is determined as the target pixel.
12. The method according to any one of claims 5-11, characterized in that, Determining the target pixel points in the video frame whose residual information meets the residual conditions includes: The original bitstream is reconstructed to obtain the reconstructed image corresponding to the video frame; The reconstructed image is then subjected to image enhancement processing to obtain an enhanced reconstructed image; The enhanced reconstructed image and the reconstructed image are subjected to a difference operation to obtain a difference image; the difference image contains target pixels whose residual information meets the residual conditions.
13. The method according to any one of claims 5-12, characterized in that, The step of identifying a designated region in the video frame based on the position of the target pixel in the video frame to obtain a detailed image of the video frame includes: The video frame is divided into regions to obtain at least two first regions corresponding to the video frame; If any one of the at least two first regions contains the target pixel, then that first region is identified as a designated region in the video frame.
14. The method according to any one of claims 5-13, characterized in that, The step of identifying any of the first regions as a designated region in the video frame includes: Divide any first region including the target pixel into regions to obtain at least two second regions corresponding to any first region; the area size of the second region is smaller than the area size of the first region. If any of the at least two second regions contains the target pixel, then that second region is identified as a designated region in the video frame.
15. A decoding device, characterized in that, The decoding device is mounted on a computer device, and the decoding device includes: The acquisition unit is used to acquire the bitstream data of a video frame, the bitstream data including the original bitstream and the detail bitstream of the video frame; the original bitstream is obtained by encoding the video frame; the detail bitstream is obtained by encoding a specified region in the video frame, the specified region being the region in the video frame to be enhanced in detail; The processing unit is used to perform image reconstruction processing on the detail stream to obtain the reconstructed detail image of the video frame; the reconstructed detail image contains the reconstructed pixel information of the pixels in the specified region; The processing unit is also used to decode the original bitstream to obtain the reconstructed original image of the video frame; The processing unit is further configured to perform detail enhancement processing on the specified region in the reconstructed original image based on the reconstructed detail image, so as to obtain the reconstructed frame image of the video frame.
16. An encoding device, characterized in that, The encoding device is mounted on a computer device, and the encoding device includes: An acquisition unit is configured to acquire a video frame to be encoded, and to encode the video frame to obtain the original bitstream of the video frame; and, A processing unit is configured to determine a specified region in the video frame and obtain a detail image of the video frame; the specified region refers to the region in the video frame to be enhanced for detail. The processing unit is further configured to encode the detail image to obtain the detail bitstream of the video frame; The processing unit is further configured to send the original bitstream and the detail bitstream of the video frame to the decoding end for joint image decoding.
17. A computer device, characterized in that, include: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the decoding method as described in any one of claims 1-4, or the encoding method as described in any one of claims 5-14.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as the decoding method as described in any one of claims 1-4, or to implement the encoding method as described in any one of claims 5-14.
19. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the decoding method as described in any one of claims 1-4, or the encoding method as described in any one of claims 5-14.
Citation Information
Patent Citations
Coding and decoding methods and related devices
CN114913249A
Video encoding and decoding method, device, equipment, storage medium and computer program
CN116132684A
Residual coding method and device, video coding method and device, and storage medium
CN117136540A
Image encoding / decoding method, electronic equipment and computer readable storage medium
CN118381930A
Coding apparatus, coding method, decoding apparatus, and decoding method
US20150208097A1