Image processing method and device, equipment and storage medium
By encoding and decoding based on the scale of the target block during the video frame encoding and decoding process, the problem of difficult performance and complexity in the prior art is solved, and more efficient video frame encoding and decoding is achieved.
Patent Information
- Application Number
- CN202311691334.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
How to improve the compromise between performance and complexity in the video frame encoding and decoding process, and solve the problem of performance improvement in the prior art with increased complexity.
The encoded data of the video frame is decoded based on the scale of the target block, and the encoding and decoding information is determined using the scale of the target block during the encoding process, and the encoding and decoding process is optimized.
The performance and complexity ratio in the video frame encoding and decoding process is improved, the encoding and decoding process is optimized, and the dependence on the target block division flag bits is reduced.
Smart Images

Figure CN120128704A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to an image processing method, apparatus, and computer-readable storage medium. Background Art
[0002] With the progress of scientific and technological research, a vast amount of video resources have emerged on the Internet. Since the transmission of many videos (such as live broadcasts, online videos, etc.) is real-time, technicians are committed to improving the performance of encoding and decoding to meet the ever-increasing video (such as clarity) requirements of video consumers. It has been found that when improving the performance of encoding and decoding, the complexity of encoding and decoding usually also increases. How to improve the trade-off ratio between performance and complexity in the video frame encoding and decoding process has become a hot issue in current research. Summary of the Invention
[0003] Embodiments of this application provide an image processing method, apparatus, device, and computer-readable storage medium, which can improve the trade-off ratio between performance and complexity in the video frame encoding and decoding process.
[0004] On the one hand, an embodiment of this application provides an image processing method, including:
[0005] Obtain bitstream data of a video frame, where the bitstream data includes encoded data of the video frame;
[0006] Decode the encoded data of the video frame based on the scale of a target block, and present the video frame; the target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0007] In an embodiment of this application, bitstream data of a video frame is obtained, the bitstream data includes encoded data of the video frame, the encoded data of the video frame is decoded based on the scale of a target block, and the video frame is presented. The target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. It can be seen that during the decoding process, by using the scale of the target block to determine the encoding and decoding information of the target block (such as the division method of the target block and whether there is a residual), the decoding process of the video frame can be optimized (such as without parsing the division flag bit of the target block), thereby improving the trade-off ratio between performance and complexity in the video frame encoding and decoding process.
[0008] On the one hand, an embodiment of this application provides an image processing method, including:
[0009] Obtain a video frame to be encoded;
[0010] Encode the video frame based on the scale of a target block to obtain bitstream data of the video frame; the target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0011] In the embodiments of the present application, a video frame to be encoded is obtained, and encoding processing is performed on the video frame based on the scale of a target block to obtain bitstream data of the video frame. The target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. It can be seen that during the encoding process, by using the scale of the target block to determine the encoding and decoding information of the target block (such as the partitioning method of the target block and whether there is a residual), the encoding process of the video frame can be optimized (such as not needing to configure the partitioning flag bit of the target block), thereby improving the trade-off ratio between performance and complexity during the video frame encoding and decoding process.
[0012] On the one hand, the embodiments of the present application provide an image processing apparatus, which includes:
[0013] An acquisition unit, configured to acquire bitstream data of a video frame, where the bitstream data includes the encoded data of the video frame;
[0014] A processing unit, configured to perform decoding processing on the encoded data of the video frame based on the scale of a target block and present the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0015] On the one hand, the embodiments of the present application provide an image processing apparatus, which includes:
[0016] An acquisition unit, configured to acquire a video frame to be encoded;
[0017] A processing unit, configured to perform encoding processing on the video frame based on the scale of a target block to obtain bitstream data of the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0018] Correspondingly, the present application provides a computer device, which includes:
[0019] A memory, in which a computer program is stored;
[0020] A processor, configured to load the computer program to implement the above image processing method.
[0021] Correspondingly, the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded and executed by a processor to implement the above image processing method.
[0022] Accordingly, the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image processing method.
[0023] In an embodiment of the present application, an encoding end obtains a video frame to be encoded, performs encoding processing on the video frame based on the scale of a target block, and obtains bitstream data of the video frame. The target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. A decoding end obtains the bitstream data of the video frame. The bitstream data includes the encoded data of the video frame, performs decoding processing on the encoded data of the video frame based on the scale of the target block, and presents the video frame. It can be seen that during the encoding and decoding process, by using the scale of the target block to determine the encoding and decoding information of the target block (such as the division method of the target block and whether there is a residual), the encoding and decoding process of the video frame can be optimized (such as no need to configure / parse the division flag bit of the target block), thereby improving the trade-off ratio between performance and complexity during the video frame encoding and decoding process. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1a It is a schematic diagram of a video encoding and decoding framework provided by an embodiment of the present application;
[0026] Figure 1b It is a schematic diagram of a block division method in the fourth-generation audio and video coding standard provided by an embodiment of the present application;
[0027] Figure 1c It is a schematic diagram of the division result of a coding tree unit provided by an embodiment of the present application;
[0028] Figure 1d It is an image processing scenario diagram provided by an embodiment of the present application;
[0029] Figure 2 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0030] Figure 3 It is a schematic diagram of a right boundary block and a lower boundary block provided by an embodiment of the present application;
[0031] Figure 4Flow chart of another image processing method provided by an embodiment of the present application;
[0032] Figure 5 Structural schematic diagram of an image processing apparatus provided by an embodiment of the present application;
[0033] Figure 6 Structural schematic diagram of another image processing apparatus provided by an embodiment of the present application;
[0034] Figure 7 Structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0036] The present application relates to technologies related to encoding and decoding. The following briefly introduces the technologies related to encoding and decoding:
[0037] Coding Unit (CU): It refers to the basic unit when encoding a video frame. During the encoding process, the coding unit can refer to the entire video frame (when the video frame is not divided), or a part of the area in the video frame (when the video frame is divided).
[0038] Intra prediction: It means that when encoding a coding unit, the information of other video frames except the video frame to which the coding unit belongs in the video is not referred to.
[0039] Inter prediction: It means that when encoding a coding unit, the information of the video frames adjacent to the video frame to which the coding unit belongs in the video is referred to.
[0040] Video signal: The video signal can be captured by a camera or generated by a computer device. During the encoding and decoding process, due to different statistical methods of the characteristics of the video signal, the corresponding encoding and decoding methods may also be different.
[0041] Taking the modern mainstream video coding technologies, such as the international video coding standard HEVC (High Efficiency Video Coding), the international video coding standard VVC (Versatile Video Coding), and the Chinese national video coding standard AVS (Audio Video Coding Standard) as examples, they adopt a hybrid coding framework.Figure 1a A schematic diagram of a video encoding and decoding framework provided in an embodiment of the present application. Figure 1a As shown in the figure, during the encoding and decoding process, the input original video signal is subjected to the following series of operations and processing:
[0042] 1) Block partition structure: The input image is processed into several non-overlapping processing units according to the size of the processing unit, and similar compression operations are performed on each processing unit. This processing unit is called a coding tree unit (CTU) or a largest coding unit (LCU). The CTU can be further divided into more refined units to obtain one or more basic coding units, called coding units (CU). Each CU is the most basic element in a coding process.
[0043] Figure 1b Schematic diagram of the block division method in the fourth generation audio and video codec standard provided in the embodiment of the present application. Figure 1b As shown in the figure, there are three types of partition trees in the fourth generation audio and video codec standard (AVS4): quadtree, binary tree (horizontal, vertical) and extended quadtree (horizontal, vertical). Through the recursive partitioning of the three partition trees, the entire CTU can be divided into a state that is more suitable for prediction. Different partitioning methods can be indicated by the partition flag bit. For example, the flag bit corresponding to the quad partition is 1, and the flag bit corresponding to the vertical binary partition is 0101.
[0044] Figure 1c A schematic diagram of the division result of a coding tree unit provided in an embodiment of the present application. Figure 1c As shown in FIG. 1 , the entire CTU is divided into 25 coding blocks. The coding device writes the block division result into the bitstream; for example, the coding device can write the block division result into the bitstream according to Figure 1b The partition flag bit and Figure 1c The coding tree corresponding to the CTU in the bitstream is used to indicate the division method of the CTU. The decoding device can obtain the division method of the CTU by parsing the division flag bit in the bitstream.
[0045] 2) Predictive Coding: This includes intra-frame prediction and inter-frame prediction. The original video signal is predicted by the selected reconstructed video signal to obtain a residual video signal. The content production device needs to select the most suitable predictive coding mode for the current CU from among many possible predictive coding modes and inform the content playback device.
[0046] a. Intra-frame prediction: The predicted signal comes from the already encoded and reconstructed area in the same image
[0047] b. Inter-frame prediction: The predicted signal comes from other images that have been encoded and are different from the current image (referred to as reference images).
[0048] 3) Transform coding and quantization: The residual video signal undergoes transform operations such as the Discrete Fourier Transform (DFT) and the Discrete Cosine Transform (DCT) to convert the signal into the transform domain, and the resulting coefficients are called transform coefficients. The signal in the transform domain is further subjected to a lossy quantization operation, losing some information, so that the quantized signal is conducive to compressed representation. In some video coding standards, there may be more than one transform method to choose from. Therefore, the content production device also needs to select one of the transforms for the current encoded CU and inform the content playback device. The fineness of quantization is usually determined by the Quantization Parameter (QP). A larger QP value means that a larger range of coefficients will be quantized to the same output, usually resulting in greater distortion and a lower bit rate; on the contrary, a smaller QP value means that a smaller range of coefficients will be quantized to the same output, usually resulting in less distortion and a corresponding higher bit rate.
[0049] 4) Entropy coding or statistical coding: The quantized transform domain signal will be statistically compressed encoded according to the frequency of each value, and finally a binary (0 or 1) compressed bitstream is output. At the same time, other information generated by the encoding, such as the selected mode, motion vectors, etc., also needs to be entropy encoded to reduce the bit rate. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. Common statistical coding methods include Variable Length Coding (VLC) or Context-Adaptive Binary Arithmetic Coding (CABAC).
[0050] 5) Loop Filtering: For an already encoded image, through inverse quantization, inverse transformation, and prediction compensation operations (the reverse operations of the above 2-4), a reconstructed decoded image can be obtained. Specifically, at the decoding end, for each CU, after the decoding device obtains the compressed bitstream, it first performs entropy decoding on the compressed bitstream to obtain the prediction coding mode information and the quantized transform coefficients. On the one hand, the decoding device inverse quantizes and inverse transforms each of the quantized transform coefficients to obtain the residual signal; on the other hand, the decoding device determines the prediction signal corresponding to the current CU according to the prediction coding mode information. Based on the residual signal and the prediction signal of the CU, the reconstructed signal of the CU can be obtained. Finally, the reconstructed value of the decoded image needs to undergo loop filtering operations to generate the final output signal. Compared with the original image, due to the influence of quantization, some information is different from the original image, resulting in distortion. Filtering operations on the reconstructed image, such as deblocking, Sample Adaptive Offset (SAO) filter, or Adaptive Loop Filter (ALF), etc., can effectively reduce the degree of distortion caused by quantization. Since these filtered reconstructed images will be used as references for subsequent encoded images to predict future signals, the above filtering operations are also called loop filtering, that is, filtering operations within the encoding loop.
[0051] Based on the above technologies related to encoding and decoding, an embodiment of the present application provides an image processing solution, which can improve the trade-off ratio between performance and complexity in the video frame encoding and decoding process. Figure 1d An image processing scenario diagram provided by an embodiment of the present application is shown in Figure 1dAs shown in the figure, the image processing scenario provided by this application includes a terminal device 101 and a server 102. The image processing solution provided by this application can be executed by the terminal device 101 or by the server 102. Among them, the terminal device may include, but is not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (abbreviated as MID), intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable devices, etc. This application embodiment does not make any limitations in this regard; the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This application embodiment does not make any limitations in this regard.
[0052] It should be noted that Figure 1d the numbers of the terminal device and the server in the figure are only for illustration and do not constitute an actual limitation of this application. The terminal device 101 and the server 102 can be connected by wired or wireless means, and this application does not make any restrictions in this regard.
[0053] The general process of the image processing solution provided by this application is as follows: The server 102 obtains a video frame to be encoded, and encodes the video frame based on the scale of the target block in the video frame to obtain the bitstream data of the video frame; among them, the target block is obtained by dividing the video frame; specifically, the target block can be a block to be divided (a block that will continue to be divided), or a coding unit (a block that will no longer be divided). The scale of the target block is used to determine the encoding and decoding information of the target block. In one implementation, the server 102 obtains the scale parameter of the video frame, and the scale parameter is used to optimize the encoding and decoding process of the target block (such as the division method of the target block, and determine whether there is a residual in the target block). The scale parameter of the video frame includes at least one of the following: the coding tree unit parameter of the video frame, the maximum inter-frame transformation parameter of the video frame, and the maximum intra-frame prediction parameter of the video frame. In the video to which the video frame belongs, the scale parameters of different types of video frames (such as key frames and non-key frames) are different, and the scale parameters of the same type of video frame can be the same or different. The server 102 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame. In another implementation, the target block is the lower boundary block or the right boundary block of the video frame, and the server 102 encodes the video frame based on the aspect ratio of the target block to obtain the bitstream data of the video frame.
[0054] Accordingly, the terminal device 101 obtains the bitstream data of the video frame, and the bitstream data includes the encoded data of the video frame. The terminal device 101 decodes the encoded data of the video frame based on the scale of the target block in the video frame and presents the video frame. In one implementation, the bitstream data further includes the scale parameter of the video frame. The terminal device 101 decodes the encoded data of the video frame based on the scale of the target block and the scale parameter and presents the video frame according to the decoding result. In another implementation, the target block is the lower boundary block or the right boundary block of the video frame. The terminal device 101 decodes the video frame based on the aspect ratio of the target block and presents the video frame according to the decoding result.
[0055] In the embodiments of the present application, the encoding end obtains the video frame to be encoded, encodes the video frame based on the scale of the target block, and obtains the bitstream data of the video frame. The target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. The decoding end obtains the bitstream data of the video frame, and the bitstream data includes the encoded data of the video frame. The decoding end decodes the encoded data of the video frame based on the scale of the target block and presents the video frame. It can be seen that in the encoding and decoding process, by using the scale of the target block to determine the encoding and decoding information of the target block (such as the division method of the target block and whether there is a residual), the encoding and decoding process of the video frame can be optimized (such as without configuring / parsing the division flag bit of the target block), thereby improving the trade-off ratio between performance and complexity in the video frame encoding and decoding process.
[0056] Based on the above image processing solution, the embodiments of the present application propose a more detailed image processing method. The image processing method proposed in the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.
[0057] Please refer to Figure 2 , Figure 2 which is a flowchart of an image processing method provided by the embodiments of the present application. This image processing method can be executed by a computer device; for example, it is executed by the Figure 1d terminal device 101 shown in
[0058] As Figure 2 shown, the image processing method may include the following steps S201 and S202:
[0059] S201. Obtain the bitstream data of the video frame.
[0060] The bitstream data includes the encoded data of a video frame. The encoded data may include the encoding and decoding information of a target block in the video frame. The target block is obtained by partitioning the video frame. The target block may be a block to be partitioned or a coding unit. The encoding and decoding information of the target block may include at least one of the following: the partitioning method of the target block, the residual flag bit, and the predictive coding mode. For example, when the target block is a block to be partitioned, the encoding and decoding information of the target block may include the partitioning method of the target block; when the target block is a coding unit, the encoding and decoding information of the target block may include the residual flag bit and the predictive coding mode of the target block.
[0061] S202. Decode the encoded data of the video frame based on the scale of the target block, and present the video frame.
[0062] The scale of the target block is used to determine the encoding and decoding information of the target block.
[0063] In one implementation, the bitstream data further includes the scale parameter of the video frame. The scale parameter is used to optimize the encoding and decoding process of the target block (such as the partitioning method of the target block and determining whether there is a residual in the target block). In the video to which the video frame belongs, the scale parameters of different types of video frames (such as key frames and non-key frames) may be the same or different. For example, the scale parameter is the coding tree unit parameter, and the coding tree unit parameters of key frames and non-key frames in the video are different. The scale parameters of the same type of video frame may be the same or different. The scale parameter of the video frame includes at least one of the following: the coding tree unit parameter of the video frame, the maximum inter-frame transform parameter of the video frame, and the maximum intra-prediction parameter of the video frame. The computer device decodes the encoded data of the video frame based on the scale of the target block and the scale parameter, and presents the video frame according to the decoding result.
[0064] The bitstream data includes at least one video sequence of the video to which the video frame belongs. Each video sequence includes the encoded data of at least one video frame. The encoded data of any video frame includes the picture header of the video frame. The scale parameter of the video frame may be a default value or configured based on the video frame. If the scale parameter of the video frame is a default value, the computer device decodes the encoded data of the video frame based on the default value (default scale parameter) and the scale of the target block, and presents the video frame. The default value is a parameter pre-agreed by the encoding device and the decoding device (computer device), or a parameter specified in the encoding and decoding standard. If the scale parameter of the video is configured based on the video frame, the scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs, or carried in the picture header of the video frame.
[0065] It can be understood that the scale parameters of the video frame can also be partially default values and partially based on the video frame configuration. In one embodiment, the scale parameters include the coding tree unit parameters of the video frame, the maximum inter-frame transformation parameter of the video frame, and the maximum intra-prediction parameter of the video frame; among them, the coding tree unit parameters of the video frame are based on the video frame configuration (included in the bitstream data), and the maximum inter-frame transformation parameter of the video frame and the maximum intra-prediction parameter of the video frame adopt default values; for example, the coding device can configure the coding tree unit parameter (CTU) of the video frame to be 256, and the coding tree unit parameter is carried in the bitstream data; in this case, the encoding and decoding device defaults that the maximum intra-prediction parameter of the key frame is 64, the maximum intra-prediction parameter of the non-key frame is 128 (that is, the maximum intra-prediction parameters of the key frame and the non-key frame are different), and the maximum inter-frame transformation parameter of the non-key frame is 128. When it is the default value, the maximum intra-prediction parameter of the key frame, the maximum intra-prediction parameter of the non-key frame, and the maximum inter-frame transformation parameter of the non-key frame can be not carried in the bitstream data; for another example, the coding device can configure the coding tree unit parameter (CTU) of the video frame to be 256, and the coding tree unit parameter is carried in the bitstream data; in this case, the encoding and decoding device defaults that the maximum intra-prediction parameter of the key frame is 64, the maximum intra-prediction parameter of the non-key frame is 64 (that is, the maximum intra-prediction parameters of the key frame and the non-key frame are the same), and the maximum inter-frame transformation parameter of the non-key frame is 128. When it is the default value, the maximum intra-prediction parameter of the key frame, the maximum intra-prediction parameter of the non-key frame, and the maximum inter-frame transformation parameter of the non-key frame can be not carried in the bitstream data.
[0066] It should be noted that when decoding the encoded data of the video frame by using the default value (default scale parameter) and the scale of the target block, compared with decoding the encoded data of the video frame by using the configured scale parameter and the scale of the target block, it is possible to not indicate the scale parameter (or reduce the data volume of the scale parameter) in the bitstream data of the video frame, and further compress the data volume of the bitstream data of the video frame. When the scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs, the scale parameter can be used to indicate the scale of the video frames in the video sequence to which the video frame belongs, and compared with carrying it in the picture header of the video frame, it can further compress the data volume of the bitstream data. When the scale parameter is carried in the picture header of the video frame, it can indicate the scale of a single view frame, which is more flexible compared with carrying it in the sequence header.
[0067] In one embodiment, the scale parameter includes the coding tree unit parameter corresponding to the video frame, and the encoded data of the video frame is derived based on the coding tree unit parameter. The computer device determines the predictive coding parameter of the video frame through the coding tree unit parameter and a preset ratio; wherein, the predictive coding parameter includes at least one of the maximum inter-frame transformation parameter of the video frame and the maximum intra-frame prediction parameter of the video frame, and the preset ratio is 1 / N, where N is a positive integer. For example, assume N = 2 and the CTU is 256. If the predictive coding parameter includes the maximum inter-frame transformation parameter of the video frame and the maximum intra-frame prediction parameter of the video frame, then the maximum inter-frame transformation parameter of the video frame is: 256 * 1 / 2 = 128, and the maximum intra-frame prediction parameter of the video frame is: 256 * 1 / 2 = 128. After obtaining the predictive coding parameter of the video frame, the computer device decodes the encoded data of the video frame based on the scale of the target block and the predictive coding parameter of the video frame, and presents the video frame according to the decoding result.
[0068] In another embodiment, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit. The process by which the computer device decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, then the computer device determines that the predictive coding mode of the target block is non-intra-frame prediction (such as using inter-frame prediction), or decodes the target block using the first context model. It can be understood that when the predictive coding mode only includes intra-frame prediction and inter-frame prediction, the computer device can directly determine that the predictive coding mode of the target block is inter-frame prediction. The first context model is dedicated to decoding the coding unit in the non-key frame whose scale is greater than the scale indicated by the maximum intra-frame prediction parameter. It should be noted that by decoding the target block using the first context model dedicated to the target block, the prediction accuracy of the target block can be further improved, and the decoding complexity can be reduced.
[0069] For example, assume the video frame is a non-key frame, the target block is a coding unit, the scale of the target block is 256 * 256, and the scale indicated by the maximum intra-frame prediction parameter is 128 (or 64) * 128 (or 64 * 64), then the computer device determines that the predictive coding mode of the target block is non-intra-frame prediction, or decodes the target block using the first context model.
[0070] In yet another embodiment, the scale parameter includes a maximum inter-frame transformation parameter, the video frame is a non-key frame, and the target block is a block to be divided (a block that needs to be further divided, which can be indicated by a division flag bit). The process by which the computer device decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the height of the target block is greater than the height indicated by the maximum inter-frame transformation parameter and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transformation parameter, then the computer device performs a horizontal binary division process on the target block, or the computer device can parse the division flag bit of the target block and perform a decoding process on the target block using a second context model. If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, then the computer device performs a vertical binary division process on the target block, or the computer device can parse the division flag bit of the target block and perform a decoding process on the target block using a third context model. The second context model is dedicated to decoding a block to be divided in a non-key frame whose height is greater than the height indicated by the maximum inter-frame transformation parameter and whose width is less than or equal to the width indicated by the maximum inter-frame transformation parameter; the third context model is dedicated to decoding a block to be divided in a non-key frame whose height is less than or equal to the height indicated by the maximum inter-frame transformation parameter and whose width is greater than the width indicated by the maximum inter-frame transformation parameter. It should be noted that when the target block meets the conditions, decoding the target block using the dedicated second context model or third context model can further improve the prediction accuracy of the target block and reduce the decoding complexity.
[0071] For example, assume that the target block is a block to be divided, the video frame is a non-key frame, the scale (width * height) of the target block is 256 * 128, the maximum inter-frame transformation parameter is 128, and the scale (width * height) indicated by the maximum inter-frame transformation parameter is 128 * 128, that is, the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter and the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter. The computer device performs a vertical binary division process on the target block, or performs a decoding process on the target block using a third context model.
[0072] In another embodiment, the scale parameter includes a maximum inter-frame transform parameter, and the target block is a coding unit (which does not need to be further divided and can be indicated by a division flag bit). The process by which the computer device decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum inter-frame transform parameter, it is determined that there is no residual in the target block, or the computer device can parse the residual flag bit of the target block and perform decoding processing on the target block using a fourth context model. The fourth context model is dedicated to decoding coding units in the video frame whose scale is greater than the scale indicated by the maximum inter-frame transform parameter. It should be noted that when the target block meets the conditions, decoding the target block using the dedicated fourth context model can further improve the prediction accuracy of the target block and reduce the decoding complexity.
[0073] For example, assume that the target block is a coding unit, the scale of the target block is 256*256, and the maximum inter-frame transform parameter is 128 (the indicated scale is 128*128), then the computer device determines that there is no residual in the target block, or performs decoding processing on the target block using the fourth context model.
[0074] In another embodiment, the scale parameter includes a maximum intra-prediction parameter, and the video frame is a key frame. The process by which the computer device decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum intra-prediction parameter, the computer device performs quadtree partitioning on the target block, or the computer device can parse the partitioning flag bit of the target block and perform decoding processing on the target block using a fifth context model. The fifth context model is dedicated to decoding blocks in the key frame whose scale is greater than the scale indicated by the maximum intra-prediction parameter. It should be noted that decoding the target block using the fifth context model specifically for the target block can further improve the prediction accuracy of the target block and reduce the decoding complexity.
[0075] For example, assume that the video frame is a key frame, the scale of the target block is 256*256, and the maximum intra-prediction parameter is 128 (or 64) and the indicated scale is 128*128 (or 64*64), then the computer device performs quadtree partitioning on the target block, or performs decoding processing on the target block using the fifth context model.
[0076] In another implementation, the target block is a lower boundary block or a right boundary block of the video frame. The computer device decodes the video frame based on the aspect ratio of the target block and presents the video frame according to the decoding result. The right boundary block refers to the block at the right boundary of the video frame, and the lower boundary block refers to the block at the lower boundary of the video frame. Figure 3Schematic diagram of the right boundary block and the lower boundary block provided by the embodiments of the present application. As Figure 3 shown, there is an overlapping part between the right boundary block and the right side boundary of the video frame, and there is an overlapping part between the lower boundary block and the lower side boundary of the video frame.
[0077] In one embodiment, the target block is the right boundary block of the video frame. The process of the computer device decoding the encoded data of the video frame includes: if the ratio of the height to the width of the target block is less than the first ratio threshold, the computer device performs vertical binary partitioning on the target block, or the computer device can parse the partitioning flag bit of the target block and perform decoding processing on the target block using the sixth context model; correspondingly, if the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, the computer device performs horizontal binary partitioning on the target block, or the computer device can parse the partitioning flag bit of the target block and perform decoding processing on the target block using the seventh context model. The sixth context model is dedicated to decoding the right boundary blocks in the video frame whose ratio of height to width is less than the first ratio threshold, and the seventh context model is dedicated to decoding the right boundary blocks in the video frame whose ratio of height to width is greater than or equal to the first ratio threshold. It should be noted that when the target block meets the conditions, decoding the target block through the dedicated sixth context model or seventh context model can further improve the prediction accuracy of the target block and reduce the decoding complexity.
[0078] For example, let the first ratio threshold be 8, the target block be the right boundary block of the video frame, and the scale (width * height) of the target block be 64 * 128. Then the ratio of the height to the width of the target block is 128 / 64 = 2 < 8, and the computer device performs vertical binary partitioning on the target block or performs decoding processing on the target block using the sixth context model.
[0079] In another embodiment, the target block is the lower boundary block of a video frame. The process of the computer device decoding the encoded data of the video frame includes: if the ratio of the width to the height of the target block is less than the second ratio threshold, the computer device performs a horizontal binary partitioning process on the target block, or the computer device can parse the partitioning flag bit of the target block and perform decoding processing on the target block using the eighth context model; correspondingly, if the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, the computer device performs a vertical binary partitioning process on the target block, or the computer device can parse the partitioning flag bit of the target block and perform decoding processing on the target block using the ninth context model. The eighth context model is dedicated to decoding the lower boundary blocks in the video frame whose ratio of width to height is less than the second ratio threshold, and the ninth context model is dedicated to decoding the lower boundary blocks in the video frame whose ratio of width to height is greater than or equal to the second ratio threshold. It should be noted that when the target block meets the conditions, decoding the target block using the dedicated eighth context model or ninth context model can further improve the prediction accuracy of the target block and reduce the decoding complexity.
[0080] For example, let the second ratio threshold be 8, the target block be the lower boundary block of the video frame, and the scale (width * height) of the target block be 128 * 16. Then the ratio of the width to the height of the target block is 128 / 16 = 8 ≥ 8, and the computer device performs a vertical binary partitioning process on the target block or decodes the target block using the ninth context model.
[0081] In the embodiments of the present application, the bitstream data of the video frame is obtained. The bitstream data includes the encoded data of the video frame. Based on the scale of the target block, the encoded data of the video frame is decoded and the video frame is presented. The target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. It can be seen that during the decoding process, determining the encoding and decoding information of the target block (such as the partitioning method of the target block and whether there is a residual) through the scale of the target block can optimize the decoding process of the video frame (such as without parsing the partitioning flag bit of the target block), thereby improving the trade-off ratio between performance and complexity during the video frame encoding and decoding process.
[0082] Please refer to Figure 4 , Figure 4 which is a flowchart of another image processing method provided by the embodiments of the present application. This image processing method can be executed by a computer device; for example, it is executed by the Figure 1d server 102 shown in
[0083] As Figure 4 shown, this image processing method may include the following steps S401 and S402:
[0084] S401. Obtain the video frame to be encoded.
[0085] S402. Encode the video frame based on the scale of the target block to obtain the bitstream data of the video frame.
[0086] The target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. The target block can be a block to be partitioned (a block that still needs to be further partitioned), or a coding unit (a block that does not need to be further partitioned). The encoding and decoding information of the target block can include at least one of the following: the partitioning method of the target block, the residual flag bit, the predictive coding mode; for example, when the target block is a block to be partitioned, the encoding and decoding information of the target block can include the partitioning method of the target block; when the target block is a coding unit, the encoding and decoding information of the target block can include the residual flag bit and the predictive coding mode of the target block.
[0087] In one implementation, the computer device obtains the scale parameter of the video frame, and the scale parameter is used to optimize the encoding and decoding process of the target block (such as the partitioning method of the target block, and determining whether there is a residual in the target block). The scale parameter of the video frame can be a default value, or can be configured based on the video frame. In the video to which the video frame belongs, the scale parameters of different types of video frames (such as key frames and non-key frames) can be the same or different; for example, the scale parameter is the coding tree unit parameter, and the coding tree unit parameters of key frames and non-key frames in the video are different. The scale parameters of the same type of video frame can be the same or different. The scale parameter of the video frame includes at least one of the following: the coding tree unit parameter of the video frame, the maximum inter-frame transform parameter of the video frame, the maximum intra-prediction parameter of the video frame. After obtaining the scale parameter of the video frame, the computer device encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame.
[0088] If the scale parameter of the video frame is a default value, the computer device encodes the video frame based on the scale of the target block and the default value (default scale parameter) to obtain the bitstream data of the video frame. The default value is a parameter agreed in advance by the encoding device and the decoding device (computer device), or a parameter specified in the encoding and decoding standard, and does not need to be carried in the bitstream data of the video frame.
[0089] If the scale parameter of a video frame is configured based on the video frame, the computer device encodes the video frame based on the scale of the target block and the scale parameter to obtain the encoded data of the video frame, and the encoded data of the video frame includes the picture header of the video frame. After obtaining the encoded data of the video frame, on the one hand, the computer device adds the encoded data of the video frame to the video sequence to which the video frame belongs; on the other hand, the computer device adds the scale parameter of the video frame to the picture header of the video frame, or configures the sequence header of the video sequence to which the video frame belongs based on the scale parameter of the video frame to obtain the bitstream data of the video frame. In one implementation, the scale parameters of video frames belonging to the same video sequence are the same. In this case, the sequence header can be configured once based on the scale parameter of the video frames in any one video sequence.
[0090] The scale parameter of the video frame can also be partially default values and partially configured based on the video frame. In one embodiment, the scale parameter includes the coding tree unit parameter of the video frame, the maximum inter-frame transformation parameter of the video frame, and the maximum intra-prediction parameter of the video frame; among them, the coding tree unit parameter of the video frame is configured based on the video frame (included in the bitstream data), and the maximum inter-frame transformation parameter and the maximum intra-prediction parameter of the video frame are default values; for example, the coding device can configure the coding tree unit parameter (CTU) of the video frame to be 256, and the coding tree unit parameter is carried in the bitstream data; in this case, the codec device defaults that the maximum intra-prediction parameter of the key frame is 64, the maximum intra-prediction parameter of the non-key frame is 128 (that is, the maximum intra-prediction parameters of the key frame and the non-key frame are different), and the maximum inter-frame transformation parameter of the non-key frame is 128. When they are default values, the maximum intra-prediction parameter of the key frame, the maximum intra-prediction parameter of the non-key frame, and the maximum inter-frame transformation parameter of the non-key frame may not be carried in the bitstream data; for another example, the coding device can configure the coding tree unit parameter (CTU) of the video frame to be 256, and the coding tree unit parameter is carried in the bitstream data; in this case, the codec device defaults that the maximum intra-prediction parameter of the key frame is 64, the maximum intra-prediction parameter of the non-key frame is 64 (that is, the maximum intra-prediction parameters of the key frame and the non-key frame are the same), and the maximum inter-frame transformation parameter of the non-key frame is 128. When they are default values, the maximum intra-prediction parameter of the key frame, the maximum intra-prediction parameter of the non-key frame, and the maximum inter-frame transformation parameter of the non-key frame may not be carried in the bitstream data.
[0091] It should be noted that when decoding the encoded data of a video frame by using the default value (default scale parameter) and the scale of the target block, compared with decoding the encoded data of the video frame by using the configured scale parameter and the scale of the target block, it is possible to avoid indicating the scale parameter in the bitstream data of the video frame (or reducing the data volume of the scale parameter), thereby further compressing the data volume of the bitstream data of the video frame. When the scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs, the scale parameter can be used to indicate the scale of the video frames in the video sequence to which the video frame belongs, and compared with carrying it in the picture header of the video frame, the data volume of the bitstream data can be further compressed. When the scale parameter is carried in the picture header of the video frame, it can indicate the scale of a single view frame, which is more flexible than carrying it in the sequence header.
[0092] In one embodiment, the scale parameter includes the coding tree unit parameter corresponding to the video frame, and the coding tree unit parameter is used to derive the encoded data of the video. The computer device determines the prediction coding parameter of the video frame based on the coding tree unit parameter and a preset ratio; wherein, the prediction coding parameter includes at least one of the maximum inter-frame transformation parameter of the video frame and the maximum intra-frame prediction parameter of the video frame, and the preset ratio is 1 / N, where N is a positive integer. For key frames and non-key frames, the preset ratios corresponding to the maximum inter-frame transformation parameters can be different (that is, the maximum inter-frame transformation parameters of key frames and non-key frames obtained based on the coding tree unit parameter and the preset ratio are different).
[0093] For example, let N = 4, the CTU be 256, the scale indicated by the CTU be 256*256, the prediction coding parameter include the maximum inter-frame transformation parameter of the video frame and the maximum intra-frame prediction parameter of the video frame, the preset ratio corresponding to the key frame be 1 / 4, and the preset ratio corresponding to the non-key frame be 1 / 2; then the maximum intra-frame prediction parameter of the key frame is 256*1 / 4 = 64, and the scale indicated by the maximum intra-frame prediction parameter of the key frame is: 64*64; the maximum intra-frame prediction parameter of the non-key frame is 256*1 / 2 = 128, the scale indicated by the maximum intra-frame prediction parameter of the non-key frame is: 128*128, the maximum inter-frame transformation parameter of the non-key frame is: 256*1 / 2 = 128, and the scale indicated by the maximum inter-frame transformation parameter of the non-key frame is: 128*128. After obtaining the prediction coding parameter of the video frame, the computer device encodes the encoded data of the video frame based on the scale of the target block and the prediction coding parameter of the video frame to obtain the bitstream data of the video frame.
[0094] In another embodiment, the scale parameter includes the maximum intra prediction parameter, the video frame is a non-key frame, and the target block is a coding unit. The process by which the computer device encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: if the scale of the target block is greater than the scale indicated by the maximum intra prediction parameter, the computer device determines that the prediction coding mode of the target block is non-intra prediction (such as using inter prediction), or encodes the target block using the first context model. It can be understood that when the prediction coding mode only includes intra prediction and inter prediction, the computer device can directly determine that the prediction coding mode of the target block is inter prediction. The first context model is dedicated to encoding coding units in non-key frames whose scale is greater than the scale indicated by the maximum intra prediction parameter. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated first context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0095] For example, assume that the video frame is a non-key frame, the target block is a coding unit, the scale of the target block is 256*256, the maximum intra prediction parameter is 128, and the scale indicated by the maximum intra prediction parameter is 128*128. Then the computer device determines that the prediction coding mode of the target block is non-intra prediction, or encodes the target block using the first context model.
[0096] In another embodiment, the scale parameter includes the maximum inter transform parameter, the video frame is a non-key frame, and the target block is a block to be partitioned. The process by which the computer device encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: if the height of the target block is greater than the height indicated by the maximum inter transform parameter and the width of the target block is less than or equal to the width indicated by the maximum inter transform parameter, the computer device performs horizontal binary partitioning on the target block, or encodes the target block using the second context model. If the height of the target block is less than or equal to the height indicated by the maximum inter transform parameter and the width of the target block is greater than the width indicated by the maximum inter transform parameter, the computer device performs vertical binary partitioning on the target block, or encodes the target block using the third context model. The second context model is dedicated to encoding blocks to be partitioned whose height is greater than the height indicated by the maximum inter transform parameter and whose width is less than or equal to the width indicated by the maximum inter transform parameter; the third context model is dedicated to encoding blocks to be partitioned whose height is less than the height indicated by the maximum inter transform parameter and whose width is greater than the width indicated by the maximum inter transform parameter. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated second context model or third context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0097] For example, assume that the target block is the block to be partitioned, the video frame is a non-key frame, the scale (width * height) of the target block is 256 * 128, the maximum inter-frame transformation parameter is 128, and the scale (width * height) indicated by the maximum inter-frame transformation parameter is 128 * 128. That is, the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, and the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter. The computer device performs vertical binary partitioning on the target block, or encodes the target block using the third context model.
[0098] In yet another embodiment, the scale parameter includes the maximum inter-frame transformation parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction. The process of the computer device encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: if the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, the computer device determines that there is no residual in the target block, or encodes the target block using the fourth context model. In this case, the computer device can configure the residual flag bit of the target block to be empty (or not configure the residual flag bit of the target block). The fourth context model is dedicated to encoding coding units in the video frame whose predictive coding mode is inter-frame prediction and whose scale is greater than the scale indicated by the maximum inter-frame transformation parameter. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated fourth context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0099] For example, assume that the target block is a coding unit, the predictive coding mode of the coding unit is inter-frame prediction, the scale of the target block is 256 * 256, the maximum inter-frame transformation parameter is 128, and the scale indicated by the maximum inter-frame transformation parameter is 128 * 128. Then the computer device determines that there is no residual in the target block, or encodes the target block using the fourth context model.
[0100] In another embodiment, the scale parameter includes the maximum intra-frame prediction parameter, and the video frame is a key frame. The process of the computer device encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: if the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, the computer device partitions the target block, or encodes the target block using the fifth context model; where the partitioning method includes at least one of binary partitioning and quaternary partitioning. The fifth context model is dedicated to encoding blocks in the key frame whose scale is greater than the scale indicated by the maximum inter-frame transformation parameter. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated fifth context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0101] For example, assume that the video frame is a key frame, the scale of the target block is 256*256, the maximum intra-prediction parameter is 128 (or 64), and the scale indicated by the maximum intra-prediction parameter is 128*128 (or 64*64). Then, the computer device continues to divide the target block until the scale of the target block is not greater than the scale indicated by the maximum intra-prediction parameter, or encodes the target block using the fifth context model.
[0102] In another embodiment, the target block is a right boundary block or a bottom boundary block of the video frame. The computer device encodes the video frame based on the aspect ratio of the target block to obtain the bitstream data of the video frame. The right boundary block refers to the block located at the right boundary of the video frame, and the bottom boundary block refers to the block located at the bottom boundary of the video frame. For details, please refer to Figure 3 .
[0103] In one embodiment, when the target block is a right boundary block of the video frame, the process of the computer device encoding the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes: if the ratio of the height to the width of the target block is less than the first ratio threshold, the computer device performs vertical binary partitioning on the target block, or encodes the target block using the sixth context model; correspondingly, if the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, the computer device performs horizontal binary partitioning on the target block, or encodes the target block using the seventh context model. The first ratio threshold can be 2 M , where M is a positive integer. The sixth context model is dedicated to encoding the right boundary blocks in the video frame whose ratio of height to width is less than the first ratio threshold, and the seventh context model is dedicated to encoding the right boundary blocks in the video frame whose ratio of height to width is greater than or equal to the first ratio threshold. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated sixth context model or seventh context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0104] For example, assume that the first ratio threshold is 8, the target block is a right boundary block of the video frame, and the scale (height * width) of the target block is 128*64. Then, the ratio of the height to the width of the target block is 128 / 64 = 2 < 8, and the computer device performs vertical binary partitioning on the target block, or encodes the target block using the sixth context model.
[0105] In another embodiment, the target block is the lower boundary block of a video frame. The process by which the computer device encodes the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes: if the ratio of the width to the height of the target block is less than a second ratio threshold, the computer device performs a horizontal binary partitioning process on the target block, or encodes the target block using an eighth context model; correspondingly, if the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, the computer device performs a vertical binary partitioning process on the target block, or encodes the target block using a ninth context model. The second ratio threshold may be 2 M , where M is a positive integer. The first ratio threshold and the second ratio threshold may be the same or different. The eighth context model is dedicated to encoding the lower boundary blocks in the video frame whose ratio of width to height is less than the second ratio threshold, and the ninth context model is dedicated to encoding the lower boundary blocks in the video frame whose ratio of width to height is greater than or equal to the second ratio threshold. It should be noted that when the target block meets the conditions, encoding the target block using the dedicated eighth context model or ninth context model can further improve the prediction accuracy of the target block and reduce the encoding complexity.
[0106] For example, let the second ratio threshold be 8, the target block be the lower boundary block of a video frame, and the scale (width * height) of the target block be 128 * 16. Then the ratio of the width to the height of the target block is 128 / 16 = 8 ≥ 8, and the computer device performs a vertical binary partitioning process on the target block, or encodes the target block using the ninth context model.
[0107] In yet another embodiment, the target block is a basic coding unit in a video frame. The coding process of the video frame involves the calculation of high-frequency coefficients and low-frequency coefficients of each basic coding unit. The high-frequency coefficients are the coefficients of the high-frequency part (frequency greater than the frequency threshold) obtained after the signal or image undergoes discrete cosine transform or wavelet transform or other transforms. The low-frequency coefficients are the coefficients of the low-frequency part (frequency less than or equal to the frequency threshold) obtained after the signal or image undergoes wavelet decomposition or other transforms. The computer device encodes the video frame based on the scale and scale parameters of the target block. The process of obtaining the bitstream data of the video frame includes: setting the high-frequency coefficients greater than the high-frequency coefficient threshold to a preset value (such as 0); for example, setting the high-frequency coefficient threshold to 32. If the transform coefficient matrix is 64*64, the computer device can set the coefficients in the 32nd to 63rd columns and the 32nd to 63rd rows of the transform coefficient matrix to 0 (that is, only retain the coefficients in the 0th to 31st rows and 0th to 31st columns of the transform coefficient matrix). It should be noted that during the process of calculating the transform coefficients, skipping the calculation of the high-frequency coefficients greater than the high-frequency coefficient threshold (that is, setting the high-frequency coefficients greater than the high-frequency coefficient threshold to a preset value) can reduce the coding complexity of the coding device (computer device) and have little impact on the image quality of the video frame.
[0108] In the embodiments of the present application, a video frame to be encoded is obtained, and the video frame is encoded based on the scale of the target block to obtain the bitstream data of the video frame. The target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block. It can be seen that during the encoding process, determining the encoding and decoding information of the target block (such as the division method of the target block and whether there is a residual) through the scale of the target block can optimize the encoding process of the video frame (such as not needing to configure the division flag bit of the target block), thereby improving the trade-off ratio between performance and complexity during the video frame encoding and decoding process.
[0109] The above has elaborated in detail the method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the following provides the device of the embodiments of the present application.
[0110] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an image processing device provided by an embodiment of the present application. Figure 5 The shown image processing device can be installed in a computer device, and the computer device can specifically be Figure 1d the terminal device 101 shown in Figure 5 The shown image processing device can be used to execute some or all of the functions described in the method embodiment above. Please refer to Figure 2 , Figure 5 This image processing device includes:
[0111] An acquisition unit 501, configured to acquire bitstream data of a video frame, where the bitstream data includes encoded data of the video frame;
[0112] A processing unit 502, configured to perform decoding processing on the encoded data of the video frame based on a scale of a target block and present the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine encoding and decoding information of the target block.
[0113] In one implementation, the processing unit 502 is configured to perform decoding processing on the encoded data of the video frame based on the scale of the target block, specifically:
[0114] Perform decoding processing on the encoded data of the video frame based on the scale of the target block and scale parameters;
[0115] Wherein, the scale parameters are used to optimize the encoding and decoding process of the target block, the scale parameters are default values, or are carried in the bitstream data; the scale parameters include at least one of the following: coding tree unit parameters, maximum inter-frame transformation parameters, and maximum intra-frame prediction parameters.
[0116] In one implementation, the scale parameters include coding tree unit parameters corresponding to the video frame, and the encoded data of the video frame is derived based on the coding tree unit parameters; the processing unit 502 is configured to perform decoding processing on the encoded data of the video frame based on the scale of the target block and the scale parameters, specifically:
[0117] Determine prediction coding parameters of the video frame through the coding tree unit parameters and a preset ratio;
[0118] Perform decoding processing on the encoded data of the video frame based on the scale of the target block and the prediction coding parameters of the video frame;
[0119] Wherein, the prediction coding parameters include at least one of the maximum inter-frame transformation parameters of the video frame and the maximum intra-frame prediction parameters of the video frame.
[0120] In one implementation, the scale parameters include maximum intra-frame prediction parameters, the video frame is a non-key frame, and the target block is a coding unit; the process of the processing unit 502 performing decoding processing on the encoded data of the video frame based on the scale of the target block and the scale parameters includes:
[0121] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameters, determine that the prediction coding mode of the target block is non-intra-frame prediction, or perform decoding processing on the target block using a first context model.
[0122] In one implementation, the scale parameters include maximum inter-frame transformation parameters, the video frame is a non-key frame, and the target block is a block to be partitioned; the process of the processing unit 502 performing decoding processing on the encoded data of the video frame based on the scale of the target block and the scale parameters includes:
[0123] If the height of the target block is greater than the height indicated by the maximum inter-frame transformation parameter, and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transformation parameter, then perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using a second context model;
[0124] If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter, and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, then perform vertical binary partitioning on the target block, or perform decoding processing on the target block using a third context model.
[0125] In one implementation, the scale parameter includes the maximum inter-frame transformation parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction; the process of the processing unit 502 decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0126] If the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, then determine that there is no residual in the target block, or perform decoding processing on the target block using a fourth context model.
[0127] In one implementation, the scale parameter includes the maximum intra-frame prediction parameter, and the video frame is a key frame; the process of the processing unit 502 decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0128] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, then perform quadtree partitioning on the target block, or perform decoding processing on the target block using a fifth context model.
[0129] In one implementation, the bitstream data includes at least one video sequence of the video to which the video frame belongs, each video sequence includes the encoded data of at least one video frame, and the encoded data of any video frame includes the picture header of the video frame;
[0130] The scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs, or carried in the picture header of the video frame;
[0131] The video includes key frames and non-key frames, and the scale parameters of the key frames and non-key frames are different.
[0132] In one implementation, the target block is the right boundary block of the video frame, and the process of the processing unit 502 decoding the encoded data of the video frame based on the scale of the target block includes:
[0133] If the ratio of the height to the width of the target block is less than the first ratio threshold, perform vertical binary partitioning on the target block, or perform decoding processing on the target block using the sixth context model;
[0134] If the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using the seventh context model.
[0135] In one implementation, the target block is the lower boundary block of a video frame. The process of the processing unit 502 performing decoding processing on the encoded data of the video frame based on the attribute information of the target block includes:
[0136] If the ratio of the width to the height of the target block is less than the second ratio threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using the eighth context model;
[0137] If the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, perform vertical binary partitioning on the target block, or perform decoding processing on the target block using the ninth context model.
[0138] According to an embodiment of the present application, Figure 2 Some of the steps involved in the image processing method shown can be performed by Figure 5 each unit in the image processing device shown. For example, Figure 2 the step S201 shown in Figure 5 can be performed by the obtaining unit 501 shown, and the step S202 can be performed by Figure 5 the processing unit 502 shown. Figure 5 Each unit in the image processing device shown can be separately or all combined into one or several other units to form, or some of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units are realized by one unit. In other embodiments of the present application, the image processing device may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by multiple units collaborating.
[0139] According to another embodiment of the present application, it can be achieved by running, on a general computing device such as a computer device including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), software that can execute as Figure 2A computer program (including program code) for each step involved in the corresponding method shown is used to construct an image processing apparatus as shown in Figure 5 and to implement the image processing method of the embodiment of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above-mentioned computing device through the computer-readable recording medium, and run therein.
[0140] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the image processing apparatus provided in the embodiment of the present application are similar to those of the image processing method in the method embodiment of the present application. For the principle of problem-solving and the beneficial effects of the method, reference can be made to the method embodiment. For the sake of brevity, it will not be elaborated here.
[0141] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of another image processing apparatus provided in the embodiment of the present application. Figure 6 The image processing apparatus shown can be mounted in a computer device, and the computer device can specifically be Figure 1d the server 102 shown in Figure 6 The image processing apparatus shown can be used to execute some or all of the functions in the method embodiment described above. Please refer to Figure 4 ,and the image processing apparatus includes: Figure 6 An acquisition unit 601, configured to acquire a video frame to be encoded;
[0142] A processing unit 602, configured to perform encoding processing on the video frame based on the scale of the target block to obtain bitstream data of the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0143] In one implementation, the processing unit 602 is configured to perform encoding processing on the video frame based on the scale of the target block to obtain bitstream data of the video frame, and specifically includes:
[0144] Obtain a scale parameter of the video frame, where the scale parameter is used to optimize the encoding and decoding process of the target block, and the scale parameter is a default value or is configured based on the video frame; the scale parameter includes at least one of the following: coding tree unit parameter, maximum inter-frame transform parameter, maximum intra-frame prediction parameter;
[0145] Perform encoding processing on the video frame based on the scale of the target block and the scale parameter to obtain bitstream data of the video frame.
[0146]
[0147] In one embodiment, the scale parameter includes the coding tree unit parameter corresponding to the video frame, and the coding tree unit parameter is used to derive the coding data of the video; the process of the processing unit 602 encoding the video frame based on the scale of the target block and the scale parameter includes:
[0148] Determine the predictive coding parameters of the video frame based on the coding tree unit parameter and a preset ratio;
[0149] Among them, the predictive coding parameters include at least one of the maximum inter-frame transform parameter of the video frame and the maximum intra-frame prediction parameter of the video frame.
[0150] In one embodiment, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit; the process of the processing unit 602 encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0151] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, then determine that the predictive coding mode of the target block is non-intra-frame prediction, or encode the target block using the first context model.
[0152] In one embodiment, the scale parameter includes the maximum inter-frame transform parameter, the video frame is a non-key frame, and the target block is a block to be partitioned; the process of the processing unit 602 encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0153] If the height of the target block is greater than the height indicated by the maximum inter-frame transform parameter and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transform parameter, then perform horizontal binary partitioning on the target block, or encode the target block using the second context model;
[0154] If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transform parameter and the width of the target block is greater than the width indicated by the maximum inter-frame transform parameter, then perform vertical binary partitioning on the target block, or encode the target block using the third context model.
[0155] In one embodiment, the scale parameter includes the maximum inter-frame transform parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction; the process of the processing unit 602 encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0156] If the scale of the target block is greater than the scale indicated by the maximum inter-frame transform parameter, then determine that there is no residual in the target block, or encode the target block using the fourth context model.
[0157] In one embodiment, the scale parameter includes the maximum intra prediction parameter, and the video frame is a key frame; the process by which the processing unit 602 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0158] If the scale of the target block is greater than the scale indicated by the maximum intra prediction parameter, the target block is partitioned, or the target block is encoded using the fifth context model; wherein, the partitioning method includes at least one of binary partitioning and quadtree partitioning.
[0159] In one embodiment, the scale parameter is configured based on the video frame; the process by which the processing unit 602 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0160] Configuring the sequence header of the video sequence to which the video frame belongs based on the scale parameter, or adding the scale parameter to the picture header of the video frame;
[0161] Wherein, the video includes key frames and non-key frames, and the scale parameters of the key frames and non-key frames are different.
[0162] In one embodiment, the target block is the right boundary block of the video frame, and the process by which the processing unit 602 encodes the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes:
[0163] If the ratio of the height to the width of the target block is less than the first ratio threshold, the target block is vertically binary partitioned, or the target block is encoded using the sixth context model;
[0164] If the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, the target block is horizontally binary partitioned, or the target block is encoded using the seventh context model.
[0165] In one embodiment, the target block is the bottom boundary block of the video frame, and the process by which the processing unit 602 encodes the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes:
[0166] If the ratio of the width to the height of the target block is less than the second ratio threshold, the target block is horizontally binary partitioned, or the target block is encoded using the eighth context model;
[0167] If the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, the target block is vertically binary partitioned, or the target block is encoded using the ninth context model.
[0168] In one embodiment, the process of the processing unit 602 encoding a video frame based on the scale of a target block to obtain the bitstream data of the video frame includes:
[0169] Set high-frequency coefficients greater than the high-frequency coefficient threshold to a preset value.
[0170] According to an embodiment of the present application, Figure 4 Some of the steps involved in the image processing method shown can be performed by Figure 6 each unit in the image processing device shown. For example, Figure 4 the step S401 shown in can be performed by Figure 6 the acquisition unit 601 shown, and the step S402 can be performed by Figure 6 the processing unit 602 shown. Figure 6 Each unit in the image processing device shown can be separately or all combined into one or several other units to form, or some of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the image processing device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by multiple units collaborating.
[0171] According to another embodiment of the present application, it is possible to construct an image processing device as shown in by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in on a general computing device such as a computer device including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), and to implement the image processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein. Figure 4 in, and to Figure 6 implement the image processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0172] Based on the same inventive concept, the principle of solving problems and the beneficial effects of the image processing device provided in the embodiments of the present application are similar to the principle of solving problems and the beneficial effects of the image processing method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity of description, they will not be elaborated here.
[0173] Please refer to Figure 7 , Figure 7The following is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device may be a terminal device or a server. As Figure 7 shown, the computer device at least includes a processor 701, a communication interface 702, and a memory 703. Among them, the processor 701, the communication interface 702, and the memory 703 can be connected through a bus or other means. Among them, the processor 701 (or Central Processing Unit (CPU)) is the computing core and control core of the computer device. It can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power-on and power-off instructions sent by an object to the computer device and control the computer device to perform power-on and power-off operations; Another example is that the CPU can transmit various interactive data between the internal structures of the computer device, and so on. The communication interface 702 may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be controlled by the processor 701 to be used for receiving and transmitting data; the communication interface 702 can also be used for the transmission and interaction of internal data of the computer device. The memory 703 (Memory) is the memory device in the computer device and is used to store programs and data. It can be understood that the memory 703 here can include both the built-in memory of the computer device and, of course, the extended memory supported by the computer device. The memory 703 provides a storage space, and the operating system of the computer device is stored in this storage space, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.
[0174] An embodiment of the present application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the processing system of the computer device is stored in this storage space. And, a computer program suitable for being loaded and executed by the processor 701 is also stored in this storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0175] In one embodiment, the computer device is a decoding device, and the processor 701 performs the following operations by running the computer program in the memory 703:
[0176] Obtain the bitstream data of the video frame, where the bitstream data includes the encoded data of the video frame;
[0177] Decode the encoded data of the video frame based on the scale of the target block and present the video frame; the target block is obtained by dividing the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0178] As an optional embodiment, a specific embodiment in which the processor 701 decodes the encoded data of the video frame based on the scale of the target block is:
[0179] Decode the encoded data of the video frame based on the scale of the target block and the scale parameter;
[0180] Among them, the scale parameter is used to optimize the encoding and decoding process of the target block, and the scale parameter is a default value or carried in the bitstream data; the scale parameter includes at least one of the following: coding tree unit parameter, maximum inter-frame transform parameter, maximum intra-frame prediction parameter.
[0181] As an optional embodiment, the scale parameter includes the coding tree unit parameter corresponding to the video frame, and the encoded data of the video frame is derived based on the coding tree unit parameter; a specific embodiment in which the processor 701 decodes the encoded data of the video frame based on the scale of the target block and the scale parameter is:
[0182] Determine the predictive coding parameters of the video frame through the coding tree unit parameter and a preset ratio;
[0183] Decode the encoded data of the video frame based on the scale of the target block and the predictive coding parameters of the video frame;
[0184] Among them, the predictive coding parameters include at least one of the maximum inter-frame transform parameter of the video frame and the maximum intra-frame prediction parameter of the video frame.
[0185] As an optional embodiment, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit; the process of the processor 701 decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0186] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, it is determined that the predictive coding mode of the target block is non-intra-frame prediction, or the target block is decoded using the first context model.
[0187] As an optional embodiment, the scale parameter includes the maximum inter-frame transform parameter, the video frame is a non-key frame, and the target block is a block to be divided; the process of the processor 701 decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0188] If the height of the target block is greater than the height indicated by the maximum inter-frame transformation parameter and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transformation parameter, perform horizontal binary partitioning on the target block or decode the target block using a second context model;
[0189] If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, perform vertical binary partitioning on the target block or decode the target block using a third context model.
[0190] As an alternative embodiment, the scale parameter includes the maximum inter-frame transformation parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction; the process by which the processor 701 decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0191] If the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, determine that there is no residual in the target block or decode the target block using a fourth context model.
[0192] As an alternative embodiment, the scale parameter includes the maximum intra-frame prediction parameter, and the video frame is a key frame; the process by which the processor 701 decodes the encoded data of the video frame based on the scale of the target block and the scale parameter includes:
[0193] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, perform quadtree partitioning on the target block or decode the target block using a fifth context model.
[0194] As an alternative embodiment, the bitstream data includes at least one video sequence of the video to which the video frame belongs, each video sequence includes the encoded data of at least one video frame, and the encoded data of any video frame includes the picture header of the video frame;
[0195] The scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs or in the picture header of the video frame;
[0196] The video includes key frames and non-key frames, and the scale parameters of the key frames and non-key frames are different.
[0197] As an alternative embodiment, the target block is the right boundary block of the video frame, and the process by which the processor 701 decodes the encoded data of the video frame based on the scale of the target block includes:
[0198] If the ratio of the height to the width of the target block is less than the first ratio threshold, perform vertical binary partitioning on the target block or decode the target block using a sixth context model;
[0199] If the ratio of the height to the width of the target block is greater than or equal to the first proportional threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using the seventh context model.
[0200] As an alternative embodiment, the target block is the lower boundary block of the video frame. The process by which the processor 701 decodes the encoded data of the video frame based on the attribute information of the target block includes:
[0201] If the ratio of the width to the height of the target block is less than the second proportional threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using the eighth context model;
[0202] If the ratio of the width to the height of the target block is greater than or equal to the second proportional threshold, perform vertical binary partitioning on the target block, or perform decoding processing on the target block using the ninth context model.
[0203] In one embodiment, the computer device is an encoding device. The processor 701 performs the following operations by running the computer program in the memory 703:
[0204] Obtain the video frame to be encoded;
[0205] Perform encoding processing on the video frame based on the scale of the target block to obtain the bitstream data of the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
[0206] As an alternative embodiment, the specific embodiment in which the processor 701 performs encoding processing on the video frame based on the scale of the target block to obtain the bitstream data of the video frame is:
[0207] Obtain the scale parameter of the video frame. The scale parameter is used to optimize the encoding and decoding process of the target block. The scale parameter is a default value or is configured based on the video frame; the scale parameter includes at least one of the following: coding tree unit parameter, maximum inter-frame transform parameter, maximum intra-frame prediction parameter;
[0208] Perform encoding processing on the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame.
[0209] As an alternative embodiment, the scale parameter includes the coding tree unit parameter corresponding to the video frame. The coding tree unit parameter is used to derive the encoded data of the video. The process by which the processor 701 performs encoding processing on the video frame based on the scale of the target block and the scale parameter includes:
[0210] Determine the predictive coding parameters of the video frame based on the coding tree unit parameter and the preset ratio;
[0211] Among them, the prediction coding parameters include at least one of the maximum inter-frame transformation parameter of the video frame and the maximum intra-frame prediction parameter of the video frame.
[0212] As an alternative embodiment, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit; the process in which the processor 701 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0213] If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, it is determined that the prediction coding mode of the target block is non-intra-frame prediction, or the target block is encoded using the first context model.
[0214] As an alternative embodiment, the scale parameter includes the maximum inter-frame transformation parameter, the video frame is a non-key frame, and the target block is a block to be partitioned; the process in which the processor 701 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0215] If the height of the target block is greater than the height indicated by the maximum inter-frame transformation parameter and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transformation parameter, the target block is subjected to horizontal binary partitioning processing, or the target block is encoded using the second context model;
[0216] If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, the target block is subjected to vertical binary partitioning processing, or the target block is encoded using the third context model.
[0217] As an alternative embodiment, the scale parameter includes the maximum inter-frame transformation parameter, the target block is a coding unit, and the prediction coding mode of the coding unit is inter-frame prediction; the process in which the processor 701 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0218] If the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, it is determined that there is no residual in the target block, or the target block is encoded using the fourth context model.
[0219] As an alternative embodiment, the scale parameter includes the maximum intra-frame prediction parameter, and the video frame is a key frame; the process in which the processor 701 encodes the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0220] If the size of the target block is greater than the size indicated by the maximum intra prediction parameter, the target block is partitioned, or the target block is encoded using the fifth context model; wherein, the partitioning method includes at least one of binary partitioning and quadtree partitioning.
[0221] As an alternative embodiment, the scale parameter is configured based on the video frame; the process of the processor 701 encoding the video frame based on the size of the target block and the scale parameter to obtain the bitstream data of the video frame includes:
[0222] Configuring the sequence header of the video sequence to which the video frame belongs based on the scale parameter, or adding the scale parameter to the picture header of the video frame;
[0223] Wherein, the video includes key frames and non-key frames, and the scale parameters of the key frames and non-key frames are different.
[0224] As an alternative embodiment, the target block is the right boundary block of the video frame, and the process of the processor 701 encoding the video frame based on the size of the target block to obtain the bitstream data of the video frame includes:
[0225] If the ratio of the height to the width of the target block is less than the first ratio threshold, perform vertical binary partitioning on the target block, or encode the target block using the sixth context model;
[0226] If the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, perform horizontal binary partitioning on the target block, or encode the target block using the seventh context model.
[0227] As an alternative embodiment, the target block is the lower boundary block of the video frame, and the process of the processor 701 encoding the video frame based on the size of the target block to obtain the bitstream data of the video frame includes:
[0228] If the ratio of the width to the height of the target block is less than the second ratio threshold, perform horizontal binary partitioning on the target block, or encode the target block using the eighth context model;
[0229] If the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, perform vertical binary partitioning on the target block, or encode the target block using the ninth context model.
[0230] As an alternative embodiment, the process of the processor 701 encoding the video frame based on the size of the target block to obtain the bitstream data of the video frame includes:
[0231] Set the high-frequency coefficients greater than the high-frequency coefficient threshold to a preset value.
[0232] Based on the same inventive concept, the principle and beneficial effects of the computer device provided in the embodiments of the present application for solving problems are similar to those of the image processing method in the method embodiments of the present application. One can refer to the principle and beneficial effects of the method implementation. For the sake of brevity, it will not be elaborated here.
[0233] The embodiments of the present application further provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by a processor to perform the image processing method in the above method embodiments.
[0234] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above image processing method.
[0235] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0236] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0237] In the embodiments of the present application, the "module" or "unit" involved refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0238] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0239] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand the implementation of all or part of the above embodiments and the equivalent changes made according to the claims of the present application still fall within the scope covered by the application.
Claims
1. An image processing method, characterized in that, the method includes: obtaining bitstream data of a video frame, where the bitstream data includes encoded data of the video frame; decoding the encoded data of the video frame based on the scale of a target block, and presenting the video frame; the target block is obtained by dividing the video frame, and the scale of the target block is used to determine encoding and decoding information of the target block.
2. The method according to claim 1, characterized in that, the decoding the encoded data of the video frame based on the scale of the target block includes: decoding the encoded data of the video frame based on the scale of the target block and the scale parameter; wherein, the scale parameter is used to optimize the encoding and decoding process of the target block, the scale parameter is a default value or carried in the bitstream data; the scale parameter includes at least one of the following: coding tree unit parameter, maximum inter-frame transform parameter, maximum intra-frame prediction parameter.
3. The method according to claim 2, characterized in that, the scale parameter includes the coding tree unit parameter corresponding to the video frame, and the encoded data of the video frame is derived based on the coding tree unit parameter; the decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes: determining prediction coding parameters of the video frame through the coding tree unit parameter and a preset ratio; decoding the encoded data of the video frame based on the scale of the target block and the prediction coding parameters of the video frame; wherein, the prediction coding parameters include at least one of the maximum inter-frame transform parameter of the video frame and the maximum intra-frame prediction parameter of the video frame.
4. The method according to claim 2, characterized in that, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit; the process of decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, it is determined that the prediction coding mode of the target block is non-intra-frame prediction, or the target block is decoded using a first context model.
5. The method according to claim 2, characterized in that, the scale parameter includes the maximum inter-frame transform parameter, the video frame is a non-key frame, and the target block is a block to be divided; the process of decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the height of the target block is greater than the height indicated by the maximum inter-frame transform parameter, and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transform parameter, then perform horizontal binary division processing on the target block, or decode the target block using a second context model. If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter, and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, perform vertical binary partitioning on the target block, or perform decoding processing on the target block using a third context model.
6. The method according to claim 2, wherein, the scale parameter includes a maximum inter-frame transformation parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction; the process of decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, determine that there is no residual in the target block, or perform decoding processing on the target block using a fourth context model.
7. The method according to claim 2, wherein, the scale parameter includes a maximum intra-frame prediction parameter, and the video frame is a key frame; the process of decoding the encoded data of the video frame based on the scale of the target block and the scale parameter includes: if the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, perform quadtree partitioning on the target block, or perform decoding processing on the target block using a fifth context model.
8. The method according to claim 2, wherein, the bitstream data includes at least one video sequence of the video to which the video frame belongs, each video sequence includes encoded data of at least one video frame, and the encoded data of any video frame includes the picture header of the video frame; the scale parameter of the video frame is carried in the sequence header of the video sequence to which the video frame belongs, or carried in the picture header of the video frame; the video includes key frames and non-key frames, and the scale parameters of the key frames and the non-key frames are different.
9. The method according to claim 1, wherein, the target block is the right boundary block of the video frame, and the process of decoding the encoded data of the video frame based on the scale of the target block includes: if the ratio of the height to the width of the target block is less than a first ratio threshold, perform vertical binary partitioning on the target block, or perform decoding processing on the target block using a sixth context model; if the ratio of the height to the width of the target block is greater than or equal to the first ratio threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using a seventh context model.
10. The method according to claim 1, wherein, the target block is the lower boundary block of the video frame, and the process of decoding the encoded data of the video frame based on the attribute information of the target block includes: if the ratio of the width to the height of the target block is less than a second ratio threshold, perform horizontal binary partitioning on the target block, or perform decoding processing on the target block using an eighth context model; If the ratio of the width to the height of the target block is greater than or equal to the second ratio threshold, perform vertical binary partitioning on the target block, or perform decoding on the target block using the ninth context model.
11. An image processing method, characterized in that, the method includes: Obtain a video frame to be encoded; Based on the scale of the target block, perform encoding on the video frame to obtain the bitstream data of the video frame; the target block is obtained by partitioning the video frame, and the scale of the target block is used to determine the encoding and decoding information of the target block.
12. The method according to claim 11, characterized in that, the performing encoding on the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes: Obtain the scale parameter of the video frame, the scale parameter is used to optimize the encoding and decoding process of the target block, the scale parameter is a default value, or is configured based on the video frame; the scale parameter includes at least one of the following: coding tree unit parameter, maximum inter-frame transform parameter, maximum intra-frame prediction parameter; Based on the scale of the target block and the scale parameter, perform encoding on the video frame to obtain the bitstream data of the video frame.
13. The method according to claim 12, characterized in that, the scale parameter includes the coding tree unit parameter corresponding to the video frame, the coding tree unit parameter is used to derive the encoded data of the video, and the process of performing encoding on the video frame based on the scale of the target block and the scale parameter includes: Based on the coding tree unit parameter and a preset ratio, determine the prediction coding parameter of the video frame; wherein, the prediction coding parameter includes at least one of the maximum inter-frame transform parameter of the video frame and the maximum intra-frame prediction parameter of the video frame.
14. The method according to claim 12, characterized in that, the scale parameter includes the maximum intra-frame prediction parameter, the video frame is a non-key frame, and the target block is a coding unit; the process of performing encoding on the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, determine that the prediction coding mode of the target block is non-intra-frame prediction, or perform encoding on the target block using the first context model.
15. The method according to claim 12, characterized in that, the scale parameter includes the maximum inter-frame transform parameter, the video frame is a non-key frame, and the target block is a block to be partitioned; the process of performing encoding on the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: If the height of the target block is greater than the height indicated by the maximum inter-frame transform parameter, and the width of the target block is less than or equal to the width indicated by the maximum inter-frame transform parameter, perform horizontal binary partitioning on the target block, or perform encoding on the target block using the second context model; If the height of the target block is less than or equal to the height indicated by the maximum inter-frame transformation parameter, and the width of the target block is greater than the width indicated by the maximum inter-frame transformation parameter, perform vertical binary partitioning on the target block, or perform encoding processing on the target block using a third context model.
16. The method according to claim 12, wherein, the scale parameter includes a maximum inter-frame transformation parameter, the target block is a coding unit, and the predictive coding mode of the coding unit is inter-frame prediction; the process of encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: If the scale of the target block is greater than the scale indicated by the maximum inter-frame transformation parameter, determine that there is no residual in the target block, or perform encoding processing on the target block using a fourth context model.
17. The method according to claim 12, wherein, the scale parameter includes a maximum intra-frame prediction parameter, and the video frame is a key frame; the process of encoding the video frame based on the scale of the target block and the scale parameter to obtain the bitstream data of the video frame includes: If the scale of the target block is greater than the scale indicated by the maximum intra-frame prediction parameter, divide the target block, or perform encoding processing on the target block using a fifth context model; wherein, the division method includes at least one of binary partitioning and quadtree partitioning.
18. The method according to claim 11, wherein, the process of encoding the video frame based on the scale of the target block to obtain the bitstream data of the video frame includes: Set the high-frequency coefficients greater than the high-frequency coefficient threshold to a preset value.
19. A computer device, wherein, comprising: a memory storing a computer program therein; a processor for loading the computer program to implement the image processing method according to any one of claims 1-10, or for loading the computer program to implement the image processing method according to any one of claims 11-18.
20. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to implement the image processing method according to any one of claims 1-10, or is suitable for being loaded and executed by a processor to implement the image processing method according to any one of claims 11-18.
Citation Information
Cited By
Image processing method and apparatus, and computer-readable storage medium
EP4770086A1