Video encoding method, video playback method, related devices and media

By performing abnormal distortion point detection and mode cost calibration during the video encoding process, and selecting appropriate prediction modes, the image block distortion problem caused by improper prediction mode selection is solved, and the subjective quality of the image block is improved.

CN111629206BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010452023.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-25
Publication Date
2025-07-18
Estimated Expiration
2040-05-25

AI Technical Summary

Technical Problem

In the existing video encoding technology, inappropriate selection of prediction mode can easily lead to greater distortion of the image block after encoding, affecting the subjective quality of the image block.

Method used

By performing abnormal distortion point detection in the candidate prediction mode in the mode information set, the mode cost is calibrated, and the target prediction mode is selected according to the calibration mode cost for prediction processing, the distortion probability of the image block after encoding is reduced.

Benefits of technology

Without affecting the compression efficiency and coding complexity, the image compression quality is effectively improved and the subjective quality of the image block is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111629206B_ABST
    Figure CN111629206B_ABST
Patent Text Reader

Abstract

The present application discloses a video encoding method, a video playing method, related devices and media; the video encoding method includes: obtaining a target prediction unit and a set of mode information in a target image block; performing abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain detection results corresponding to at least one candidate prediction mode; calibrating the mode cost of at least one candidate prediction mode in the set of mode information according to the detection results corresponding to at least one candidate prediction mode to obtain a calibrated set of mode information; selecting a target prediction mode from a plurality of candidate prediction modes according to the mode cost of each candidate prediction mode in the calibrated set of mode information; and performing prediction processing on the target prediction unit by using the target prediction mode to obtain encoded data of the target image block. The present application can reduce the probability of distortion occurring in the image block after encoding and improve the subjective quality of the image block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, specifically to the field of image processing technologies, and particularly to a video encoding method, a video playing method, a video encoding device, a video playing device, a video encoding apparatus, a video playing apparatus, and a computer storage medium. Background Art

[0002] Video encoding typically divides an image to be encoded into multiple image blocks, and obtains the bitstream data of the image to be encoded by encoding each image block. During the encoding process of any image block, it is usually necessary to first perform prediction on each prediction unit corresponding to the image block using a prediction mode to obtain the residual block of the image block; then perform subsequent processing such as transform quantization on the residual block to obtain the encoded data of the image block. Research shows that during the encoding process of an image block, if the selected prediction mode for the prediction unit is inappropriate, it is likely to cause significant distortion in the encoded image block, resulting in low subjective quality of the image block. Summary of the Invention

[0003] Embodiments of the present invention provide a video encoding method, a video playing method, related devices, and a medium, which can reduce the probability of distortion occurring in an encoded image block, thereby improving the subjective quality of the image block.

[0004] On the one hand, embodiments of the present invention provide a video encoding method, which includes:

[0005] Obtain a target prediction unit in a target image block and a set of mode information of the target prediction unit, where the set of mode information includes multiple candidate prediction modes and the mode costs of various candidate prediction modes;

[0006] Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain the detection results corresponding to the at least one candidate prediction mode;

[0007] According to the detection results corresponding to the at least one candidate prediction mode, calibrate the mode costs of the at least one candidate prediction mode in the set of mode information to obtain a calibrated set of mode information;

[0008] According to the mode costs of the candidate prediction modes in the calibrated set of mode information, select a target prediction mode from the multiple candidate prediction modes;

[0009] Perform prediction processing on the target prediction unit using the target prediction mode to obtain the encoded data of the target image block.

[0010] On the other hand, an embodiment of the present invention provides a video encoding device, which includes:

[0011] An acquisition unit, configured to acquire a target prediction unit in a target image block and a set of mode information of the target prediction unit, where the set of mode information includes multiple candidate prediction modes and mode costs of various candidate prediction modes;

[0012] An encoding unit, configured to perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode;

[0013] The encoding unit is further configured to calibrate the mode costs of the at least one candidate prediction mode in the set of mode information according to the detection results corresponding to the at least one candidate prediction mode to obtain a calibrated set of mode information;

[0014] The encoding unit is further configured to select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated set of mode information;

[0015] The encoding unit is further configured to perform prediction processing on the target prediction unit by using the target prediction mode to obtain encoded data of the target image block.

[0016] In an implementation manner, the at least one candidate prediction mode is an inter prediction mode, and the inter prediction mode at least includes the following modes: a first prediction mode, a second prediction mode, and a third prediction mode;

[0017] The first prediction mode refers to a mode that needs to transmit index information of a reference image block related to the target image block;

[0018] The second prediction mode refers to a mode that needs to transmit residual information of the target image block and index information of a reference image block related to the target image block;

[0019] The third prediction mode refers to a mode that needs to transmit residual information of the target image block, motion vector data of the target image block, and index information of a reference image block related to the target image block.

[0020] In still another implementation manner, when the encoding unit is configured to perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode, it may specifically be configured to:

[0021] Predict the pixel values of each pixel point in the target prediction unit using a reference prediction mode to obtain the predicted values of each pixel point; the reference prediction mode is any mode in the inter-frame prediction mode;

[0022] Calculate the absolute value of the residual between the pixel value and the predicted value of each pixel point in the target prediction unit;

[0023] If there is a pixel point in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there is an abnormal distortion point in the target prediction unit under the reference prediction mode;

[0024] If there is no pixel point in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there is no such abnormal distortion point in the target prediction unit under the reference prediction mode.

[0025] In another implementation, the target threshold is associated with the reference prediction mode;

[0026] If the reference prediction mode is the first prediction mode in the inter-frame prediction mode, the target threshold is equal to the first threshold; the first threshold is greater than the invalid value and less than the maximum value of the pixel value range;

[0027] If any of the candidate prediction modes is the second prediction mode or the third prediction mode in the inter-frame prediction mode, the target threshold is equal to the second threshold; the second threshold is greater than or equal to the first threshold and less than the maximum value of the pixel value range.

[0028] In another implementation, when the coding unit is used to calibrate the mode cost of at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain a calibrated mode information set, it can be specifically used for:

[0029] If the detection result corresponding to the reference prediction mode indicates that there is no abnormal distortion point in the target prediction unit under the reference prediction mode, keep the mode cost of the reference prediction mode in the mode information set unchanged to obtain a calibrated mode information set; the reference prediction mode is any mode in the inter-frame prediction mode;

[0030] If the detection result corresponding to the reference prediction mode indicates that there is an abnormal distortion point in the target prediction unit under the reference prediction mode, adjust the mode cost of the reference prediction mode in the mode information set using the cost adjustment strategy of the reference prediction mode to obtain a calibrated mode information set.

[0031] In another implementation, when the encoding unit is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it is specifically used for:

[0032] If the reference prediction mode is the second prediction mode or the third prediction mode, the mode cost of the reference prediction mode is amplified by using a penalty factor to obtain the calibrated mode cost of the reference prediction mode.

[0033] In another implementation, when the encoding unit is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it can also be used for:

[0034] If the reference prediction mode is the first prediction mode, a preset cost is obtained; the preset cost is greater than the mode costs of the candidate prediction modes other than the first prediction mode in the calibrated mode information set and greater than the mode cost of the first prediction mode in the mode information set;

[0035] In the mode information set, the mode cost of the reference prediction mode is adjusted to the preset cost.

[0036] In another implementation, when the encoding unit is used to select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set, it can be specifically used for:

[0037] From the multiple candidate prediction modes, the candidate prediction mode with the smallest mode cost in the calibrated mode information is selected as the target prediction mode.

[0038] In another implementation, when the encoding unit is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it can also be used for:

[0039] If the reference prediction mode is the first prediction mode, the mode cost of the first prediction mode in the mode information set remains unchanged;

[0040] A disable flag is added to the first prediction mode, and the disable flag indicates that the first prediction mode is prohibited from being used to perform prediction processing on the target prediction unit.

[0041] In another implementation, when the encoding unit is used to select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set, it can be specifically used for:

[0042] If the candidate prediction mode with the minimum mode cost in the calibrated mode information set is not the first prediction mode, or the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode does not have the disable flag, then use the candidate prediction mode with the minimum mode cost as the target prediction mode;

[0043] If the candidate prediction mode with the minimum mode cost is the first prediction mode, and the first prediction mode has the disable flag, then select the candidate prediction mode with the second smallest mode cost in the calibrated mode information set as the target prediction mode.

[0044] In another implementation manner, the multiple candidate prediction modes include: an intra prediction mode and an inter prediction mode; the coding unit can also be used to:

[0045] Perform complexity analysis on the target prediction unit to obtain the prediction complexity of the target prediction unit;

[0046] If it is determined according to the prediction complexity that the target prediction unit meets the preset conditions, then use the intra prediction mode to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block; wherein, the preset conditions include: the prediction complexity is less than or equal to the complexity threshold, and there are abnormal distortion points in at least one mode of the target prediction unit in the inter prediction mode;

[0047] If it is determined according to the prediction complexity that the target prediction unit does not meet the preset conditions, then perform the step of calibrating the mode costs of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain the calibrated mode information set.

[0048] On the other hand, an embodiment of the present invention provides a video encoding device, the video encoding device includes an input interface and an output interface, and the video encoding device further includes:

[0049] A processor, adapted to implement one or more instructions; and,

[0050] A computer storage medium, the computer storage medium stores one or more first instructions, and the one or more first instructions are adapted to be loaded and executed by the processor to perform the following steps:

[0051] Obtain a target prediction unit in a target image block and a mode information set of the target prediction unit, the mode information set includes multiple candidate prediction modes and the mode costs of various candidate prediction modes;

[0052] Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the mode information set, and obtain the detection results corresponding to the at least one candidate prediction mode;

[0053] According to the detection results corresponding to the at least one candidate prediction mode, calibrate the mode costs of the at least one candidate prediction mode in the mode information set to obtain a calibrated mode information set;

[0054] Select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set;

[0055] Perform prediction processing on the target prediction unit using the target prediction mode to obtain the encoded data of the target image block.

[0056] On the other hand, an embodiment of the present invention provides a computer storage medium, which stores one or more first instructions, and the one or more first instructions are adapted to be loaded and executed by a processor to perform the following steps:

[0057] Obtain a target prediction unit in a target image block and the mode information set of the target prediction unit, where the mode information set includes multiple candidate prediction modes and the mode costs of various candidate prediction modes;

[0058] Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the mode information set, and obtain the detection results corresponding to the at least one candidate prediction mode;

[0059] According to the detection results corresponding to the at least one candidate prediction mode, calibrate the mode costs of the at least one candidate prediction mode in the mode information set to obtain a calibrated mode information set;

[0060] Select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set;

[0061] Perform prediction processing on the target prediction unit using the target prediction mode to obtain the encoded data of the target image block.

[0062] In the encoding process of the embodiments of the present invention, abnormal distortion points of a target prediction unit may first be detected under at least one candidate prediction mode in a mode information set to obtain detection results corresponding to at least one candidate prediction mode. Secondly, according to the detection results corresponding to at least one candidate prediction mode, the mode costs of at least one candidate prediction mode in the mode information set may be calibrated; so that the mode costs of the candidate prediction modes in the calibrated mode information set can more accurately reflect the corresponding bitrates and distortions of the candidate prediction modes, thereby enabling the selection of a target prediction mode more suitable for the target prediction unit from multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set. Then, a suitable target prediction mode may be used to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block, which can reduce the probability of distortion of the target image block after encoding to a certain extent; and since the embodiments of the present invention mainly reduce the distortion probability by modifying the mode decision process to select a suitable target prediction mode, it is possible to effectively improve the image compression quality and enhance the subjective quality of the target image block without substantially affecting the compression efficiency and encoding complexity.

[0063] On the other hand, the embodiments of the present invention provide a video playback method, which includes:

[0064] Obtaining bitstream data of each frame image in an image frame sequence corresponding to a target video, where the bitstream data of each frame image includes encoded data of multiple image blocks; the encoded data of each image block in other frame images except the first frame image in the image frame sequence is encoded by the above video encoding method;

[0065] Decoding the bitstream data of each frame image to obtain each frame image;

[0066] Sequentially displaying each frame image on a playback interface.

[0067] On the other hand, the embodiments of the present invention provide a video playback device, which includes:

[0068] An obtaining unit, configured to obtain bitstream data of each frame image in an image frame sequence corresponding to a target video, where the bitstream data of each frame image includes encoded data of multiple image blocks; the encoded data of each image block in other frame images except the first frame image in the image frame sequence is encoded by the above video encoding method;

[0069] A decoding unit, configured to decode the bitstream data of each frame image to obtain each frame image;

[0070] A display unit, configured to sequentially display each frame image on a playback interface.

[0071] In another aspect, an embodiment of the present invention provides a video playback device, which includes an input interface and an output interface. The video playback device further includes:

[0072] a processor, adapted to implement one or more instructions; and,

[0073] a computer storage medium storing one or more second instructions, the one or more second instructions being adapted to be loaded and executed by the processor to perform the following steps:

[0074] Obtain the bitstream data of each frame image in the image frame sequence corresponding to the target video. The bitstream data of each frame image includes the encoded data of multiple image blocks; the encoded data of each image block in the other frame images except the first frame image in the image frame sequence is encoded by the above video encoding method;

[0075] Decode the bitstream data of each frame image to obtain each frame image;

[0076] Display each frame image in sequence on the playback interface.

[0077] In another aspect, an embodiment of the present invention provides a computer storage medium storing one or more second instructions, the one or more second instructions being adapted to be loaded and executed by a processor to perform the following steps:

[0078] Obtain the bitstream data of each frame image in the image frame sequence corresponding to the target video. The bitstream data of each frame image includes the encoded data of multiple image blocks; the encoded data of each image block in the other frame images except the first frame image in the image frame sequence is encoded by the above video encoding method;

[0079] Decode the bitstream data of each frame image to obtain each frame image;

[0080] Display each frame image in sequence on the playback interface.

[0081] In the embodiment of the present invention, the bitstream data of each frame image in the image frame sequence corresponding to the target video can be obtained first. The bitstream data of each frame image includes the encoded data of multiple image blocks. Secondly, the bitstream data of each frame image can be decoded to obtain each frame image; and each frame image can be displayed in sequence on the playback interface. Since the encoded data of each image block in the other frame images except the first frame image in the image frame sequence corresponding to the target video is encoded by the above video encoding method; therefore, the probability of distortion of each image block can be effectively reduced, so that when each frame image is displayed on the playback interface, the probability of dirty points in each frame image can be reduced to a certain extent, and the subjective quality of each frame image can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0083] Figure 1a It is a schematic diagram of the architecture of an image processing system provided by an embodiment of the present invention;

[0084] Figure 1b It is a schematic diagram of an image processing flow provided by an embodiment of the present invention;

[0085] Figure 1c It is a schematic diagram of dividing a frame image into image blocks provided by an embodiment of the present invention;

[0086] Figure 1d It is a schematic diagram of dividing an inter-frame prediction mode provided by an embodiment of the present invention;

[0087] Figure 1e It is a schematic diagram of encoding a frame image provided by an embodiment of the present invention;

[0088] Figure 1f It is another schematic diagram of encoding a frame image provided by an embodiment of the present invention;

[0089] Figure 2 It is a schematic diagram of the flow of a video encoding method provided by an embodiment of the present invention;

[0090] Figure 3 It is a schematic diagram of the flow of a video encoding method provided by an embodiment of the present invention;

[0091] Figure 4 It is a schematic diagram of the flow of a video playback method provided by an embodiment of the present invention;

[0092] Figure 5 It is an application scenario diagram of a video encoding method and a video playback method provided by an embodiment of the present invention;

[0093] Figure 6 It is a schematic diagram of the structure of a video encoding device provided by an embodiment of the present invention;

[0094] Figure 7 It is a schematic diagram of the structure of a video encoding device provided by an embodiment of the present invention;

[0095] Figure 8 It is a schematic diagram of the structure of a video playback device provided by an embodiment of the present invention;

[0096] Figure 9 It is a schematic structural diagram of a video playback device provided by an embodiment of the present invention. Detailed implementation manners

[0097] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0098] In the embodiments of the present invention, an image processing system is involved; see Figure 1a As shown, the image processing system at least includes: a video encoding device 11, a video playback device 12, and a transmission medium 13. Among them, the video encoding device 11 can be a server or a terminal; the server here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.; the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. The interior of the video encoding device 11 can at least include an encoder, and the encoder is used to execute a series of encoding processes. The video playback device 12 can be any device with a video playback function, such as a terminal such as a smart phone, a tablet computer, a laptop computer, a smart watch, or a device such as a projector or a projector that can project video images onto a screen for playback. The video playback device 12 can at least include a decoder, and the decoder is used to execute a series of decoding processes. The transmission medium 13 refers to the space or entity through which data is transmitted, and is mainly used to transmit data between the video encoding device 11 and the video decoding device 12; specifically, it can include but is not limited to: network media such as mobile networks, wireless networks, and wired networks, or removable hardware media with read and write functions such as USB flash drives (Universal Serial Bus) and external hard drives. It should be understood that Figure 1a It only exemplarily represents the architecture of the image processing system involved in the embodiments of the present invention, and does not limit the specific architecture of the image processing system. For example, in other embodiments, the video encoding device 11 and the video playback device 12 can also be the same device; again, in other embodiments, the number of video playback devices 12 is not limited to one, etc.

[0099] In the above image processing system, the processing flow for any frame image in the image frame sequence can be seen together in Figure 1b As shown, it generally includes the following several stages:

[0100] (1) Encoding stage:

[0101] After the encoder in the video encoding device 11 obtains the current frame image to be encoded, it can encode the current frame image based on the mainstream video encoding standards to obtain the bitstream data of the current frame image. Among them, the mainstream video encoding standards may include, but are not limited to: H.264, H.265, VVC (Versatile Video Coding), AVS3 (Audio Video coding Standard 3), etc.; here, H.265 can also be referred to as HEVC (High Efficiency Video Coding). Taking H.265 as an example, its general encoding process is as follows:

[0102] ① Divide the current frame image to be encoded into several image blocks (or called CU (Coding Unit)), where the image block here refers to the basic unit of video encoding. In specific implementation, the current frame image to be encoded can be first divided into several non-overlapping LCU (Largest Coding Unit). Then, several CUs can be further obtained by dividing the corresponding LCU according to the characteristics of each LCU, as Figure 1c shown; it should be understood that Figure 1c is only an exemplary representation of the division method of LCU and does not limit it; as Figure 1c shows that the LCU is evenly divided into multiple CUs, but in fact, the LCU can also be unevenly divided into multiple CUs. Each CU can correspond to a prediction mode of a certain mode type, such as the inter prediction mode (Inter mode) and the intra prediction mode (Intra mode). Among them, the intra prediction mode mainly searches for reference image blocks from the already encoded image blocks in the current frame image and uses the decoded information of the reference image blocks for prediction; the intra prediction mode needs to transmit the corresponding prediction mode and residual information to the decoder. The inter prediction mode mainly searches for the reference image block that matches the current image block from the already encoded frame images in the image frame sequence according to the motion estimation (ME) and motion compensation (MC) of the current image block, and uses the decoded information of the reference image block for prediction based on the MV (Motion Vector).

[0103] See Figure 1dAs shown, the inter-frame prediction mode may at least include the AMVP (Advanced Motion Vector Prediction) mode and the Merge mode. Among them, the AMVP mode needs to transmit the motion vector data (MVD) of the current image block, the index information of the reference image block, and the residual information of the current image block to the decoder; here, the MVD refers to the difference between the MVP (predicted motion vector) and the motion vector obtained based on motion estimation (ME). The Merge mode, on the other hand, does not need to transmit the motion vector data (MVD) of the current image block, and it can be further divided into the normal Merge mode and the SKIP mode. The SKIP mode here is a special case of the Merge mode. The difference between the SKIP mode and the normal Merge mode is that: the normal Merge mode needs to transmit the index information of the reference image block and the residual information of the current image block to the decoder, while the SKIP mode only needs to transmit the index information of the reference image block to the decoder and does not need to transmit the residual information of the current image block.

[0104] ②For the current image block to be encoded, the pixel values of each pixel in the current image block are predicted using a prediction mode to obtain the prediction block corresponding to the current image block; the prediction block includes the predicted values of each pixel. In a specific implementation, the current image block can be further divided into one or more prediction units (PU), and mode decision is performed to dynamically decide the prediction mode of each prediction unit corresponding to the current image block according to the characteristics of the input signal. Specifically, the mode type can be determined first according to the characteristics of the current image block, such as the inter-frame prediction type or the intra-frame prediction type; then, the corresponding prediction mode is selected from the prediction modes of this mode type according to the characteristics of each prediction unit. If the determined mode type is the intra-frame prediction type, the prediction mode of each prediction unit corresponding to the current image block is the intra-frame prediction mode; if the determined mode type is the inter-frame prediction type, the prediction mode of each prediction unit corresponding to the current image block can be the AMVP mode, the normal Merge mode, or the SKIP mode; in this case, the prediction modes of each prediction unit corresponding to the current image block can be the same or different. After determining the prediction mode of each prediction unit, the corresponding prediction mode can be used to perform prediction processing on each prediction unit to obtain the prediction result of each prediction unit. Then, the prediction results of each prediction unit are combined to obtain the prediction block corresponding to the current image block.

[0105] ③Calculate the residual block of the current image block based on the prediction block and the current image block. Here, the residual block includes the difference between the predicted value and the actual pixel value of each pixel in the current image block. Then, perform transformation, quantization, and entropy coding processing on the residual block in sequence to obtain the encoded data of the current image block. Iteratively execute steps ②-③ involved in the above encoding process until all image blocks in the current frame image are encoded. At this time, the encoded data of each image block included in the current frame image can be obtained, thereby obtaining the bitstream data of the current frame image.

[0106] (2) Transmission stage:

[0107] After the video encoding device 11 goes through the above encoding stage, it obtains the bitstream data and encoding information of the current frame image and transmits them to the video display device 12, so that the video display device 12 decodes the encoded data of each image block using the encoding information through the decoding stage to obtain the current frame image. Among them, the bitstream data includes the encoded data of each image block in the current frame image; the encoding information at least includes the transmission information specified by the prediction mode used when predicting the prediction units of each image block in the current frame image, such as the motion vector data of the current image block, the index information of the reference image block, and the residual information of the current image block specified by the AMVP mode, the index information of the reference image block specified by the SKIP mode, and other transmission information.

[0108] (3) Decoding stage:

[0109] After the video display device 12 receives the bitstream data and encoding information of the current frame image, it can decode the encoded data of each image block in the bitstream data according to the encoding information in sequence. The decoding process for any image block is as follows: perform decoding, inverse quantization, and inverse transformation processing on the encoded data of the current image block in sequence to obtain the residual block of the current image block. Then, the prediction mode used by each prediction unit of the current image block can be determined according to the transmission information corresponding to the current image block in the encoding information, and the image block can be obtained based on the determined prediction mode and the residual block. Iteratively execute each step involved in the above decoding process, and the individual image blocks of the current frame image can be obtained, thereby obtaining the current frame image. After obtaining the current frame image, the video display device 12 can display the current frame image on the playback interface.

[0110] As can be seen from the above image processing flow, the mode decision process involved in the encoding stage usually involves multiple prediction modes. If the selected prediction mode for the prediction unit during the mode decision process is inappropriate and the quantization parameter (QP) involved in the transformation quantization of the residual block is large, then some particularly abnormal distortion points are likely to occur after the image block is encoded, such as pixel points with a distortion as high as 100+; which in turn causes some dirty points to appear in the image block obtained by decoding through the decoding stage, affecting the subjective quality of the image block and the frame image, such as Figure 1e shown. Based on this, an embodiment of the present invention proposes a video coding scheme; this video coding scheme is mainly used to guide the encoder in mode decision during the encoding process, and reduce the probability of abnormal distortion points occurring after the image block is encoded by selecting an appropriate prediction mode for the prediction unit; the abnormal distortion points here refer to pixel points whose absolute value of the difference between the decoded pixel value and the pixel value before encoding is greater than a certain threshold. In specific implementation, the principle of this video coding scheme is roughly as follows:

[0111] For the target frame image to be encoded, the target frame image can be divided into one or more image blocks; and an image block can be selected from the target frame image as the target image block to be encoded, and then the target image block is further divided into one or more prediction units. For any prediction unit in the target image block, when selecting a prediction mode for the prediction unit through mode decision, the mode cost of each prediction mode can be obtained first; and it is detected whether there are abnormal distortion points in the prediction unit under at least one prediction mode. If not, a prediction mode is selected for the prediction unit from multiple prediction modes using the mode decision algorithm; the mode decision algorithm here is used to indicate selecting the prediction mode with the minimum mode cost. If there are, the mode decision algorithm is adjusted; the adjustment of the mode decision algorithm here means: first calibrating the mode cost of the prediction mode corresponding to the abnormal distortion point, and then selecting a prediction mode for the prediction unit according to the calibrated mode cost of the calibrated prediction mode and the mode cost of the uncalibrated prediction mode. After a prediction mode is selected for the prediction unit, the selected prediction mode can be used to perform prediction processing on the prediction unit; iterate the above steps to perform prediction processing on each prediction unit in the target image block, so as to obtain the encoded data of the target image block. After obtaining the encoded data of the target image block, an image block can be reselected from the target frame image as the new target image block, and the above steps are executed to obtain the encoded data of the new target image block; after all the image blocks in the target frame image are encoded, the bitstream data of the target frame image can be obtained.

[0112] To more clearly illustrate the beneficial effects of the video coding scheme proposed by the embodiment of the present invention, still taking the target frame image as Figure 1eTaking the original frame image shown in the upper middle figure as an example, the video encoding scheme of the embodiment of the present invention is used to encode it, and the following can be obtained: Figure 1f The encoded frame image is shown in the lower middle figure. Figure 1e The lower side of the figure and Figure 1f As can be seen from the lower figure in the figure, the encoding scheme proposed in the embodiment of the present invention can correct the mode decision process according to the abnormal distortion point detection result by adding an abnormal distortion point detection mechanism in the mode decision process, thereby effectively reducing the number of abnormal distortion points generated by the target image frame after encoding and improving the subjective quality of the target image frame.

[0113] Based on the description of the above video encoding scheme, an embodiment of the present invention proposes a video encoding method; the video encoding method can be executed by the above-mentioned video encoding device, specifically, by an encoder in the video encoding device. Figure 2 , the video encoding method may include the following steps S201-S205:

[0114] S201, obtaining a target prediction unit in a target image block and a mode information set of the target prediction unit.

[0115] In an embodiment of the present invention, the target prediction unit may be any prediction unit in the target image block; the mode information set of the target prediction unit may include multiple candidate prediction modes and mode costs of various candidate prediction modes, and the mode cost here may be used to reflect the code rate and distortion caused by using the candidate prediction mode to predict the target prediction unit, which may include but is not limited to rate distortion cost. Among them, the multiple candidate prediction modes may at least include: intra-frame prediction mode and inter-frame prediction mode; the inter-frame prediction mode may at least include the following modes: first prediction mode, second prediction mode and third prediction mode. The so-called first prediction mode refers to a mode in which the index information of the reference image block related to the target image block needs to be transmitted, which may specifically be the SKIP mode mentioned above; the second prediction mode refers to a mode in which the residual information of the target image block and the index information of the reference image block related to the target image block need to be transmitted, which may specifically be the ordinary merge mode mentioned above; the third prediction mode refers to a mode in which the residual information of the target image block, the motion vector data of the target image block, and the index information of the reference image block related to the target image block need to be transmitted, which may specifically be the AMVP mode mentioned above.

[0116] S202: Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the mode information set to obtain a detection result corresponding to the at least one candidate prediction mode.

[0117] In a specific implementation, abnormal distortion points can be detected for a target prediction unit under each candidate prediction mode in a mode information set to obtain detection results corresponding to each candidate prediction mode. That is, in this specific implementation, at least one candidate prediction mode may include an intra prediction mode and an inter prediction mode. Among them, the detection result of each candidate prediction mode can be used to indicate whether there are abnormal distortion points in the target prediction unit under this candidate prediction mode. In another specific implementation, research shows that the probability of generating abnormal distortion points when using the intra prediction mode for prediction is relatively small; the probability of generating abnormal distortion points when using the inter prediction mode for prediction is relatively large, especially the SKIP mode in the inter prediction mode. Since when using the SKIP mode for prediction, the motion vector (MV) is derived from other reference image blocks and the residual information is not transmitted, although the SKIP mode can greatly save the bit rate and improve the coding efficiency, in some special scenarios (such as screen sharing scenarios, video live broadcast scenarios, etc.), it is easy to cause excessive local point distortion, making the probability of generating abnormal distortion points in the target image block relatively large. Based on this research result, in the embodiments of the present invention, abnormal distortion points can be detected for the target prediction unit only under each mode included in the inter prediction mode to obtain detection results corresponding to each mode included in the inter prediction mode. That is, in this specific implementation, at least one candidate prediction mode may be the inter prediction mode. In this specific implementation, since no detection process is performed on the intra prediction mode, by reducing the operation of detecting abnormal distortion points for the target prediction unit in the intra prediction mode, processing resources can be effectively saved and the coding speed can be improved.

[0118] As can be seen from the foregoing, an abnormal distortion point refers to a pixel point whose absolute value of the difference between the decoded pixel value and the pixel value before encoding is greater than a certain threshold. Therefore, in one embodiment, the embodiments of the present invention can use the difference between the predicted value and the actual pixel value (i.e., the pixel value before encoding) of the pixel point to determine whether the pixel point is an abnormal pixel point. Based on this, the detection principle of detecting abnormal distortion points for the target prediction unit under any candidate prediction mode is as follows: The pixel values of each pixel point in the target prediction unit can be predicted using this any candidate prediction mode. If there is at least one pixel point with a relatively large difference between the predicted value and the actual pixel value, it can be determined that there are abnormal distortion points in the target prediction unit under this any candidate prediction mode; if the differences between the predicted values and the actual pixel values of each pixel point are relatively small, it can be determined that there are no abnormal distortion points in the target prediction unit under this any candidate prediction mode.

[0119] In another implementation, when at least one candidate prediction mode is an inter-frame prediction mode, the difference between the motion compensation value of a pixel and the actual pixel value (i.e., the pixel value before encoding) can also be used to determine whether the pixel is an abnormal pixel, so as to improve the accuracy of the detection result; here, the motion compensation value is equal to the sum of the predicted value of the pixel and the residual obtained after inverse transformation and inverse quantization of the residual information. It should be noted that since the first prediction mode does not transmit residual information, the motion compensation value and the predicted value of the pixel in the first prediction mode are equal. Based on this, the detection principle for detecting abnormal distortion points in a target prediction unit in any mode of the inter-frame prediction mode can also be as follows: Any mode can be used to predict the pixel values of each pixel in the target prediction unit, and the motion compensation value of each pixel can be calculated according to the predicted value of each pixel and the residual information; if there is a large difference between the motion compensation value of at least one pixel and the actual pixel value, it can be determined that there are abnormal distortion points in the target prediction unit in any mode; if the differences between the motion compensation values of each pixel and the actual pixel values are all small, it can be determined that there are no abnormal distortion points in the target prediction unit in any mode.

[0120] S203. According to the detection results corresponding to at least one candidate prediction mode, calibrate the mode costs of at least one candidate prediction mode in the mode information set to obtain a calibrated mode information set.

[0121] In a specific implementation, the detection results of each candidate prediction mode in at least one candidate prediction mode to be detected can be traversed in sequence. In each traversal process, the mode cost of the currently traversed candidate prediction mode in the mode information set can be calibrated according to the detection result of the currently traversed candidate prediction mode. Specifically, if the detection result of the currently traversed candidate prediction mode indicates that there is no abnormal distortion point in the target prediction unit under the currently traversed candidate prediction mode, the mode cost of the currently traversed candidate prediction mode in the mode information is kept unchanged; that is, in this case, the calibrated mode cost of the currently traversed candidate prediction mode is the same as the mode cost before calibration. If the detection result of the currently traversed candidate prediction mode indicates that there is an abnormal distortion point in the target prediction unit under the currently traversed candidate prediction mode, at least one of the following penalty processes can be performed on the mode cost of the currently traversed candidate prediction mode in the mode information set: magnifying the mode cost of the currently traversed candidate prediction mode, and adding a disabled flag to the currently traversed candidate prediction mode; that is, in this case, the calibrated mode cost of the currently traversed candidate prediction mode may be the same as or different from the mode cost before calibration. Iterate the above traversal steps until all the candidate prediction modes detected in step S202 are traversed, and a calibrated mode information set can be obtained; the calibrated mode information set includes the calibrated mode costs of each candidate prediction mode detected in step S202, and the mode costs of the candidate prediction modes not detected in step S202.

[0122] S204. Select a target prediction mode from multiple candidate prediction modes according to the mode costs of each candidate prediction mode in the calibrated mode information set.

[0123] In a specific implementation, the candidate prediction mode with the smallest mode cost in the calibrated mode information set can be selected from multiple candidate prediction modes as the target prediction mode. Optionally, if the candidate prediction mode with the smallest mode cost in the calibrated mode information set has a disabled flag, the candidate prediction mode with the second smallest (i.e., the second smallest) mode cost in the calibrated mode information set can be selected as the target prediction mode. Further, if the candidate prediction mode with the second smallest mode cost in the calibrated mode information set also has a disabled flag, the candidate prediction mode with the third smallest mode cost in the calibrated mode information set can be selected as the target prediction mode, and so on. In another specific implementation, the standby prediction modes can be screened out from multiple candidate prediction modes according to the mode costs of each candidate prediction mode in the calibrated mode information set; the standby prediction modes here refer to the candidate prediction modes with mode costs greater than the cost threshold in the calibrated mode information set, and the cost threshold can be set according to empirical values. Then, a standby prediction mode can be randomly selected from the screened standby prediction modes as the target prediction mode.

[0124] S205. Perform prediction processing on the target prediction unit in the target prediction mode to obtain the encoded data of the target image block.

[0125] In a specific implementation, the pixel value prediction can be performed on each pixel point in the target prediction unit in the target prediction mode to obtain the prediction result of the target prediction unit. The prediction result of the target prediction unit here may include the predicted values of each pixel point in the target prediction unit. Repeat the above steps S201 - S205 iteratively to obtain the prediction results of each prediction unit in the target image block. Then, the prediction results of each prediction unit can be combined to obtain the prediction block corresponding to the target image block, and the residual block can be obtained based on the target image block and the prediction block. Finally, perform transformation, quantization, and entropy coding processing on the residual block in sequence to obtain the encoded data of the target image block.

[0126] In the embodiment of the present invention during the encoding process, the abnormal distortion points can be detected for the target prediction unit under at least one candidate prediction mode in the mode information set first to obtain the detection results corresponding to at least one candidate prediction mode. Secondly, the mode cost of at least one candidate prediction mode in the mode information set can be calibrated according to the detection results corresponding to at least one candidate prediction mode, so that the mode costs of each candidate prediction mode in the calibrated mode information set can more accurately reflect the corresponding code rate and distortion of the candidate prediction mode, thereby enabling the selection of a target prediction mode more suitable for the target prediction unit from multiple candidate prediction modes according to the mode costs of each candidate prediction mode in the calibrated mode information set. Then, a suitable target prediction mode can be used to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block, which can reduce the probability of distortion of the target image block after encoding to a certain extent. And since the embodiment of the present invention mainly reduces the distortion probability by modifying the mode decision process to select a suitable target prediction mode, it can effectively improve the image compression quality and enhance the subjective quality of the target image block without substantially affecting the compression efficiency and encoding complexity.

[0127] Please refer to Figure 3 , which is a schematic flowchart of another video encoding method provided by the embodiment of the present invention. This video encoding method can be executed by the above-mentioned video encoding device, specifically by the encoder in the video encoding device. Please refer to Figure 3 , this video encoding method may include the following steps S301 - S307:

[0128] S301. Obtain the target prediction unit in the target image block and the mode information set of the target prediction unit.

[0129] In a specific implementation, the target image block can be divided into at least one prediction unit; then, an unprocessed prediction unit is selected from the at least one prediction unit as the target prediction unit. After determining the target prediction unit, a set of pattern information matching the target prediction unit can also be obtained. As described above, the set of pattern information of the target prediction unit includes multiple candidate prediction patterns and the pattern costs of various candidate prediction patterns; accordingly, the specific implementation of obtaining the set of pattern information matching the target prediction unit can include the following steps:

[0130] First, multiple candidate prediction patterns matching the target prediction unit can be determined. Specifically, it can be detected whether the target image block belongs to an I Slice (Intra Slice). Since an I Slice usually only includes I macroblocks, and I macroblocks can only use the encoded pixel points in the current frame image as a reference for intra-frame prediction; therefore, if the target image block belongs to an I Slice, the intra-frame prediction mode can be directly used to perform prediction processing on the target prediction unit, and the subsequent steps are not executed. If the target image block does not belong to an I Slice, it means that either the intra-frame prediction mode or the inter-frame prediction mode can be used to predict the target prediction unit; therefore, each mode in the intra-frame prediction mode and the inter-frame prediction mode can be selected as multiple candidate prediction patterns matching the target prediction unit. After determining the multiple candidate prediction patterns, the pattern costs of each candidate prediction pattern can be calculated respectively using a cost function; the cost function here can include but is not limited to: the cost function of the RDO (Rate-distortion Optimized) mode, such as the cost function shown in the following formula 1.1; the cost function of the non-RDO mode, such as the cost function shown in the following formula 1.2, etc. Then, the calculated pattern costs of each candidate prediction pattern and the corresponding candidate prediction patterns can be added to the set of pattern information.

[0131] cost = HAD + λ·R Formula 1.1

[0132] cost = SAD + 4R*λ(QP) Formula 1.2

[0133] In Equation 1.1, cost represents the mode cost of a candidate prediction mode; HAD represents the sum of the absolute values of the coefficients after the residual signal of the target prediction unit is subjected to Hadamard transform; λ represents the Lagrangian coefficient, and R represents the number of bits required to encode the candidate prediction mode (i.e., the coding rate). In Equation 1.2, cost still represents the mode cost of the candidate prediction mode; SAD represents the sum of the absolute differences between the prediction results obtained by using the candidate prediction mode to predict the pixel values of each pixel point in the target prediction unit and the target prediction unit; 4R represents the number of bits estimated after using the candidate prediction mode; λ(QP) represents an exponential function related to the quantization parameter (QP). It should be noted that the above Equations 1.1 and 1.2 are only used to illustrate the cost function and are not exhaustive.

[0134] S302. Detect abnormal distortion points for the target prediction unit under at least one candidate prediction mode in the mode information set, and obtain detection results corresponding to at least one candidate prediction mode.

[0135] In the embodiments of the present invention, at least one candidate prediction mode is mainly taken as an example of an inter-frame prediction mode for illustration; the inter-frame prediction mode at least includes the following modes: a first prediction mode (SKIP mode), a second prediction mode (ordinary Merge mode), and a third prediction mode (AMVP mode). Since the principle of detecting abnormal distortion points for the target prediction unit under each mode in the inter-frame prediction mode is similar, for the convenience of description, the implementation manner of step S302 is described below by taking a reference prediction mode as an example; wherein, the reference prediction mode is any mode in the inter-frame prediction mode; that is, the reference prediction mode can be the first prediction mode, the second prediction mode, or the third prediction mode.

[0136] In a specific implementation, the following Equation 1.3 can be used to detect abnormal distortion points:

[0137] ABS(DIFF(x, y)) > TH Equation 1.3

[0138] In the above formula 1.3, DIFF(x, y) represents the residual (i.e., the difference) between the actual pixel value and the predicted value of the pixel at position (x, y) in the target prediction unit; ABS represents taking the absolute value, and TH represents the target threshold. If the absolute value of the residual between the actual pixel value and the predicted value of the pixel at position (x, y) in the target prediction unit is greater than the target threshold, then it can be determined that this pixel is an abnormal pixel; otherwise, it can be determined that this pixel is a normal pixel. Based on this, during the specific execution of step S302, the pixel values of each pixel in the target prediction unit can be predicted using the reference prediction mode first to obtain the predicted values of each pixel. Secondly, the absolute value of the residual between the pixel value and the predicted value of each pixel in the target prediction unit can be calculated. If there is a pixel in the target prediction unit whose absolute value of the residual is greater than the target threshold, then it is determined that the detection result corresponding to the reference prediction mode indicates that there is an abnormal distortion point in the target prediction unit in the reference prediction mode; if there is no pixel in the target prediction unit whose absolute value of the residual is greater than the target threshold, then it can be determined that the detection result corresponding to the reference prediction mode indicates that there is no abnormal distortion point in the target prediction unit in the reference prediction mode.

[0139] Among them, the target threshold can at least include the following two value-taking methods: In one implementation manner, the above-mentioned target threshold can be set to a unified fixed value according to empirical values; that is, in this case, regardless of whether the reference prediction mode is the first prediction mode, the second prediction mode, or the third prediction mode, the target threshold for abnormal distortion point judgment is the same. In another implementation manner, since the second prediction mode and the third prediction mode need to transmit residual information, the detection criteria for abnormal distortion points in the second prediction mode and the third prediction mode can be relaxed; in this case, the target threshold can be set specifically for different reference target thresholds according to empirical values; that is, the target threshold can be associated with the reference prediction mode. Specifically, if the reference prediction mode is the first prediction mode in the inter-frame prediction mode, the target threshold can be equal to the first threshold. Among them, the first threshold is greater than the invalid value and less than the maximum value of the pixel value range; the invalid value here can be set according to empirical values, for example, set to 0; the maximum value of the pixel value range can be determined according to the pixel bit width. The so-called pixel bit width refers to the number of pixels transmitted or displayed at one time. That is to say, the value range of the first threshold (TH1) can be: 0 < TH1 < (2 << (BITDEPTH)); where, BITDEPTH represents the pixel bit width, << represents the exponentiation operation, and 2 << (BITDEPTH) represents the maximum value of the pixel value range; for example, if the pixel bit width is 8, then the maximum value of the pixel value range is equal to 2 to the 8th power (i.e., 256); then the first threshold can be any value between 0 and 256, for example, the first threshold can be set to 30. If any candidate prediction mode is the second prediction mode or the third prediction mode in the inter-frame prediction mode, the target threshold can be equal to the second threshold; the second threshold is greater than or equal to the first threshold and less than the maximum value of the pixel value range. That is to say, the value range of the second threshold (TH2) can be: TH1 <= TH2 < (2 << (BITDEPTH)).

[0140] In another specific implementation, during the specific execution of step S302, the pixel values of each pixel point in the target prediction unit can be predicted first using the reference prediction mode to obtain the predicted values of each pixel point; and the motion compensation value of each pixel point can be calculated based on the predicted value and residual information of each pixel point. Secondly, the absolute value of the difference between the pixel value and the motion compensation value of each pixel point in the target prediction unit can be calculated. If there are pixel points in the target prediction unit whose absolute value of the difference is greater than the target threshold, it can be determined that the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit in the reference prediction mode; if there are no pixel points in the target prediction unit whose absolute value of the difference is greater than the target threshold, it can be determined that the detection result corresponding to the reference prediction mode indicates that there are no abnormal distortion points in the target prediction unit in the reference prediction mode. Among them, the target threshold in this implementation manner can be set to a unified fixed value.

[0141] S303. Perform complexity analysis on the target prediction unit to obtain the prediction complexity of the target prediction unit.

[0142] In a specific implementation, gradient operations can be performed on the pixel values included in the target prediction unit to obtain the image gradient value of the target prediction unit; and this image gradient value can be used as the prediction complexity of the target prediction unit. In another specific implementation, variance operations or mean operations can be performed on the pixel values included in the target prediction unit, and the obtained variance or mean can be used as the prediction complexity of the target prediction unit. It should be understood that the embodiments of the present invention only exemplarily list two specific implementation manners of complexity analysis and are not exhaustive. After analyzing and obtaining the prediction complexity of the target prediction unit, it can be detected whether the target prediction unit meets a preset condition according to the prediction complexity; wherein, the preset condition may at least include: the prediction complexity is less than or equal to a complexity threshold, and there are abnormal distortion points in at least one mode of the target prediction unit in the inter-frame prediction mode. If it is satisfied, step S304 can be executed; if it is not satisfied, steps S305 - S308 can be executed. Thus, it can be seen that by adding prior information such as prediction complexity in the embodiments of the present invention, when it is determined according to the prediction complexity that the target prediction unit meets the preset condition, the decision-making process of other modes can be skipped and intra-frame prediction can be directly performed; in this way, the mode decision-making process can be effectively accelerated, thereby improving the encoding speed.

[0143] S304. If it is determined according to the prediction complexity that the target prediction unit meets the preset condition, then use the intra-frame prediction mode to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block.

[0144] S305. If it is determined according to the prediction complexity that the target prediction unit does not meet the preset condition, then calibrate the mode cost of at least one candidate prediction mode in the mode information set according to the detection results corresponding to at least one candidate prediction mode to obtain a calibrated mode information set.

[0145] Since the principle of calibrating the mode cost of each candidate prediction mode according to the detection result of each candidate prediction mode is similar; for the convenience of description, the implementation manner of step S305 will be described below by taking the reference prediction mode as an example; wherein, the reference prediction mode is any mode in the inter-frame prediction mode; that is, the reference prediction mode can be the first prediction mode, the second prediction mode or the third prediction mode. In a specific implementation, if the detection result corresponding to the reference prediction mode indicates that there is no abnormal distortion point in the target prediction unit in the reference prediction mode, the mode cost of the reference prediction mode in the mode information set can be kept unchanged to obtain the calibrated mode information set. If the detection result corresponding to the reference prediction mode indicates that there is an abnormal distortion point in the target prediction unit in the reference prediction mode, the mode cost of the reference prediction mode in the mode information set can be adjusted by using the cost adjustment strategy of the reference prediction mode to obtain the calibrated mode information set.

[0146] Specifically, if the reference prediction mode is the second prediction mode or the third prediction mode, the step of adjusting the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode may include: amplifying the mode cost of the reference prediction mode by using a penalty factor to obtain the calibrated mode cost of the reference prediction mode. Wherein, the penalty factor is any value greater than 0; its specific value can be set according to empirical values. It should be noted that in this case, if the target prediction mode selected through step S306 later is the second prediction mode or the third prediction mode, the video coding device (or encoder) can also be forced to transmit the residual values of the pixels in the target image block that have not been transformed and quantized to the decoder, so that the decoder can perform decoding according to the residual values. If the reference prediction mode is the first prediction mode, the step of adjusting the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode may further include the following implementation manners:

[0147] The first implementation manner: a preset cost can be obtained; the preset cost is greater than the mode costs of the candidate prediction modes other than the first prediction mode in the calibrated mode information set, and greater than the mode cost of the first prediction mode in the mode information set (that is, greater than the mode cost of the first prediction mode before calibration). Then, in the mode information set, the mode cost of the reference prediction mode is adjusted to the preset cost. In this implementation manner, the specific implementation manner of the following step S306 can be: directly selecting the candidate prediction mode with the smallest mode cost in the calibrated mode information from multiple candidate prediction modes as the target prediction mode. It can be seen that in this implementation manner, if there is an abnormal distortion point in the target prediction unit in the first prediction mode, the mode cost of the first prediction mode can be set to an infinite value, so that the first prediction mode will not be selected when the target prediction mode is selected in ascending order of the mode cost later.

[0148] In the second implementation manner, the mode cost of the first prediction mode in the mode information set can be kept unchanged; and a disable flag is added to the first prediction mode, and the disable flag indicates that it is prohibited to use the first prediction mode to perform prediction processing on the target prediction unit. In this implementation manner, the specific implementation manner of the following step S306 can be: If the candidate prediction mode with the smallest mode cost in the calibrated mode information set is not the first prediction mode, or the candidate prediction mode with the smallest mode cost is the first prediction mode and the first prediction mode does not have a disable flag, then the candidate prediction mode with the smallest mode cost can be used as the target prediction mode. If the candidate prediction mode with the smallest mode cost is the first prediction mode and the first prediction mode has a disable flag, then the candidate prediction mode with the second smallest mode cost in the calibrated mode information set can be selected as the target prediction mode. Thus, it can be seen that in this implementation manner, if there are abnormal distortion points in the target prediction unit in the first prediction mode, the first prediction mode can be prevented from being selected as the target prediction mode in the subsequent selection of the target prediction mode by adding a disable flag.

[0149] In the third implementation manner, the mode cost of the first prediction mode in the mode information set can be kept unchanged; and no processing is performed on the first prediction mode. In this implementation manner, the specific implementation manner of the following step S306 can be: If the candidate prediction mode with the smallest mode cost in the calibrated mode information set is not the first prediction mode, then the candidate prediction mode with the smallest mode cost is directly used as the target prediction mode. If the candidate prediction mode with the smallest mode cost in the calibrated mode information set is not the first prediction mode, then the detection result of the first prediction mode can be queried again. If the detection result of the first prediction mode indicates that there are abnormal distortion points in the target prediction unit in the first prediction mode, then the candidate prediction mode with the second smallest mode cost in the calibrated mode information set can be selected as the target prediction mode; if the detection result of the first prediction mode indicates that there are no abnormal distortion points in the target prediction unit in the first prediction mode, then the first prediction mode can be used as the target prediction mode.

[0150] S306. Select a target prediction mode from multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set.

[0151] S307. Perform prediction processing on the target prediction unit using the target prediction mode to obtain the encoded data of the target image block.

[0152] In the encoding process of the embodiments of the present invention, abnormal distortion points can be detected for a target prediction unit under at least one candidate prediction mode in a mode information set to obtain detection results corresponding to at least one candidate prediction mode. Secondly, according to the detection results corresponding to at least one candidate prediction mode, the mode costs of at least one candidate prediction mode in the mode information set can be calibrated; so that the mode costs of each candidate prediction mode in the calibrated mode information set can more accurately reflect the corresponding bit rate and distortion of the corresponding candidate prediction mode, thereby enabling the selection of a target prediction mode more suitable for the target prediction unit from multiple candidate prediction modes according to the mode costs of each candidate prediction mode in the calibrated mode information set. Then, a suitable target prediction mode can be used to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block, which can reduce the probability of distortion of the target image block after encoding to a certain extent; and, since the embodiments of the present invention mainly reduce the distortion probability by correcting the mode decision process to select a suitable target prediction mode, it is possible to effectively improve the image compression quality and enhance the subjective quality of the target image block without substantially affecting the compression efficiency and encoding complexity.

[0153] Based on the related descriptions of the embodiments of the above video encoding method, the embodiments of the present invention also propose a video playback method; this video playback method can be executed by the above-mentioned video playback device. Please refer to Figure 4 , the video playback method may include the following steps S401 - S403:

[0154] S401, Obtain the bitstream data of each frame image in the image frame sequence corresponding to the target video.

[0155] Among them, the target video may include but is not limited to: screen sharing video, web conference video, webcast video, movie and TV drama video, short video, etc. The bitstream data of each frame image in the image frame sequence corresponding to the target video may include the encoded data of multiple image blocks; and, the encoded data of each image block in other frame images except the first frame image in the image frame sequence can be encoded by the video encoding method as Figure 2 or Figure 3 shown.

[0156] In a specific implementation, the video playback device can obtain the bitstream data of each frame image in the image frame sequence corresponding to the target video from the video encoding device. In one implementation, the bitstream data of each frame image in the image frame sequence corresponding to the target video can be obtained through real-time encoding; in this case, the video playback device can obtain the bitstream data of each frame image from the video encoding device in real time. That is to say, in this implementation, every time the video encoding device encodes the bitstream data of a frame image, it can transmit the bitstream data of this frame image to the video playback device for decoding and playback. In yet another implementation, the bitstream data of each frame image in the image frame sequence corresponding to the target video can also be obtained through pre-offline encoding; in this case, the video playback device can also obtain the bitstream data of each frame image in the image frame sequence from the video encoding device at once. That is to say, in this implementation, after the video encoding device encodes the bitstream data of all frame images, it can then transmit the bitstream data of all frame images to the video playback device for decoding and playback.

[0157] S402. Decode the bitstream data of each frame image to obtain each frame image.

[0158] S403. Display each frame image in sequence on the playback interface.

[0159] It should be noted that the specific implementation manners of steps S402 - S403 can refer to the relevant content in the decoding stage mentioned in the foregoing image processing flow, which will not be elaborated here. And it should also be noted that if the bitstream data of each frame image in the image frame sequence corresponding to the target video is obtained through real-time encoding and transmitted to the video playback device in real time, then every time the video playback device receives the bitstream data of a frame image, it can execute steps S402 - S403 to achieve real-time display of the frame images, thereby realizing real-time playback of the target video.

[0160] In the embodiment of the present invention, first, the bitstream data of each frame image in the image frame sequence corresponding to the target video can be obtained. The bitstream data of each frame image includes the encoding data of multiple image blocks. Secondly, the bitstream data of each frame image can be decoded to obtain each frame image; and each frame image can be displayed in sequence on the playback interface. Since the encoding data of each image block in other frame images except the first frame image in the image frame sequence corresponding to the target video is encoded by using the above-mentioned video encoding method; thus, the probability of distortion of each image block can be effectively reduced, so that when each frame image is displayed on the playback interface, the probability of each frame image having dirty points can be reduced to a certain extent, and the subjective quality of each frame image can be improved.

[0161] It should be understood that the video encoding method and video playback method proposed in the embodiments of the present invention can be applied to various application scenarios; such as screen sharing scenarios in video conferencing, video live streaming scenarios, movie and TV drama video playback scenarios, and so on. Taking the screen sharing scenario in video conferencing as an example, the specific applications of the video encoding method and video playback method proposed in the embodiments of the present invention are described below:

[0162] During the process of multiple users using communication clients with video conferencing functions (such as enterprise WeChat clients, Tencent Meeting clients, etc.) to conduct video conferencing, if user A wants to share their screen content with other users, the screen sharing function can be enabled. After the first communication client used by user A detects that the screen sharing function is turned on, it can obtain the display content in the terminal screen corresponding to user A in real time, and generate the current frame image of the screen sharing video based on the real-time obtained display content. Then, the first communication client can encode the current frame image to obtain the bitstream data of the current frame image. Specifically, the current frame image can be divided into multiple image blocks, and the video encoding method shown in Figure 2 or Figure 3 is used to encode each image block to obtain the encoded data of each image block; and the encoded data of each image block is combined to obtain the bitstream data of the current frame image. After obtaining the bitstream data of the current frame image, the first communication client can transmit the bitstream data of the current frame image to the second communication client used by other users through the server. Correspondingly, after the second communication client used by other users receives the bitstream data of the current frame image sent by the first communication client, it can decode the bitstream data of the current frame image using the video playback method shown in Figure 4 to obtain the current frame image; and display the current frame image in the user interface, as shown in Figure 5

[0163] It can be seen that by using the video encoding method and video playback method proposed in the embodiments of the present invention, the probability of abnormal distortion points in the screen sharing scenario can be effectively reduced, the compressed video quality of the screen sharing video can be effectively improved, and the subjective quality of the screen sharing video can be enhanced.

[0164] Based on the description of the above video encoding method embodiments, the embodiments of the present invention also disclose a video encoding device, and the video encoding device can be a computer program (including program code) running in a video encoding device. This video encoding device can execute Figures 2 - 3 the method shown in. Please refer to Figure 6 , and the video encoding device can run the following units:

[0165] ​An acquisition unit 601, configured to acquire a target prediction unit in a target image block and a set of mode information of the target prediction unit, where the set of mode information includes a plurality of candidate prediction modes and mode costs of various candidate prediction modes;

[0166] An encoding unit 602, configured to perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information, so as to obtain a detection result corresponding to the at least one candidate prediction mode;

[0167] The encoding unit 602 is further configured to calibrate the mode costs of the at least one candidate prediction mode in the set of mode information according to the detection results corresponding to the at least one candidate prediction mode, so as to obtain a calibrated set of mode information;

[0168] The encoding unit 602 is further configured to select a target prediction mode from the plurality of candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated set of mode information;

[0169] The encoding unit 602 is further configured to perform prediction processing on the target prediction unit by using the target prediction mode, so as to obtain encoded data of the target image block.

[0170] In an implementation manner, the at least one candidate prediction mode is an inter prediction mode, and the inter prediction mode at least includes the following modes: a first prediction mode, a second prediction mode, and a third prediction mode;

[0171] The first prediction mode refers to a mode that needs to transmit index information of a reference image block related to the target image block;

[0172] The second prediction mode refers to a mode that needs to transmit residual information of the target image block and index information of a reference image block related to the target image block;

[0173] The third prediction mode refers to a mode that needs to transmit residual information of the target image block, motion vector data of the target image block, and index information of a reference image block related to the target image block.

[0174] In still another implementation manner, when the encoding unit 602 is configured to perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode, it may be specifically configured to:

[0175] Perform pixel value prediction on each pixel point in the target prediction unit by using a reference prediction mode to obtain a predicted value of each pixel point; the reference prediction mode is any mode in the inter prediction mode;

[0176] Calculate the absolute value of the residual between the pixel value and the predicted value of each pixel point in the target prediction unit;

[0177] If there are pixel points in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit in the reference prediction mode;

[0178] If there are no pixel points in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there are no such abnormal distortion points in the target prediction unit in the reference prediction mode.

[0179] In another implementation, the target threshold is associated with the reference prediction mode;

[0180] If the reference prediction mode is the first prediction mode in the inter-frame prediction mode, the target threshold is equal to the first threshold; the first threshold is greater than the invalid value and less than the maximum value of the pixel value range;

[0181] If any of the candidate prediction modes is the second prediction mode or the third prediction mode in the inter-frame prediction mode, the target threshold is equal to the second threshold; the second threshold is greater than or equal to the first threshold and less than the maximum value of the pixel value range.

[0182] In another implementation, when the encoding unit 602 is used to calibrate the mode cost of at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain a calibrated mode information set, it can be specifically used for:

[0183] If the detection result corresponding to the reference prediction mode indicates that there are no abnormal distortion points in the target prediction unit in the reference prediction mode, keep the mode cost of the reference prediction mode in the mode information set unchanged to obtain a calibrated mode information set; the reference prediction mode is any mode in the inter-frame prediction mode;

[0184] If the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit in the reference prediction mode, adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode to obtain a calibrated mode information set.

[0185] In another implementation, when the encoding unit 602 is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it can be specifically used for:

[0186] If the reference prediction mode is the second prediction mode or the third prediction mode, a penalty factor is used to amplify the mode cost of the reference prediction mode to obtain the calibrated mode cost of the reference prediction mode.

[0187] In another implementation, when the encoding unit 602 is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it can also be used for:

[0188] If the reference prediction mode is the first prediction mode, obtain a preset cost; the preset cost is greater than the mode costs of each candidate prediction mode other than the first prediction mode in the calibrated mode information set, and greater than the mode cost of the first prediction mode in the mode information set;

[0189] In the mode information set, adjust the mode cost of the reference prediction mode to the preset cost.

[0190] In another implementation, when the encoding unit 602 is used to select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set, it can specifically be used for:

[0191] Select, from the multiple candidate prediction modes, the candidate prediction mode with the smallest mode cost in the calibrated mode information as the target prediction mode.

[0192] In another implementation, when the encoding unit 602 is used to adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode, it can also be used for:

[0193] If the reference prediction mode is the first prediction mode, keep the mode cost of the first prediction mode in the mode information set unchanged;

[0194] Add a disabled flag to the first prediction mode, and the disabled flag indicates that it is prohibited to use the first prediction mode to perform prediction processing on the target prediction unit.

[0195] In another implementation, when the encoding unit 602 is used to select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set, it can specifically be used for:

[0196] If the candidate prediction mode with the minimum mode cost in the calibrated mode information set is not the first prediction mode, or if the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode does not have the disable flag, then use the candidate prediction mode with the minimum mode cost as the target prediction mode;

[0197] If the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode has the disable flag, then select the candidate prediction mode with the second smallest mode cost in the calibrated mode information set as the target prediction mode.

[0198] In another implementation, the multiple candidate prediction modes include: intra prediction mode and inter prediction mode; correspondingly, the encoding unit 602 can also be used to:

[0199] Perform complexity analysis on the target prediction unit to obtain the prediction complexity of the target prediction unit;

[0200] If it is determined according to the prediction complexity that the target prediction unit meets the preset conditions, then use the intra prediction mode to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block; wherein, the preset conditions include: the prediction complexity is less than or equal to the complexity threshold, and there are abnormal distortion points in at least one mode of the target prediction unit in the inter prediction mode;

[0201] If it is determined according to the prediction complexity that the target prediction unit does not meet the preset conditions, then perform the step of calibrating the mode costs of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain the calibrated mode information set.

[0202] According to an embodiment of the present invention, Figures 2 - 3 Each step involved in the method shown can be performed by Figure 6 each unit in the video encoding device shown. For example, Figure 2 the step S201 shown in Figure 6 can be performed by the obtaining unit 601 shown in Figure 6 , and the steps S202 - S205 can be performed by Figure 3 the encoding unit 602 shown in Figure 6 ; again, Figure 6 the step S301 shown in

[0203] can be performed by the obtaining unit 601 shown in Figure 6Each unit in the video encoding device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operations without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, based on the video encoding device, other units may also be included. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0204] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in Figures 2 - 3 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a video encoding device as shown in Figure 6 and to implement the video encoding method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0205] In the encoding process of the embodiments of the present invention, abnormal distortion point detection can be first performed on a target prediction unit under at least one candidate prediction mode in the mode information set to obtain detection results corresponding to at least one candidate prediction mode. Secondly, the mode cost of at least one candidate prediction mode in the mode information set can be calibrated according to the detection results corresponding to at least one candidate prediction mode; so that the mode costs of the candidate prediction modes in the calibrated mode information set can more accurately reflect the corresponding bit rate and distortion of the candidate prediction modes, thereby enabling the selection of a target prediction mode more suitable for the target prediction unit from multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set. Then, an appropriate target prediction mode can be used to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block, which can reduce the probability of distortion of the target image block after encoding to a certain extent; and since the embodiments of the present invention mainly reduce the distortion probability by correcting the mode decision process to select an appropriate target prediction mode, it is possible to effectively improve the image compression quality and enhance the subjective quality of the target image block without substantially affecting the compression efficiency and encoding complexity.

[0206] Based on the descriptions of the above video encoding method embodiments and video encoding device embodiments, the embodiments of the present invention also provide a video encoding device. Please refer toFigure 7 , the video encoding device at least includes a processor 701, an input interface 702, an output interface 703, a computer storage medium 704, and an encoder 705. Among them, the computer storage medium 704 can be stored in the memory of the video encoding device. The computer storage medium 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the video encoding device, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; in one embodiment, the processor 701 described in the embodiment of the present invention can be used to perform a series of video encodings on a target image block, including: obtaining a target prediction unit in the target image block and a set of mode information of the target prediction unit, the set of mode information including multiple candidate prediction modes and the mode costs of various candidate prediction modes; performing abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode; calibrating the mode costs of the at least one candidate prediction mode in the set of mode information according to the detection result corresponding to the at least one candidate prediction mode to obtain a calibrated set of mode information; selecting a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated set of mode information; using the target prediction mode to perform prediction processing on the target prediction unit to obtain encoded data of the target image block, and so on.

[0207] The embodiment of the present invention also provides a computer storage medium (Memory). The computer storage medium is a memory device in the video encoding device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the video encoding device and, of course, the extended storage medium supported by the video encoding device. The computer storage medium provides a storage space, and the operating system of the video encoding device is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor 701 are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.

[0208] In one embodiment, one or more first instructions stored in a computer storage medium may be loaded and executed by a processor 701 to implement the corresponding steps of the method in the above-described embodiment of the video encoding method; specifically, one or more first instructions in the computer storage medium are loaded and executed by the processor 701 to perform the following steps:

[0209] Obtain a target prediction unit in a target image block and a set of mode information of the target prediction unit, where the set of mode information includes a plurality of candidate prediction modes and mode costs of various candidate prediction modes;

[0210] Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode;

[0211] According to the detection results corresponding to the at least one candidate prediction mode, calibrate the mode costs of the at least one candidate prediction mode in the set of mode information to obtain a calibrated set of mode information;

[0212] Select a target prediction mode from the plurality of candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated set of mode information;

[0213] Perform prediction processing on the target prediction unit using the target prediction mode to obtain encoded data of the target image block.

[0214] In one implementation manner, the at least one candidate prediction mode is an inter prediction mode, and the inter prediction mode at least includes the following modes: a first prediction mode, a second prediction mode, and a third prediction mode;

[0215] The first prediction mode refers to a mode that needs to transmit index information of a reference image block related to the target image block;

[0216] The second prediction mode refers to a mode that needs to transmit residual information of the target image block and index information of a reference image block related to the target image block;

[0217] The third prediction mode refers to a mode that needs to transmit residual information of the target image block, motion vector data of the target image block, and index information of a reference image block related to the target image block.

[0218] In yet another implementation manner, when performing abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information to obtain a detection result corresponding to the at least one candidate prediction mode, the one or more first instructions are loaded and specifically executed by the processor 701:

[0219] Predict the pixel values of each pixel point in the target prediction unit using a reference prediction mode to obtain the predicted values of each pixel point; the reference prediction mode is any mode in the inter-frame prediction mode;

[0220] Calculate the absolute value of the residual between the pixel value and the predicted value of each pixel point in the target prediction unit;

[0221] If there are pixel points in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit in the reference prediction mode;

[0222] If there are no pixel points in the target prediction unit whose absolute value of the residual is greater than the target threshold, determine that the detection result corresponding to the reference prediction mode indicates that there are no such abnormal distortion points in the target prediction unit in the reference prediction mode.

[0223] In another implementation, the target threshold is associated with the reference prediction mode;

[0224] If the reference prediction mode is the first prediction mode in the inter-frame prediction mode, the target threshold is equal to the first threshold; the first threshold is greater than the invalid value and less than the maximum value of the pixel value range;

[0225] If any of the candidate prediction modes is the second prediction mode or the third prediction mode in the inter-frame prediction mode, the target threshold is equal to the second threshold; the second threshold is greater than or equal to the first threshold and less than the maximum value of the pixel value range.

[0226] In another implementation, when calibrating the mode costs of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain the calibrated mode information set, the one or more first instructions can be loaded and specifically executed by the processor 701:

[0227] If the detection result corresponding to the reference prediction mode indicates that there are no abnormal distortion points in the target prediction unit in the reference prediction mode, keep the mode cost of the reference prediction mode in the mode information set unchanged to obtain the calibrated mode information set; the reference prediction mode is any mode in the inter-frame prediction mode;

[0228] If the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit in the reference prediction mode, adjust the mode cost of the reference prediction mode in the mode information set using the cost adjustment strategy of the reference prediction mode to obtain the calibrated mode information set.

[0229] In another implementation, when adjusting the pattern cost of the reference prediction pattern in the pattern information set by using the cost adjustment strategy of the reference prediction pattern, the one or more first instructions are loaded and specifically executed by the processor 701:

[0230] If the reference prediction pattern is the second prediction pattern or the third prediction pattern, the pattern cost of the reference prediction pattern is amplified by using a penalty factor to obtain the calibrated pattern cost of the reference prediction pattern.

[0231] In another implementation, when adjusting the pattern cost of the reference prediction pattern in the pattern information set by using the cost adjustment strategy of the reference prediction pattern, the one or more first instructions are loaded and specifically executed by the processor 701:

[0232] If the reference prediction pattern is the first prediction pattern, a preset cost is obtained; the preset cost is greater than the pattern costs of the candidate prediction patterns other than the first prediction pattern in the calibrated pattern information set, and greater than the pattern cost of the first prediction pattern in the pattern information set;

[0233] In the pattern information set, the pattern cost of the reference prediction pattern is adjusted to the preset cost.

[0234] In another implementation, when selecting a target prediction pattern from the multiple candidate prediction patterns according to the pattern costs of the candidate prediction patterns in the calibrated pattern information set, the one or more first instructions are loaded and specifically executed by the processor 701:

[0235] From the multiple candidate prediction patterns, the candidate prediction pattern with the smallest pattern cost in the calibrated pattern information is selected as the target prediction pattern.

[0236] In another implementation, when adjusting the pattern cost of the reference prediction pattern in the pattern information set by using the cost adjustment strategy of the reference prediction pattern, the one or more first instructions are loaded and specifically executed by the processor 701:

[0237] If the reference prediction pattern is the first prediction pattern, the pattern cost of the first prediction pattern in the pattern information set remains unchanged;

[0238] A disable flag is added to the first prediction pattern, and the disable flag indicates that the first prediction pattern is prohibited from being used to perform prediction processing on the target prediction unit.

[0239] In another implementation, when selecting a target prediction mode from the multiple candidate prediction modes according to the mode cost of each candidate prediction mode in the calibrated mode information set, the one or more first instructions are loaded and specifically executed by the processor 701:

[0240] If the candidate prediction mode with the minimum mode cost in the calibrated mode information set is not the first prediction mode, or the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode does not have the disable flag, then the candidate prediction mode with the minimum mode cost is used as the target prediction mode;

[0241] If the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode has the disable flag, then the candidate prediction mode with the second smallest mode cost in the calibrated mode information set is selected as the target prediction mode.

[0242] In another implementation, the multiple candidate prediction modes include: intra prediction modes and inter prediction modes; correspondingly, the one or more first instructions can also be loaded and specifically executed by the processor 701:

[0243] Perform complexity analysis on the target prediction unit to obtain the prediction complexity of the target prediction unit;

[0244] If it is determined according to the prediction complexity that the target prediction unit meets the preset conditions, then the intra prediction mode is used to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block; wherein, the preset conditions include: the prediction complexity is less than or equal to the complexity threshold, and there are abnormal distortion points in at least one mode of the target prediction unit in the inter prediction mode;

[0245] If it is determined according to the prediction complexity that the target prediction unit does not meet the preset conditions, then perform the step of calibrating the mode cost of the at least one candidate prediction mode in the mode information set according to the detection result corresponding to the at least one candidate prediction mode to obtain the calibrated mode information set.

[0246] In the encoding process of the embodiments of the present invention, abnormal distortion points of a target prediction unit may be detected under at least one candidate prediction mode in a mode information set first to obtain detection results corresponding to at least one candidate prediction mode. Secondly, according to the detection results corresponding to at least one candidate prediction mode, the mode costs of at least one candidate prediction mode in the mode information set may be calibrated, so that the mode costs of the candidate prediction modes in the calibrated mode information set can more accurately reflect the corresponding code rates and distortions of the candidate prediction modes, thereby enabling the selection of a target prediction mode more suitable for the target prediction unit from multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set. Then, a suitable target prediction mode may be used to perform prediction processing on the target prediction unit to obtain encoded data of the target image block, which can reduce the probability of distortion of the target image block after encoding to a certain extent. And since the embodiments of the present invention mainly reduce the distortion probability by modifying the mode decision process to select a suitable target prediction mode, the image compression quality can be effectively improved and the subjective quality of the target image block can be enhanced while basically not affecting the compression efficiency and encoding complexity.

[0247] Based on the description of the above embodiments of the video playing method, embodiments of the present invention also disclose a video playing device, and the video playing device may be a computer program (including program code) running in a video playing device. The video playing device can execute Figure 4 the method shown. Please refer to Figure 8 , and the video playing device may run the following units:

[0248] An obtaining unit 801, configured to obtain bitstream data of each frame of image in an image frame sequence corresponding to a target video, and the bitstream data of each frame of image includes encoded data of multiple image blocks; the encoded data of each image block in other frames of image in the image frame sequence except the first frame of image is encoded by using Figure 2 or Figure 3 the video encoding method shown;

[0249] A decoding unit 802, configured to decode the bitstream data of each frame of image to obtain each frame of image;

[0250] A display unit 803, configured to sequentially display each frame of image on a playing interface.

[0251] According to an embodiment of the present invention, Figure 4 each step involved in the method shown may be executed by each unit in the Figure 8 video playing device shown. For example, Figure 4 steps S401 - S403 shown in Figure 8It is executed by the acquisition unit 801, decoding unit 802, and display unit 803 shown in. According to another embodiment of the present invention, Figure 8 Each unit in the video playback encoding device shown can be respectively or wholly combined into one or several other units to form, or a certain (some) unit among them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, based on the video playback device, other units may also be included. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.

[0252] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in Figure 4 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), random access storage medium (RAM), read-only storage medium (ROM), etc., to construct a video playback device as shown in Figure 8 and to implement the video playback method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0253] Based on the description of the above video playback method embodiments and video playback device embodiments, the embodiments of the present invention also provide a video playback device. Please refer to Figure 9, the video playback device at least includes a processor 901, an input interface 902, an output interface 903, a computer storage medium 904, and a decoder 905. Among them, the computer storage medium 904 can be stored in the memory of the video playback device. The computer storage medium 904 is used to store a computer program, and the computer program includes program instructions. The processor 901 is used to execute the program instructions stored in the computer storage medium 904. The processor 901 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the video playback device, and is adapted to implement one or more instructions. Specifically, it is adapted to load and execute one or more instructions to implement the corresponding method flow or corresponding function. In one embodiment, the processor 901 described in the embodiments of the present invention can be used to perform a series of video playbacks on a target video, including: obtaining the bitstream data of each frame image in the image frame sequence corresponding to the target video, and the bitstream data of each frame image includes the encoded data of multiple image blocks; the encoded data of each image block in other frame images in the image frame sequence except the first frame image is encoded by the video encoding method shown in Figure 2 or Figure 3 ; decoding the bitstream data of each frame image to obtain each frame image; sequentially displaying each frame image on the playback interface, and so on.

[0254] The embodiments of the present invention also provide a computer storage medium (Memory). The computer storage medium is a memory device in the video playback device and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the video playback device and, of course, the extended storage medium supported by the video playback device. The computer storage medium provides a storage space, and the operating system of the video playback device is stored in this storage space. And, one or more instructions adapted to be loaded and executed by the processor 901 are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.

[0255] In one embodiment, one or more second instructions stored in the computer storage medium can be loaded and executed by the processor 901 to implement the corresponding steps of the method in the above embodiments of the video playback method; specifically, one or more second instructions in the computer storage medium are loaded and executed by the processor 901 as follows:

[0256] Obtain the bitstream data of each frame image in the image frame sequence corresponding to the target video. The bitstream data of each frame image includes the encoded data of multiple image blocks; the encoded data of each image block in other frame images except the first frame image in the image frame sequence is encoded by Figure 2 or Figure 3 the video encoding method shown;

[0257] Decode the bitstream data of each frame image to obtain each frame image;

[0258] Display each frame image in sequence on the playback interface.

[0259] In the embodiment of the present invention, first, the bitstream data of each frame image in the image frame sequence corresponding to the target video can be obtained. The bitstream data of each frame image includes the encoded data of multiple image blocks. Secondly, the bitstream data of each frame image can be decoded to obtain each frame image; and each frame image is displayed in sequence on the playback interface. Since the encoded data of each image block in other frame images except the first frame image in the image frame sequence corresponding to the target video is encoded by the above video encoding method; therefore, the probability of distortion of each image block can be effectively reduced, so that when each frame image is displayed on the playback interface, the probability of dirty points appearing in each frame image can be reduced to a certain extent, and the subjective quality of each frame image is improved.

[0260] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A video encoding method, characterized in that, Including: Obtain a target prediction unit in a target image block and a set of mode information of the target prediction unit, where the set of mode information includes multiple candidate prediction modes and mode costs of various candidate prediction modes; the multiple candidate prediction modes include intra prediction modes and inter prediction modes; Perform abnormal distortion point detection on the target prediction unit under at least one candidate prediction mode in the set of mode information, and obtain a detection result corresponding to the at least one candidate prediction mode; the at least one candidate prediction mode is each mode in the inter prediction modes. When using a reference prediction mode to predict the pixel values of each pixel point in the target prediction unit to obtain the predicted values of each pixel point, calculate the absolute value of the residual between the pixel value and the predicted value of each pixel point in the target prediction unit. When there is a pixel point in the target prediction unit whose absolute value of the residual is greater than a target threshold, the detection result corresponding to the reference prediction mode is used to indicate that there is an abnormal distortion point in the target prediction unit under the reference prediction mode. The reference prediction mode is any mode in the inter prediction modes; the target threshold is related to the candidate prediction mode. The target threshold corresponding to a candidate prediction mode that needs to transmit the residual information of the target image block is greater than the target threshold corresponding to a candidate prediction mode that does not need to transmit the residual information; According to the detection results corresponding to the at least one candidate prediction mode, calibrate the mode costs of the at least one candidate prediction mode in the set of mode information to obtain a calibrated set of mode information; Select a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated set of mode information; Perform prediction processing on the target prediction unit using the target prediction mode to obtain encoded data of the target image block.

2. The method according to claim 1, characterized in that, The inter prediction mode at least includes the following modes: a first prediction mode, a second prediction mode, and a third prediction mode; The first prediction mode refers to a mode that needs to transmit index information of a reference image block related to the target image block; The second prediction mode refers to a mode that needs to transmit the residual information of the target image block and index information of a reference image block related to the target image block; The third prediction mode refers to a mode that needs to transmit the residual information of the target image block, the motion vector data of the target image block, and index information of a reference image block related to the target image block.

3. The method according to claim 2, characterized in that, If there is no pixel point in the target prediction unit whose absolute value of the residual is greater than the target threshold, the detection result corresponding to the reference prediction mode indicates that there is no abnormal distortion point in the target prediction unit under the reference prediction mode.

4. The method according to claim 3, characterized in that, The target threshold is associated with the reference prediction mode; If the reference prediction mode is the first prediction mode in the inter prediction modes, the target threshold is equal to a first threshold; the first threshold is greater than zero and less than the maximum value of the pixel value range; If any of the candidate prediction modes is the second prediction mode or the third prediction mode among the inter-frame prediction modes, the target threshold is equal to the second threshold; the second threshold is greater than or equal to the first threshold and less than the maximum value of the pixel value range.

5. The method according to claim 2, characterized in that, Calibrating the mode costs of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain a calibrated mode information set includes: If the detection result corresponding to the reference prediction mode indicates that there are no abnormal distortion points in the target prediction unit under the reference prediction mode, keep the mode cost of the reference prediction mode in the mode information set unchanged to obtain a calibrated mode information set; the reference prediction mode is any mode among the inter-frame prediction modes; If the detection result corresponding to the reference prediction mode indicates that there are abnormal distortion points in the target prediction unit under the reference prediction mode, adjust the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode to obtain a calibrated mode information set.

6. The method according to claim 5, wherein Adjusting the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode includes: If the reference prediction mode is the second prediction mode or the third prediction mode, use a penalty factor to amplify the mode cost of the reference prediction mode to obtain the calibrated mode cost of the reference prediction mode.

7. The method according to claim 6, wherein Adjusting the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode further includes: If the reference prediction mode is the first prediction mode, obtain a preset cost; the preset cost is greater than the mode costs of the candidate prediction modes other than the first prediction mode in the calibrated mode information set and greater than the mode cost of the first prediction mode in the mode information set; In the mode information set, adjust the mode cost of the reference prediction mode to the preset cost.

8. The method according to claim 7, wherein Selecting a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set includes: Select the candidate prediction mode with the minimum mode cost in the calibrated mode information as the target prediction mode from the multiple candidate prediction modes.

9. The method according to claim 6, wherein Adjusting the mode cost of the reference prediction mode in the mode information set by using the cost adjustment strategy of the reference prediction mode further includes: If the reference prediction mode is the first prediction mode, keep the mode cost of the first prediction mode in the mode information set unchanged; Add a disabled flag to the first prediction mode, and the disabled flag indicates that it is prohibited to use the first prediction mode to perform prediction processing on the target prediction unit.

10. The method according to claim 9, characterized in that, Selecting a target prediction mode from the multiple candidate prediction modes according to the mode costs of the candidate prediction modes in the calibrated mode information set includes: If the candidate prediction mode with the minimum mode cost in the calibrated mode information set is not the first prediction mode, or if the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode does not have the disable flag, then use the candidate prediction mode with the minimum mode cost as the target prediction mode; If the candidate prediction mode with the minimum mode cost is the first prediction mode and the first prediction mode has the disable flag, then select the candidate prediction mode with the second smallest mode cost in the calibrated mode information set as the target prediction mode.

11. The method according to any one of claims 2 to 10, characterized in that, The method further includes: Performing complexity analysis on the target prediction unit to obtain the prediction complexity of the target prediction unit; If it is determined according to the prediction complexity that the target prediction unit meets the preset conditions, then use the intra prediction mode to perform prediction processing on the target prediction unit to obtain the encoded data of the target image block; wherein, the preset conditions include: the prediction complexity is less than or equal to a complexity threshold, and there are abnormal distortion points in at least one mode of the target prediction unit in the inter prediction mode; If it is determined according to the prediction complexity that the target prediction unit does not meet the preset conditions, then perform the step of calibrating the mode costs of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode to obtain the calibrated mode information set.

12. A video playing method, characterized in that, Includes: Obtaining the bitstream data of each frame of image in the image frame sequence corresponding to the target video, and the bitstream data of each frame of image includes the encoded data of multiple image blocks; the encoded data of each image block in other frames of the image frame sequence except the first frame image is encoded by using the video encoding method according to any one of claims 1-11; Decoding the bitstream data of each frame of image to obtain each frame of image; Sequentially displaying each frame of image on the playback interface.

13. A video encoding device, characterized in that, Includes: An obtaining unit, configured to obtain a target prediction unit in a target image block and a mode information set of the target prediction unit, where the mode information set includes multiple candidate prediction modes and the mode costs of various candidate prediction modes; The multiple candidate prediction modes include an intra prediction mode and an inter prediction mode; A coding unit is configured to perform abnormal distortion point detection on a target prediction unit under at least one candidate prediction mode in the mode information set, so as to obtain detection results corresponding to the at least one candidate prediction mode; the at least one candidate prediction mode is each mode in the inter-frame prediction mode. When pixel value prediction is performed on each pixel point in the target prediction unit by using a reference prediction mode to obtain predicted values of each pixel point, the absolute value of the residual between the pixel value and the predicted value of each pixel point in the target prediction unit is calculated. When there is a pixel point in the target prediction unit whose absolute value of the residual is greater than a target threshold, the detection result corresponding to the reference prediction mode is used to indicate that there is an abnormal distortion point in the target prediction unit under the reference prediction mode, and the reference prediction mode is any mode in the inter-frame prediction mode; the target threshold is related to the candidate prediction mode, and the target threshold corresponding to the candidate prediction mode that needs to transmit the residual information of the target image block is greater than the target threshold corresponding to the candidate prediction mode that does not need to transmit the residual information. The coding unit is further configured to calibrate the mode cost of the at least one candidate prediction mode in the mode information set according to the detection results corresponding to the at least one candidate prediction mode, so as to obtain a calibrated mode information set. The coding unit is further configured to select a target prediction mode from the multiple candidate prediction modes according to the mode cost of each candidate prediction mode in the calibrated mode information set. The coding unit is further configured to perform prediction processing on the target prediction unit by using the target prediction mode, so as to obtain the encoded data of the target image block.

14. A video playback device, characterized in that, It includes: An acquisition unit is configured to acquire bitstream data of each frame image in an image frame sequence corresponding to a target video, and the bitstream data of each frame image includes encoded data of multiple image blocks; the encoded data of each image block in other frame images except the first frame image in the image frame sequence is encoded by using the video coding method according to any one of claims 1-11. A decoding unit is configured to decode the bitstream data of each frame image to obtain each frame image. A display unit is configured to sequentially display each frame image on a playback interface.

15. A computer storage medium, characterized in that, The computer storage medium stores one or more first instructions, and the one or more first instructions are suitable for being loaded and executed by a processor to execute the video coding method according to any one of claims 1-11; alternatively, the computer storage medium stores one or more second instructions, and the one or more second instructions are suitable for being loaded and executed by a processor to execute the video playback method according to claim 12.

Citation Information

Patent Citations

  • Fast video encoding method with block partitioning

    US20160261870A1