Video encoding method, apparatus, device, and computer-readable storage medium
By dividing the encoding blocks and judging the texture complexity of the video frames, skipping the intra-frame mode, and using the palette mode to encode the encoding blocks, the problem of high computational complexity in efficient video encoding is solved, and the encoding speed and efficiency are improved.
Patent Information
- Application Number
- CN202010818441.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-08-14
AI Technical Summary
During the efficient video encoding process, the parallel computing complexity of intra-block copying and palette encoding tools is high, resulting in increased computing burden and it is difficult to improve the efficiency of screen content video compression.
By dividing the coded blocks of the video frames to be encoded, the texture complexity of each coded block is determined, and when the mode skipping condition is met, the intra mode is skipped, and the coded blocks are encoded using the palette mode to optimize the encoding process.
While ensuring the quality of video frame encoding, the video frame encoding speed is improved, the calculation complexity is reduced, and the encoding efficiency is improved.
Smart Images

Figure CN114079769B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video coding technologies, and in particular, to a video coding method, apparatus, device, and computer-readable storage medium. Background Art
[0002] During the encoding process of High Efficiency Video Coding (HEVC), there are many mode decision processes. Rate Distortion Optimization (RDO) is the most important video coding decision technology in the current industry, which can effectively select the optimal coding mode for each link of video coding, thereby improving the compression efficiency.
[0003] In HEVC Screen Content Coding (SCC), through Intra Block Copy (IBC) and Palette (PLT) coding tools, the compression efficiency of screen content videos can be greatly improved. However, in its main calculation process, the vast majority of calculations are mainly serial calculations. Since it is difficult to perform large-scale parallel calculations on the PLT module through Single Instruction Multiple Data (SIMD), directly placing the PLT module in the existing HEVC420 encoder will cause a significant increase in computational complexity. Summary of the Invention
[0004] Embodiments of the present application provide a video coding method, apparatus, device, and computer-readable storage medium, which can improve the encoding speed of video frames while ensuring the encoding quality of video frames.
[0005] The technical solution of the embodiments of the present application is implemented as follows:
[0006] Embodiments of the present application provide a video coding method, including:
[0007] Performing encoding block partitioning on a target video frame to be encoded to obtain a plurality of encoding blocks corresponding to the target video frame;
[0008] Respectively determining the texture complexity of each of the encoding blocks;
[0009] Based on the texture complexity of each of the encoding blocks, when it is determined that the corresponding encoding block satisfies the mode skip condition, for the intra mode and the palette mode executed sequentially, skipping the intra mode, and
[0010] Encoding the encoding block through the palette mode to implement encoding of the target video frame when the encoding of the plurality of encoding blocks is completed.
[0011] An embodiment of the present application provides a video encoding device, which is characterized in that the device includes:
[0012] A partitioning module, configured to perform encoding block partitioning on a target video frame to be encoded, and obtain a plurality of encoding blocks corresponding to the target video frame;
[0013] A determining module, configured to respectively determine the texture complexity of each of the encoding blocks;
[0014] An encoding module, configured to, when it is determined that the corresponding encoding block satisfies the mode skip condition based on the texture complexity of each encoding block, for the intra mode and the palette mode that are sequentially executed, skip the intra mode, and
[0015] encode the encoding block through the palette mode, so as to implement encoding of the target video frame when encoding of the plurality of encoding blocks is completed.
[0016] In the above solution, the determining module is further configured to respectively obtain the gradient value of each encoding block, and use the gradient value as the texture complexity of the corresponding encoding block;
[0017] respectively compare the gradient value of each encoding block with a gradient threshold to obtain a comparison result;
[0018] When the comparison result indicates that the gradient value reaches the gradient threshold, it is determined that the encoding block corresponding to the gradient value satisfies the mode skip condition.
[0019] In the above solution, the determining module is further configured to respectively perform the following operations on each encoding block:
[0020] partition the encoding block into a plurality of encoding sub-blocks;
[0021] obtain the gradient values of the plurality of encoding sub-blocks corresponding to the encoding block;
[0022] use the sum of the gradient values of the plurality of encoding sub-blocks corresponding to the encoding block as the gradient value of the encoding block.
[0023] In the above solution, the determining module is further configured to, when the target encoding frame satisfies a first condition, encode each encoding sub-block through the intra block copy (IBC) mode, and
[0024] determine the gradient value of the corresponding encoding sub-block during the process of encoding the encoding sub-block through the IBC mode.
[0025] In the above solution, the determining module is further configured to obtain the type of the target encoding frame;
[0026] When the type of the target encoded frame indicates that the target encoded frame is an intra-coded frame, it is determined that the target encoded frame satisfies the first condition;
[0027] When the type of the target encoded frame indicates that the target encoded frame is a forward predicted frame, obtain the current coding mode of the target coding block, and
[0028] when the current coding mode of the coding block is not the skip mode and the residual quantization value obtained by encoding the coding block using the inter-frame mode is not zero, it is determined that the target encoded frame satisfies the first condition.
[0029] In the above solution, the determining module is further configured to perform the following processing for each coding sub-block:
[0030] Determine the horizontal gradient and vertical gradient of each pixel in the coding sub-block;
[0031] Obtain the gradient average value of the horizontal gradient and vertical gradient corresponding to each pixel in the coding sub-block;
[0032] Use the sum of the gradient average values corresponding to each pixel included in the coding sub-block as the gradient value of the coding sub-block.
[0033] In the above solution, the encoding module is further configured to obtain the type of the target encoded frame;
[0034] When the type of the target encoded frame indicates that the target encoded frame is a forward predicted frame, obtain the current coding mode of the coding block; when the current coding mode of the coding block is the skip mode, end the processing for the coding block; or,
[0035] When the current coding mode of the coding block is the inter-frame mode and the residual quantization value obtained by encoding the coding block using the inter-frame mode is zero, end the processing for the coding block.
[0036] In the above solution, the encoding module is further configured to obtain the first rate-distortion cost when encoding the coding block using the skip mode;
[0037] When the first rate-distortion cost is not greater than the rate-distortion threshold, determine that the current coding mode of the coding block is the skip mode.
[0038] In the above solution, the encoding module is further configured to obtain the predicted value of the coding block and the actual value of the coding block when encoding the coding block using the inter-frame mode;
[0039] Compare the actual value of the coding block with the predicted value to obtain the residual between the actual value and the predicted value of the coding block;
[0040] Encode the residual to obtain a residual quantization value.
[0041] In the above solution, the encoding module is further configured to generate a color palette, where the color palette is a table including at least one color value in the encoding block;
[0042] Respectively obtain the indexes of each pixel in the encoding block in the table, and use the indexes as the encoding result of the encoding block.
[0043] In the above solution, the encoding module is further configured to generate a color palette, where the color palette is a table including at least one color value in the encoding block;
[0044] Sample the pixels in the encoding block according to a preset pixel sampling interval to obtain a plurality of sampled pixels;
[0045] Respectively obtain the indexes of each of the sampled pixels in the encoding block in the table, and use the indexes as the encoding result of the encoding block.
[0046] In the above solution, when the encoding module is further configured to determine that the corresponding encoding block does not meet the mode skip condition based on the texture complexity, encode the encoding block through an intra-frame mode;
[0047] Obtain a second rate-distortion cost when encoding the encoding block through the intra-frame mode;
[0048] When the second rate-distortion cost is greater than the rate-distortion threshold, encode the encoding block through the color palette mode.
[0049] In the above solution, the encoding module is further configured to encode the encoding block through the vertical mode, horizontal mode, direct current mode, and plane mode included in the intra-frame mode respectively;
[0050] The obtaining of the second rate-distortion cost when encoding the encoding block through the intra-frame mode includes:
[0051] Respectively obtain the sum of absolute transform errors when encoding through the vertical mode, horizontal mode, direct current mode, and plane mode;
[0052] According to the obtained sum of absolute transform errors, use the mode corresponding to the minimum value in the obtained sum of absolute transform errors as the optimal intra-frame mode;
[0053] Obtain the second rate-distortion cost when encoding the encoding block through the optimal intra-frame mode.
[0054] An embodiment of the present application provides a computer device, including:
[0055] A memory for storing executable instructions;
[0056] A processor, when executing the executable instructions stored in the memory, implements the video encoding method provided by the embodiments of the present application.
[0057] The embodiments of the present application provide a computer-readable storage medium storing executable instructions, which are used to cause a processor to implement the video encoding method provided by the embodiments of the present application when executed.
[0058] The embodiments of the present application have the following beneficial effects:
[0059] In the present application, the texture complexity of each encoding block is determined respectively; when it is determined that the corresponding encoding block satisfies the mode skipping condition based on the texture complexity of each encoding block, for the intra-mode and palette mode executed in sequence, the intra-mode is skipped, and the encoding block is encoded through the palette mode, so as to implement the encoding of the target video frame when the encoding of the multiple encoding blocks is completed; in this way, the intra-mode can be skipped when the texture complexity satisfies the mode skipping condition, so as to improve the encoding speed of the video frame while ensuring the encoding quality of the video frame. Description of the Drawings
[0060] Figure 1 is a schematic structural diagram of an encoding system 100 for a video frame provided by the embodiments of the present application;
[0061] Figure 2 is a schematic structural diagram of a computer device provided by the embodiments of the present application;
[0062] Figure 3 is a schematic flowchart of a video frame encoding method provided by the embodiments of the present invention;
[0063] Figure 4 is a schematic diagram of encoding block division provided by the embodiments of the present application;
[0064] Figure 5 is a schematic diagram of executing the intra-block copy mode provided by the embodiments of the present application;
[0065] Figure 6 is a schematic diagram of executing the palette mode provided by the embodiments of the present application;
[0066] Figure 7 is a schematic flowchart of a video frame encoding method provided by the embodiments of the present application;
[0067] Figure 8 is a schematic diagram of the composition structure of a video frame encoding device provided by the embodiments of the present application. Detailed Embodiments
[0068] To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0069] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0070] In the following description, the terms "first", "second", and "third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0072] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0073] 1) Coding block, which is a basic concept in video coding technology. By dividing a video frame into blocks of different sizes, different compression strategies are selected for different positions.
[0074] 2) Texture, which is a visual feature reflecting the homogeneous phenomenon in an image. It reflects the surface structure organization arrangement attribute of an object's surface with slow changes or periodic changes.
[0075] 3) I-frame, that is, an intra-coded frame, represents a key frame. Usually, it is the first frame of each group of pictures (GOP, Group of Pictures). After moderate compression, it serves as a random access reference point and can be regarded as an image.
[0076] 4) P-frame, a forward prediction coded frame, represents the difference between this frame and a previous I-frame (or P-frame). When decoding, the previously cached picture needs to be superimposed with the difference defined in this frame to generate the final picture.
[0077] See Figure 1 , Figure 1It is a schematic architecture diagram of the video encoding system 100 provided by an embodiment of the present application. To support an exemplary application, the terminal 400 (exemplarily showing the terminal 400-1 and the terminal 400-2) is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0078] The terminal 400-1 is used to collect video in real time, and use the currently collected video frame as the target video frame to be encoded; perform encoding block division on the target video frame to be encoded to obtain a plurality of encoding blocks corresponding to the target video frame; respectively determine the texture complexity of each of the encoding blocks; when it is determined that the corresponding encoding block satisfies the mode skip condition based on the texture complexity of each of the encoding blocks, for the intra-frame mode and the palette mode executed in sequence, skip the intra-frame mode, and encode the encoding block through the palette mode, so as to complete the encoding of the target video frame when the encoding of the plurality of encoding blocks is completed; send the encoded target video frame to the server 200.
[0079] The server 200 is used to send the encoded target video frame to 400-2;
[0080] The terminal 400-2 is used to decode the encoded target video frame and display the decoded video frame.
[0081] In some embodiments, the server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and are not limited in the embodiments of the present invention.
[0082] The following will make a detailed description of the hardware structure of the computer device for the video encoding method provided by the embodiment of the present invention. The computer device includes, but is not limited to, a server or a terminal. See Figure 2 , Figure 2 It is a schematic structure diagram of the computer device provided by an embodiment of the present application. Figure 2The computer device shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to achieve connection and communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0083] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0084] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0085] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.
[0086] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0087] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0088] An operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0089] A network communication module 452 for reaching other computing devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0090] A presentation module 453 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.);
[0091] An input processing module 454 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432.
[0092] In some embodiments, the video encoding device provided in the embodiments of the present application can be implemented in software. Figure 2 Shown is a video encoding device 455 stored in the memory 450, which can be software in the form of a program and a plug-in, etc., including the following software modules: a partitioning module 4551, a determination module 4552, and an encoding module 4553. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented.
[0093] The functions of each module will be described below.
[0094] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video encoding method provided in the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0095] Based on the above description of the video encoding system and computer device in the embodiments of the present invention, the video encoding method provided in the embodiments of the present invention will be described below. Refer to Figure 3 , Figure 3It is a schematic flowchart of the video encoding method provided by an embodiment of the present invention; in some embodiments, the video encoding method can be implemented independently by a server or a terminal, or jointly implemented by a server and a terminal. Taking the terminal implementation as an example, the video encoding method provided by an embodiment of the present invention includes:
[0096] Step 301: The terminal performs encoding block partitioning on the target video frame to be encoded, obtaining a plurality of encoding blocks corresponding to the target video frame.
[0097] In actual implementation, in video encoding, a video frame is usually divided into several encoding blocks. If the encoding blocks are arranged in a sheet form, the video encoding algorithm encodes each encoding block one by one and organizes them into a continuous video bitstream.
[0098] In some embodiments, a client is set on the terminal, such as a desktop sharing client, a network conference client, an instant messaging client, etc. The terminal can collect video data through the client to encode the collected video data.
[0099] Exemplarily, taking the application to a desktop sharing client as an example, the screen content image is captured in real time from the image display unit of the terminal, and the currently obtained screen content image is used as the target video frame to be encoded, so as to encode the target video frame in real time, and then transmit the encoded target video frame to the terminal for desktop sharing with the current terminal, so that the terminal for desktop sharing with the current terminal can obtain it.
[0100] Exemplarily, taking the application to a network conference client as an example, here, the first terminal and the second terminal hold a video conference through the network conference client. During the video conference, each of the first terminal and the second terminal can collect video images in real time through the camera, and use the currently obtained screen content image as the target video frame to be encoded, so as to encode the target video frame in real time, and then transmit the encoded target video frame to the other terminal. Each of the first terminal and the second terminal can also receive the encoded target video frame transmitted by the other terminal, and can decode the encoded target video frame to restore the target video frame, and can display the target video frame on the accessible display device according to the restored target video frame.
[0101] Step 302: Determine the texture complexity of each encoding block respectively.
[0102] Here, texture is a visual feature reflecting the homogeneous phenomenon in an image, which reflects the surface structure organization arrangement attribute of the object surface with slow changes or periodic changes.
[0103] In actual implementation, the texture complexity can be described by the second moment (variance) of the grayscale histogram of the coding block, or by the gradient of the coding block, or by other means.
[0104] In some embodiments, before determining the texture complexity of each coding block respectively, the method further includes: obtaining the type of the target coding frame; when the type of the target coding frame indicates that the target coding frame is a forward prediction frame, obtaining the current coding mode of the coding block; when the current coding mode of the coding block is the skip mode, ending the processing for the coding block; or, when the current coding mode of the coding block is the inter mode and the residual quantization value obtained by encoding the coding block in the inter mode is zero, ending the processing for the coding block.
[0105] In actual implementation, obtain the type of the target coding frame, such as an I-frame (intra-coded frame), a P-frame (forward prediction coded frame). When the target coding frame is a P-frame, it is necessary to determine whether the current coding mode of the coding block is the skip mode, and whether the residual quantization value obtained in the inter mode is zero; when the current coding mode of the coding block is the skip mode, or the residual quantization value obtained in the inter mode is zero, end the processing for the coding block; otherwise, obtain the texture complexity of the coding block.
[0106] Here, since there is no skip mode and inter mode in the coding mode for I-frames, when the target coding frame is an I-frame, directly obtain the texture complexity of the coding block.
[0107] In some embodiments, the terminal can determine that the current coding mode is the skip mode in the following way: obtain the first rate-distortion cost when encoding the coding block in the skip mode; when the first rate-distortion cost is not greater than the rate-distortion threshold, determine that the current coding mode of the coding block is the skip mode.
[0108] In actual implementation, encode the coding block in the skip mode to obtain the first rate-distortion cost when encoding the coding block in the skip mode. If the first rate-distortion cost is not greater than the rate-distortion threshold, it means that the effect of encoding the coding block in the skip mode meets the standard. Based on this, take the skip mode as the optimal coding mode for the corresponding coding block, and there is no need to encode the coding block in other modes.
[0109] In some embodiments, the residual quantization value when encoding a coding block in an inter-frame mode can be obtained in the following manner: obtain the predicted value and the actual value of the coding block when encoding the coding block in the inter-frame mode; compare the actual value of the coding block with the predicted value to obtain the residual between the actual value and the predicted value of the coding block; encode the residual to obtain the residual quantization value.
[0110] In actual implementation, the inter-frame (Inter) mode refers to using the temporal correlation of the video, and predicting the pixels of the target video frame using the pixels of adjacent encoded video frames, so as to effectively remove the temporal redundancy of the video. Since the video sequence usually includes strong temporal correlation, the predicted residual value is close to 0. Using the residual signal as the input of the subsequent module for transformation, quantization, scanning, and entropy coding can achieve efficient compression of the video signal.
[0111] Here, the residual quantization value can be the coded block pattern (CBP) information. The CBP information is a syntax element used to reflect the residual situation in the coding of the coding block. In this embodiment, the coding block size is 16x16, and the CBP information has a total of 6 bits. The first 2 bits represent the UV components; the last 4 bits are the Y components, representing 4 8x8 coding sub-blocks within the coding block respectively. If any bit is 0, it indicates that all the transform coefficient levels (the values in the matrix after transforming and quantizing the pixel residuals, hereinafter collectively referred to as levels) in the corresponding 8x8 block are all 0, otherwise it indicates that the transform coefficient levels in the corresponding 8x8 block are not all 0.
[0112] Step 303: When it is determined that the corresponding coding block meets the mode skip condition based on the texture complexity of each coding block, for the intra-frame mode and the palette mode that are sequentially executed, skip the intra-frame mode and encode the coding block through the palette mode, so as to complete the encoding of the target video frame when the encoding of multiple coding blocks is completed.
[0113] The inventors of the present application found in the process of implementing the present invention that when the texture complexity of the coding block is low, that is, the texture information of the coding block is relatively simple, it is necessary to calculate the intra-frame mode (standard Intra mode); while when the texture complexity of the coding block is high, the rate-distortion cost of the palette (PLT) mode is usually better than that of the standard Intra mode. Therefore, the standard Intra mode can be skipped and the PLT mode can be directly executed.
[0114] In some embodiments, the terminal may determine the texture complexity of each coding block in the following manner: obtain the gradient value of each coding block respectively, and use the gradient value as the texture complexity of the corresponding coding block; correspondingly, the following manner may be adopted to determine that the coding block corresponding to the gradient value satisfies the mode skip condition: compare the gradient value of each coding block with the gradient threshold respectively to obtain a comparison result; when the comparison result indicates that the gradient value reaches the gradient threshold, determine that the coding block corresponding to the gradient value satisfies the mode skip condition.
[0115] In actual implementation, the complexity may be described by the gradient value of the coding block, and a gradient threshold is preset. By comparing the gradient value with the preset gradient threshold, it is determined whether the gradient value reaches the preset gradient threshold. If so, the standard Intra mode is skipped and the PLT mode is directly executed; otherwise, it is determined that the standard Intra mode needs to be executed.
[0116] Here, the smaller the set gradient threshold is, the lower the probability of needing to execute the standard Intra mode. When there is no need to execute the Intra mode, the computational complexity can be reduced and the encoding speed can be improved.
[0117] In some embodiments, the gradient value of each coding block may be obtained in the following manner:
[0118] The following operations are respectively performed for each coding block: divide the coding block into multiple coding sub-blocks; obtain the gradient values of the multiple coding sub-blocks corresponding to the coding block; and use the sum of the gradient values of the multiple coding sub-blocks corresponding to the coding block as the gradient value of the coding block.
[0119] In actual implementation, the coding block may be divided into multiple coding sub-blocks of the same size to obtain the gradient values of the respective coding sub-blocks; after obtaining the gradient values of the respective coding sub-blocks, sum the gradient values of the respective coding sub-blocks to obtain the gradient value of the coding block.
[0120] For example, Figure 4 is a schematic diagram of the coding block division provided by the embodiments of the present application. Refer to Figure 4 , the target video frame may be divided into 4 16×16 coding blocks, and each 16×16 coding block may be divided into 4 8×8 coding sub-blocks.
[0121] In some embodiments, the gradient value of each coding sub-block may be obtained in the following manner: when the target coding frame meets the first condition, each coding sub-block is encoded through the Intra Block Copy (IBC) mode, and during the process of encoding the coding sub-block through the IBC mode, the gradient value of the corresponding coding sub-block is determined.
[0122] Here, when the IBC mode is adopted, the encoding sub-block searches for the most similar block among the already reconstructed encoding sub-blocks in the target video frame as the prediction sub-block, and calculates the block vector (BV, Block Bector), where the block vector indicates the positional relationship between the current block and the prediction sub-block.
[0123] In actual implementation, Figure 5 is a schematic diagram of the execution of the intra-frame block copy mode provided by the embodiments of the present application. Refer to Figure 5 , search for the prediction sub-block 502 similar to the encoding sub-block 501 in the reconstructed part of the target video frame, and use the vector from the position where the encoding sub-block 501 is located to the position where the prediction sub-block 502 is located as the block vector 503. Among them, the current prediction sub-block and the current encoding sub-block must be in the same slice and the same tile to avoid the dependency relationships of different slices or tiles and affect the parallel processing capabilities of slices and tiles; and, the prediction sub-block needs to be restricted within Figure 5 the gray area in to avoid affecting the wavefront parallel processing capabilities.
[0124] In actual implementation, a hash-based search method can be used to find the optimal block vector. For example, for an 8x8 encoding sub-block, each node in the Hash table represents the position of a BV candidate in the target video frame, and only the BV candidates with the same Hash value as the current block are checked, and the length of the Hash value is 16 bits.
[0125] Here, the 16-bit Hash value can be calculated by the following formula:
[0126] H = msb(dc0,3) << 13 + msb(dc1,1) << 10 + msb(dc2,3) << 7 + msb(dc3,3) << 4 + msb(grad BLK ,4) (1),
[0127] where, msb(X, n) represents the highest n significant bits of X, dc0, dc1, dc2, and dc3 respectively represent the DC values of the 4 4x4 sub-blocks of the 8x8 encoding sub-block, and grad BLK represents the gradient value of the 8x8 encoding sub-block.
[0128] That is, during the execution of IBC, the gradient value of the encoding sub-block needs to be calculated. Based on this, the gradient value of the encoding block can be obtained using the gradient value of this encoding sub-block to avoid recalculating the gradient value of the encoding sub-block. In this way, additional computational effort can be avoided and the encoding efficiency can be improved.
[0129] In some embodiments, the target encoded frame can be determined to meet the first condition in the following manner: obtain the type of the target encoded frame; when the type of the target encoded frame indicates that the target encoded frame is an intra-coded frame, determine that the target encoded frame meets the first condition; when the type of the target encoded frame indicates that the target encoded frame is a forward-predicted frame, obtain the current coding mode of the target coding block, and when the current coding mode of the coding block is not the skip mode and the residual quantization value obtained by encoding the coding block using the inter-frame mode is not zero, determine that the target encoded frame meets the first condition.
[0130] In actual implementation, obtain the type of the target encoded frame, such as an I-frame (intra-coded frame) or a P-frame (forward-predicted coded frame). When the target encoded frame is an I-frame, determine that the target encoded frame meets the first condition and directly encode each coding sub-block through the IBC mode; when the target encoded frame is a P-frame, it is necessary to determine whether the current coding mode of the coding block is the skip mode and whether the residual quantization value obtained in the inter-frame (Inter) mode is zero; when the current coding mode of the coding block is not the skip mode and the residual quantization value obtained in the Inter mode is zero, determine that the target encoded frame meets the first condition and encode each coding sub-block through the IBC mode.
[0131] In some embodiments, the gradient values of multiple coding sub-blocks corresponding to a coding block can be obtained in the following manner:
[0132] For each coding sub-block, perform the following processing: determine the horizontal gradient and vertical gradient of each pixel in the coding sub-block; obtain the gradient average value of the horizontal gradient and vertical gradient corresponding to each pixel in the coding sub-block; take the sum of the gradient average values corresponding to each pixel included in the coding sub-block as the gradient value of the coding sub-block.
[0133] In actual implementation, for all pixels in the coding sub-block except for the first row and the first column, calculate the horizontal gradient and vertical gradient corresponding to the pixel, and then take the average value of the horizontal gradient and vertical gradient as the gradient value of the pixel; sum up the gradient values corresponding to all pixels in the coding sub-block to obtain the gradient value of the coding sub-block.
[0134] Here, for all pixels in the coding sub-block except for the first row and the first column, the absolute value of the pixel difference between the pixel and the adjacent left pixel can be obtained, and the absolute value of the pixel difference between the pixel and the adjacent left pixel is used as the horizontal gradient corresponding to the pixel; and the absolute value of the pixel difference between the pixel and the adjacent upper pixel can be obtained, and the absolute value of the pixel difference between the pixel and the adjacent upper pixel is used as the vertical gradient corresponding to the pixel.
[0135] For example, for an 8x8 coding sub-block, the gradient value of the coding sub-block can be obtained through formula (2) and formula (3):
[0136] g (i,j) = (|p (i,j-1) - p (i,j) | + |p (i-1,j) - p (i,j) |) / 2 (2)
[0137]
[0138] where g (i,j) represents the gradient value corresponding to the pixel at the i-th row and j-th column in the coding sub-block, p (i,j) represents the pixel at the i-th row and j-th column in the coding sub-block, and g represents the gradient value of the coding sub-block.
[0139] For a 16x16 coding block, after calculating the gradient values of the 4 8x8 coding sub-blocks included in the 16x16 coding block, the gradient value of the 16x16 coding block can be obtained through formula (4):
[0140]
[0141] where g 16×16 represents the gradient value of the 16x16 coding block, g sub8×8_1 、g sub8×8_2 、g sub8×8_3 represent the gradient values of the four coding sub-blocks of the 16x16 coding block.
[0142] In some embodiments, the coding block can be encoded in the PLT mode in the following manner: generating a palette, where the palette is a table containing at least one color value in the coding block; respectively obtaining the indexes of each pixel in the coding block in the table, and using the indexes as the encoding result of the coding block.
[0143] Here, the palette is a table containing at least one color value in the coding block, and each entry in the table contains three components. For a 4:2:0 or 4:2:2 color format, if the current position pixel has no chrominance component, only the first component is used to reconstruct the pixel. There is a special index, escape index, in the color palette for pixels that have no corresponding value in the table. In this case, in addition to transmitting the index es cape index of the pixel in the bitstream, the quantization value of the pixel also needs to be transmitted.
[0144] Figure 6 is a schematic diagram of implementing the palette mode provided by the embodiments of the present application. See Figure 6, where the size of the coding block is 4×4 and the size of the palette is 4. According to the color of the pixel, the index of the pixel in the palette is obtained. For example, the index of the pixel in the first row and the first column is 2; if there is no value corresponding to a certain pixel in the palette, its index is the index corresponding to escape. For example, the pixel in the second row and the second column corresponds to the escape index 4.
[0145] In some embodiments, the coding block can be encoded in the PLT mode in the following manner: generate a palette, where the palette is a table containing at least one color value in the coding block; sample the pixels in the coding block according to a preset pixel sampling interval to obtain a plurality of sampled pixels; respectively obtain the indexes of the sampled pixels in the coding block in the table, and use the indexes as the encoding result of the coding block.
[0146] Here, the selection of PLT and the actual PLT entropy coding complexity can be simplified by reducing the sampling accuracy of the PLT mode, thereby reducing the computational complexity of the PLT mode.
[0147] In practical applications, if the Intra mode is skipped and the PLT mode is directly executed, since the Intra mode is not executed, the computational complexity is reduced. Here, the accuracy of the PLT mode can be appropriately increased. For example, the sampling accuracy of the PLT mode can be increased from 3x3 to 3x2 pixel regions. In this way, the quality and speed of encoding can be improved simultaneously.
[0148] In some embodiments, when the terminal further determines that the corresponding coding block does not meet the mode skip condition based on the texture complexity, the coding block is encoded in the Intra mode; obtain the second rate distortion cost when the coding block is encoded in the Intra mode; when the second rate distortion cost is greater than the rate distortion threshold, the coding block is encoded in the palette mode.
[0149] In actual implementation, the coding block is encoded in the Intra mode to obtain an encoding result; then, according to the encoding result of the Intra mode, the second rate distortion cost when the coding block is encoded in the Intra mode is calculated, and further it is determined whether the coding block needs to be encoded in the palette mode. If the second rate distortion cost is greater than the rate distortion threshold, the coding block is encoded in the palette mode; if the second rate distortion cost is less than the rate distortion threshold, it means that the Intra mode is the optimal coding mode. Therefore, there is no need to encode the coding block in the palette mode anymore.
[0150] It should be noted that the coding block is coded in the palette mode only when the second rate-distortion cost is greater than the rate-distortion threshold. In this way, the computational complexity can be reduced. That is, when the second rate-distortion cost is less than the rate-distortion threshold, even if the result of coding the coding block in the palette mode may be better than the Intra mode, since the coding quality already meets the requirements, there is no need to increase unnecessary calculations.
[0151] In some embodiments, the terminal can code the coding block in the following manner: respectively through the vertical mode, horizontal mode, DC mode, and plane mode included in the intra mode.
[0152] Correspondingly, the second rate-distortion cost when coding the coding block in the intra mode can be obtained in the following manner: respectively obtain the sum of absolute transform errors when coding through the vertical mode, horizontal mode, DC mode, and plane mode; according to the obtained sum of absolute transform errors, use the mode corresponding to the minimum value in the obtained sum of absolute transform errors as the optimal intra mode; obtain the second rate-distortion cost when coding the coding block through the optimal intra mode.
[0153] Here, the sum of absolute transformed differences (SATD) refers to the sum of absolute values after hadamard transformation.
[0154] In actual implementation, the intra mode includes 4 modes, namely the vertical mode, horizontal mode, DC mode, and plane mode. The coding block is coded through these four modes respectively to obtain the predicted values of the coding blocks corresponding to the 4 modes; for each mode, calculate the difference between the predicted value and the actual value of the coding block to obtain the residual; here, the residual is a matrix, perform hadamard transformation on this matrix, and then sum the absolute values of the matrix elements to obtain the sum of absolute transform errors corresponding to each mode.
[0155] Here, the smaller the value of the sum of absolute transform errors, the better the corresponding mode. Based on this, the sum of absolute transform errors corresponding to the four modes can be compared to select the mode corresponding to the minimum value as the optimal intra mode.
[0156] In this application, the texture complexity of each coding block is determined separately; based on the texture complexity of each coding block, when it is determined that the corresponding coding block satisfies the mode skip condition, for the intra mode and palette mode executed sequentially, the intra mode is skipped, and the coding block is encoded through the palette mode, so as to implement the encoding of the target video frame when the encoding of multiple coding blocks is completed; in this way, when the texture complexity satisfies the mode skip condition, the intra mode can be skipped to improve the encoding speed of the video frame while ensuring the encoding quality of the video frame.
[0157] Next, an exemplary application of the embodiments of this application in a practical application scenario will be described.
[0158] In the related art, since it is difficult to perform large-scale parallel computing on the PLT module using SIMD technology, directly putting the PLT module into the existing HEVC420 encoder will bring a huge increase in complexity. Therefore, it is necessary to introduce the PLT module locally. First, the PLT mode is only used for 16x16 and 32x32 blocks, the 8x8 block uses the IBC mode, Intra4x4 and transformskip. Then, the palette sampling accuracy of the PLT module is reduced from per-pixel to sampling one pixel every 3x3 pixels, which can simplify the palette selection process and reduce the actual PLT entropy coding complexity, and can greatly reduce the computational complexity of the PLT module. At the same time, some coding gains of the PLT module are obtained under the condition of increasing a certain amount of computational complexity. However, after deep optimization of the HEVC encoder and the IBC mode, the computational amount of the originally optimized PLT mode becomes relatively large, and further deep optimization is required to improve the speed of the SCC encoder, especially in the artificial intelligence (AI) mode.
[0159] Based on this, a video coding method according to an embodiment of this application is proposed. Figure 7 It is a schematic flowchart of the video coding method provided by the embodiment of this application. Refer to Figure 7 The video coding method provided by the embodiment of this application includes:
[0160] Step 701: Divide the target video frame into multiple 16x16 coding blocks.
[0161] Here, the target video frame is divided into multiple 16x16 coding blocks, and each coding block is encoded one by one in units of coding blocks.
[0162] Step 702: Obtain the type of the target video frame. When the target video frame is an I frame, execute Step 703; when the target video frame is a P frame, execute Step 709.
[0163] Step 703: For each coding block of the target video frame, divide the 16x16 coding block into 4 8x8 coding sub-blocks, and perform prediction on each coding sub-block through the IBC mode respectively.
[0164] Here, when adopting the IBC mode, the coding sub-block will search for the most similar block in the reconstructed coding sub-blocks of the target video frame as the prediction sub-block, and calculate the block vector (BV, Block Bector). Among them, the block vector indicates the positional relationship between the current block and the prediction sub-block.
[0165] In actual implementation, a hash-based search method can be used to find the optimal block vector. When calculating the hash value of the coding sub-block, the gradient value of the coding sub-block will be calculated. Therefore, the intermediate results of predicting each 8x8 coding sub-block through the IBC mode can be utilized to obtain the gradient value of each coding sub-block.
[0166] Step 704: During the process of predicting each 8x8 coding sub-block through the IBC mode, determine the gradient value of each coding sub-block.
[0167] Step 705: Sum up the gradient values of each 8x8 coding sub-block to obtain the gradient value of the 16x16 coding block.
[0168] Here, for all pixels in the 8x8 coding sub-block except for the first row and the first column, the absolute value of the pixel difference between this pixel and the adjacent left pixel can be obtained, and the absolute value of the pixel difference between this pixel and the adjacent left pixel is used as the horizontal gradient corresponding to this pixel; and the absolute value of the pixel difference between this pixel and the adjacent upper pixel is obtained, and the absolute value of the pixel difference between this pixel and the adjacent upper pixel is used as the vertical gradient corresponding to this pixel; then take the average of the horizontal gradient and the vertical gradient as the gradient value of this pixel; sum up the gradient values corresponding to all pixels in the 8x8 coding sub-block, and the gradient value of the 8x8 coding sub-block can be obtained.
[0169] That is to say, the gradient value of the 8x8 coding sub-block can be obtained through the following formula:
[0170] g (i,j) =(|p (i,j-1) -p (i,j) |+|p (i-1,j) -p (i,j) |) / 2;
[0171]
[0172] where, g (i,j) represents the gradient value corresponding to the pixel at the i-th row and the j-th column in the coding sub-block, p (i,j) represents the pixel at the i-th row and the j-th column in the coding sub-block, and g represents the gradient value of the coding sub-block. Here, the pixel refers to the luminance pixel.
[0173] After obtaining the gradient values of four 8x8 coded sub-blocks in an 8x8 coded sub-block, the gradient value of the 8x8 coded sub-block can be obtained through the following formula:
[0174]
[0175] where g 16×16 represents the gradient value of a 16x16 coded block, g sub8×8_1 、g sub8×8_2 、g sub8×8_3 represent the gradient values of four coded sub-blocks of a 16x16 coded block.
[0176] Step 706: Determine whether the gradient value of the 16x16 coded block is less than the gradient threshold. If so, execute Step 707; otherwise, execute Step 708.
[0177] Here, the gradient value is used to describe the texture complexity. When the gradient value of the 16x16 coded block is less than the gradient threshold, it indicates that the texture information of the 16x16 coded block is relatively simple, and it is necessary to calculate the standard Intra mode; when the gradient value of the 16x16 coded block is greater than or equal to the gradient threshold, it indicates that the texture information of the 16x16 coded block is relatively complex, and there is a high probability that the PLT mode can obtain a better rate-distortion cost than the standard Intra mode.
[0178] Regarding the setting of the gradient threshold, the smaller the gradient threshold, the lower the probability of executing the standard Intra mode, which can reduce the computational complexity and improve the encoding speed.
[0179] Step 707: Predict the 16x16 coded block through the Intra16x16 mode.
[0180] Step 708: Predict the 16x16 coded block through the PLT mode.
[0181] When skipping Step 707 and directly executing Step 708, since the standard Intra mode is not executed, the computational complexity is reduced. Here, the accuracy of the PLT mode can be improved. For example, the sampling accuracy of the PLT mode can be increased from 3x3 to 3x2 pixel regions. In this way, both the encoding quality and speed can be improved.
[0182] Step 709: For each 16x16 coded block corresponding to the target video frame, determine whether the coded block is a skip block. If so, execute Step 712; otherwise, execute Step 710.
[0183] Step 710: Predict the 16x16 coded block through the standard Inter mode to obtain the CBP information when using the standard Inter mode.
[0184] Step 711: Determine whether the CBP information in the standard Inter mode is zero. If so, execute Step 712; otherwise, execute Step 703.
[0185] Here, the CBP information is a syntax element used to reflect the residual situation in the encoding of the coding block. The CBP information has 6 bits, where the first 2 bits represent the UV components; the last 4 bits are the Y components, respectively representing 4 8x8 coding sub-blocks within the coding block. If any bit is 0, it indicates that all the transform coefficient levels (the values in the matrix after transforming and quantizing the pixel residuals, hereinafter collectively referred to as levels) in the corresponding 8x8 block are all 0, otherwise it indicates that the transform coefficient levels in the corresponding 8x8 block are not all 0.
[0186] Step 712: Determine the optimal coding mode for each corresponding coding block respectively, so as to perform the encoding of the target video frame based on the determination.
[0187] Here, after executing each coding mode, the optimal coding mode for the corresponding coding block can be selected according to the execution results of each coding mode, so as to determine to encode the coding block through the optimal coding mode. After determining the optimal coding mode corresponding to all coding blocks in the target video frame, the encoding of the target video frame can be achieved.
[0188] The embodiments of the present application have the following beneficial effects:
[0189] For the QQ265 SCC 444 encoder, when Tintra16x16_skip is set to 1024 and the drawing board sampling accuracy of the PLT module is set to 3x2. In the common screen scene video, the encoding speed in the AI mode is increased by 10%, and at the same time, the BD-rate is increased by -0.27%; in the long-play (LP) mode, the encoding complexity is decreased by 3.1%, and at the same time, the BD-rate is increased by -0.23%.
[0190] Next, continue to describe the exemplary structure of the implementation of the video encoding device 455 provided by the embodiments of the present application as a software module. Figure 8 It is a schematic diagram of the composition structure of the video encoding device provided by the embodiments of the present application. As Figure 8 shown, the video encoding device 455 provided by the embodiments of the present application includes:
[0191] A partitioning module 4551, configured to perform coding block partitioning on the target video frame to be encoded, and obtain a plurality of coding blocks corresponding to the target video frame;
[0192] A determination module 4552, configured to determine the texture complexity of each of the encoding blocks respectively;
[0193] An encoding module 4553, configured to, when determining that a corresponding encoding block satisfies a mode skip condition based on the texture complexity of each of the encoding blocks, skip the intra mode for the sequentially executed intra mode and palette mode, and
[0194] encode the encoding block through the palette mode, so as to implement encoding of the target video when encoding of the multiple encoding blocks is completed.
[0195] In some embodiments, the determination module 4552 is further configured to respectively obtain the gradient value of each of the encoding blocks, and use the gradient value as the texture complexity of the corresponding encoding block;
[0196] respectively compare the gradient value of each of the encoding blocks with a gradient threshold to obtain a comparison result;
[0197] when the comparison result indicates that the gradient value reaches the gradient threshold, determine that the encoding block corresponding to the gradient value satisfies the mode skip condition.
[0198] In some embodiments, the determination module 4552 is further configured to respectively perform the following operations on each of the encoding blocks:
[0199] divide the encoding block into multiple encoding sub-blocks;
[0200] obtain the gradient values of the multiple encoding sub-blocks corresponding to the encoding block;
[0201] use the sum of the gradient values of the multiple encoding sub-blocks corresponding to the encoding block as the gradient value of the encoding block.
[0202] In some embodiments, the determination module 4552 is further configured to, when the target encoding frame satisfies a first condition, encode each of the encoding sub-blocks through an intra block copy (IBC) mode, and
[0203] during the process of encoding the encoding sub-blocks through the IBC mode, determine the gradient value of the corresponding encoding sub-block.
[0204] In some embodiments, the determination module is further configured to obtain the type of the target encoding frame;
[0205] when the type of the target encoding frame indicates that the target encoding frame is an intra-coded frame, determine that the target encoding frame satisfies the first condition;
[0206] when the type of the target encoding frame indicates that the target encoding frame is a forward prediction frame, obtain the current encoding mode of the target encoding block, and
[0207] When the current coding mode of the coding block is not the skip mode and the residual quantization value obtained by coding the coding block using the inter-frame mode is not zero, it is determined that the target coding frame meets the first condition.
[0208] In some embodiments, the determining module 4552 is further configured to perform the following processing for each of the coding sub-blocks:
[0209] Determine the horizontal gradient and vertical gradient of each pixel in the coding sub-block;
[0210] Obtain the gradient average value of the horizontal gradient and vertical gradient corresponding to each pixel in the coding sub-block;
[0211] Take the sum of the gradient average values corresponding to each pixel included in the coding sub-block as the gradient value of the coding sub-block.
[0212] In some embodiments, the coding module 4553 is further configured to obtain the type of the target coding frame;
[0213] When the type of the target coding frame indicates that the target coding frame is a forward prediction frame, obtain the current coding mode of the coding block; when the current coding mode of the coding block is the skip mode, end the processing for the coding block; or,
[0214] When the current coding mode of the coding block is the inter-frame mode and the residual quantization value obtained by coding the coding block using the inter-frame mode is zero, end the processing for the coding block.
[0215] In some embodiments, the coding module 4553 is further configured to obtain the first rate-distortion cost when coding the coding block using the skip mode;
[0216] When the first rate-distortion cost is not greater than the rate-distortion threshold, determine that the current coding mode of the coding block is the skip mode.
[0217] In some embodiments, the coding module 4553 is further configured to obtain the predicted value of the coding block and the actual value of the coding block when coding the coding block using the inter-frame mode;
[0218] Compare the actual value of the coding block with the predicted value to obtain the residual between the actual value and the predicted value of the coding block;
[0219] Encode the residual to obtain a residual quantization value.
[0220] In some embodiments, the coding module 4553 is further configured to generate a color palette, where the color palette is a table including at least one color value in the coding block;
[0221] Obtain the indices of each pixel in the encoding block in the table respectively to implement encoding of the encoding block.
[0222] In some embodiments, the encoding module 4553 is further configured to generate a palette, where the palette is a table including at least one color value in the encoding block;
[0223] Sample the pixels in the encoding block according to a preset pixel sampling interval to obtain a plurality of sampled pixels;
[0224] Obtain the indices of each of the sampled pixels in the encoding block in the table respectively, and use the indices as the encoding result of the encoding block.
[0225] In some embodiments, when the encoding module 4553 is further configured to determine that the corresponding encoding block does not meet the mode skip condition based on the texture complexity, encode the encoding block through an intra mode;
[0226] Obtain a second rate-distortion cost when encoding the encoding block through the intra mode;
[0227] When the second rate-distortion cost is greater than the rate-distortion threshold, encode the encoding block through a palette mode.
[0228] In some embodiments, the encoding module 4553 is further configured to encode the encoding block through the vertical mode, horizontal mode, DC mode, and planar mode included in the intra mode respectively;
[0229] The obtaining of the second rate-distortion cost when encoding the encoding block through the intra mode includes:
[0230] Obtain the sum of absolute transform errors when encoding through the vertical mode, horizontal mode, DC mode, and planar mode respectively;
[0231] According to the obtained sum of absolute transform errors, use the mode corresponding to the minimum value in the obtained sum of absolute transform errors as the optimal intra mode;
[0232] Obtain the second rate-distortion cost when encoding the encoding block through the optimal intra mode.
[0233] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video encoding method described above in the embodiments of the present application.
[0234] Embodiments of the present application provide a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by a processor, cause the processor to execute the method provided by the embodiments of the present application. For example, as Figure 3 the method shown.
[0235] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0236] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0237] As an example, the executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file storing other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (e.g., files storing one or more modules, subroutines, or code portions).
[0238] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or, on multiple computing devices distributed at multiple locations and interconnected by a communication network.
[0239] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A video encoding method, characterized in that, The method includes: Performing encoding block partitioning on a target video frame to be encoded, obtaining a plurality of encoding blocks corresponding to the target video frame; Performing the following operations respectively for each of the encoding blocks: Dividing the encoding block into a plurality of encoding sub-blocks; When a target encoding frame satisfies a first condition, encoding each of the encoding sub-blocks respectively through an Intra Block Copy (IBC) mode, and determining a gradient value of a corresponding encoding sub-block during the process of encoding the encoding sub-block through the IBC mode; Taking the sum of the gradient values of the plurality of encoding sub-blocks corresponding to the encoding block as the gradient value of the encoding block; Taking the gradient value as the texture complexity of the corresponding encoding block; Comparing the gradient value of each of the encoding blocks with a gradient threshold respectively to obtain a comparison result; When the comparison result indicates that the gradient value reaches the gradient threshold, determining that the encoding block corresponding to the gradient value satisfies a mode skip condition; When it is determined that the corresponding encoding block satisfies the mode skip condition based on the texture complexity of each of the encoding blocks, for an intra mode and a palette mode that are sequentially executed, skipping the intra mode, and Encoding the encoding block through the palette mode to implement encoding of the target video frame when the encoding of the plurality of encoding blocks is completed.
2. The method according to claim 1, wherein, Before encoding each of the encoding sub-blocks respectively through the Intra Block Copy (IBC) mode, the method further includes: Obtaining the type of the target encoding frame; When the type of the target encoding frame indicates that the target encoding frame is an intra-coded frame, determining that the target encoding frame satisfies the first condition; When the type of the target encoding frame indicates that the target encoding frame is a forward prediction frame, obtaining the current encoding mode of a target encoding block, and When the current encoding mode of the encoding block is not a skip mode and the residual quantization value obtained by encoding the encoding block through the inter mode is not zero, determining that the target encoding frame satisfies the first condition.
3. The method according to claim 1, wherein Before performing the following operations respectively for each of the encoding blocks, the method further includes: Obtaining the type of the target encoding frame; When the type of the target encoding frame indicates that the target encoding frame is a forward prediction frame, obtaining the current encoding mode of the encoding block; When the current encoding mode of the encoding block is a skip mode, ending the processing for the encoding block; or, When the current encoding mode of the encoding block is an inter mode and the residual quantization value obtained by encoding the encoding block through the inter mode is zero, ending the processing for the encoding block.
4. The method according to claim 3, wherein The method further includes: Obtaining a first rate-distortion cost when encoding the encoding block through a skip mode; When the first rate-distortion cost is not greater than a rate-distortion threshold, determining that the current encoding mode of the encoding block is a skip mode.
5. The method according to claim 3, characterized in that, The residual quantization value obtained by encoding the encoding block through the inter mode includes: Obtaining a predicted value of the encoding block and an actual value of the encoding block when encoding the encoding block through the inter mode; Comparing the actual value of the encoding block with the predicted value to obtain a residual between the actual value and the predicted value of the encoding block; Encode the residual to obtain a residual quantization value.
6. The method according to any one of claims 1 to 5, characterized in that, Encoding the encoded block through the palette mode includes: Generate a palette, which is a table containing at least one color value in the encoded block; Obtain the index of each pixel in the encoded block in the table respectively, and use the index as the encoding result of the encoded block.
7. The method according to any one of claims 1 to 5, characterized in that Encoding the encoded block through the palette mode includes: Generate a palette, which is a table containing at least one color value in the encoded block; Sample the pixels in the encoded block according to a preset pixel sampling interval to obtain a plurality of sampled pixels; Obtain the index of each of the sampled pixels in the encoded block in the table respectively, and use the index as the encoding result of the encoded block.
8. The method according to any one of claims 1 to 5, characterized in that The method further includes: When it is determined that the corresponding encoded block does not meet the mode skip condition based on the texture complexity, encode the encoded block through the intra mode; Obtain the second rate-distortion cost when encoding the encoded block through the intra mode; When the second rate-distortion cost is greater than the rate-distortion threshold, encode the encoded block through the palette mode.
9. The method according to claim 8, characterized in that Encoding the encoded block through the intra mode includes: Encode the encoded block through the vertical mode, horizontal mode, DC mode, and planar mode included in the intra mode respectively; Obtaining the second rate-distortion cost when encoding the encoded block through the intra mode includes: Obtain the sum of the absolute transform errors when encoding through the vertical mode, horizontal mode, DC mode, and planar mode respectively; According to the obtained sum of the absolute transform errors, use the mode corresponding to the minimum value in the obtained sum of the absolute transform errors as the optimal intra mode; Obtain the second rate-distortion cost when encoding the encoded block through the optimal intra mode.
10. A video encoding device, characterized in that, The apparatus includes: A partitioning module for partitioning the target video frame to be encoded into encoded blocks corresponding to the target video frame; A determination module for performing the following operations on each of the encoded blocks respectively: partitioning the encoded block into a plurality of encoded sub-blocks; when the target encoded frame meets the first condition, encode each of the encoded sub-blocks through the intra block copy (IBC) mode, and determine the gradient value of the corresponding encoded sub-block during the encoding of the encoded sub-block through the IBC mode; use the sum of the gradient values of the plurality of encoded sub-blocks corresponding to the encoded block as the gradient value of the encoded block; Use the gradient value as the texture complexity of the corresponding encoded block; compare the gradient value of each encoded block with the gradient threshold respectively to obtain a comparison result; when the comparison result indicates that the gradient value reaches the gradient threshold, determine that the encoded block corresponding to the gradient value meets the mode skip condition; An encoding module for, when it is determined that the corresponding encoded block meets the mode skip condition based on the texture complexity of each encoded block, skipping the intra mode for the sequentially executed intra mode and palette mode, and Encoding the encoded block through the palette mode to implement encoding of the target video frame when the encoding of the multiple encoded blocks is completed.
11. A computer device, characterized in that, Comprising: A memory for storing executable instructions; A processor for implementing the video encoding method according to any one of claims 1 to 9 when executing the executable instructions stored in the memory.
12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer executable instructions or computer program are executed by the processor, the video encoding method according to any one of claims 1 to 9 is implemented.
13. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer executable instructions or computer program are executed by the processor, the video encoding method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Fast screen content coding method based on spatiotemporal correlation
CN107623850A