Video encoding method, apparatus, electronic device, and storage medium
By combining a low-complexity first interpolation method with a high-compression second coding standard in video coding, the problem of balancing compression rate and coding efficiency in video coding is solved, achieving high-efficiency video coding results.
Patent Information
- Application Number
- CN202210605670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-05-30
AI Technical Summary
Existing video coding technologies struggle to strike a balance between video compression rate and coding efficiency. High-efficiency video coding standards like HEVC employ high computational complexity in their interpolation methods, resulting in low coding efficiency.
A low-complexity first interpolation method is used to perform subpixel interpolation on the reference frame to generate the search plane. Motion estimation and video coding are then performed in conjunction with a high-compression second coding standard. By differentiating the interpolation method, the computational complexity of the motion estimation process is reduced while maintaining the video compression rate.
While maintaining the video compression rate, the complexity and computational load of the video encoding process are reduced, and the encoding efficiency is improved, achieving a balance between compression rate and computational load.
Smart Images

Figure CN115037947B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to cloud computing, video processing, and media cloud technologies, which can be applied in intelligent cloud scenarios. Specifically, it relates to a video encoding method, apparatus, electronic device, and storage medium. Background Technology
[0002] HEVC (High Efficiency Video Coding) is a next-generation video coding and compression standard widely used in video compression-related fields, such as live streaming and video-on-demand. When using the HEVC coding standard for video encoding, the compression rate or bitrate is high, resulting in good video quality. Summary of the Invention
[0003] This disclosure provides a video encoding method, apparatus, electronic device, and storage medium.
[0004] According to one aspect of this disclosure, a video coding method is provided, comprising:
[0005] Based on the first interpolation method under the first coding standard, sub-pixel interpolation is performed on the reference frame corresponding to the target frame to be encoded in the video to be encoded to obtain the search plane;
[0006] In the search plane, motion estimation is performed on each block to be encoded in the target frame to be encoded to obtain the search motion vector;
[0007] Based on the second interpolation method under the second coding standard, a prediction plane for the target frame to be encoded is generated according to the search motion vector, which is used for video encoding based on the second coding standard.
[0008] The interpolation complexity of the first interpolation method is lower than that of the second interpolation method; the video compression rate of the second encoding standard is higher than that of the first encoding standard.
[0009] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0010] At least one processor; and
[0011] A memory communicatively connected to the at least one processor; wherein,
[0012] The memory stores instructions that can be executed by the at least one processor, which, when executed, enable the at least one processor to perform the video encoding method provided in any embodiment of this disclosure.
[0013] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the video encoding method provided in any embodiment of this disclosure.
[0014] The technical solution disclosed herein improves encoding efficiency while maintaining video compression rate.
[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0016] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0017] Figure 1 This is a schematic diagram of a video encoding method provided according to an embodiment of the present disclosure;
[0018] Figure 2A This is a schematic diagram of another video encoding method provided according to an embodiment of the present disclosure;
[0019] Figure 2B This is a schematic diagram illustrating a process of performing half-pixel interpolation on a reference block based on H.264 interpolation according to an embodiment of this disclosure;
[0020] Figure 2C This is a schematic diagram illustrating a process of interpolating a reference block by 1 / 4 pixel based on H.264 interpolation according to an embodiment of this disclosure;
[0021] Figure 2D This is a schematic diagram of a reference block provided according to an embodiment of the present disclosure;
[0022] Figure 2E This is a schematic diagram illustrating a process of sub-pixel interpolation of a reference block based on HEVC interpolation according to an embodiment of this disclosure;
[0023] Figure 3 This is a schematic diagram of yet another video encoding method provided according to an embodiment of the present disclosure;
[0024] Figure 4 This is a schematic diagram of yet another video encoding method provided according to an embodiment of the present disclosure;
[0025] Figure 5 This is a structural diagram of a video encoding apparatus provided according to an embodiment of the present disclosure;
[0026] Figure 6 This is a block diagram of an electronic device used to implement the video encoding method of the embodiments of this disclosure. Detailed Implementation
[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0028] Video coding mainly includes stages such as prediction, transform, quantization, loop filtering, and entropy coding. Prediction, as a crucial component of video coding, is divided into intra-frame prediction and inter-frame prediction. Inter-frame prediction may include motion estimation and motion compensation. Sub-pixel interpolation is typically used in motion estimation and motion compensation to improve accuracy. However, interpolation methods under some high-compression-rate coding standards are computationally complex and computationally intensive, resulting in low coding efficiency and failing to achieve a balance between video compression rate and coding efficiency.
[0029] To address the aforementioned technical problems, this disclosure provides a video encoding method and a video encoding apparatus, applicable to application scenarios involving video encoding. The video encoding methods provided in this disclosure can be executed by a video encoding apparatus, which can be implemented in software and / or hardware and specifically configured within an electronic device, such as a computer and / or server, etc., and this disclosure does not impose any limitations thereon.
[0030] To facilitate understanding, we will first provide a detailed explanation of video encoding methods.
[0031] See Figure 1 The video encoding method shown includes:
[0032] S110. Based on the first interpolation method under the first coding standard, sub-pixel interpolation is performed on the reference frame corresponding to the target frame to be encoded in the video to be encoded to obtain the search plane.
[0033] In this context, encoding standards can be understood as the encoding rules that must be followed during the video encoding process. Encoding standards typically specify the interpolation methods that must be followed, using pixel-level or sub-pixel-level interpolation to ensure image quality during video encoding.
[0034] The first encoding standard can be a standard with a relatively low video compression rate but high computational efficiency, i.e., low computational complexity. The first interpolation method can be the interpolation method required to be followed in the first encoding standard. In an optional embodiment, the first encoding standard can be at least one of the H.264 (digital video compression format) standard and the MPEG-2 standard (compression standard for moving images and audio based on digital storage media); correspondingly, the first interpolation method can be at least one of the subpixel interpolation method based on the H.264 standard and the subpixel interpolation method based on the MPEG-2 standard.
[0035] The video to be encoded can be a video that needs to be compressed or encoded; the video to be encoded may include at least one frame to be encoded, and the frame to be encoded carries the content information of the video image. The target frame to be encoded can be any frame to be encoded in the video to be encoded, or a video image frame to be encoded selected from at least one frame to be encoded based on certain requirements or rules. For example, the target frame to be encoded can be manually selected by relevant technicians according to actual needs, or it can be automatically selected according to the image complexity of the frame to be encoded; this embodiment does not limit this.
[0036] The reference frame can be an adjacent encoded frame of the frame to be encoded, and can be a forward encoded frame or a backward encoded frame; this disclosure does not impose any limitations on this. The encoded frame can be the result of decoding and reconstructing an image frame after encoding it.
[0037] The search plane can be a plane whose pixel accuracy is not lower than that of the reference frame after sub-pixel interpolation of the target frame to be encoded. There can be at least one search plane, and the pixel accuracy of different search planes can be the same or different; this disclosure does not impose any limitations on this.
[0038] For example, the plane to be searched may include at least one of a whole-pixel plane, a half-pixel plane, and a quarter-pixel plane. Of course, a pixel plane with higher precision, such as a quarter-pixel plane, may also be generated as needed, and this disclosure does not limit it in any way.
[0039] To improve the efficiency of generating the search plane, in one optional embodiment, the reference frame can be pre-divided into at least one reference block. Subpixel interpolation is performed on at least one reference block, and the interpolation results of different reference blocks are spliced together according to the positional relationship of the reference blocks to obtain the search plane.
[0040] Since subsequent video encoding is performed according to the second coding standard, the reference block can be obtained based on the block division method in the second coding standard, avoiding encoding failures caused by differences between standards. The second coding standard can be a coding standard with relatively high video compression ratio but relatively low computational efficiency, i.e., high computational complexity. In an optional embodiment, the second coding standard can be the HEVC standard.
[0041] For example, a target frame to be encoded can be selected from the frames to be encoded in the video to be encoded; the encoded frames adjacent to the target frame to be encoded can be used as reference frames; based on a second encoding standard, the reference frames can be divided using existing partitioning techniques to obtain at least one reference block; for each reference block, subpixel-level interpolation can be performed on the reference block based on a first interpolation method under a first encoding standard to obtain at least one interpolation block corresponding to the reference block; the interpolation blocks can be stitched together according to the positions of the corresponding pixels to generate at least one search plane. In a specific implementation, the first encoding standard can be the H.264 standard, the first interpolation method can be a subpixel interpolation method based on the H.264 standard, and the second encoding standard can be the HEVC standard.
[0042] Specifically, in the process of subpixel level interpolation of the reference block based on the first interpolation method under the first encoding standard, the reference block can be sequentially interpolated with integer pixels, half pixels, and quarter pixels to generate corresponding integer pixel interpolation blocks, half pixel interpolation blocks, and quarter pixel interpolation blocks, which are used to generate integer pixel planes, half pixel planes, and quarter pixel planes as important components of the plane to be searched.
[0043] S120. In the search plane, perform motion estimation on each block to be encoded in the target frame to be encoded to obtain the search motion vector.
[0044] The block to be encoded can be an image region divided into parts based on a second coding standard and using existing partitioning techniques or methods. Motion estimation can be performed by searching the search plane for blocks that closely match (e.g., the best match) each block to be encoded. The search motion vector is used to characterize the motion offset between the block to be encoded and the matched block, and can be a motion vector obtained based on the positional relationship between the block to be encoded and the block that closely matches (e.g., the best match).
[0045] For example, according to the second coding standard, using existing partitioning methods, each search plane can be divided into at least one search block, and the target frame to be encoded can be divided into at least one block to be encoded. Within each search block, the search block with the highest matching degree (e.g., the highest) with each block to be encoded is searched and used as the matching block that matches each block to be encoded in the target frame. A search motion vector is generated based on the positional relationship between the matching block and the corresponding block to be encoded. This disclosure does not limit the search method used in the search block matching process.
[0046] It should be noted that during the search process of a search block, the overall cost of the search block also needs to be considered, and the search block with the minimum overall cost is selected as the matching block. The overall cost may include at least one of the following: the number of bits consumed during the motion vector encoding process, inter-block differences, and block distortion. The overall cost can be determined using at least one cost determination method in the prior art, and this disclosure does not impose any limitations on the specific method for determining the overall cost.
[0047] S130. Based on the second interpolation method under the second coding standard, a prediction plane for the target frame to be encoded is generated according to the search motion vector, which is used for video coding based on the second coding standard.
[0048] The second coding standard can be a video standard with a higher video compression rate than the first compression standard; correspondingly, the second interpolation method is the interpolation method that needs to be followed under the second coding standard.
[0049] In a specific example, if the first encoding standard is H.264 or MPEG-2, then the second encoding standard can be HEVC. Correspondingly, the second interpolation method can be a subpixel interpolation method based on the HEVC standard.
[0050] The prediction plane is used to generate a prediction frame for the target frame to be encoded. Subsequently, the difference data between the frame to be encoded and the prediction frame can be determined by subtraction and other methods to obtain a residual frame. By performing transformation, quantization and entropy coding on the residual frame in sequence, the video encoding of the target frame to be encoded can be achieved.
[0051] It should be noted that the interpolation complexity of the first interpolation method is lower than that of the second interpolation method, in order to reduce computational complexity and workload during motion estimation in video encoding; the video compression rate of the second encoding standard is higher than that of the first encoding standard, in order to ensure that while reducing computational complexity and workload during video encoding, the video compression rate is guaranteed.
[0052] After determining the search motion vector, motion compensation can be performed based on the second interpolation method under the second coding standard, and a prediction plane can be generated based on the motion compensation result.
[0053] For example, subpixel interpolation can be performed on the target frame to be encoded based on the interpolation method of the second coding standard to obtain a compensation reference plane; in the compensation reference plane, the compensation block corresponding to each matching block is determined according to the search motion vector, and the compensation blocks are combined according to the pixel position to generate the prediction plane of the target frame to be encoded.
[0054] Accordingly, the current block to be encoded can be determined from each block to be encoded, and the predicted motion vector can be determined based on the search motion vectors of other blocks to be encoded around the current block to be encoded; the difference between the predicted motion vector and the search motion vector is taken as the motion vector difference; and video coding is performed based on the second coding standard according to the motion vector difference and the prediction plane.
[0055] The number of compensation reference planes is at least one, and the pixel precision of different compensation reference planes can be the same or different. To ensure matching accuracy and reduce the amount of data computation during subsequent encoding, the pixel precision of the compensation reference plane usually corresponds to the pixel precision of the search plane used to generate the search motion vector. In a specific example, the compensation reference plane may include at least one of a whole pixel plane, a half pixel plane, and a quarter pixel plane.
[0056] The technical solution of this disclosure adopts a first interpolation method with lower interpolation complexity and less computational load to replace the second interpolation method with higher interpolation complexity and more computational load in the motion estimation process of video encoding under the second coding standard. This reduces the computational complexity in the motion estimation process, thereby reducing the amount of data computation in the motion estimation process. At the same time, the second interpolation method with higher compression ratio under the second coding standard is used in the motion compensation process of video encoding, ensuring the video compression ratio. This achieves a balance between compression ratio and computational load in the video encoding process. While ensuring that the video compression ratio or bitrate effect remains unchanged, the complexity and computational load of the video encoding process are reduced, thereby improving the video encoding efficiency.
[0057] Based on the above technical solutions, this disclosure also provides an optional embodiment in which the generation process of the search screen is optimized to achieve comprehensive generation of the search plane based on different interpolation methods. It should be noted that for parts not described in detail in this disclosure's embodiments, please refer to the relevant descriptions in other embodiments, which will not be repeated here.
[0058] See Figure 2A The video encoding method shown includes:
[0059] S210. Divide the reference frame corresponding to the target frame to be encoded into at least one reference block; the reference block includes a first reference block and a second reference block.
[0060] For example, based on the second coding standard, existing block partitioning techniques for frames to be encoded can be used to partition the reference frame corresponding to the target frame to be encoded, thereby obtaining at least one reference block. The sizes of different reference blocks can be the same or different, and this can be determined during the partitioning process.
[0061] The reference block may include a first reference block and a second reference block. The first and second reference blocks may be determined by random or uniform partitioning, or by partitioning based on preset rules; this disclosure does not impose any limitations on this.
[0062] Because the compression rate of the second coding standard is higher than that of the first coding standard, the video frames reconstructed after compression using the second coding standard carry richer detail information than those reconstructed using the first coding standard. To avoid the loss of significant detail information during video compression, reference blocks can be divided into a first reference block and a second reference block based on the richness of detail information they carry. The first reference block can be a reference block with a simpler pixel structure, meaning it carries less detail information; the second reference block can be a reference block with a more complex pixel structure, meaning it carries more detail information.
[0063] In one optional embodiment, at least one reference block can be manually divided by relevant technical personnel according to the complexity of the pixel structure corresponding to each reference block to obtain a first reference block and a second reference block.
[0064] In another alternative embodiment, the degree of deviation of the first pixel of the reference block can be determined based on the pixel value of each pixel in the reference block; and the reference block can be divided into a first reference block and a second reference block based on the degree of deviation of the first pixel.
[0065] The deviation of the first pixel is used to characterize the complexity of the content information carried by the reference block. It can be quantified numerically by the pixel variance or pixel standard deviation of the reference block. The larger the deviation of the first pixel, the more complex the content information carried by the reference block; the smaller the deviation of the first pixel, the simpler the content information carried by the reference block.
[0066] For example, the pixel values corresponding to the pixels of the reference block can be determined; based on the pixel values of each pixel in the reference block, the pixel variance or pixel standard deviation of the reference block can be determined, and the determination result can be used as the first pixel deviation degree of the reference block.
[0067] Based on the degree of deviation of the first pixel of each reference block, the reference block with a lower degree of deviation of the first pixel is divided into the first reference block, and the reference block with a higher degree of deviation of the first pixel is divided into the second reference block.
[0068] In an optional embodiment, a first deviation threshold may be introduced as a measure of the degree of deviation of the first pixel; accordingly, the reference block is divided into a first reference block and a second reference block according to the degree of deviation of the first pixel, including: taking the reference block with the degree of deviation of the first pixel less than the first deviation threshold as the first reference block; and taking the reference block with the degree of deviation of the first pixel not less than the first deviation threshold as the second reference block.
[0069] The first deviation threshold can be preset or adjusted by relevant technical personnel, or determined through a large number of experiments.
[0070] Specifically, the deviation degree of the first pixel corresponding to each reference block and the magnitude of the first deviation threshold are determined; the reference block in each reference block whose deviation degree of the first pixel is less than the first deviation threshold is taken as the first reference block; the reference block in each reference block whose deviation degree of the first pixel is not less than the first deviation threshold is taken as the second reference block.
[0071] For example, at least one reference block obtained from the division of the reference frame includes reference block A, reference block B, reference block C, and reference block D. The first pixel deviation degrees corresponding to reference blocks A, B, C, and D are a1, a2, a3, and a4, respectively. Where a1 and a2 are less than a first deviation threshold a0, then reference blocks A and B corresponding to the first pixel deviation degrees a1 and a2 are the first reference blocks; where a3 and a4 are not less than the first deviation threshold a0, then reference blocks C and D corresponding to the first pixel deviation degrees a3 and a4 are the second reference blocks.
[0072] This optional embodiment uses a reference block whose first pixel deviation is less than a first deviation threshold as the first reference block and a reference block whose first pixel deviation is not less than the first deviation threshold as the second reference block. This method is relatively simple and can ensure the efficiency of reference block division while taking into account the accuracy of video encoding.
[0073] In one specific embodiment, at least one reference block corresponding to a reference frame is determined, and the pixel value of each pixel in each reference block is determined; for each reference block, the pixel standard deviation or pixel variance corresponding to the reference block is determined based on the pixel value of each pixel in the reference block to characterize the first pixel deviation degree corresponding to the reference block; a reference block whose first pixel deviation degree is less than a first deviation threshold is designated as a first reference block, and a reference block whose first pixel deviation degree is not less than the first deviation threshold is designated as a second reference block.
[0074] The above technical solution characterizes the complexity of the content information carried by the reference block by introducing the degree of pixel deviation, and divides each reference block into a first reference block and a second reference block based on the degree of pixel deviation. Thus, different interpolation methods are selected according to the complexity of the content information carried, realizing differentiated interpolation of the reference block according to the complexity of the content. This avoids the loss of image details caused by rashly adopting an interpolation method with low interpolation complexity, and helps to improve the richness and comprehensiveness of the information carried by the video coding result.
[0075] S220. Based on the first interpolation method, sub-pixel interpolation is performed on the first reference block to obtain the first interpolation block of the corresponding reference block.
[0076] The first interpolation block can be the result obtained by performing sub-pixel interpolation, such as whole pixel, half pixel, and 1 / 4 pixel, on the first reference block.
[0077] The following section uses the first interpolation method, which is the subpixel interpolation method based on the H.264 standard, as an example to explain the interpolation process in detail.
[0078] For example, such as Figure 2B The diagram illustrates a process of half-pixel interpolation of a reference block based on H.264 interpolation. Here, AT represents the position of an integer pixel in the first reference block, and ae and aa-hh represent half-pixel positions. The position between two integer pixels can be considered as one half-pixel position.
[0079] Under the H.264 standard, half-pixel precision interpolation uses a 6-tap filter with tap coefficients of (1, -5, 20, 20, -5, 1). The pixel values at half-pixel positions a, b, and e can be calculated as follows:
[0080] a=(E-5F+20G+20H-5I+J+16)>>5;
[0081] b=(A-5C+20G+20M-5Q+S+16)>>5;
[0082] e=(cc-5dd+20b+20c-5ee+ff+16)>>5;
[0083] Here, ">>5" means moving 5 positions to the right, which is equivalent to dividing the result by 2. 5 .
[0084] For example, such as Figure 2CThe diagram illustrates a process of interpolating a reference block into quarter-pixel positions using H.264 interpolation. Here, AD represents the position of an integer pixel in the first reference block, aa-ee represents the position of a half-pixel, and ae represents the position of a quarter-pixel. The positions between two half-pixels, and between an integer pixel and a half-pixel, can be used to generate the positions of the quarter-pixels. The pixel values of positions a, b, and e at the quarter-pixel positions can be calculated as follows:
[0085] a = (A + aa + 1) >> 1;
[0086] b = (A + bb + 1) >> 1;
[0087] e = (aa + bb + 1) >> 1;
[0088] Here, ">>1" means moving one position to the right, which is equivalent to dividing the result by 2. 1 It should be noted that the pixel value at position e is calculated from its neighboring pair of half-pixels, aa and bb, rather than from the whole pixel A and the half-pixel ee. Similarly, the pixel values of other half-pixels and quarter-pixels can be obtained, which will not be elaborated upon in this embodiment.
[0089] S230. Based on the second interpolation method, sub-pixel interpolation is performed on the second reference block to obtain the second interpolation block of the corresponding reference block.
[0090] The second interpolation block can be the result obtained by performing sub-pixel interpolation, such as whole pixel points, half pixel points, and 1 / 4 pixel points, on the second reference block.
[0091] The following section uses the second interpolation method, which is a sub-pixel interpolation method based on the HEVC standard, as an example to explain the interpolation process in detail.
[0092] Among them, the interpolation complexity of the second interpolation method is higher than that of the first interpolation method; the video compression rate of the second coding standard is higher than that of the first coding standard.
[0093] For example, such as Figure 2D The diagram shows a reference block, wherein A -1,-1 A 0,-1 A 1,-1 A 2,-1 A -1,0 A 0,0 A 1,0 A 2,0 A -1,1 A 0,1 A 1,1 A 2,1 A -1,2 A0,2 A 1,2 and A 2,2 These represent the integer pixels corresponding to the reference block. If Figure 2D The reference block shown is the second reference block that needs to be processed using the second interpolation method. Therefore, based on the HEVC interpolation method, sub-pixel interpolation is performed on this second reference block to obtain the corresponding second interpolated block. It should be noted that... Figure 2D The second reference block is described by way of example only and should not be construed as a specific limitation on the shape and size of the second reference block.
[0094] A schematic diagram illustrating the process of sub-pixel interpolation of the reference block based on HEVC interpolation can be seen as follows: Figure 2E As shown in the diagram. The pixels in the gray area represent the integer pixels corresponding to the second reference block. Subpixel interpolation is performed on the rows and columns containing the integer pixels corresponding to the second reference block to obtain the interpolated half-pixels and quarter-pixels.
[0095] It is understandable that interpolating between integer pixels can result in a half-pixel. Taking A as an example... 0,0 Taking sub-pixel points near point b as an example, where b 0,0 It can be A 0,0 and A 1,0 Half-pixel between, h 0,0 It can be A 0,0 and A 0,1 Half-pixel between, h 1,0 It can be A 1,0 and A 1,1 between half-pixels, b 0,1 It can be A 0,1 and A 1,1 The half-pixel between, j 0,0 It can be A 0,0 and A 1,1 Half-pixels between.
[0096] Interpolating between whole pixels and half pixels results in a pixel that is 1 / 4 of a pixel. For example, a 0,0 It could be A 0,0 and b 0,0 1 / 4 pixel between, d 0,0 It could be A 0,0 and h 0,0 1 / 4 pixel between, e 0,0 It could be A 0,0 and j 0,0 The remaining 1 / 4 pixels are not described in detail in this embodiment.
[0097] A pixel obtained by interpolating between a half-pixel and a whole pixel can be a 3 / 4 pixel. For example, c 0,0 It can be b 0,0 and A 1,0 Between 3 / 4 pixels, n 0,0 It can be h 0,0 and A 0,1 The remaining 3 / 4 pixels are not described in detail in this embodiment. It should be noted that in this embodiment, the 3 / 4 pixels obtained after subpixel interpolation of the second reference block can all be regarded as 1 / 4 pixels. Therefore, after subpixel interpolation of the second reference block, the corresponding integer pixels, half pixels, and 1 / 4 pixels of the reference block are obtained respectively.
[0098] It should be noted that the HEVC standard specifies that the values for the luminance component at the half-pixel position are generated by an 8-tap filter, while the values at the 1 / 4 and 3 / 4 pixel positions are generated by a 7-tap filter. The tap coefficients are shown in Table 1.
[0099] Table 1
[0100] Subpixel position Tap coefficient 1 / 4 {-1,4,-10,58,17,-5,1} 1 / 2 {-1,4,-11,40,40,-11,4,-1} 3 / 4 {1,-5,17,58,10,4,-1}
[0101] Interpolate the row or column containing integer pixels. Let A... 0,0 Taking nearby sub-pixels as an example, the following conditions are met: 0,0 b 0,0 c 0,0 d 0,0 h 0,0 and n 0,0 Among them, a 0,0 b 0,0 and c 0,0 The value can be calculated using integer pixels in the horizontal direction, d 0,0 j 0,0 and n 0,0 The value can be calculated using integer pixels in the vertical direction. Based on the tap coefficient corresponding to the sub-pixel position of the HEVC standard, the calculation is as follows:
[0102] a 0,0 =-A -3,0 +4A -2,0 -10A -1,0 +58A 0,0 +17A 1,0 -5A 2,0 +A 3,0 ;
[0103] h 0,0 =-A 0,-3 +4A 0,-2 -11A 0,-1 +40A 0,0+40A 0,1 -11A 0,2 +4A 0,3 -A 0,4 ;
[0104] Pixels at other positions b 0,0 c 0,0 d 0,0 and n 0,0 The values can be calculated using the corresponding filter tap coefficients, which will not be elaborated further in this embodiment.
[0105] Interpolate for the remaining sub-pixel positions that are not in the same row or column as the integer pixel. For pixels not in the same row or column, such as e... 0,0 f 0,0 g 0,0 i 0,0 j 0,0 k 0,0 p 0,0 q 0,0 and r 0,0 'a' is required 0,0 b 0,0 c 0,0 d 0,0 h 0,0 and n 0,0 The calculation is as follows:
[0106] e 0,0 =(-a 0,-3 +4a 0,-2 -10a 0,-1 +58a 0,0 +17a 0,1 -5a 0,2 +a 0,3 )>>6;
[0107] j 0,0 = (-b 0,-3 +4b 0,-2 -11b 0,-1 +40b 0,0 +40b 0,1 -11b 0,2 +4b 0,3 -b 0,4 )>>6;
[0108] r 0,0 =(-c 0,-2 -5c 0,-1 +17c 0,0 +58c 0,1 -10c 0,2 +4c 0,3 -c 0,4 )>>6;
[0109] Here, ">>6" means moving 6 positions to the right, which is equivalent to dividing the result by 2. 6 .
[0110] The remaining pixels f 0,0 g 0,0 i 0,0 k 0,0 p 0,0 and q 0,0 The calculation requires the use of appropriate filters, which will not be elaborated upon in this embodiment.
[0111] The sub-pixel points (including whole pixels, half pixels, 1 / 4 pixels, and 3 / 4 pixels (equivalent to 1 / 4 pixels) after sub-pixel interpolation of the second reference block are combined according to their pixel positions to obtain interpolation blocks including whole pixels, half pixels, and 1 / 4 pixels.
[0112] For example, continuing from the previous example, the interpolated A -1,-1 A 0,-1 A 1,-1 A 2,-1 A -1,0 A 0,0 A 1,0 A 2,0 A -1,1 A 0,1 A 1,1 A 2,1 A -1,2 A 0,2 A 1,2 and A 2,2 By combining pixels according to their positions, an interpolation block of integer pixels is obtained; the half-pixel points b corresponding to each integer pixel are then... 0,-1 b 0,0 ... b 0,1 By combining pixels according to their positions, an interpolation block of half a pixel is obtained; 1 / 4 pixel a is then used as an interpolation block. 0,-1 a 0,0 ... a 0,1 By combining these pixels according to their positions, an interpolation block of 1 / 4 pixel is obtained. Other half-pixels, such as h... 0,0 or j 0,0 And other 1 / 4 pixel points such as d 0,0 or c 0,0 The same method is used to combine the interpolation blocks for other half-pixels and other 1 / 4-pixel interpolation blocks, which will not be described in detail in this embodiment.
[0113] It is understandable that if the pixel size corresponding to the second reference block is a 4*4 pixel block, then the size of the second interpolation block obtained after subpixel interpolation is also 4*4, and each 4*4 reference block can obtain 16 interpolation blocks, including interpolation blocks of whole pixels, interpolation blocks of half pixels and interpolation blocks of 1 / 4 pixels.
[0114] It should be noted that this embodiment does not limit the execution order of S220 and S230. For example, S220 can be executed first and then S230 can be executed, or S230 can be executed first and then S220 can be executed, or S220 and S230 can be executed simultaneously or in an alternating manner.
[0115] Understandably, based on the aforementioned sub-pixel interpolation processes using the HEVC standard and the H.264 standard, the HEVC interpolation process is significantly more complex and computationally intensive than the H.264 interpolation process. If the entire motion estimation process for video coding were to use the more complex HEVC standard for interpolation, the computational load would be high, resulting in low efficiency and ultimately lower video coding efficiency.
[0116] S240. Generate the search plane based on the first interpolation block and the second interpolation block.
[0117] The first interpolation block and the second interpolation block can be combined according to pixel position to obtain the plane to be searched. The plane to be searched may include at least one of the following: a whole pixel plane, a half pixel plane, and a quarter pixel plane.
[0118] For example, a reference frame is divided into two first reference blocks and two second reference blocks, namely first reference block A1, first reference block A2, second reference block B1, and second reference block B2, and the size of the two first reference blocks and the two second reference blocks is 4*4 pixels. Based on the first interpolation method, after sub-pixel interpolation of any first reference block, the number of first interpolation blocks of the corresponding reference block is 16. Similarly, based on the second interpolation method, after sub-pixel interpolation of any second reference block, the number of second interpolation blocks of the corresponding reference block is also 16, and the size of both the first and second interpolation blocks is 4*4 pixels. Therefore, by concatenating the 16 first interpolation blocks corresponding to the first reference block A1, the 16 first interpolation blocks corresponding to the first reference block A2, the 16 second interpolation blocks corresponding to the second reference block B1, and the 16 second interpolation blocks corresponding to the second reference block B2 according to the pixel positions of the interpolation blocks, 16 search planes of size 16*16 pixels can be obtained.
[0119] S250. In the search plane, perform motion estimation on each block to be encoded in the target frame to be encoded to obtain the search motion vector.
[0120] S260. Based on the second interpolation method under the second coding standard, a prediction plane for the target frame to be encoded is generated according to the search motion vector, which is used for video coding based on the second coding standard.
[0121] The present invention divides the reference frame corresponding to the target frame to be encoded into a first reference block and a second reference block. By using different interpolation methods, sub-pixel interpolation is performed on the first and second reference blocks respectively to obtain the interpolation blocks of the corresponding reference blocks for generating the search plane. Thus, in the video encoding process corresponding to the second encoding standard, at least part of the interpolation process in the motion estimation process is replaced with the first interpolation method with lower computational complexity. Different interpolation methods can be selected and used as needed, which improves the flexibility of the motion estimation process while taking into account the video compression rate and computational complexity of video encoding.
[0122] Based on the above technical solutions, this disclosure also provides an optional embodiment, in which the partitioning mechanism of the reference block into a first reference block and a second reference block is improved. It should be noted that for parts not described in detail in the embodiments of this disclosure, please refer to the relevant descriptions in other embodiments, which will not be repeated here.
[0123] See Figure 3 The video encoding method shown includes:
[0124] S310. Divide the reference frame corresponding to the target frame to be encoded into at least one reference block.
[0125] S320. Based on the first interpolation method, perform sub-pixel interpolation on each reference block to obtain the interpolation block of the corresponding reference block.
[0126] For example, based on the interpolation method under the first coding standard, sub-pixel interpolation is performed on each reference block obtained by dividing the reference frame to obtain at least one interpolation block corresponding to each reference block.
[0127] S330. Determine the degree of deviation of the second pixel of the interpolation block based on the pixel values of each pixel in the interpolation block.
[0128] The deviation of the second pixel is used to characterize the complexity of the content information carried by the interpolation block. It can be quantified numerically by the pixel variance or pixel standard deviation of the interpolation block. The larger the deviation of the second pixel, the more complex the content information carried by the interpolation block; the smaller the deviation of the second pixel, the simpler the content information carried by the interpolation block.
[0129] For example, the pixel value corresponding to the pixel point of the interpolation block can be determined; based on the pixel value of each pixel point in the interpolation block, the pixel variance or pixel standard deviation of the interpolation block can be determined, and the determination result can be used as the second pixel deviation degree of the interpolation block.
[0130] S340. Based on the degree of deviation of the second pixel, the corresponding reference block is divided into a first reference block and a second reference block.
[0131] For example, the reference block corresponding to the interpolation block with a lower degree of deviation of the second pixel can be used as the first reference block; and the reference block corresponding to the interpolation block with a higher degree of deviation of the second pixel can be used as the second reference block.
[0132] In an optional embodiment, a second deviation threshold may be introduced as a criterion for measuring the degree of deviation of the second pixel; accordingly, based on the degree of deviation of the second pixel, the corresponding reference block is divided into a first reference block and a second reference block, including: taking the reference block corresponding to the interpolation block whose degree of deviation of the second pixel is less than the second deviation threshold as the first reference block; and taking the reference block corresponding to the interpolation block whose degree of deviation of the second pixel is not less than the second deviation threshold as the second reference block.
[0133] The second deviation threshold can be preset or adjusted by relevant technical personnel, or determined through extensive testing. It should be noted that the second deviation threshold can be the same as or different from the aforementioned first deviation threshold, and this disclosure does not impose any limitations on their numerical values.
[0134] Specifically, the degree of second pixel deviation and the magnitude of the second deviation threshold corresponding to at least one interpolation block obtained by each reference block based on the first interpolation method are determined; the reference block corresponding to the interpolation block with the degree of second pixel deviation less than the second deviation threshold is taken as the first reference block; the reference block corresponding to the interpolation block with the degree of second pixel deviation not less than the second deviation threshold is taken as the second reference block.
[0135] It should be noted that after interpolating each reference block based on the first interpolation method, each reference block can correspond to at least one interpolation block. Each interpolation block corresponds to the degree of deviation of its second pixel.
[0136] This optional embodiment uses the interpolation block corresponding to the second pixel deviation less than the first deviation threshold as the first reference block, and the interpolation block corresponding to the second pixel deviation not less than the second deviation threshold as the second reference block. The partitioning method is relatively simple and can ensure the partitioning efficiency of the reference block while taking into account the video coding accuracy.
[0137] In an optional embodiment, if the second pixel deviation of any interpolation block in at least one interpolation block corresponding to the reference block is less than the second deviation threshold, then the reference block can be used as the first reference block; otherwise, the reference block can be used as the second reference block.
[0138] In another optional embodiment, if, among at least one interpolation block corresponding to a reference block, there exists an interpolation block whose second pixel deviation is less than a second deviation threshold that satisfies a preset first quantity threshold, then that reference block can be used as the first reference block; otherwise, that reference block is used as the second reference block. The preset first quantity threshold can be set or adjusted by relevant technical personnel according to needs or experience, or determined based on extensive experimentation. For example, if the number of interpolation blocks corresponding to reference block A is 16, and the preset first quantity threshold is 9, then there are 9 interpolation blocks whose second pixel deviation is less than the second deviation threshold; in this case, the reference block can be used as the first reference block; otherwise, that reference block is used as the second reference block.
[0139] Understandably, by introducing a second pixel offset degree and filtering interpolation blocks carrying relatively complex information, it can be determined that the corresponding region in the reference block to which the interpolation block belongs carries relatively rich content information. Insisting on using interpolation blocks obtained using the first interpolation method will result in the loss of content details during the encoding process. Therefore, such reference blocks can be used as second reference blocks, and subsequent subpixel interpolation will revert to using the second interpolation method to avoid the loss of content details.
[0140] S350, The interpolation block corresponding to the first reference block is taken as the first interpolation block.
[0141] Since S320 interpolates each reference block based on the first interpolation method to obtain the corresponding interpolation block, the interpolation block corresponding to the first reference block can be directly used as the first interpolation block, without repeating the interpolation operation based on the first interpolation method.
[0142] S360. Based on the second interpolation method, sub-pixel interpolation is performed on the second reference block to obtain the second interpolation block of the corresponding reference block.
[0143] Subpixel interpolation is performed on the second reference block using only the second interpolation method, instead of the entire reference block. This reduces the computational complexity of the motion estimation process and improves the interpolation efficiency during the interpolation of the first reference block.
[0144] S370. Generate the search plane based on the first interpolation block and the second interpolation block.
[0145] S380. In the search plane, perform motion estimation on each block to be encoded in the target frame to be encoded to obtain the search motion vector.
[0146] S390. Based on the second interpolation method under the second coding standard, a prediction plane for the target frame to be encoded is generated according to the search motion vector, which is used for video coding based on the second coding standard.
[0147] The present invention discloses an embodiment that interpolates each reference block based on a first interpolation method to obtain an interpolation block for the corresponding reference block. It uses a method to determine the second pixel deviation degree of the interpolation block to characterize the complexity of the content information carried by the interpolation, indirectly characterizing the complexity of the content information in the reference block to which the interpolation block belongs. Based on the pixel deviation degree of the interpolation block, each reference block is divided to obtain a first reference block and a second reference block. This achieves a fallback from the first interpolation method to the second interpolation method in the second reference block interpolation process. That is, the motion trajectory process only uses the second interpolation method for the second reference block, instead of performing sub-pixel interpolation on the entire reference block. This reduces the computational complexity of the motion estimation process and improves the interpolation efficiency during the interpolation of the first reference block, thus facilitating subsequent targeted interpolation of reference blocks with different levels of complexity.
[0148] Based on the above technical solutions, this disclosure also provides an optional embodiment to improve the selection mechanism of the target frame to be encoded. It should be noted that for parts not described in detail in this disclosure, please refer to the relevant descriptions in other embodiments, which will not be repeated here.
[0149] See Figure 4 The video encoding method shown includes:
[0150] S410. Obtain the frames to be encoded from the video to be encoded.
[0151] The video to be encoded may include at least one frame to be encoded; the frame to be encoded carries the content information of the video frame.
[0152] For example, existing video frame extraction techniques can be used to extract video frames from the video to be encoded, thereby obtaining the frames to be encoded. This disclosure does not limit the specific video frame extraction method.
[0153] S420. Determine the deviation degree of the third pixel of the corresponding block to be encoded based on the pixel values of each pixel in the block to be encoded in the frame to be encoded.
[0154] The block to be encoded can be the result of dividing the frame to be encoded based on the aforementioned second encoding standard and using existing partitioning techniques or methods.
[0155] The deviation of the third pixel is used to characterize the complexity of the content information carried by the block to be encoded. It can be quantified numerically by the pixel variance or pixel standard deviation of the block to be encoded. The larger the deviation of the third pixel, the more complex the content information carried by the block to be encoded; the smaller the deviation of the third pixel, the simpler the content information carried by the block to be encoded.
[0156] For example, the pixel value corresponding to the pixel point of the block to be encoded can be determined; based on the pixel value of each pixel point in the block to be encoded, the pixel variance or pixel standard deviation of the block to be encoded can be determined, and the determination result can be used as the degree of deviation of the third pixel of the block to be encoded.
[0157] S430. Based on the degree of deviation of each third pixel, determine whether to use the frame to be encoded as the target frame to be encoded.
[0158] For example, the frame to which the block to be encoded belongs with a relatively low deviation of the third pixel can be used as the target frame to be encoded.
[0159] It is understandable that the frame to be encoded can be divided into at least one block to be encoded, and each block to be encoded corresponds to the degree of deviation of its third pixel.
[0160] In an optional embodiment, a third deviation threshold may be introduced as a measure of the degree of deviation of the third pixel; accordingly, based on the degree of deviation of each third pixel, it is determined whether to use the frame to be encoded as the target frame to be encoded, including: if the degree of deviation of each third pixel is less than the third deviation threshold, then the frame to be encoded is used as the target frame to be encoded.
[0161] The third deviation threshold can be preset or adjusted by relevant technical personnel, or determined through extensive testing. It should be noted that the third deviation threshold may be the same as or different from the aforementioned first or second deviation threshold, and this disclosure does not impose any limitations on the numerical values of the three.
[0162] Specifically, the deviation degree of the third pixel and the magnitude of the third deviation threshold are determined for each block to be encoded corresponding to the frame to be encoded. If the deviation degree of the third pixel for each block to be encoded is less than the third deviation threshold, then the frame to be encoded is designated as the target frame to be encoded. If the deviation degree of the third pixel for any block to be encoded is not less than the third deviation threshold, then that frame to be encoded can be designated as another frame to be encoded. The other frames to be encoded can be any other frame to be encoded besides the target frame.
[0163] It should be noted that since the deviation of the third pixel of each block to be encoded in the target frame is less than the third deviation threshold, it indicates that the content information carried by each block to be encoded in the target frame is not complex. The subsequent use of the first interpolation method with lower computational complexity for subpixel interpolation as the basis for video encoding will not result in the loss of a large amount of content details. This reduces the amount of data computation in the motion estimation process while taking into account the quality of video encoding.
[0164] This optional embodiment selects target frames with less content information by using frames whose deviation of each third pixel is less than the third deviation threshold as target frames to be encoded. This allows for the selection of target frames to be encoded only for these frames with less content information. In the motion estimation stage, a first interpolation method with lower computational complexity is used for subpixel interpolation, avoiding the degradation of video encoding quality caused by rashly using the first interpolation method for motion estimation on all frames to be encoded. Thus, while maintaining video encoding quality, the amount of data computation in the motion estimation process is reduced.
[0165] In another optional embodiment, if the number of coding blocks in each coding block corresponding to the frame to be encoded where the deviation of the third pixel is less than the third deviation threshold reaches a preset second quantity threshold, then the frame to be encoded can be used as the target frame to be encoded.
[0166] The preset second quantity threshold can be set or adjusted by relevant technical personnel according to needs or experience, or determined based on a large number of experiments. It should be noted that the second quantity threshold and the aforementioned first quantity threshold can be the same or different, and this disclosure does not impose any limitation on their numerical values. For example, if the number of blocks to be encoded corresponding to frame A is 16, and the preset second quantity threshold is 9, then there are 9 blocks whose third pixel deviation is less than the third deviation threshold; therefore, this frame to be encoded can be used as the target frame to be encoded.
[0167] S440. Based on the first interpolation method under the first coding standard, sub-pixel interpolation is performed on the reference frame corresponding to the target frame to be encoded in the video to be encoded to obtain the search plane.
[0168] S450. In the search plane, perform motion estimation on each block to be encoded in the target frame to be encoded to obtain the search motion vector.
[0169] S460. Based on the second interpolation method under the second coding standard, a prediction plane for the target frame to be encoded is generated according to the search motion vector, which is used for video coding based on the second coding standard.
[0170] Among them, the interpolation complexity of the first interpolation method is lower than that of the second interpolation method; the video compression rate of the second coding standard is higher than that of the first coding standard.
[0171] The present invention selects a target frame to be encoded from the frame to be encoded by determining the degree of deviation of the third pixel corresponding to each block to be encoded in the frame to be encoded. This achieves accurate determination of the target frame to be encoded, which facilitates the subsequent targeted use of the first interpolation method under the first encoding standard to process the target frame to be encoded, which carries simple content information. This balances computational efficiency with ensuring the quality of the video encoding result.
[0172] Figure 5 This is a schematic diagram of a video encoding apparatus according to an embodiment of the present disclosure. This embodiment is applicable to application scenarios involving video encoding. The apparatus can be configured in an electronic device and can implement the video encoding method described in any embodiment of the present disclosure. (Reference) Figure 5 The video encoding device 500 specifically includes the following:
[0173] The search plane determination module 501 is used to perform sub-pixel interpolation on the reference frame corresponding to the target frame to be encoded in the video to be encoded based on the first interpolation method under the first encoding standard, so as to obtain the search plane.
[0174] The motion vector determination module 502 is used to perform motion estimation on each block to be encoded in the target frame to be encoded in the search plane to obtain the search motion vector;
[0175] The video encoding module 503 is used to generate a prediction plane of the target frame to be encoded based on the search motion vector according to the second interpolation method under the second encoding standard, and is used to perform video encoding based on the second encoding standard.
[0176] The interpolation complexity of the first interpolation method is lower than that of the second interpolation method; the video compression rate of the second encoding standard is higher than that of the first encoding standard.
[0177] The technical solution of this disclosure adopts a first interpolation method with lower interpolation complexity and less computational load to replace the second interpolation method with higher interpolation complexity and more computational load in the motion estimation process of video encoding under the second coding standard. This reduces the computational complexity in the motion estimation process, thereby reducing the amount of data computation in the motion estimation process. At the same time, the second interpolation method with higher compression ratio under the second coding standard is used in the motion compensation process of video encoding, ensuring the video compression ratio. This achieves a balance between compression ratio and computational load in the video encoding process. While ensuring that the video compression ratio or bitrate effect remains unchanged, the complexity and computational load of the video encoding process are reduced, thereby improving the video encoding efficiency.
[0178] In one optional implementation, the plane to be searched determination module includes:
[0179] A reference block partitioning unit is used to partition the reference frame corresponding to the target frame to be encoded into at least one reference block; the reference block includes a first reference block and a second reference block.
[0180] The first interpolation block determination unit is configured to perform sub-pixel interpolation on the first reference block based on the first interpolation method to obtain a first interpolation block of the corresponding reference block; and
[0181] The second interpolation block determination unit is used to perform sub-pixel interpolation on the second reference block based on the second interpolation method to obtain the second interpolation block of the corresponding reference block;
[0182] The search plane generation unit is used to generate the search plane based on the first interpolation block and the second interpolation block.
[0183] In an optional embodiment, the video encoding device 500 further includes:
[0184] The deviation degree determination module is used to determine the first pixel deviation degree of the reference block based on the pixel value of each pixel in the reference block;
[0185] The reference block division module is used to divide the reference block into a first reference block and a second reference block according to the degree of deviation of the first pixel.
[0186] In one optional implementation, the reference block partitioning module includes:
[0187] The first reference block determination unit is configured to designate reference blocks whose deviation from the first pixel is less than a first deviation threshold as first reference blocks; and...
[0188] The second reference block determination unit is used to determine the reference block whose deviation degree of the first pixel is not less than the first deviation threshold as the second reference block.
[0189] In one optional implementation, the first interpolation block determining unit includes:
[0190] An interpolation block determination subunit is used to perform sub-pixel interpolation on each of the reference blocks based on the first interpolation method to obtain the interpolation block of the corresponding reference block;
[0191] The deviation degree determination subunit is used to determine the second pixel deviation degree of the interpolation block based on the pixel value of each pixel in the interpolation block;
[0192] A reference block division subunit is used to divide a corresponding reference block into a first reference block and a second reference block according to the degree of deviation of the second pixel;
[0193] The first interpolation block determination subunit is used to take the interpolation block corresponding to the first reference block as the first interpolation block.
[0194] In one optional implementation, the reference block partitioning subunit includes:
[0195] The first reference block is determined from the unit, used to take the interpolation block corresponding to the second pixel deviation degree being less than the second deviation threshold as the first reference block; and...
[0196] The second reference block determination unit is used to identify the interpolation block corresponding to the second pixel with a deviation degree not less than the second deviation threshold as the second reference block.
[0197] In an optional embodiment, the video encoding device 500 further includes:
[0198] The frame to be encoded acquisition module is used to acquire the frame to be encoded in the video to be encoded.
[0199] The third deviation degree determination module is used to determine the third pixel deviation degree of the corresponding block to be encoded based on the pixel value of each pixel point of the block to be encoded in the frame to be encoded.
[0200] The target frame to be encoded module is used to determine whether to use the frame to be encoded as the target frame to be encoded based on the degree of deviation of each of the third pixels.
[0201] In an optional implementation, the target frame to be encoded determination module includes:
[0202] The target frame to be encoded is determined by a unit that determines the frame to be encoded as the target frame to be encoded if the deviation of each of the third pixels is less than the third deviation threshold.
[0203] The video encoding apparatus described above can execute the various video encoding methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing each video encoding method.
[0204] The technical solutions disclosed herein involve the collection, storage, use, processing, transmission, provision, and disclosure of the video to be encoded, all of which comply with relevant laws and regulations and do not violate public order and good morals.
[0205] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0206] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0207] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0208] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0209] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as video encoding methods. For example, in some embodiments, the video encoding method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the video encoding method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform video encoding methods by any other suitable means (e.g., by means of firmware).
[0210] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0211] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0212] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0213] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0214] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0215] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem that addresses the management difficulties and weak business scalability inherent in traditional physical hosting and VPS services. Servers can also be servers for distributed systems or servers integrated with blockchain technology.
[0216] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0217] Cloud computing refers to a technology system that enables access to a shared pool of physical or virtual resources via a network. These resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for applications such as artificial intelligence and blockchain, as well as for model training.
[0218] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0219] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for video coding, comprising: performing sub-pixel interpolation on a reference frame corresponding to a target frame to be coded in a video to be coded based on a first interpolation manner under a first coding standard to obtain a search plane; performing motion estimation on each block to be coded in the target frame to be coded in the search plane to obtain a search motion vector; and generating a prediction plane of the target frame to be coded based on the search motion vector according to a second interpolation manner under a second coding standard, wherein the interpolation complexity of the first interpolation manner is lower than that of the second interpolation manner, and the video compression rate of the second coding standard is higher than that of the first coding standard; wherein the performing sub-pixel interpolation on the reference frame corresponding to the target frame to be coded based on the first interpolation manner under the first coding standard to obtain the search plane comprises: dividing the reference frame corresponding to the target frame to be coded into at least one reference block based on the second coding standard, wherein the reference block comprises a first reference block and a second reference block, and the first reference block and the second reference block are obtained by dividing the at least one reference block based on the complexity of the pixel structure corresponding to each reference block; performing sub-pixel interpolation on the first reference block based on the first interpolation manner to obtain a first interpolation block of the corresponding reference block; and performing sub-pixel interpolation on the second reference block based on the second interpolation manner to obtain a second interpolation block of the corresponding reference block; and generating the search plane according to the first interpolation block and the second interpolation block. 2.The method of claim 1, further comprising: determining a first pixel deviation degree of each pixel point in the reference block according to the pixel value of the pixel point; and dividing the reference block into the first reference block and the second reference block according to the first pixel deviation degree. The dividing the reference block into the first reference block and the second reference block according to the first pixel deviation degree comprises: taking the reference block with the first pixel deviation degree less than a first deviation threshold as the first reference block; and taking the reference block with the first pixel deviation degree not less than the first deviation threshold as the second reference block. The performing sub-pixel interpolation on the first reference block based on the first interpolation manner to obtain the first interpolation block of the corresponding reference block comprises: performing sub-pixel interpolation on each reference block based on the first interpolation manner to obtain an interpolation block of the corresponding reference block; determining a second pixel deviation degree of each pixel point in the interpolation block according to the pixel value of the pixel point; dividing the corresponding reference block into the first reference block and the second reference block according to the second pixel deviation degree; and taking the interpolation block corresponding to the first reference block as the first interpolation block. The dividing the corresponding reference block into the first reference block and the second reference block according to the second pixel deviation degree comprises: taking the interpolation block corresponding reference block with the second pixel deviation degree less than a second deviation threshold as the first reference block; and taking the interpolation block corresponding reference block with the second pixel deviation degree not less than the second deviation threshold as the second reference block. 3. The method of claim 2, wherein, 4. The method of claim 1, wherein, 5. The method of claim 4, wherein, The second reference block corresponding to the interpolation block whose second pixel deviation degree is not less than a second deviation threshold is taken as a second reference block.
6. The method of any one of claims 1-5, further comprising: obtaining a to-be-encoded frame in the to-be-encoded video; determining a third pixel deviation degree of each to-be-encoded block in the to-be-encoded frame according to pixel values of each pixel point in the to-be-encoded block; determining whether to take the to-be-encoded frame as the target to-be-encoded frame according to the third pixel deviation degrees.
7. The method of claim 6, wherein, The determining whether to take the to-be-encoded frame as the target to-be-encoded frame according to the third pixel deviation degrees comprises: if each third pixel deviation degree is less than a third deviation threshold, taking the to-be-encoded frame as the target to-be-encoded frame.
8. A video encoding apparatus, comprising: a to-be-searched plane determination module configured to perform sub-pixel interpolation on a reference frame corresponding to a target to-be-encoded frame in a to-be-encoded video based on a first interpolation manner under a first encoding standard to obtain a to-be-searched plane; a motion vector determination module configured to perform motion estimation on each to-be-encoded block in the target to-be-encoded frame in the to-be-searched plane to obtain a search motion vector; a video encoding module configured to generate a prediction plane of the target to-be-encoded frame based on a second interpolation manner under a second encoding standard according to the search motion vector, and perform video encoding based on the second encoding standard; wherein an interpolation complexity of the first interpolation manner is lower than an interpolation complexity of the second interpolation manner, and a video compression rate of the second encoding standard is higher than a video compression rate of the first encoding standard; wherein the to-be-searched plane determination module comprises: a reference block division unit configured to divide the reference frame corresponding to the target to-be-encoded frame into at least one reference block based on the second encoding standard, wherein the reference block comprises a first reference block and a second reference block, and the first reference block and the second reference block are obtained by dividing at least one reference block based on a pixel structure complexity of each reference block; a first interpolation block determination unit configured to perform sub-pixel interpolation on the first reference block based on the first interpolation manner to obtain a first interpolation block of the corresponding reference block; and a second interpolation block determination unit configured to perform sub-pixel interpolation on the second reference block based on the second interpolation manner to obtain a second interpolation block of the corresponding reference block; a to-be-searched plane generation unit configured to generate the to-be-searched plane according to the first interpolation block and the second interpolation block.
9. The apparatus of claim 8, further comprising: a deviation degree determination module configured to determine a first pixel deviation degree of each reference block in the reference block according to pixel values of each pixel point in the reference block; a reference block division module configured to divide the reference block into the first reference block and the second reference block according to the first pixel deviation degree.
10. The apparatus of claim 9, wherein, The reference block division module comprises: a first reference block determination unit configured to take a reference block whose first pixel deviation degree is less than a first deviation threshold as the first reference block; and The second reference block determination unit is configured to determine, as the second reference block, a reference block whose second pixel deviation degree is not less than a second deviation threshold.
11. The apparatus of claim 8, wherein, The first interpolation block determination unit comprises: The interpolation block determination subunit is configured to perform sub-pixel interpolation on each reference block based on the first interpolation mode to obtain an interpolation block of the corresponding reference block; The deviation degree determination subunit is configured to determine the second pixel deviation degree of the interpolation block according to pixel values of each pixel point in the interpolation block; The reference block division subunit is configured to divide the corresponding reference block into the first reference block and the second reference block according to the second pixel deviation degree; The first interpolation block determination subunit is configured to determine, as the first interpolation block, a corresponding interpolation block of the first reference block.
12. The apparatus of claim 11, wherein, The reference block division subunit comprises: The first reference block determination subunit is configured to determine, as the first reference block, a reference block corresponding to an interpolation block whose second pixel deviation degree is less than a second deviation threshold; and The second reference block determination subunit is configured to determine, as the second reference block, a reference block corresponding to an interpolation block whose second pixel deviation degree is not less than the second deviation threshold.
13. The apparatus according to any one of claims 8-12, further comprising: The to-be-encoded frame acquisition module is configured to acquire a to-be-encoded frame in the to-be-encoded video; The third deviation degree determination module is configured to determine a third pixel deviation degree of each to-be-encoded block in the to-be-encoded frame according to pixel values of each pixel point in the to-be-encoded block; The target to-be-encoded frame determination module is configured to determine whether to determine the to-be-encoded frame as the target to-be-encoded frame according to each third pixel deviation degree.
14. The apparatus of claim 13, wherein, The target to-be-encoded frame determination module comprises: The target to-be-encoded frame determination unit is configured to determine the to-be-encoded frame as the target to-be-encoded frame if each third pixel deviation degree is less than a third deviation threshold.
15. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the video encoding method in any one of claims 1-7.
16. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable a computer to perform the video encoding method in any one of claims 1-7.
17. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the video encoding method in any one of claims 1-7.
Citation Information
Patent Citations
Motion estimation engine with parallel interpolation and search hardware
US20040120401A1