ERP panoramic video VVC intra-frame mode fast prediction method and storage medium
By dividing panoramic video frames into different latitude areas and combining MPM lists and texture calculations to optimize intra-frame mode selection, the problem that the VVC fast CU partitioning algorithm fails to effectively combine the panoramic video characteristics of the ERP format is solved, thereby reducing coding complexity and improving coding efficiency.
Patent Information
- Application Number
- CN202410734794.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-06-07
AI Technical Summary
The existing VVC fast CU partitioning algorithm fails to effectively combine the characteristics of ERP format panoramic video, resulting in high complexity of panoramic video encoding and serious redundant calculation, which affects the encoding efficiency.
Based on the sampling characteristics of ERP panoramic video, the video frames are divided into different latitude regions. By statistically analyzing the distribution of intra-frame mode selection, the MPM list and texture calculation are used to filter redundant modes and optimize the intra-frame prediction mode selection.
It significantly reduces the encoding time of ERP panoramic videos, reduces encoding complexity and cost, while keeping the encoding quality loss to a minimum, and is suitable for efficient encoding of panoramic videos.
Smart Images

Figure CN118694967B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video coding, and more specifically, provides a fast intra-frame mode prediction partitioning method suitable for Versatile Video Coding (VVC), which can be applied to panoramic video coding scenarios in equirectangular projection (ERP) format. Background Art
[0002] With the growing demand for high-quality video applications, high-definition video, such as 4K and 8K, has become mainstream. Panoramic video, a type of dynamic video captured in all directions using a 3D camera, is widely used in smart tourism, live sports broadcasts, and large-scale entertainment events. Panoramic video offers a highly immersive and realistic experience, but it also requires a large amount of data, significantly increasing the transmission and storage costs of its applications. To further improve video coding efficiency, in July 2020, the International Organization for Standardization / International Electrotechnical Commission and the International Telecommunication Union's Telecommunication Standardization Sector jointly developed the next-generation Versatile Video Coding (VVC) standard. VVC utilizes advanced coding techniques, including a nested multitree quadtree partitioning structure, 67 directional intra-frame prediction modes, a merge mode with motion vectors, geometric partitioning inter-frame prediction, and joint intra-frame and inter-frame prediction, significantly improving video coding efficiency. This makes it ideal for the compression of high-definition video, especially panoramic video. Compared to the previous generation coding standard, High Efficiency Video Coding (HEVC), VVC improves coding efficiency by nearly 50%, but also significantly increases coding complexity. Researching fast algorithms that can significantly reduce VVC's computational complexity is currently a key focus for many researchers.
[0003] With the advancement of computer vision technology, 360° panoramic virtual videos have emerged. Traditional videos, due to video format size limitations, can only display a single image within a single frame. Panoramic videos, on the other hand, break the boundaries of this frame, enveloping the viewer in the entire image, allowing them to observe the captured scene from all angles. The panoramic video mapping process generates a large amount of image data, necessitating compression of the panoramic images to conserve storage space and transmission bandwidth. Before transmission, panoramic videos must be converted into rectangular planar videos—that is, 3D images are mapped onto a 2D plane. These videos are then encoded using video coding techniques before data transmission. With the promulgation of the next-generation video coding standard (VVC), the Video Coding Experts Group and the Moving Picture Experts Group released a Call for Proposals (CFP) for next-generation video compression technologies. This document lists several mapping methods for converting panoramic videos into planar videos, such as equirectangular projection (ERP), cubic projection, and icosahedron projection. ERP is the most widely used projection format and currently the primary projection format for 360° panoramic videos. ERP maps the longitudes of a panoramic video sphere into equally spaced vertical lines and the latitudes into equally spaced horizontal lines, resulting in an aspect ratio of 2:1. It uses a completely linear transformation formula, making the mapping simple and easy to implement. The resulting two-dimensional image is also relatively intuitive and easy to observe. However, the original spherical video uses the same number of sampling points on each latitude, resulting in more redundant sampling points closer to the poles. At the poles, where only a small number of sampling points are required, the same number of sampling points as at the equator are used, causing pixels in the polar regions to be stretched infinitely, resulting in oversampling. This uneven pixel density distribution leads to severe image distortion and a large amount of redundant data. Compared to traditional video, panoramic video has a higher resolution. While high-resolution panoramic video provides a breathtaking visual experience, it also consumes a large amount of data, placing a significant burden on its transmission and storage. Effectively compressing this massive amount of data, along with the significant amount of redundant data introduced by the mapping, is a primary challenge in the field of panoramic video applications.
[0004] The VVC fast CU partitioning algorithm has not been effectively combined with the ERP format panoramic video characteristics, and there is room for further reduction in coding complexity. VVC is the latest video coding standard with higher coding efficiency. However, panoramic videos contain a large amount of data due to their ultra-high-definition resolution. Therefore, it is necessary to effectively combine the ERP format panoramic video characteristics with the VVC fast CU partitioning algorithm to achieve fast coding of panoramic videos. ERP format panoramic videos have the characteristics of low sampling density in the central area of the image and high sampling density in the polar areas. In view of this characteristic, removing the redundant parts of the mapping with minimal loss of video quality and saving its coding complexity are problems that need to be solved in current panoramic video coding research.
[0005] To address the problem of redundant computational complexity in intra-frame prediction of panoramic videos, the present invention proposes a fast intra-frame mode selection algorithm based on the characteristics of different latitude regions. First, the selection distribution of intra-frame modes in the polar regions is statistically analyzed, and the number of mode lists is determined based on the modes in the MPM list. For the mid-latitude region, the selection distribution of horizontal and vertical angle mode predictions is statistically analyzed, and redundant vertical angle prediction modes are pre-screened through texture calculation and candidate mode list information. For the equatorial region, the last intra-frame angle prediction mode in the candidate mode list is screened out by comparing texture features in four directions.
[0006] The closest prior art is a prior application, publication number CN117041736A, by the same inventor as the present application. This application discloses a method and storage medium for fast CU partitioning of ERP panoramic video (VVC), belonging to the field of video coding. The method comprises the following steps: utilizing the sampling characteristics of ERP panoramic video to divide the coded frame into different latitudinal regions; making an early termination decision on the current CU partitioning mode based on the distribution characteristics of the CU quadtree depths in different latitudinal regions and the correlation between adjacent CUs; for CUs to be further partitioned, using gradient differences to evaluate the current CU texture characteristics, skipping redundant horizontal or vertical partitioning modes; and for CUs with blurred textures, using a weighted quadratic comparison based on latitudinal sampling weights to determine whether to skip the vertical partitioning mode; finally, using two-dimensional Haar wavelet transform coefficients to evaluate the differences between sub-CUs and determine whether to skip the ternary tree partitioning mode. This method utilizes the characteristics of ERP panoramic video to rapidly partition VVC intra-frame coded CUs. The present invention, however, utilizes these characteristics to rapidly determine the intra-frame prediction mode of a CU. Combining these two methods can significantly reduce the complexity of ERP panoramic video VVC intra-frame coding and lower the implementation and operating costs of video encoders. Summary of the Invention
[0007] The present invention aims to solve the above problems in the prior art. A method and storage medium for fast intra-frame mode prediction of ERP panoramic video VVC is proposed. The technical solution of the present invention is as follows:
[0008] A method for fast prediction of intra-frame modes of ERP panoramic video VVC includes the following steps:
[0009] S1. Based on the sampling characteristics of ERP panoramic video, the video encoding frames are divided into five regions according to their altitude: polar region, mid-latitude region and equatorial region;
[0010] S2. Determine the region to which the current CU (coding unit) belongs; if it is the polar region, execute step S3; if it is the mid-latitude region, execute step S6; if it is the equatorial region, execute step S9;
[0011] S3. Determine whether the first two prediction modes in the current CU traditional prediction angle mode list include DC (direct current) or Planar (planar) mode; if yes, execute step S4; if no, execute step S12;
[0012] S4. Determine whether the first 1 / 2 modes in the MPM (most probable mode) list of the current CU include DC or Planar mode; if yes, go to step S5; if no, go to step S12;
[0013] S5. Clear the rate-distortion cost candidate mode list, add the modes included in the first 1 / 2 of the MPM and the first two digits of the CU candidate mode list, and execute step S11.
[0014] S6. Calculate the horizontal and vertical gradient values of the current CU and compare them to determine whether the CU is approaching the horizontal texture; if so, proceed to step S7; otherwise, execute step S12;
[0015] S7, determine whether the first intra-frame angle prediction mode in the current CU candidate list belongs to the horizontal area; if so, go to step S8; otherwise, go to step S12;
[0016] S8, traverse the rate-distortion cost list in sequence, eliminate the angle prediction mode belonging to the vertical area, and execute step S11;
[0017] S9. Calculate the gradients of the four directions of the current CU, record the area where the minimum directional gradient is located, and execute step S10.
[0018] S10, after the cost calculation of the two rounds of coarse selection mode and the merging of the MPM list, determine whether the last angle prediction mode in the list is in the area where the minimum directional gradient is located; if not, remove it from the candidate list mode and execute step S11; if it is, execute step S12;
[0019] S11, calculating the rate-distortion cost of the filtered list and selecting the best intra-frame prediction mode;
[0020] S12. Perform intra-frame prediction according to the VTM (VVC Test Reference Model) original platform algorithm to obtain the optimal intra-frame prediction mode for the current CU.
[0021] Furthermore, the step S1 divides the video encoding frames into five regions, namely, polar regions, mid-latitude regions, and equatorial regions, according to their altitude based on the sampling characteristics of the ERP panoramic video, specifically including:
[0022] The video frames are divided into five regions according to their altitude: the polar regions, mid-latitude regions, and the equatorial region. The 3D panoramic video is converted into a 2D video encoding format through ERP mapping. Based on the different sampling and stretching degrees of different regions of the video frame, the video frames are divided into five regions from top to bottom: the polar regions, mid-latitude regions, equatorial region, mid-latitude regions, and the polar regions. The altitudes of the North and South Poles, the two mid-latitude regions, and the equatorial region account for 24%, 36%, and 40% of the vertical resolution of the video frame, respectively.
[0023] Furthermore, in the step S3, S3 determines whether the first two prediction modes in the current CU traditional prediction angle mode list include DC (direct current) or Planar (planar) mode, specifically: the CU traditional prediction angle mode is divided into 65 angle prediction modes and 2 non-angle prediction modes, among which the non-angle prediction mode is divided into Planar mode and DC mode. The Planar mode is suitable for areas where pixel values change slowly, and its predicted pixels are regarded as the average values of the prediction values in the horizontal and vertical directions; the DC mode calculates the average value of the reference pixels on the left or above the current CU, and selects different average reference pixel rows as predicted pixel values according to their different sizes, which is more suitable for large flat areas.
[0024] Furthermore, in step S4, the MPM (most probable mode) list of the current CU is constructed using the correlation between adjacent block modes. The MPM list technology utilizes the following steps: first, the 35 prediction modes identical to those in HEVC are traversed, and the candidate list is updated by comparing the Hadamard costs; then, the modes in the candidate list are traversed, and their left and right adjacent modes are obtained, and the Hadamard costs are calculated and further compared with the previously selected mode. The Hadamard cost is calculated using a Lagrangian optimization method, and its calculation formula is shown in (1).
[0025] J=D SAD +λ×Bit mode (1)
[0026] J in formula (1) represents the Hadamard value calculated using the current prediction model; D SADRepresents the absolute error sum between the original block and the reconstructed block; Bit mode Represents the bit rate required to encode the current mode.
[0027] After two rounds of Hadamard cost calculations to determine the lowest-cost candidate angular prediction modes, VVC constructs the MPM list. In H.266 / VVC, the MPM list contains six prediction modes, with the selection of the prediction mode primarily based on the CU above and to the left of the current CU.
[0028] Furthermore, in step S6, the horizontal and vertical gradient values of the current CU are calculated by using two Scharr operators with 3×3 convolution kernels to calculate the vertical and horizontal gradients of the CU pixels respectively. The calculation formula is as follows:
[0029]
[0030] G in formula (2) h (x,y) and G v (x,y) represents the horizontal gradient and vertical gradient of the pixel at coordinate (x,y), respectively; A(x,y) represents the pixel matrix centered at coordinate (x,y).
[0031] Furthermore, in step S6, whether the CU approaches the horizontal texture is determined by calculating the average gradient of the current CU pixel in the horizontal and vertical directions using formula (3).
[0032]
[0033] |G in formula (3) h (x,y)| and|G h (x,y)|respectively the absolute value of the horizontal gradient and vertical gradient of the pixel at the (i,j) coordinate; Gradient h and Gradient v are the horizontal average gradient and vertical average gradient of the current CU respectively; if Gradient h / Gradient v ≤0.9, the horizontal gradient is significantly smaller than the vertical gradient, indicating that the current CU approaches the horizontal texture and the CU tends to use the horizontal area angle prediction mode as the optimal intra prediction mode.
[0034] Furthermore, in step S7, whether it belongs to a horizontal area is determined by determining that the CU whose optimal angle prediction mode is selected within the range of 10-26 is a horizontal area;
[0035] In step S8, the CU whose optimal angle prediction mode is selected within the range of 42-58 is determined to be a vertical area.
[0036] Furthermore, in step S9, the gradients in the four directions of the current CU are calculated, that is, the gradients in the horizontal, vertical, 45° and 135° directions of the current CU are calculated; the gradients in the horizontal and vertical directions of the current CU are calculated using formulas (2) and (3), and the gradient operators in the 45° and 135° directions are obtained by rotating the horizontal and vertical convolution kernel operators, and their calculation is shown in formula (4).
[0037]
[0038] G in formula (4) 45 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 45° direction; A(x,y) represents the pixel matrix centered at coordinate (x,y); G 135 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 135° direction; the gradient calculation of the 45° and 135° directions of the current CU is shown in formula (5);
[0039]
[0040] |G in formula (5) 45 (x,y)| and|G 135 (x, y)| is the absolute value of the 45° and 135° gradient of the CU at the (i, j) coordinate; W and H are the width and height of the CU respectively; Gradient 45 and Gradient 135 are the average gradients of the current CU in the 45° and 135° directions respectively.
[0041] Furthermore, in step S10, it is determined whether it is in the area where the minimum directional gradient is located, where the area includes four areas: horizontal, vertical, 45° and 135°; the 65 angle prediction modes of VVC are divided into the above four areas; among them, the angle prediction modes of 10-26 are divided into the horizontal area; the angle prediction modes of 27-41 are divided into the 135° area; the angle prediction modes of 42-58 are divided into the vertical area; and the angle prediction modes of 2-9 and 59-66 are divided into the 45° area.
[0042] A storage medium stores a computer program therein, wherein when the computer program is read by a processor, any of the above-mentioned ERP panoramic video VVC intra-frame mode fast prediction methods is executed.
[0043] The advantages and beneficial effects of the present invention are as follows:
[0044] To address the problem of redundant computational complexity in intra-frame prediction of VVC panoramic videos, the present invention proposes a fast intra-frame mode selection algorithm based on the characteristics of different latitude regions. First, the intra-frame mode selection distribution in the polar regions is statistically analyzed, and the mode selection distribution is judged based on the modes in the MPM list to determine whether to reduce the number of mode lists; for the mid-latitude region, the selection distribution of horizontal and vertical angle mode predictions is statistically analyzed, and redundant vertical angle prediction modes are screened out in advance through texture calculation and candidate mode list information; for the equatorial region, the last intra-frame angle prediction mode in the candidate mode list is screened out by comparing the texture features in four directions.
[0045] In order to verify the encoding acceleration performance of the proposed method, the proposed method was compared with the encoding acceleration comprehensive performance of the VTM-8.0-360Lib-10.1 platform. The peak signal-to-noise ratio (WS-PSNR) based on spherical weights, time saving TS (Time Saving) and is a performance evaluation indicator. BDBR represents the bitrate savings compared to the original platform at the same quality. Its value also indirectly reflects the accuracy of the CU division of the fast algorithm. The larger the BDBR value, the greater the coding efficiency loss. TS represents the percentage improvement in encoding time using the proposed method compared to the encoding time of the original platform. Its calculation formula is as follows:
[0046]
[0047] Where T org (QP i ) indicates that the original platform is at QP = QP i The encoding time is T pro (QP i ) indicates that the algorithm proposed in this paper is used when QP=QP i The encoding time is .
[0048] Table 1 shows the encoding performance of 14 panoramic video test sequences. The data in Table 1 shows that compared to the original VVC algorithm, the fast algorithm proposed in this invention can reduce encoding time by an average of 9.47%, increase BDBR by only 0.15%, and reduce WS-PSNR by only 0.007dB. Among them, the test sequence "Harbor" saves the least encoding time, 7%. The test sequence "Broadway" saves the most encoding time, reaching 11.1%. The difference in encoding time savings between the two sequences is not significant, demonstrating that the algorithm proposed in this invention has good robustness and a relatively stable acceleration effect on all sequences. The sequence with the largest bit rate increase is "AerialCity", with an increase of only 0.21%, indicating that compared to the original platform algorithm, the algorithm proposed in this invention has little loss in encoding efficiency. The present invention can be used for accelerated encoding under the VVC intra-frame coding configuration. While ensuring minimal reduction in coding efficiency and coding quality, it can effectively reduce the VVC encoding time and reduce the implementation and operating costs of the video encoder. It can be applied to application scenarios with high real-time requirements for ERP format panoramic videos.
[0049] Table 1
[0050]
[0051] The proposed method is combined with the fast CU partitioning method (proposed by the inventors in their prior patent application publication number CN117041736A) for intra-frame encoding of ERP panoramic videos using VVC. The acceleration effect is shown in Table 2. The average encoding time is reduced by 49.26%, almost halving the encoding time. The average bitrate increases by only 1.02%, and the WS-PSNR loss is only 0.049 dB, demonstrating that the cost of using the acceleration algorithm for the video encoder is minimal.
[0052] Table 2
[0053]
[0054] BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 The present invention provides a flowchart of a fast algorithm for VVC intra-frame mode prediction based on ERP panoramic video latitude and texture characteristics in a preferred embodiment;
[0056] Figure 2 This is an example of dividing a video frame into different latitude regions based on the sampling characteristics of the ERP panoramic video;
[0057] Figure 3 This is an example diagram of VVC intra-frame prediction mode. DETAILED DESCRIPTION
[0058] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.
[0059] The technical solution of the present invention to solve the above technical problems is:
[0060] Figure 1 The present invention provides a fast VVC intra-frame mode prediction algorithm and storage medium based on ERP panoramic video latitude and texture characteristics. The method of the present invention includes the following steps:
[0061] S1. Based on the sampling characteristics of ERP panoramic video, the video encoding frames are divided into five regions according to their altitude: polar region, mid-latitude region and equatorial region.
[0062] S2. Determine the region to which the current CU (coding unit) belongs. If it is a polar region, execute step S3; if it is a mid-latitude region, execute step S6; if it is an equatorial region, execute step S9.
[0063] S3. Determine whether the first two prediction modes in the current CU's traditional prediction angle mode list include DC (direct current) or Planar (planar) mode. If so, proceed to step S4; if not, proceed to step S12.
[0064] S4. Determine whether the first half of the MPM list of the current CU contains DC or Planar mode. If so, proceed to step S5; if not, proceed to step S12.
[0065] S5. Clear the rate-distortion cost candidate mode list, add the modes included in the first 1 / 2 modes of the MPM and the first two digits of the CU candidate mode list, and execute step S11.
[0066] S6. Calculate the horizontal and vertical gradient values of the current CU and compare them to determine whether the CU is approaching horizontal texture. If so, proceed to step S7; otherwise, proceed to step S12.
[0067] S7: Determine whether the first intra-frame angle prediction mode in the current CU candidate list belongs to the horizontal region. If so, proceed to step S8; otherwise, proceed to step S12.
[0068] S8. Traverse the rate-distortion cost list in sequence, eliminate the angle prediction modes belonging to the vertical area, and execute step S11.
[0069] S9. Calculate the gradients of the four directions of the current CU, record the area where the minimum directional gradient is located, and execute step S10.
[0070] S10. After calculating the costs of the two rounds of coarse selection modes and merging the MPM list, determine whether the last angle prediction mode in the list is in the region of minimum directional gradient. If not, remove it from the candidate list and proceed to step S11. If so, proceed to step S12.
[0071] S11. Calculate the rate-distortion cost for the filtered list and select the best intra-frame prediction mode.
[0072] S12. Perform intra-frame prediction according to the VTM (VVC Test Reference Model) original platform algorithm to obtain the optimal intra-frame prediction mode for the current CU.
[0073] Preferably, in step S1, the video frame is divided into five regions according to its altitude: the polar region, the mid-latitude region, and the equatorial region. The 3D panoramic video is converted into a 2D video encoding format through ERP mapping. Based on the different sampling and stretching degrees of different regions of the video frame, the video frame is divided into five regions from top to bottom: the polar region, the mid-latitude region, the equatorial region, the mid-latitude region, and the polar region. The altitudes of the North and South Poles, the two mid-latitude regions, and the equatorial region account for 24%, 36%, and 40% of the vertical resolution of the video frame, respectively.
[0074] Preferably, in step S3, the CU traditional prediction angle mode is divided into 65 angle prediction modes and 2 non-angle prediction modes, wherein the non-angle prediction mode is divided into Planar mode and DC mode. The Planar mode is suitable for areas where pixel values change slowly, and its predicted pixels can be regarded as the average value of the predicted values in the horizontal and vertical directions. The DC mode calculates the average value of the reference pixels to the left or above the current CU, and selects different average reference pixel rows as the predicted pixel values according to their different sizes, which is more suitable for large flat areas.
[0075] Preferably, in step S4, the MPM (most probable mode) list of the current CU is constructed by using the correlation between adjacent block modes. The MPM list technology utilizes the following steps: first, traverse the same 35 prediction modes as HEVC, and update the candidate list by comparing the Hadamard cost; then, traverse the modes in the candidate list and obtain their left and right adjacent modes, calculate the Hadamard cost, and further compare it with the previously selected mode. The Hadamard cost is calculated using a Lagrangian optimization method, and its calculation formula is shown in (1).
[0076] J=D SAD +λ×Bit mode (1)
[0077] J in formula (1) represents the Hadamard value calculated using the current prediction model; D SADRepresents the absolute error sum between the original block and the reconstructed block; Bit mode Represents the bit rate required to encode the current mode.
[0078] After two rounds of Hadamard cost calculations to determine the lowest-cost candidate angular prediction modes, VVC constructs the MPM list. In H.266 / VVC, the MPM list contains six prediction modes, with the selection of the prediction mode primarily based on the CU above and to the left of the current CU.
[0079] Preferably, in step S6, the horizontal and vertical gradient values of the current CU are calculated by using two Scharr operators with 3×3 convolution kernels to calculate the vertical and horizontal gradients of the CU pixels respectively, and the calculation formula is as follows:
[0080]
[0081] G in formula (2) h (x,y) and G v (x,y) represents the horizontal gradient and vertical gradient of the pixel at coordinate (x,y), respectively; A(x,y) represents the pixel matrix centered at coordinate (x,y).
[0082] Preferably, in step S6, whether the CU approaches the horizontal texture is determined by calculating the average gradient of the current CU pixel in the horizontal and vertical directions using formula (3).
[0083]
[0084] |G in formula (3) h (x,y)| and|G h (x,y)|respectively the absolute value of the horizontal gradient and vertical gradient of the pixel at the (i,j) coordinate; Gradient h and Gradient v are the horizontal average gradient and vertical average gradient of the current CU respectively. h / Gradient v ≤0.9, the horizontal gradient is significantly smaller than the vertical gradient, indicating that the current CU approaches the horizontal texture and the CU tends to use the horizontal area angle prediction mode as the optimal intra prediction mode.
[0085] Preferably, in step S7, whether it belongs to the horizontal area is determined by determining that the CU whose optimal angle prediction mode is selected within the range of 10-26 is the horizontal area.
[0086] Preferably, in step S8, the CU that has the best angle prediction mode selected within the range of 42-58 is determined as a vertical area.
[0087] Preferably, in step S9, the gradients in the four directions of the current CU are calculated, that is, the gradients in the horizontal, vertical, 45°, and 135° directions of the current CU are calculated. Formulas (2) and (3) are used to calculate the gradients in the horizontal and vertical directions of the current CU, and the gradient operators in the 45° and 135° directions are obtained by rotating the horizontal and vertical convolution kernel operators, and their calculation is shown in formula (4).
[0088]
[0089] G in formula (4) 45 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 45° direction; A(x,y) represents the pixel matrix centered at coordinate (x,y); G 135 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 135° direction. The gradient calculation of the 45° and 135° directions of the current CU is shown in formula (5).
[0090]
[0091] |G in formula (5) 45 (x,y)| and|G 135 (x, y)| is the absolute value of the 45° and 135° gradient of the CU at the (i, j) coordinate; W and H are the width and height of the CU respectively; Gradient 45 and Gradient 135 are the average gradients of the current CU in the 45° and 135° directions respectively.
[0092] Preferably, in step S10, it is determined whether the angle is in the region where the minimum directional gradient is located. Here, the region includes four regions: horizontal, vertical, 45°, and 135°. The 65 angle prediction modes of VVC are divided into the above four regions. Among them, the angle prediction modes of 10-26 are divided into the horizontal region; the angle prediction modes of 27-41 are divided into the 135° region; the angle prediction modes of 42-58 are divided into the vertical region; and the angle prediction modes of 2-9 and 59-66 are divided into the 45° region.
[0093] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0094] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0096] The above embodiments should be understood as merely illustrating the present invention and not as limiting the scope of protection of the present invention. After reading the contents of the present invention, technicians may make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A fast prediction method for intra-frame mode of ERP panoramic video VVC, characterized by: The following steps are involved: S1. Based on the sampling characteristics of ERP panoramic video, the video encoding frames are divided into five regions according to their altitude: polar region, mid-latitude region and equatorial region; S2. Determine the region to which the current CU coding unit belongs; if it is the polar region, execute step S3; if it is the mid-latitude region, execute step S6; if it is the equatorial region, execute step S9; S3. Determine whether the first two prediction modes in the current CU traditional prediction angle mode list include DC or Planar mode; if yes, execute step S4; if no, execute step S12; S4. Determine whether the first 1 / 2 modes of the most probable MPM mode list of the current CU include DC or Planar mode; if so, execute step S5; if not, execute step S12; S5. Clear the rate-distortion cost candidate mode list, add the modes included in the first 1 / 2 of the MPM and the first two digits of the CU candidate mode list, and execute step S11. S6. Calculate the horizontal and vertical gradient values of the current CU and compare them to determine whether the CU is approaching the horizontal texture; if so, proceed to step S7; otherwise, execute step S12; S7, determine whether the first intra-frame angle prediction mode in the current CU candidate list belongs to the horizontal area; if so, go to step S8; otherwise, go to step S12; S8, traverse the rate-distortion cost list in sequence, eliminate the angle prediction mode belonging to the vertical area, and execute step S11; S9. Calculate the gradients of the four directions of the current CU, record the area where the minimum directional gradient is located, and execute step S10. S10, after the cost calculation of the two rounds of coarse selection modes and the merging of the MPM list, determining whether the last angle prediction mode in the list is in the area where the minimum directional gradient is located; If it does not belong, then remove it from the candidate list mode and execute step S11; If yes, go to step S12; S11, calculating the rate-distortion cost of the filtered list and selecting the best intra-frame prediction mode; S12. Perform intra-frame prediction according to the VTM (VVC Test Reference Model) original platform algorithm to obtain the optimal intra-frame prediction mode for the current CU.
2. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1 is characterized in that: The step S1 divides the video encoding frames into five regions, namely the polar region, the mid-latitude region and the equatorial region, according to their altitude based on the sampling characteristics of the ERP panoramic video. Specifically, the steps include: The video frames are divided into five regions according to their altitude: the polar regions, mid-latitude regions, and the equatorial region. The 3D panoramic video is converted into a 2D video encoding format through ERP mapping. Based on the different sampling and stretching degrees of different regions of the video frame, the video frames are divided into five regions from top to bottom: the polar regions, mid-latitude regions, equatorial region, mid-latitude regions, and the polar regions. The altitudes of the North and South Poles, the two mid-latitude regions, and the equatorial region account for 24%, 36%, and 40% of the vertical resolution of the video frame, respectively.
3. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1 is characterized in that: In the step S3, S3 determines whether the first two prediction modes in the current CU traditional prediction angle mode list include DC or Planar mode. Specifically, the CU traditional prediction angle mode is divided into 65 angle prediction modes and 2 non-angle prediction modes, wherein the non-angle prediction mode is divided into Planar mode and DC mode. The Planar mode is suitable for areas where pixel values change slowly, and its predicted pixels are regarded as the average values of the prediction values in the horizontal and vertical directions; the DC mode calculates the average value of the reference pixels on the left or above the current CU, and selects different average reference pixel rows as the predicted pixel values according to their different sizes, which is more suitable for large flat areas.
4. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1 is characterized in that: In step S4, the MPM most probable mode list of the current CU is constructed. The MPM list technology utilizes the correlation between adjacent block modes. The construction steps are as follows: first, traverse the 35 prediction modes that are the same as HEVC, and update the candidate list by comparing the Hadamard cost; then, traverse the modes in the candidate list, obtain their left and right adjacent modes, calculate the Hadamard cost and further compare it with the previously selected mode; the Hadamard cost adopts the calculation method based on the Lagrangian optimization method, and its calculation formula is shown in (1); J=D SAD +λ×Bit mode (1) J in formula (1) represents the Hadamard value calculated using the current prediction model; D SAD Represents the absolute error sum between the original block and the reconstructed block; Bit mode Represents the bit rate required to encode the current mode; After two rounds of Hadamard cost calculation, several candidate angle prediction modes with the lowest cost are obtained, and VVC constructs the MPM list; in H.266 / VVC, the MPM list contains 6 prediction modes, and the selection of the prediction mode is mainly based on the CU above and the CU to the left of the current CU.
5. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1 is characterized in that: In step S6, the horizontal and vertical gradient values of the current CU are calculated by using two Scharr operators with 3×3 convolution kernels to calculate the vertical and horizontal gradients of the CU pixels respectively. The calculation formula is as follows: G in formula (2) h (x,y) and G v (x,y) represents the horizontal gradient and vertical gradient of the pixel at coordinate (x,y), respectively; A(x,y) represents the pixel matrix centered at coordinate (x,y).
6. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 5 is characterized in that: In step S6, whether the CU approaches the horizontal texture is determined by calculating the average gradient of the current CU pixel in the horizontal and vertical directions using formula (3): |G in formula (3) h (x,y)| and|G h (x,y)|respectively the absolute value of the horizontal gradient and vertical gradient of the pixel at the (i,j) coordinate; Gradient h and Gradient v are the horizontal average gradient and vertical average gradient of the current CU respectively; if Gradient h / Gradient v ≤0.9, the horizontal gradient is significantly smaller than the vertical gradient, indicating that the current CU approaches the horizontal texture and the CU tends to use the horizontal area angle prediction mode as the optimal intra prediction mode.
7. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1, characterized in that: In step S7, whether it belongs to a horizontal area is determined by determining that the CU whose optimal angle prediction mode is selected within the range of 10-26 is a horizontal area; In step S8, the CU whose optimal angle prediction mode is selected within the range of 42-58 is determined to be a vertical area.
8. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 6 is characterized in that: In step S9, the gradients of the four directions of the current CU are calculated, that is, the gradients of the horizontal, vertical, 45° and 135° directions of the current CU are calculated; the gradients of the horizontal and vertical directions of the current CU are calculated using formulas (2) and (3), and the gradient operators in the 45° and 135° directions are obtained by rotating the horizontal and vertical convolution kernel operators, and their calculation is shown in formula (4): G in formula (4) 45 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 45° direction; A(x,y) represents the pixel matrix centered at coordinate (x,y); G 135 (x,y) represents the gradient of the pixel at coordinate (x,y) in the 135° direction; the gradient calculation of the 45° and 135° directions of the current CU is shown in formula (5); |G in formula (5) 45 (x,y)| and|G 135 (x, y)| is the absolute value of the 45° and 135° gradient of the CU at the (i, j) coordinate; W and H are the width and height of the CU respectively; Gradient 45 and Gradient 135 are the average gradients of the current CU in the 45° and 135° directions respectively.
9. The ERP panoramic video VVC intra-frame mode fast prediction method according to claim 1, characterized in that: In step S10, it is determined whether the angle is in the region where the minimum directional gradient is located. The regions here include horizontal, vertical, 45° and 135°, totaling four regions; the 65 angle prediction modes of VVC are divided into the above four regions; among them, the angle prediction modes 10-26 are divided into the horizontal region; the angle prediction modes 27-41 are divided into the 135° region; the angle prediction modes 42-58 are divided into the vertical region; and the angle prediction modes 2-9 and 59-66 are divided into the 45° region.
10. A storage medium storing a computer program, characterized in that: When the computer program is read by a processor, the method for fast prediction of intra-mode of ERP panoramic video VVC according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
CU division decision based on texture similarity
CN111683245A
Rapid CU division method for VVC of ERP panoramic video and storage medium
CN117041736A