A recording encoding method for ultra-high-definition video

By constructing partitioning mode decision values ​​and optimizing the horizontal and vertical partitioning modes of CU units, the problem of uneven regional sampling in ultra-high-definition 360° panoramic video encoding is solved, thereby improving encoding efficiency and encoding quality.

CN120881283BActive Publication Date: 2026-01-23GUANGDONG TUSHENG ULTRA HD INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511222730.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-01-23
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing H.266/VVC video coding methods suffer from uneven regional sampling in ultra-high-definition 360° panoramic video coding, leading to redundant partitioning and traversal, which increases coding time and resource consumption.

Method used

By analyzing the grayscale homogeneity, texture prominence, gradient significance ratio, and ROI attention weight of panoramic encoded frames, a partitioning mode decision value is constructed to optimize the horizontal and vertical partitioning modes of CU units and skip redundant traversal.

Benefits of technology

It reduces encoding time, improves the recording and encoding efficiency of ultra-high-definition panoramic video, and ensures the encoding quality of key details and features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120881283B_ABST
    Figure CN120881283B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video coding, in particular to a recording and coding method for ultra-high-definition video, which comprises the following steps: acquiring all panoramic coding frames in the ultra-high-definition video, and then acquiring target CU units; acquiring horizontal stretching evaluation values of the target CU units according to the discrete degrees of the gray values of all the pixel points in each target CU unit and the contrast of the gray level co-occurrence matrix in the neighborhood window of each key point in each target CU unit; dividing the panoramic coding frames into two-pole regions, mid-latitude regions and equatorial regions; acquiring reservation evaluation values of the target CU units according to the gradient characteristics of the horizontal and vertical directions of each pixel point and the ROI attention weight of each pixel point, and then acquiring division mode decision values of the target CU units, and then acquiring the division modes of the target CU units. The recording and coding efficiency of the ultra-high-definition panoramic video is improved by adaptively evaluating the division modes of the target CU units.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of video coding, in particular to a recording and coding method for ultra-high-definition video. BACKGROUND

[0002] With the gradual transformation of audio and video technology systems from SDI baseband signal architecture to IP network architecture, the wide application of ultra-high-definition (4K / 8K) panoramic video in sports live broadcast, virtual reality, smart tourism and other IP production and broadcast scenarios, the data volume of ultra-high-definition 360° panoramic video is several times that of ultra-high-definition flat video, which not only needs to ensure ultra-high-definition quality, but also needs to be compressed by efficient coding to adapt to the limited bandwidth resources in the IP production and broadcast process.

[0003] The recording and coding of ultra-high-definition 360° panoramic video need to process massive data, and the ERP projection has a significant problem of uneven regional sampling. The H.266 / VVC video coding method (Versatile Video Coding) does not evaluate the horizontal division mode and vertical division mode of each coding unit according to the principle of "brute force recursion and selection of the optimal", determines the optimal division mode of each CU unit through a depth-first traversal process, and has many redundant division traversals, which greatly increases the time consumed by video coding. SUMMARY

[0004] In order to solve the above technical problems, the application provides a recording and coding method for ultra-high-definition video to solve the existing problems.

[0005] The recording and coding method for ultra-high-definition video provided by the application adopts the following technical scheme:

[0006] An embodiment of the application provides a recording and coding method for ultra-high-definition video, which comprises the following steps:

[0007] All panoramic coding frames in the ultra-high-definition video are obtained, and then all panoramic coding frames are subjected to VVC intra-frame coding; all CU units that need to be divided into CU units in the VVC intra-frame coding recursive process are recorded as target CU units;

[0008] The in-row gray homogeneity of each target CU unit is obtained according to the discrete degree of the gray values of the pixel points in each target CU unit; all key points in each panoramic coding frame are obtained, and the horizontal texture prominence of each key point is obtained according to the contrast of the gray co-occurrence matrix in the neighborhood window of the key point; the horizontal stretching evaluation value of each target CU unit is obtained according to the in-row gray homogeneity of each target CU unit and the horizontal texture prominence of all key points in the target CU unit;

[0009] The panoramic coding frame is divided into polar region, middle latitude region and equator region; gradient significant proportions of each pixel point are obtained according to differences between horizontal gradients of each pixel point and average horizontal gradients of regions to which the pixel points belong, and differences between vertical gradients of each pixel point and average vertical gradients of regions to which the pixel points belong; matching points of each key point in other panoramic coding frames are obtained; when each pixel point is a key point, ROI attention weights of each pixel point are obtained according to matching point numbers of each pixel point and proportions of matching points same as regions to which the pixel points belong in all matching points corresponding to the pixel points; when each pixel point is a non-key point, the ROI attention weight of each pixel point is a preset constant;

[0010] According to the gradient significant proportions and the ROI attention weights of all pixel points in each target CU unit, a reservation evaluation value of each target CU unit is obtained, and in combination with a horizontal stretching evaluation value of each target CU unit, a division mode decision value of each target CU unit is obtained, and then a division mode of each target CU unit is obtained.

[0011] Preferably, a calculation formula of the in-line gray homogeneity of each target CU unit is: ; in the formula, is the in-line gray homogeneity of the target CU unit, is a standard deviation of the gray value of the pth row of pixel points in the target CU unit, is the number of rows of the target CU unit.

[0012] Preferably, a calculation formula of the horizontal texture prominence of each key point is: ; in the formula, is the horizontal texture prominence of the jth key point, is a contrast of a GLCM matrix with a direction of 0° and a step distance of d in a neighborhood window of the jth key point, and M is a gray level number of the panoramic coding frame; wherein the neighborhood window of the jth key point refers to a window with a size of h x h centered on the jth key point, and h is a preset side length.

[0013] Preferably, the horizontal stretching evaluation value of each target CU unit refers to a ratio of the in-line gray homogeneity of each target CU unit to a cumulative sum of the horizontal texture prominences of all key points inside the target CU unit.

[0014] Preferably, a specific process of dividing the panoramic coding frame into the polar region, the middle latitude region and the equator region is as follows: taking a left-bottom pixel point of the panoramic coding frame as an origin, taking a horizontal direction as an X axis and taking a vertical direction as a Y axis, a rectangular coordinate system of the panoramic coding frame is constructed; a pixel region in the panoramic coding frame with a longitudinal coordinate between and is recorded as the polar region, and a pixel region in the panoramic coding frame with a longitudinal coordinate between and The pixel region between the equator and the north pole is recorded as a middle latitude region, and the pixel region between the equator and the south pole is recorded as a middle latitude region. The pixel region between the equator and the north pole is recorded as a middle latitude region, and the pixel region between the equator and the south pole is recorded as a middle latitude region. is the height of the panoramic coded frame.

[0015] Preferably, the method for obtaining the gradient significant proportion of each pixel point is: the absolute difference value between the horizontal gradient of each pixel point and the average horizontal gradient of the region to which the pixel point belongs in the panoramic coded frame, and the absolute difference value between the vertical gradient of each pixel point and the average vertical gradient of the region to which the pixel point belongs in the panoramic coded frame are counted, and recorded as the horizontal gradient significant value and the vertical gradient significant value of each pixel point, respectively. The ratio of the horizontal gradient significant value to the vertical gradient significant value is taken as the gradient significant proportion of each pixel point.

[0016] Preferably, the calculation formula of the ROI attention weight of each pixel point is: ; in the formula, is the ROI attention weight of the i-th pixel point in the target CU unit, is the set of all key points in the panoramic coded frame where the target CU unit is located, is the sample matching proportion of the i-th pixel point in the target CU unit, is the distribution probability of the i-th pixel point in the target CU unit, is the preset matching score of the region to which the i-th pixel point in the target CU unit belongs, and exp() is an exponential function with natural constant e as the base; wherein the sample matching proportion of the i-th pixel point refers to the ratio of the total number of matching points corresponding to the i-th pixel point in all sample frames to the total number of all sample frames; the sample frame refers to a preset proportion of panoramic coded frames uniformly sampled in ascending order of time from all panoramic coded frames; the distribution probability of the i-th pixel point refers to the proportion of the matching points corresponding to the i-th pixel point in the same region as the i-th pixel point.

[0017] Preferably, the calculation formula of the reserved evaluation value of each target CU unit is: ; in the formula, is the reserved evaluation value of the target CU unit, I is the total number of pixel points in the target CU unit, is the gradient significant proportion of the i-th pixel point in the target CU unit, is the ROI attention weight of the i-th pixel point in the target CU unit.

[0018] Preferably, the division mode decision value of each target CU unit is the normalized value of the ratio of the reserved evaluation value of the target CU unit to the horizontal stretching evaluation value.

[0019] Preferably, the specific process of obtaining the partition mode of each target CU unit is: when the partition mode decision value of the target CU unit is greater than or equal to a preset decision threshold, a vertical partition mode is adopted for the target CU unit; otherwise, a horizontal partition mode is adopted for the target CU unit.

[0020] The present application has at least the following beneficial effects:

[0021] 1. The present application is aimed at the serious problem of horizontal stretching in the two-pole area of ERP projection. By analyzing the texture detail features after stretching, the horizontal stretching evaluation value and texture saliency are constructed, the unobvious degree of horizontal texture features is evaluated, and then it is beneficial to judge whether the horizontal partition mode should be adopted;

[0022] 2. The present application uses key point matching and attention weighting mechanism to identify the frequently appearing ROI area in the video, and gives the corresponding weight to the pixel points of the ROI area, constructs the retention evaluation value of each target CU unit, so that the ROI area is taken as an important reference area when evaluating the partition mode of each target CU unit, and the coding quality of key detail features is guaranteed;

[0023] 3. In the scenario of huge amount of super high definition panoramic video data, the present application directly determines the vertical / horizontal partition mode of the target CU unit by constructing the partition mode decision value of each target CU unit, skips the redundant partition mode traversal in the VVC standard, reduces the coding time, optimizes the calculation resource consumption, and improves the efficiency of recording and coding of super high definition panoramic video. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0025] Figure 1 A step flowchart of a recording and coding method for super high definition video provided by the present application is provided.

[0026] Figure 2 A flowchart of obtaining the partition mode decision value of each target CU unit provided by the present application is provided. DETAILED DESCRIPTION

[0027] For further elaboration of the technical means and effects taken by the present application to achieve the predetermined object of the application, the specific implementation, structure, features and effects of a recording and encoding method for ultra-high-definition video according to the present application are described in detail as follows in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0029] The specific scheme of the recording and encoding method for ultra-high-definition video provided by the present application is described in detail below in combination with the drawings.

[0030] One embodiment of the present application provides a recording and encoding method for ultra-high-definition video. Specifically, the following recording and encoding method for ultra-high-definition video is provided. Please refer to Figure 1 The method comprises the following steps:

[0031] Step one: obtain all panoramic encoded frames in the ultra-high-definition video, and then perform VVC intra-frame encoding on all panoramic encoded frames; all CU units that need to be divided in the VVC intra-frame encoding recursive process are recorded as target CU units.

[0032] The present application uses a 3D camera to take a full range of shots and obtain an ultra-high-definition 360° panoramic video, which contains all visual information in space. Any video frame is a 3D spherical panoramic image, and any point on the 3D spherical panoramic image contains a corresponding yaw angle and pitch angle. Since the equirectangular projection (ERP) does not require cropping and splicing operations, the 3D spherical panoramic image can be mapped to a 2D plane by stretching. Therefore, the present application projects the ultra-high-definition panoramic video in an ERP manner to obtain a 2D plane image according to the yaw angle and pitch angle. The size of the 2D plane image is , The equirectangular projection (ERP) algorithm is a known technology, and the specific process is not described again.

[0033] In the VVC video coding standard, the CU (Coding Unit) partition process is a key technology, and its efficiency directly affects the quality and speed of encoding. In this application, all 2D planar images obtained by ERP projection of ultra-high-definition panoramic video are used as panoramic encoding frames for VVC encoding of ultra-high-definition panoramic video. Each panoramic encoding frame is divided into several CTUs (Coding Tree Units), and the size of the CTU is 128x128. For any 128x128 CTU unit, a quadtree (QT) recursive partitioning is used to divide it into 16 initial CUs with a size of 32x32. Then, the VVC intra-frame coding method is used to recursively divide all initial CUs in the panoramic encoding frame into smaller CU units. The VVC intra-frame coding CU partitioning order is a depth-first Z-shaped scan. In order to facilitate subsequent description, all CU units that need to be partitioned in the VVC intra-frame coding recursive process are referred to as target CU units. The size of the target CU unit can be up to 32x32 or as small as 4x4.

[0034] Step two: obtain the row-wise gray level homogeneity of each target CU unit according to the dispersion degree of the gray level values of the pixels in each row of the target CU unit; obtain all key points in each panoramic encoding frame, and obtain the horizontal texture prominence of each key point according to the contrast of the gray level co-occurrence matrix in the neighborhood window of the key point; obtain the horizontal stretching evaluation value of each target CU unit according to the row-wise gray level homogeneity of the target CU unit and the horizontal texture prominence of all key points inside the target CU unit.

[0035] Currently, the CU unit partition structure in VVC video coding has been expanded from QT (Quad Tree) partition structure to QTMT (Quadtree with nested Multi-type Tree) partition structure. The QTMT partition structure includes quadtree, binary tree and ternary tree partitioning, and the binary tree and ternary tree partitioning further includes horizontal mode and vertical mode, which can be further divided into: vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning. Although these flexible partitioning methods can improve the adaptability of video texture, they also cause the coding complexity of the CU partitioning step to increase dramatically.

[0036] In the VVC video coding, the depth-first traversal rate-distortion optimization process is used, and the encoder does not determine the horizontal / vertical partition mode of the target CU unit when partitioning the CU unit. It needs to traverse all partitioning modes of the current CU, and recursively traverse the partitioning modes of the sub-CU of the current CU and the combination of the sub-CU. As the number of partitioning layers increases, the number of branches that need to be traversed increases dramatically, resulting in a significant increase in coding complexity and consuming a large amount of encoding time.

[0037] Since the ERP projection can map each parallel circle on the 360° spherical panoramic image to a horizontal line with equal length to the equator in the equirectangular frame, the 360° spherical panoramic image at the two poles is stretched into a very long horizontal strip occupying the whole row of pixels, and the texture details are also significantly elongated in the horizontal direction. When the pixel values of each row of pixels in the target CU unit of the equirectangular frame tend to be consistent, it means that the texture in the horizontal direction is less significant. In order to prevent the stretched details in the horizontal direction from being destroyed, the horizontal mode partition should be used more.

[0038] As a preferred embodiment, the in-row gray level homogeneity of the target CU unit is obtained according to the dispersion degree of the gray level values of each row of pixels in the target CU unit, which is used to represent the insignificant degree of the horizontal direction texture in the target CU unit.

[0039] In this embodiment, the in-row gray level homogeneity of the target CU unit is denoted as , and the specific expression is: ; in the formula, is the in-row gray level homogeneity of the target CU unit, is the standard deviation of the gray level value of the pth row of pixels in the target CU unit, is the number of rows of the target CU unit.

[0040] is used to reflect the difference of the pixel gray level values in the horizontal direction of the target CU unit, the greater the in-row gray level homogeneity is, the smaller the difference of the pixel gray level values in the horizontal direction is, the higher the gray level homogeneity of the target CU unit as a whole is, and the target CU unit should be partitioned in the horizontal mode more.

[0041] Further, in the process of recording and encoding the ultra-high-definition 360° panoramic video, in order to ensure the visual quality and immersive experience, more attention is often paid to the local detail part of the video frame, which usually corresponds to the edge, texture or structure mutation region of the video object, and is the most sensitive area in vision.

[0042] In this application, any panoramic encoding frame to be encoded is taken as the input of the SURF (Speeded-Up Robust Features) algorithm, the coordinate positions and descriptor vectors of all key points in the panoramic encoding frame are obtained, and the set of all key points in the panoramic encoding frame is denoted as ​The panoramic coding frame is gray quantized, and after quantization, there are M gray levels. Specifically, the original pixel gray value of the panoramic coding frame can be directly divided by (256 / M), and then rounded down. In this embodiment, M is 8. Taking the jth key point in the gray quantized panoramic coding frame as an example, a h x h (h is a preset side length, and in this embodiment, h is 15) window centered on the jth key point is taken as the neighborhood window of the jth key point, and a GLCM (gray level co-occurrence matrix) matrix with a direction of 0° and a step distance of d in the neighborhood window of the jth key point is obtained. The GLCM matrix is a known technology, and the specific process will not be described here.

[0043] In this embodiment, the horizontal texture prominence of the jth key point is denoted as , and the specific expression is: ; in the formula, , the horizontal texture prominence of the jth key point is , the contrast of the GLCM matrix with a direction of 0° and a step distance of d in the neighborhood window of the jth key point, and M is the gray level number of the panoramic coding frame.

[0044] The upper limit of the step distance d is set to , so as to avoid too large step distance leading to too few GLCM matrix data elements, which makes it difficult to fully reflect the horizontal texture features of the neighborhood window and reduces unnecessary calculation amount. The horizontal texture prominence is used to reflect the prominent features of the horizontal texture under the multi-scale step distance, and the smaller the value is, the lower the definition of the neighborhood window of the corresponding key point is, and the shallower the horizontal texture is, and the horizontal division mode should be used more.

[0045] As a preferred embodiment, the horizontal stretch evaluation value of the target CU unit is obtained according to the in-line gray homogeneity of the target CU unit and the horizontal texture prominence of all key points in the target CU unit, and is used to reflect the unobvious degree of the texture features in the horizontal direction of the target CU unit.

[0046] In this embodiment, the ratio of the in-line gray homogeneity of each target CU unit to the cumulative sum of the horizontal texture prominences of all key points in the target CU unit is taken as the horizontal stretch evaluation value of the target CU unit. The greater the horizontal stretch evaluation value of the target CU unit is, the more unobvious the texture features in the horizontal direction of the target CU unit are, and then the horizontal division mode should be used to divide the target CU unit.

[0047] Step three: dividing the panoramic coded frame into polar region, mid-latitude region and equatorial region; obtaining gradient significant proportion of each pixel point according to the difference between horizontal gradient of each pixel point and average horizontal gradient of the region to which the pixel point belongs, and the difference between vertical gradient of each pixel point and average vertical gradient of the region to which the pixel point belongs; obtaining matching points of each key point in other panoramic coded frames; when each pixel point is a key point, obtaining ROI attention weight of each pixel point according to the number of matching points of each pixel point and the proportion of matching points same as the region to which the pixel point belongs in all matching points corresponding to the pixel point; when each pixel point is a non-key point, setting ROI attention weight of each pixel point as a preset constant.

[0048] In this application, the left-bottom pixel point of the panoramic coded frame is taken as the origin, the horizontal direction is taken as the X axis, and the vertical direction is taken as the Y axis to construct the rectangular coordinate system of the panoramic coded frame. Since the panoramic coded frame is a planar image after ERP mapping, the area of the north and south poles (upper and lower parts of the panoramic coded frame) is more seriously oversampled. In this application, the pixel region between the longitudinal coordinates of and is recorded as the polar region of the panoramic coded frame, the pixel region between the longitudinal coordinates of and is recorded as the mid-latitude region of the panoramic coded frame, and the pixel region between the longitudinal coordinates of is recorded as the equatorial region of the panoramic coded frame.

[0049] In the panoramic coded frame, the texture of different regions is different, the texture features of the equatorial region are highly preserved after ERP projection, and the texture features of the polar region are severely distorted and have high detail loss after ERP projection.

[0050] In this application, the panoramic coded frame is taken as the input of the Sobel operator to obtain the horizontal gradient and vertical gradient of any pixel point in the panoramic coded frame, and the average horizontal gradient and average vertical gradient of the polar region, mid-latitude region and equatorial region of the panoramic coded frame are calculated respectively. The absolute difference between the horizontal gradient of each pixel point and the average horizontal gradient of the region to which the pixel point belongs in the panoramic coded frame is taken as the horizontal gradient significant value of each pixel point, and the absolute difference between the vertical gradient of each pixel point and the average vertical gradient of the region to which the pixel point belongs in the panoramic coded frame is taken as the vertical gradient significant value of each pixel point.

[0051] In this embodiment, the gradient significant proportion of the i-th pixel point in the target CU unit is recorded as , and the specific expression is: ; in the formula, is the gradient significant proportion of the i-th pixel point in the target CU unit, is the horizontal gradient significant value of the i-th pixel point in the target CU unit, is a vertical gradient significant value of the i-th pixel point in the target CU unit.

[0052] The greater, the higher the horizontal gradient significant value of the i-th pixel point than the vertical gradient significant value, in order to prevent the vertical detail features of the i-th pixel point divided by the CU unit from being damaged, the target CU unit at the i-th pixel point tends to adopt a vertical division mode.

[0053] Further, the 360° panoramic video has the characteristics of immersion and interactivity, and can provide personalized visual experience for users, and is generally used in scenes such as film and television production, education and training, cultural tourism, etc. There are regions that need special attention in the 360° panoramic video, i.e. Region of Interest (ROI), which usually appears frequently in the 360° panoramic video stream and contains key information that VR users want to watch.

[0054] Taking the R-th panoramic encoding frame as an example, 10% of the panoramic encoding frames are uniformly sampled in ascending order of time from all panoramic encoding frames of the panoramic video stream, and all the sampled panoramic encoding frames are denoted as sample frames. All sample frames are input into the SURF (Speeded-Up Robust Features) algorithm to obtain all key points and their corresponding descriptor vectors in the sample frames. Taking the k-th sample frame as an example, according to the R-th panoramic encoding frame and all key points and their descriptor vectors in the k-th sample frame, the FLANN (Fast Library for Approximate Nearest Neighbors) library is used to match the key points of the R-th panoramic encoding frame and the k-th sample frame, and the key point that matches successfully with the i-th key point in the R-th panoramic encoding frame is denoted as the matching point of the i-th key point in the k-th sample frame. The use of the SURF algorithm and the FLANN library is a known technology, and the specific process will not be described again.

[0055] Taking any key point in the R-th panoramic encoding frame as an example, the ratio of the total number of matching points corresponding to the key point in all sample frames to the total number of all sample frames is taken as the sample matching proportion of the key point, and the proportion of the matching points belonging to the same region as the key point in all matching points corresponding to the key point is taken as the distribution probability of the key point. For example, the i-th key point is located in the polar region, the total number of matching points of the i-th key point is 30, and the number of matching points belonging to the polar region is 5, then the distribution probability of the i-th key point is .

[0056] Through the above analysis, the application obtains the ROI attention weight of each pixel point in the target CU unit: ; in the formula, a ROI attention weight of the i-th pixel point in the target CU unit, a set of all key points in the panoramic encoding frame where the target CU unit is located, is a sample matching proportion of the i-th pixel point in the target CU unit, is a distribution probability of the i-th pixel point in the target CU unit, is a preset matching score of the region to which the i-th pixel point in the target CU unit belongs, and the preset matching scores of the two-pole region, the middle latitude region and the equatorial region are 1, 2 and 3 respectively in this embodiment, and exp() is an exponential function with the natural constant e as the base.

[0057] is used to reflect the frequency of the key point i in the 360° panoramic video stream, The greater the value is, the more the key point is paid attention to by the video shooting, and the more likely the region where the key point i is located is the ROI region in the 360° panoramic video stream. is used to give different matching weights to key points in different regions. The high attention region of the 360° panoramic video stream is often contained in the equatorial region, and the information amount of each sampling point in the equatorial region is more than that in the two-pole region. The key point matching credibility is higher, in order to improve the visual quality of the video decoding, the key point should be paid more attention, and the ROI attention weight The greater the value is.

[0058] Step four: according to the gradient significant proportion of all pixel points in each target CU unit and the ROI attention weight, the retention evaluation value of each target CU unit is obtained, and the division mode decision value of each target CU unit is obtained by combining the horizontal stretching evaluation value of each target CU unit, and then the division mode of each target CU unit is obtained.

[0059] Further, the retention evaluation value of the target CU unit is obtained through the gradient significant proportion of all pixel points in the target CU unit and the ROI attention weight, which is used to evaluate the degree of inclination of the target CU unit to the vertical division mode.

[0060] In this embodiment, the retention evaluation value of the target CU unit is denoted as , and the specific expression is: ; in the formula, is the retention evaluation value of the target CU unit, I is the total number of pixel points in the target CU unit, is the gradient significant proportion of the i-th pixel point in the target CU unit, is the ROI attention weight of the i-th pixel point in the target CU unit.

[0061] The larger the value, the stronger the significance of the horizontal gradient. Therefore, to prevent the destruction of vertical detail features, the more vertical texture details in the region containing the i-th pixel should be preserved, and thus, a vertical partitioning mode should be used. The gradient significance ratio is weighted by ROI attention weights, so that when... When greater than 0, for It has an expanding effect; when When less than 0, for It has a shrinking effect. The larger the value, the more the pixels in the target CU unit tend to be divided into vertical segments.

[0062] In a preferred embodiment, based on the retention evaluation value and horizontal stretching evaluation value of each target CU element, a partitioning mode decision value is obtained for each target CU element. This value reflects the tendency of each target CU element to be partitioned using a vertical partitioning mode. The flowchart for obtaining the partitioning mode decision value of each target CU element is shown below. Figure 2 As shown.

[0063] In this embodiment, the normalized value of the ratio of the retention evaluation value to the horizontal stretching evaluation value of the target CU unit is used as the partitioning mode decision value of the target CU unit.

[0064] A preset decision threshold Th is set, and the horizontal / vertical partitioning mode of the target CU unit during the VVC intra-coding recursion is determined as follows: When the partitioning mode decision value of the target CU unit is greater than or equal to the preset decision threshold, it indicates that the horizontal texture saliency is high and the vertical texture saliency is low within the target CU unit. Therefore, to prevent the vertical features from being destroyed, the vertical texture should be preserved as much as possible, and the vertical partitioning mode is adopted for the target CU unit. When the partitioning mode decision value of the target CU unit is less than the preset decision threshold, it indicates that the vertical texture saliency is high and the horizontal texture saliency is low within the target CU unit. To prevent the horizontal features from being destroyed, the horizontal texture should be preserved as much as possible, and the horizontal partitioning mode is adopted for the target CU unit. In this embodiment, the preset decision threshold is set to 0.5.

[0065] The above method obtains the partitioning patterns of all target CU units during the VVC intra-frame coding recursion process, skipping redundant pattern partitioning traversal in VVC coding, thereby improving the recording and coding efficiency of ultra-high-definition video.

[0066] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0067] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the difference from other embodiments.

[0068] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; the technical solutions recorded in the foregoing embodiments are modified, or some technical features are replaced equivalently, and the essence of the corresponding technical solutions does not deviate from the scope of the technical solutions of the embodiments of the present application, which should be included in the protection scope of the present application.

Claims

1. A recording and encoding method for ultra-high-definition video, characterized in that, The method includes the following steps: All panoramic encoded frames in the ultra-high-definition video are acquired, and then VVC intra-frame coding is performed on all panoramic encoded frames; all CU units that need to be divided into CU units during the VVC intra-frame coding recursive process are denoted as target CU units. The intra-row grayscale homogeneity of each target CU unit is obtained based on the dispersion of grayscale values ​​of pixels in each row of each target CU unit; all keypoints in each panoramic encoded frame are obtained, and the horizontal texture prominence of each keypoint is obtained based on the contrast of the grayscale co-occurrence matrix in the neighborhood window of each keypoint; the horizontal stretching evaluation value of each target CU unit is obtained based on the intra-row grayscale homogeneity of each target CU unit and the horizontal texture prominence of all keypoints within it. The panoramic encoded frames are divided into polar, mid-latitude, and equatorial regions. The gradient significance ratio of each pixel is obtained based on the difference between its horizontal gradient and the average gradient of its region, and the difference between its vertical gradient and the average vertical gradient of its region. Matching points for each keypoint in other panoramic encoded frames are obtained. When a pixel is a keypoint, its ROI attention weight is obtained based on the number of matching points for that pixel and the proportion of matching points in the same region as that pixel. When a pixel is not a keypoint, its ROI attention weight is set to a preset constant. Based on the gradient significance ratio and ROI attention weight of all pixels in each target CU unit, the retention evaluation value of each target CU unit is obtained. Combined with the horizontal stretching evaluation value of each target CU unit, the partitioning mode decision value of each target CU unit is obtained, and then the partitioning mode of each target CU unit is obtained.

2. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The formula for calculating the in-row grayscale homogeneity of each target CU unit is as follows: In the formula, The grayscale homogeneity within the target CU cell. It is the standard deviation of the grayscale values ​​of the p-th row of pixels in the target CU unit. It is the row number of the target CU unit.

3. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The formula for calculating the horizontal texture prominence of each key point is as follows: In the formula, Let the horizontal texture saliency of the j-th keypoint be , is the contrast of the GLCM matrix with a direction of 0° and a step size of d within the neighborhood window of the j-th keypoint, and M is the gray level of the panoramic encoded frame; where the neighborhood window of the j-th keypoint refers to an h×h window centered on the j-th keypoint, and h is the preset side length.

4. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The horizontal stretching evaluation value of each target CU unit refers to the ratio of the inline grayscale homogeneity of each target CU unit to the cumulative sum of the horizontal texture prominence of all key points within it.

5. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The specific process of dividing the panoramic encoded frame into polar regions, mid-latitude regions and equatorial regions is as follows: taking the bottom left corner pixel of the panoramic encoded frame as the origin, the horizontal direction as the X-axis, and the vertical direction as the Y-axis, a rectangular coordinate system of the panoramic encoded frame is constructed. The vertical coordinate in the panoramic encoded frame and The pixel regions between these points are denoted as the two extreme regions. The vertical coordinates of the panoramic encoded frame are then... and The pixel region between these points is denoted as the mid-latitude region, and the vertical coordinates of the panoramic encoded frame are... The pixel region between these points is denoted as the equatorial region; where... The height of the panoramic encoded frame.

6. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The method for obtaining the gradient significance ratio of each pixel is as follows: the absolute difference between the horizontal gradient of each pixel and the average level gradient of the region to which it belongs in the panoramic coding frame, and the absolute difference between the vertical gradient of each pixel and the average vertical gradient of the region to which it belongs in the panoramic coding frame, are recorded as the horizontal and vertical gradient significance values ​​of each pixel, respectively. The ratio of the horizontal gradient significance value to the vertical gradient significance value is taken as the gradient significance ratio of each pixel.

7. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The formula for calculating the ROI attention weight of each pixel is as follows: In the formula, The ROI attention weight is assigned to the i-th pixel in the target CU unit. This is the set of all keypoints in the panoramic encoded frame containing the target CU unit. It represents the sample matching percentage of the i-th pixel in the target CU unit. It is the probability distribution of the i-th pixel in the target CU unit. exp() is the preset matching score of the region to which the i-th pixel belongs in the target CU unit, and exp() is an exponential function with the natural constant e as the base; where the sample matching ratio of the i-th pixel is the ratio of the total number of matching points corresponding to the i-th pixel in all sample frames to the total number of all sample frames; the sample frame refers to a preset proportion of panoramic encoded frames uniformly sampled in ascending chronological order from all panoramic encoded frames; the distribution probability of the i-th pixel is the proportion of matching points in the same region as the i-th pixel among all matching points corresponding to the i-th pixel.

8. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The formula for calculating the retention evaluation value of each target CU unit is as follows: In the formula, The retained evaluation value for the target CU unit, where I is the total number of pixels in the target CU unit. The gradient significance ratio of the i-th pixel in the target CU unit. The ROI attention weight is assigned to the i-th pixel in the target CU unit.

9. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The partitioning mode decision value of each target CU unit refers to the normalized value of the ratio of the retention evaluation value to the horizontal stretching evaluation value of the target CU unit.

10. The recording and encoding method for ultra-high-definition video as described in claim 1, characterized in that, The specific process for obtaining the partitioning mode of each target CU unit is as follows: when the partitioning mode decision value of the target CU unit is greater than or equal to the preset decision threshold, the target CU unit adopts the vertical partitioning mode; otherwise, the target CU unit adopts the horizontal partitioning mode.

Citation Information

Patent Citations

  • Rapid CU division method for VVC of ERP panoramic video and storage medium

    CN117041736A

  • 360-degree video fast coding method and device

    CN119815030A