A Low-Complexity Feature Alignment Method for HEVC Video Frames Based on Encoded Information
By using a feature alignment method based on HEVC encoding information to generate a displacement field for feature alignment, the problems of low accuracy and high computational complexity in existing technologies for compressed video are solved, achieving fast and efficient feature alignment and video quality improvement.
Patent Information
- Application Number
- CN202511206446.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing feature alignment methods have low accuracy in compressed videos, cannot flexibly handle videos with complex motion amplitude changes, have high computational complexity, and rely on high-performance hardware devices and pre-trained models.
The feature alignment method based on HEVC encoding information extracts CU block information, CU prediction mode information, CU-level MV information, and the mapping relationship between CU reference frame list index and POC from the HEVC encoded bitstream to generate MV missing region map and pixel-level MV distribution map, and uses the encoding information to generate a displacement field for feature alignment.
It achieves fast and efficient feature alignment, reduces computational complexity, improves the quality of compressed video, adapts to videos with complex motion amplitude changes, is compatible with existing feature alignment methods, and enhances both video quality and computational efficiency.
Smart Images

Figure CN120730064B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image communication technology, and in particular to a low-complexity feature alignment method for HEVC video frames based on encoded information. Background Technology
[0002] In recent years, with the rapid development of deep learning, various low-level visual processing algorithms for compressed video have achieved better results than traditional methods. Among these deep learning algorithms, the method of using the inter-frame similarity of video to perform multi-frame joint processing has been proven to have the best effect after long-term experiments and has become the industry consensus.
[0003] Multi-frame compressed video processing algorithms typically utilize feature alignment to extract effective information from adjacent frames, thereby enhancing the information in the current frame. Therefore, the effectiveness of feature alignment directly impacts the overall quality enhancement algorithm. Existing feature alignment methods are mainly divided into two categories: optical flow-based methods and deformable convolution-based methods.
[0004] Optical flow-based methods primarily solve for the optical flow field by accurately estimating the spatial variations of pixels, and then align features based on the optical flow field. The advantage of this method is its high accuracy, but its disadvantages include the need to construct a complex optical flow estimation network, resulting in high computational complexity and a heavy reliance on high-performance hardware and pre-trained models. When processing compressed video, its accuracy is affected by compression effects such as blurring, blockiness, and color shifts.
[0005] Feature alignment methods based on deformable convolution treat the relationship between adjacent frames as a set of matrices that can be transformed by convolution. Feature alignment is achieved by estimating the local displacement fields of adjacent frames and the frame to be enhanced. However, this method is limited by the size of the convolution kernel. When the motion information in the video is too large, the algorithm cannot handle it effectively. Increasing the kernel size leads to increased computational cost, further increasing the computational burden. Furthermore, deformable convolution cannot flexibly handle videos with complex motion amplitude variations; it performs better on videos with relatively uniform motion amplitudes. Summary of the Invention
[0006] The purpose of this invention is to provide a low-complexity feature alignment method for HEVC video frames based on encoded information that can overcome the shortcomings of existing technologies, such as low accuracy of compressed video and inability to flexibly handle videos with complex motion amplitude variations. This method is fast, efficient, and low in complexity.
[0007] To achieve the above objectives, the technical solution adopted by this invention is as follows: a low-complexity feature alignment method for HEVC video frames based on encoded information, comprising the following steps:
[0008] S1, extract the encoding information of each video frame from the HEVC encoded bitstream, including CU block information, CU prediction mode information, CU-level MV information, CU reference frame list index, and the mapping relationship between the reference frame list index and POC.
[0009] S2, the t-th video frame y t and the t+ith video frame y t+i y is generated as the target frame and the frame to be aligned, respectively. t+i MV missing region map ML t+i and pixel-level MV distribution map MV t+i Including S21~S23;
[0010] S21, for y t+1 Generate CU block diagram based on CU block information. t+i Then, based on the CU prediction model information, CU t+i In the ML (Multi-Frame Prediction) mode, the coded blocks using intra-prediction mode are set to white, and the remaining coded blocks are set to black. t+i The white area is used as the area with missing MV information, and the black area is used as the area with MV information.
[0011] S22, process ML sequentially t+i Each coded block;
[0012] If the s-th encoded block y t+i (s) The coordinates are located in the black area; read its CU-level MV information. And store in ML t+i At the corresponding coordinates;
[0013] If the s-th encoded block y t+i If coordinate (s) is located in the white area, then in the video frame y at frame t+i-1... t+i-1 Find its most similar block and calculate y. t+i (s) and the motion vector MV of the most similar block t+i→t+i-1 (s), as y t+i CU-level MV information (s) Store in ML t+i At the corresponding coordinates;
[0014] S23, ML stores the CU-level MV information of all encoded blocks. t+i Marked as pixel-level MV distribution map MV t+i ;
[0015] S3, generate the CU reference frame POC distribution map for each video frame, where y t+i CU reference frame POC distribution map ND t+i ND t+i Includes y t+iThe POC sequence number of each coded block;
[0016] S4, Generate target frame y t and the frame to be aligned y t+i Displacement field M t+i→t This includes steps S41 to S42;
[0017] S41, read y t+i CU reference frame number distribution diagram ND t+i , for y t+i (s) The displacement field M of the coded block is generated according to the following formula. t+i→t (s);
[0018] ,
[0019] In the formula, recPOC t+i (s) is y t+i The s-th coded block y t+i The POC number of (s);
[0020] S42, process y sequentially according to S41. t+i From each encoded block, the displacement field M is obtained. t+i→t ;
[0021] S5, y t+i According to the displacement field M t+i→t Generate alignment frames .
[0022] As a preferred embodiment, in S23, according to the following formula, from y... t+i-1 Find the code block that satisfies Conditional encoding block y t+i-1 (u+V h ,v+V v ), as y t+i The most similar block of (s) is obtained, and MV is obtained. t+i→t+i-1 (u,v);
[0023] ,
[0024] In the formula, y t+i (u,v) is y t+i The coordinates of (s), where u and v are the x and y coordinates respectively, argmin(∙) is the argmin function, and V h V v y t+i-1 (u+V h ,v+V v (relative to y) t+i The horizontal and vertical translation coordinates of (s).
[0025] As a preferred option, y is generated in S3. t+i CU reference frame POC distribution map ND t+i Specifically:
[0026] S31, Obtain the CU reference frame list L={L0,L1,…,L...} based on the prediction direction. n ,…,L N}, where L n This is the nth reference frame in L;
[0027] S32, if y t+i (s) If the CU reference frame list index of the coded block is n, then its reference frame is L. n L is obtained based on the mapping relationship between the reference frame list index and the POC. n The corresponding POC sequence number is used as the POC sequence recPOC of this coding block. t+i (s), marked on all pixels of the coded block;
[0028] S33, process y sequentially according to S32. t+i From all coded blocks, we obtain the CU reference frame number distribution map ND. t .
[0029] Preferably, in S42, the obtained displacement field M t+i→t The correction is performed using a weighted filtering method.
[0030] As preferred methods, the methods used in S5 include feature alignment methods based on spatiotemporally deformable convolution, feature alignment methods based on feature alignment neural networks with encoded information displacement fields, and feature alignment methods based on optical flow estimation networks.
[0031] Definitions:
[0032] HEVC (High Efficiency Video Coding), also known as H.265, is a video coding standard jointly developed by the International Organization for Standardization (ISO) and the International Telecommunication Union (ITU) to improve the compression efficiency of video coding. HEVC includes intra-frame coding mode and inter-frame coding mode. HEVC frames are further divided into I-frames, P-frames, and B-frames.
[0033] Intra-frame coding: This refers to a method of predictive coding that utilizes the spatial correlation between neighboring pixels within a single image frame. This mode generates prediction blocks by performing directional predictions (such as planar, angular, and DC prediction modes) on the reconstructed pixels surrounding the current coding block, thereby reducing the amount of information required for coding. Intra-frame coding does not rely on reference information from other frames, thus providing high coding efficiency and low latency even in scenes with significant texture changes or minimal motion.
[0034] Inter-frame coding (IPC) is a predictive coding method that utilizes the temporal redundancy between different video frames. This method obtains motion vector information by searching for the best-matching candidate block in preceding and following reference frames, and then uses motion compensation to generate predictive blocks, reducing the amount of residual data encoded. IPC fully leverages the continuity and similarity in video sequences, which can significantly reduce the bitrate, but it also requires substantial computational resources for motion estimation and motion compensation.
[0035] POC (Picture Order Count) is a key parameter used to determine the correct playback order of decoded image frames on the display. Because encoders typically output frames in encoding order for compression efficiency—for example, encoding P-frames or B-frames before I-frames—this is not necessarily their playback order. POC assigns a display order number to each frame, allowing the decoder to send the frames to the display in ascending order of POC after buffering and decoding all reference frames.
[0036] The Coding Unit Reference Picture List Index (CU) indicates the position of the selected reference frame in the reference picture list L0 (List 0) or L1 (List 1) for this CU (Coding Unit, coded block) (i.e., ref_idx_l0 / ref_idx_l1 in the bitstream information). L0 is used for forward prediction or one-way prediction, and L1 is used for backward prediction or two-way prediction.
[0037] Mapping relationship between reference frame list index and POC: Map the above list index to the corresponding reference frame's POC so as to correctly obtain the reference frame sequence number used for motion compensation and decoding.
[0038] In the HEVC standard, the specific steps for encoding a P-frame are as follows: (1) The current encoded frame is first divided into coded blocks (CU) or prediction blocks (PU) according to the coding tree unit (CTU). (2) For each coded block, the encoder uses the minimum matching cost criterion (absolute error and algorithm SAD) to search for the most similar candidate block in the reference frame and records the spatial displacement information between the two. This spatial displacement information is the motion vector (MV) recorded in 4×4 pixel units. (3) If no candidate block that meets the matching criterion can be found in the reference frame, the coded block will be determined to use the intra-frame prediction mode and will not contain motion vector information.
[0039] MV information is similar to optical flow information, both reflecting the displacement of adjacent frames. However, MV information is not equal to the displacement field information of adjacent frames; it needs to be combined with other encoded information to form a displacement field reflecting adjacent frames.
[0040] Compared with the prior art, the advantages of the present invention are:
[0041] (1) Novel Approach: This invention proposes a novel low-complexity feature alignment method for HEVC video frames based on encoded information. For encoders based on the HEVC encoding standard, the method utilizes the existing encoded information in the bitstream information to achieve fast and efficient displacement field acquisition, thereby realizing low-complexity HEVC compressed frame feature alignment.
[0042] (2) Improves the quality of compressed video: The present invention tested the extent to which the method improves the quality of compressed video on the H.265 / HEVC standard test model HM16.9.
[0043] (3) Fast speed and low complexity: Compared with feature alignment methods based on optical flow, it does not require the construction of a complex optical flow estimation network, thereby reducing computational complexity and dependence on high-performance hardware and pre-trained models.
[0044] (4) Large feature alignment range: Unlike feature alignment methods based on deformable convolution, the feature alignment range that deformable convolution can handle is limited by the size of the convolution kernel. In videos with more intense motion, the displacement distance of objects is larger. In this case, the present invention can handle it more effectively.
[0045] (5) Good compatibility: This invention can be integrated with existing feature alignment methods to achieve better results, such as feature alignment based on optical flow, feature alignment based on deformable convolution, and deformable convolution based on structure from motion (SFM). Attached Figure Description
[0046] Figure 1This is a flowchart of the present invention;
[0047] Figure 2 Example diagram of CU block diagram;
[0048] Figure 3 Example image of the missing region in MV;
[0049] Figure 4 This is a schematic diagram of displacement field transformation based on reference frame number;
[0050] Figure 5 A schematic diagram of the obtained displacement field;
[0051] Figure 6 for Figure 5 A magnified view of the area within the red box;
[0052] Figure 7 This is a diagram of a feature alignment neural network structure based on the displacement field of encoded information. Detailed Implementation
[0053] The present invention will be further described below with reference to the embodiments and accompanying drawings.
[0054] Example 1: See Figures 1 to 6 A low-complexity feature alignment method for HEVC video frames based on encoded information includes the following steps:
[0055] S1, extract the encoding information of each video frame from the HEVC encoded bitstream, including CU block information, CU prediction mode information, CU-level MV information, CU reference frame list index, and the mapping relationship between the reference frame list index and POC.
[0056] S2, the t-th video frame y t and the t+ith video frame y t+i y is generated as the target frame and the frame to be aligned, respectively. t+i MV missing region map ML t+i and pixel-level MV distribution map MV t+i Including S21~S23;
[0057] S21, for y t+1 Generate CU block diagram based on CU block information. t+i Then, based on the CU prediction model information, CU t+i In the ML (Multi-Frame Prediction) mode, the coded blocks using intra-prediction mode are set to white, and the remaining coded blocks are set to black. t+i The white area is used as the area with missing MV information, and the black area is used as the area with MV information.
[0058] S22, process ML sequentially t+i Each coded block;
[0059] If the s-th encoded block y t+i (s) The coordinates are located in the black area; read its CU-level MV information. And store in ML t+i At the corresponding coordinates;
[0060] If the s-th encoded block y t+i If coordinate (s) is located in the white area, then in the video frame y at frame t+i-1... t+i-1 Find its most similar block and calculate y. t+i (s) and the motion vector MV of the most similar block t+i→t+i-1 (s), as y t+i CU-level MV information (s) Store in ML t+i At the corresponding coordinates;
[0061] S23, ML stores the CU-level MV information of all encoded blocks. t+i Marked as pixel-level MV distribution map MV t+i ;
[0062] S3, generate the CU reference frame POC distribution map for each video frame, where y t+i CU reference frame POC distribution map ND t+i ND t+i Includes y t+i The POC sequence number of each coded block;
[0063] S4, Generate target frame y t and the frame to be aligned y t+i Displacement field M t+i→t This includes steps S41 to S42;
[0064] S41, read y t+i CU reference frame number distribution diagram ND t+i , for y t+i (s) The displacement field M of the coded block is generated according to the following formula. t+i→t (s);
[0065] ,
[0066] In the formula, recPOC t+i (s) is y t+i The s-th coded block y t+i The POC number of (s);
[0067] S42, process y sequentially according to S41. t+i From each encoded block, the displacement field M is obtained. t+i→t ;
[0068] S5, y t+i According to the displacement field M t+i→t Generate alignment frames .
[0069] In S1 of this invention, among the encoding information extracted from the HEVC encoded bitstream, the CU block information refers to HEVC dividing a frame image into multiple grid units (also called encoding blocks) of different sizes, and recording the position, size, and level of each encoding block in the quadtree structure; the CU prediction mode information indicates whether each encoding block uses intra-frame prediction (Intra) or inter-frame prediction (Inter) and the specific partitioning method; the CU-level MV information (Motion Vector, MV) records the horizontal and vertical offsets of each CU using inter-frame prediction relative to the best matching block in the reference frame; the CU reference frame list index indicates the list position of the selected reference frame in the reference image list L0 or L1 (i.e., ref_idx_l0 / ref_idx_l1 in the bitstream information); the mapping relationship from the reference frame list index to the POC maps the above list index to the Picture Order Count of the corresponding reference frame so as to correctly obtain the reference image used for motion compensation and decoding.
[0070] In S21, the method for generating the CU block map based on the CU block information is as follows: The CU block information in the HEVC encoded bitstream includes the partition depth d of each coding block and the coordinates of the coding block, where the coding block coordinates are the coordinates of the top-left corner of the coding block. In the HEVC standard, the resolutions of CU blocks with partition depths of 0, 1, 2, and 3 are 64×64, 32×32, 16×16, and 8×8, respectively. Therefore, the side length of the coding block is... According to the formula Determined, where S is the size of a coding tree unit (CU) in HEVC (64). Given the top-left corner coordinates (x, y) of a coding block U, the pixel region covered by this CU can be defined as a rectangular region: In this formula, (u,v) represents the coordinates of the pixels within the coverage area. Then, the edge pixels of this rectangular region are extracted, and the edge pixels are set to white while the non-edge pixels are set to black, thus obtaining the CU block image as shown below. Figure 2 As shown.
[0071] Regarding the generation of the MV missing region map: In S21, based on the CU block map, the CU prediction mode information of each coding block is combined with black and white markings to obtain the MV missing region map ML. t+i like Figure 3 As shown, the white area represents the missing MV information area, corresponding to the intra-frame predicted coding block, while the black area represents the MV information area.
[0072] Regarding generating pixel-level MV distribution maps (MV) t+i In S22, for ML t+i The coded blocks are processed sequentially and stored in CU-level motion vector (MV) information. The black region coded blocks already contain MV information, so they are read directly. The white region coded blocks do not contain MV information, so the MV information is supplemented by calculating the motion vector by finding the most similar block in the previous frame.
[0073] Regarding the CU reference frame POC distribution map: Based on the CU reference frame list index and the mapping relationship between the reference frame list index and POC, the POC number of each coding block is obtained, thereby generating the CU reference frame POC distribution map.
[0074] Regarding the generation of the target frame y t and the frame to be aligned y t+i Displacement field M t+i→t This is achieved through the following formula of the present invention:
[0075] ,
[0076] See Figure 4 This includes 3 video frames, assuming y t+i The s-th coded block y t+i If the POC number of (s) is t, then according to the first line of the above formula, we can directly use MV. t+i (s) as y t+i Displacement field M of (s) t+i→t (s), such as Figure 4 As indicated by the short arrow. If y t+i If the POC number of (s) is not t, then it needs to be converted according to the second line of the formula, such as... Figure 4 As shown by the medium-long arrow, recPOC t+i (s) is y t+i The s-th coded block y t+i The POC number of (s) then Figure 4 middle For recPOC t+i The video frame containing (s).
[0077] See Figure 5 and Figure 6 , Figure 5 A specific displacement field diagram is provided, corresponding to the size of a video frame. The horizontal and vertical directions represent the width and height of the video frame, respectively, and each red dot represents the coordinates of a pixel. When the displacement field of a coded block is known, the displacement fields of all pixels within its range are identical to those of the coded block; therefore, see [reference needed]. Figure 6 , Figure 6 for Figure 5The image shows a magnified view of a portion of the data. The two rows of coordinates are: the first row represents the coordinates of the pixel within the video frame, which can be obtained by converting the coordinates of the coded block and the pixel coordinates within the coded block; the second row represents the displacement field of the pixel, which is the same as the displacement field of the coded block.
[0078] Example 2: See Figures 1 to 6 Based on Example 1, the specific operations of some steps are given as follows:
[0079] In S23, according to the following formula from y t+i-1 Find the code block that satisfies Conditional encoding block y t+i-1 (u+V h ,v+V v ), as y t+i The most similar block of (s) is obtained, and MV is obtained. t+i→t+i-1 (u,v);
[0080] ,
[0081] In the formula, y t+i (u,v) is y t+i The coordinates of (s), where u and v are the x and y coordinates respectively, argmin(∙) is the argmin function, and V h V v y t+i-1 (u+V h ,v+V v (relative to y) t+i The horizontal and vertical translation coordinates of (s).
[0082] y is generated in S3 t+i CU reference frame POC distribution map ND t+i Specifically:
[0083] S31, Obtain the CU reference frame list L={L0,L1,…,L...} based on the prediction direction. n ,…,L N}, where L n This is the nth reference frame in L;
[0084] S32, if y t+i (s) If the CU reference frame list index of the coded block is n, then its reference frame is L. n L is obtained based on the mapping relationship between the reference frame list index and the POC. n The corresponding POC sequence number is used as the POC sequence recPOC of this coding block. t+i (s), marked on all pixels of the coded block;
[0085] S33, process y sequentially according to S32. t+i From all coded blocks, we obtain the CU reference frame number distribution map ND. t。
[0086] In S42, the obtained displacement field M t+i→t The correction is performed using a weighted filtering method.
[0087] The methods used in S5 include feature alignment based on spatiotemporally deformable convolution, feature alignment based on a feature alignment neural network with encoded information displacement field, and feature alignment based on an optical flow estimation network.
[0088] Example 3: See Figure 7 To illustrate the effects of this invention, based on Example 1, a feature alignment neural network based on the coded information displacement field is constructed, such as... Figure 7 As shown, this includes a front-end and a back-end, and the network is used for feature alignment.
[0089] The front end consists of two parts: displacement field estimation and variation based on encoded information, as shown in the figure.
[0090] Among them, displacement field estimation based on encoded information is the method described in steps S1~S4 of Embodiment 1 of the present invention, used to estimate the t-th video frame y in the bitstream. t and the t+ith video frame y t+i Using these as the target frame and the frame to be aligned, respectively, generate their displacement fields M. t+i→t。
[0091] The change involves using a warp-based transformation module to align the frame y to be aligned. t+i With target frame y t Perform coarse alignment to generate a coarse-aligned frame.
[0092] The backend includes a deformable convolutional offset prediction module and a deformable convolution module. The deformable convolutional offset prediction module takes a coarsely aligned frame as input and outputs a deformable convolutional offset field O. t+i→t Subsequently, a deformable convolution module is used to perform fine-grained alignment of adjacent frames, forming the final aligned frames. .
[0093] Example 4: Based on the scheme in Example 1, the extent to which this method improves the quality of compressed video was tested using the H.265 / HEVC standard test model HM16.9. The test configuration and process are as follows:
[0094] Step 1. The configuration file is encoder_lowdelay_P_main.cfg. The quantization step size of the algorithm in this invention corresponds to 37 in the H.265 / HEVC standard quantization parameter QP.
[0095] Step 2. The encoding object is 18 video sequences from 5 classes of the HEVC standard test sequence, with resolutions including: 2560×1600, 1920×1080, 1280×720, 832×480, and 416×240.
[0096] Step 3. Compress each video sequence using the HM16.9 standard to obtain the encoded bitstream, then decode it to obtain the decoded video. After decoding, use the same compressed video quality enhancement network, setting four feature alignment methods before this network. Method 1: No feature alignment method. Method 2: Feature alignment method based on optical flow estimation network. Optical flow provides a one-to-one correspondence between pixel positions in different frames. By using displacement information described by the optical flow field, the features of the reference frame can be transformed to the coordinate system of the target frame to complete feature alignment. Among them, SpyNet is a lightweight optical flow estimation method, which is widely used in optical flow estimation for feature alignment. Method 3: STDF method, a feature alignment method based on spatio-temporal deformable convolution. Method 4: The method used in this invention is as follows. Figure 7 The feature alignment method shown is based on a feature alignment neural network using the displacement field of encoded information. Encoding and processing are performed using the methods described above and the algorithm framework proposed in this invention. The PSNR value ΔPSNR, representing the improvement in compressed video quality, is calculated based on the PSNR of the decoded and enhanced videos, resulting in Table 1.
[0097] Table 1. Comparison of the results of this invention with other methods in the task of improving the quality of compressed video.
[0098] method Method 1 Method 2 Method 3 Method 4 test sequence ΔPSNR ΔPSNR ΔPSNR ΔPSNR PeopleOnStreet 0.71 1.00 0.88 1.04 Traffic 0.49 1.09 1.06 1.25 BasketballDrive 0.39 0.41 0.31 0.48 BQTerrace 0.52 1.09 1.11 1.41 Cactus 0.44 0.96 0.87 0.93 Kimono 0.36 0.56 0.50 0.59 ParkScene 0.25 0.80 0.87 0.91 BasketballDrill 0.40 1.07 1.10 1.28 BQMall 0.47 1.23 1.03 1.24 PartyScene 0.25 1.57 1.40 1.55 RaceHorses 0.61 0.16 0.34 0.30 BasketballPass 0.47 1.19 1.21 1.21 BlowingBubbles 0.31 1.23 1.19 1.39 BQSquare 0.25 2.40 1.94 2.51 RaceHorses 0.59 0.69 0.58 0.71 FourPeople 0.72 1.70 1.43 1.79 Johnny 0.51 1.17 1.17 1.48 KristenAndSara 0.75 1.33 1.33 1.59 AVERAGE 0.47 1.09 1.02 1.20
[0099] According to the statistics in Table 1, the PSNR of the method of the present invention exceeds that of other alignment methods, which can effectively improve the quality of compressed video and enhance network performance.
[0100] Furthermore, the computational complexity of the method of this invention compared with the feature alignment method based on optical flow estimation networks is shown in Table 2 below:
[0101] Table 2. Comparison of computational complexity
[0102] Model Number of parameters Computational cost (FLOPs) SpyNet (Optical Flow Method) 1.2M 149.8G This invention 90.2K 34.0G
[0103] In Table 2, Parameters represents the number of model parameters, and FLOPs represents the number of floating-point operations per second.
[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A low-complexity feature alignment method for HEVC video frames based on encoded information, characterized in that, Includes the following steps: S1, extract the encoding information of each video frame from the HEVC encoded bitstream, including CU block information, CU prediction mode information, CU-level MV information, CU reference frame list index, and the mapping relationship between the reference frame list index and POC. S2, the t-th video frame y t and the t+ith video frame y t+i y is generated as the target frame and the frame to be aligned, respectively. t+i MV missing region map ML t+i and pixel-level MV distribution map MV t+i Including S21~S23; S21, for y t+1 Generate CU block diagram based on CU block information. t+i Then, based on the CU prediction model information, CU t+i In the ML (Multi-Frame Prediction) mode, the coded blocks using intra-prediction mode are set to white, and the remaining coded blocks are set to black. t+i The white area is used as the area with missing MV information, and the black area is used as the area with MV information. S22, process ML sequentially t+i Each coded block; If the s-th encoded block y t+i (s) The coordinates are located in the black area; read its CU-level MV information. And store in ML t+i At the corresponding coordinates; If the s-th encoded block y t+i If coordinate (s) is located in the white area, then in the video frame y at frame t+i-1... t+i-1 Find its most similar block and calculate y. t+i (s) and the motion vector MV of the most similar block t+i→t+i-1 (s), as y t+i CU-level MV information (s) Store in ML t+i At the corresponding coordinates; S23, ML stores the CU-level MV information of all encoded blocks. t+i Marked as pixel-level MV distribution map MV t+i ; S3, generate the CU reference frame POC distribution map for each video frame, where y t+i CU reference frame POC distribution map ND t+i ND t+i Contains y t+i The POC sequence number of each coded block; S4, Generate target frame y t and the frame to be aligned y t+i Displacement field M t+i→t This includes steps S41 to S42; S41, read y t+i CU reference frame number distribution diagram ND t+i , for y t+i (s) The displacement field M of the coded block is generated according to the following formula. t+i→t (s); , In the formula, recPOC t+i (s) is y t+i The s-th coded block y t+i The POC number of (s); S42, process y sequentially according to S41. t+i From each encoded block, the displacement field M is obtained. t+i→t ; S5, y t+i According to the displacement field M t+i→t Generate alignment frames .
2. The method for low-complexity feature alignment of HEVC video frames based on encoded information according to claim 1, characterized in that, In S23, according to the following formula from y t+i-1 Find the code block that satisfies Conditional encoding block y t+i-1 (u+V h ,v+V v ), as y t+i The most similar block of (s) is obtained, and MV is obtained. t+i→t+i-1 (u,v); , In the formula, y t+i (u,v) is y t+i The coordinates of (s), where u and v are the x and y coordinates respectively, argmin(∙) is the argmin function, and V h 、V v y t+i-1 (u+V h ,v+V v (relative to y) t+i The horizontal and vertical translation coordinates of (s).
3. The method for low-complexity feature alignment of HEVC video frames based on encoded information according to claim 1, characterized in that, y is generated in S3 t+i CU reference frame POC distribution map ND t+i Specifically: S31, Obtain the CU reference frame list L={L0,L1,…,L...} based on the prediction direction. n ,…,L N }, where L n This is the nth reference frame in L; S32, if y t+i (s) If the CU reference frame list index of the coded block is n, then its reference frame is L. n L is obtained based on the mapping relationship between the reference frame list index and the POC. n The corresponding POC sequence number is used as the POC sequence recPOC of this coding block. t+i (s), marked on all pixels of the coded block; S33, process y sequentially according to S32. t+i From all coded blocks, we obtain the CU reference frame number distribution map ND. t .
4. The method for low-complexity feature alignment of HEVC video frames based on encoded information according to claim 1, characterized in that, In S42, the obtained displacement field M t+i→t The correction is performed using a weighted filtering method.
5. The method for low-complexity feature alignment of HEVC video frames based on encoded information according to claim 1, characterized in that, The methods used in S5 include feature alignment based on spatiotemporally deformable convolution, feature alignment based on a feature alignment neural network with encoded information displacement field, and feature alignment based on an optical flow estimation network.
Citation Information
Patent Citations
Video coding method and device based on long-term reference frame, equipment and storage medium
CN111405282A
Inter-frame prediction method and device, computer equipment and storage medium
CN117615129A