A predictive coding method, device, and storage medium

By using position prediction filters in video prediction encoding, using the encoding unit position information of the current frame and the reference frame to predict the position information in the backward frame, the accuracy and calculation amount problems of the prior art in scenarios with high noise are solved, and more efficient and accurate video encoding is achieved.

CN119788877BActive Publication Date: 2025-05-30ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510282000.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-30
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing video prediction encoding method based on optical flow affine correction has reduced accuracy in scenarios with high noise, and the calculation amount is large, which affects the encoding performance.

Method used

Using a position prediction filter, the encoding unit position information in the current frame and the reference frame is used to predict the position information of the encoding unit in the backward frame, thereby guiding the encoding process of the backward frame.

Benefits of technology

It improves the accuracy and efficiency of video encoding, reduces encoding redundancy and errors caused by inaccurate location information, and improves the overall quality and performance of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788877B_ABST
    Figure CN119788877B_ABST
Patent Text Reader

Abstract

The present application discloses a prediction coding method, device, and storage medium. The method includes: inputting the position information of each coding unit in at least some coding units in the current frame and the position information of the corresponding coding unit of each coding unit in the reference frame into a position prediction filter, so that the position prediction filter predicts the position information of each coding unit in the backward frame by using the position information of each coding unit in the current frame and the corresponding position information in the reference frame; and encoding the backward frame by using the position information of at least some coding units in the backward frame. By the above method, the present application can improve the accuracy and efficiency of coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and particularly to a prediction coding method, device, and storage medium. Background Art

[0002] In the field of video prediction coding, the traditional prediction coding method based on optical flow affine correction has certain limitations. This method first assumes that the inter-frame motion of an image undergoes translation, rotation, zooming, or other irregular motions. It generates sub-block predictions based on affine motion compensation of sub-blocks, and then assumes that the motion of the sub-blocks is smooth. It corrects the difference between the affine motion vector and the original sub-block's motion vector through the optical flow gradient value to balance the motion compensation in both the spatial and temporal domains, which can improve the prediction accuracy, reduce distortion, and bit rate to a certain extent.

[0003] However, for scenarios with high noise, the accuracy of the prediction coding method based on optical flow affine correction decreases. Moreover, since additional gradient and affine calculations are required for each sub-block, the computational complexity increases significantly, seriously affecting the coding performance.

[0004] Therefore, there is an urgent need for a new prediction coding method that can reduce noise interference, improve prediction accuracy, reduce the inter-frame prediction computational complexity, and further effectively enhance the coding performance. Summary of the Invention

[0005] To solve the above technical problems, the technical solution adopted in this application is: to provide a prediction coding method, device, and storage medium to at least solve the problems of low accuracy of the prediction coding result and large computational complexity leading to reduced coding performance in the related art.

[0006] According to an embodiment of the present invention, a prediction coding method is provided, including:

[0007] Input the position information of each coding unit in at least some coding units in the current frame and the position information of the corresponding coding unit of each coding unit in the reference frame into a position prediction filter, so that the position prediction filter predicts the position information of each coding unit in the backward frame using the position information of each coding unit in the current frame and the corresponding position information in the reference frame; wherein, the reference frame is the forward frame of the current frame;

[0008] Encode the backward frame using the position information of at least some coding units in the backward frame.

[0009] To solve the above technical problems, a technical solution adopted in this application is: to provide an electronic device, including a memory and a processor. Among them, the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the prediction coding method in the above technical solution.

[0010] To solve the above technical problems, a technical solution adopted in this application is: to provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the prediction coding method in the above technical solution.

[0011] Through the above solution, the beneficial effect of this application is: the prediction coding method provided by this application can, through a position prediction filter, utilize the position information of coding units in the current frame and the reference frame to more accurately grasp the position changes of various elements in the video content, thereby more accurately predicting the position information of the coding unit in the backward frame, and then effectively guiding the coding process of the backward frame, making the encoded video closer to the original video during reconstruction, improving the accuracy and efficiency of coding, reducing coding redundancy and errors caused by inaccurate position information, enhancing the accuracy and efficiency of prediction coding, and thus improving the overall quality and performance of video coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0013] Figure 1 is a flowchart of an embodiment of the prediction coding method provided by this application;

[0014] Figure 2 is a flowchart of a working example of a Kalman filter provided by this application;

[0015] Figure 3 is a flowchart of another embodiment of the prediction coding method provided by this application;

[0016] Figure 4 is a flowchart of the coding unit division method when no target is detected provided by this application;

[0017] Figure 5 is a flowchart of the coding unit division method when a target is detected provided by this application;

[0018] Figure 6 is a schematic structural diagram of an embodiment of an electronic device provided by this application;

[0019] Figure 7 is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only partial embodiments of the present application rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0021] The mention of "embodiment" in the present application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0022] It should be noted that the terms "first", "second", etc. in the present application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0023] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the prediction coding method provided by the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to the Figure 1 flow order shown. As Figure 1 shown, this embodiment includes:

[0024] S110: Input the position information of each coding unit in at least some coding units in the current frame, and the position information of the corresponding coding unit of each coding unit in the reference frame into the position prediction filter.

[0025] In this embodiment, the current frame represents the image frame being encoded during the video encoding process. The reference frame is the forward frame of the current frame and is the encoded frame used for predicting the current frame. The correlation between the reference frame and the current frame can be used to reduce the amount of encoded data. The coding unit is the basic unit for partitioning and processing images in video encoding, and each coding unit contains a certain number of pixels.

[0026] In the prediction coding scheme provided in this application, for each image frame, it is necessary to partition the image area of the image frame to determine the image partitioning result in the image frame, so as to obtain multiple coding units in the image frame.

[0027] The methods of image partitioning include fixed-size partitioning, quadtree partitioning, and hybrid partitioning that combines multiple strategies, etc., which are not limited in this application. In one example, the image frame is partitioned into multiple image blocks according to a preset fixed size, the multiple image blocks are grouped according to the image content in the image blocks, and the image blocks with similar content are partitioned into an image block group. Each image block group is determined as a coding unit, and finally the image partitioning result of the image frame is obtained. Exemplarily, the preset fixed size is 16×16 pixels. The image frame is divided into multiple image blocks according to the fixed size, the gradient texture value of each image block is calculated respectively, and through image similarity calculation methods such as SSIM (Structural Similarity), the similarity between the image blocks is compared, and the image blocks with similar similarity are partitioned into the same image block group. Each image block group is determined as a coding unit, and finally the image partitioning result of the image frame is obtained.

[0028] For each coding unit, its position information in the current frame and the reference frame is obtained by means of coordinate calculation, etc. Exemplarily, in a pixel coordinate-based manner, the horizontal and vertical coordinates of the preset pixel point of the coding unit in the image are recorded as the position information, where the preset pixel point of the coding unit can be the diagonal pixel of the coding unit.

[0029] The obtained position information of at least some coding units in the current frame and the corresponding coding unit position information in the reference frame are input into the filter according to the input format and order specified by the position prediction filter. Among them, the position prediction filter can accurately predict the positions of these coding units in the backward frame according to the position information of the coding units in the current image frame and the reference frame. By accurately predicting the positions of the coding units, it provides strong support for subsequent coding operations, thereby improving the overall video or image coding quality; preferably, the position prediction filter is a Kalman filter, and the Kalman filter is an optimal estimation filter based on the state space model of a linear system.

[0030] After receiving the input information, the position prediction filter activates the internal algorithm module. According to the preset algorithm rules, this module analyzes and processes the input position information. For example, the algorithm module can adopt a model obtained through machine learning training, and based on the relationship between the position information learned from a large amount of existing video data and the coding mode, it matches and predicts the current input position information, and outputs the predicted coding mode and parameters of the current coding unit.

[0031] Exemplarily, a Kalman filter is used to analyze and process the input position information. Its working process mainly consists of two core steps: the prediction step and the update step. In the prediction step, based on the position estimation of the coding unit in the previous frame and the system dynamics (such as the object motion law), the position of the coding unit in the current frame is predicted. In the update step, by combining the position information of the coding unit actually observed in the current frame and the reference frame, the predicted position is corrected to obtain a more accurate position estimation of the coding unit. By continuously repeating the prediction and update steps, the Kalman filter can continuously and accurately predict the position of the coding unit between different frames. Even in the presence of system noise and observation noise, it can effectively track the position change of the coding unit and improve the accuracy and quality of coding.

[0032] In one embodiment, please refer to Figure 2 , Figure 2 which is a schematic flowchart of an example of the Kalman filter provided in this application, specifically as follows: The diagonal vertex coordinates of each coding unit in the reference frame are respectively input into the Kalman filter for initialization; the diagonal vertex coordinates of the corresponding coding unit in the current frame are input for the prediction of the backward frame; the prediction distortion between the original frame and the predicted frame of the backward frame is calculated, and it is judged whether the prediction distortion exceeds the preset threshold; the coding units with prediction distortion not exceeding the preset threshold are encoded, the backward frame is used as the current frame, and the diagonal vertex coordinates of the corresponding coding unit in the input current frame are returned for the step of predicting the backward frame; the image area where the coding unit with prediction distortion exceeding the preset threshold is located is re-divided into coding units.

[0033] After determining multiple coding units in the reference frame and the current frame, at least some coding units in the current frame and at least some coding units corresponding to the coding units (CUs) in the reference frame are determined. The position information of each coding unit in at least some coding units in the current frame and the position information of the corresponding coding unit of each coding unit in the reference frame are input into the position prediction filter, so that the position prediction filter uses the position information of each coding unit in the current frame and the corresponding position information in the reference frame to predict the position information of each coding unit in the backward frame, and applies the position information of each coding unit in the backward frame to the actual coding process of the coding units in the next frame.

[0034] S120: Encode a backward frame using the position information of at least some coding units in the backward frame.

[0035] In a video coding sequence, a backward frame is the next frame of the current frame or any subsequent picture frame of the current frame. Use the position information of at least some coding units in the current frame obtained by the above step S110 to encode the backward frame.

[0036] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another embodiment of the prediction coding method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 3 the shown process sequence. As Figure 3 shown, this embodiment includes:

[0037] S210: Input the position information of multiple coding units in a reference frame into a position prediction filter.

[0038] Determine multiple coding units in the reference frame that need to be input into the position prediction filter, and input the position information of the multiple coding units into the position prediction filter for initialization or combine with the prediction information in the position prediction filter to predict the position information of each coding unit in the backward frame (i.e., the subsequent current frame) of the reference frame. Among them, the prediction information is the information relied on by the position prediction filter when predicting the coding unit position, including historical frame coding unit position information, coding unit motion vectors, and / or filter own model parameters. By analyzing this information and combining the currently input coding units, position prediction is completed.

[0039] In one embodiment, determine a first coding unit and a second coding unit in the reference frame. Among them, the first coding unit is at least some coding units in the reference frame, and at least some coding units in the reference frame are coding units with a prediction distortion less than a preset threshold in the reference frame; the second coding unit is the coding units in the reference frame other than the first coding unit, and may include coding units with a prediction distortion not less than the preset threshold and / or coding units without prediction information in the reference frame. By dividing the coding units in the reference frame into the first coding unit and the second coding unit, different processing methods can be adopted for coding units with different characteristics to improve the accuracy of prediction coding.

[0040] If there is a first coding unit in the reference frame, input the position information of the first coding unit in the reference frame into the position prediction filter to combine with the prediction information in the filter to predict the position information of each coding unit in the backward frame.

[0041] If there are second coding units in the reference frame, input the position information of at least some of the second coding units in the reference frame into the position prediction filter to combine the prediction information in the filter to predict the position information of each coding unit in the backward frame; and / or, input the position information of at least some of the second coding units in the reference frame into the position prediction filter for initialization.

[0042] S220: Input the position information of the coding units in the current frame that have a corresponding relationship with the coding units in the reference frame into the position prediction filter.

[0043] If there is a corresponding relationship between the coding units in the current frame and the coding units in the reference frame, input the position information of the coding units in the current frame that have a corresponding relationship with the coding units in the reference frame into the position prediction filter, so that the position prediction filter can use the position information of each coding unit in the current frame and the corresponding position information in the reference frame to predict the position information of each coding unit in the backward frame.

[0044] Among them, the corresponding relationship represents a mapping relationship determined by a specific algorithm, which means that the coding units in the current frame and the reference frame have similarity or relevance in content, structure or position. In one example, an algorithm based on feature matching can be adopted to extract features of the coding units in the current frame and the reference frame, such as extracting texture features, edge features, etc. By comparing these features, the corresponding relationship between the coding units in the two frames is determined. For example, for coding units with similar texture distributions and edge contours, it is determined that they have a corresponding relationship. In another example, the corresponding relationship between the coding units in the current frame and the reference frame can also be determined based on the position information of each coding unit in the reference frame predicted in the current frame. There are various methods to determine the corresponding relationship between the coding units in the current frame and the coding units in the reference frame, which are not limited here.

[0045] In one embodiment, the coding units in the current frame that have a corresponding relationship with the coding units in the reference frame include the coding units in the current frame that correspond to the first coding units in the reference frame. Determine at least some of the coding units in the current frame that correspond to the first coding units in the reference frame as at least some of the coding units in the current frame, and input at least some of the coding units in the current frame into the position prediction filter.

[0046] If the position prediction filter predicts the position information of the first coding unit in the current frame based on the position information of the first coding unit, use the reference frame and the position information of the first coding unit in the current frame to predict the current frame, and obtain the predicted frame of the current frame, where the predicted frame includes the predicted blocks of the coding units in the current frame that correspond to the first coding unit.

[0047] Determine that the coding units with relatively accurate predicted position information in the current frame are at least part of the coding units in the current frame, and input the position information of these coding units into the position prediction filter to predict the position information in the backward frame. Specifically, a first difference threshold can be preset, and the difference degree between each coding unit in the original frame of the current frame and the corresponding coding unit in the predicted frame is calculated. If the difference degree is greater than the first difference threshold, it indicates that the predicted position information of the coding unit is inaccurate; if the difference degree is not greater than the first difference threshold, it indicates that the predicted position information of the coding unit is relatively accurate, and the coding units with relatively accurate predicted position information in the current frame are determined as at least part of the coding units in the current frame.

[0048] In one example, the original frame of the current frame is matched with the predicted frame to determine the prediction distortion of each coding unit corresponding to the first coding unit in the current frame. The coding units at the same positions in the original frame and the predicted frame of the current frame are matched, and the mean square error (MSE) is used to calculate the prediction distortion between the coding units. The difference between each pixel position of the coding units in each pair of matched original frame and predicted frame is calculated to obtain the MSE value of the coding unit. The MSE value reflects the difference at the pixel level of the coding unit, and the smaller the value, the higher the prediction accuracy. In other examples, other methods such as structural similarity (SSIM) can also be used to calculate the prediction distortion of each coding unit in the original frame and the predicted frame, which is not limited here.

[0049] After obtaining the prediction distortion of each coding unit corresponding to the first coding unit in the current frame, a preset threshold is set, and the relationship between the prediction distortion of each coding unit and the preset threshold is compared. The position information of the coding units with prediction distortion less than the preset threshold in the current frame is input into the position prediction filter, so that the position prediction filter uses the position information of the coding units with prediction distortion less than the preset distortion threshold and the position information of the corresponding coding units in the reference frame to predict the position information of the coding units with prediction distortion less than the preset distortion threshold in the backward frame.

[0050] Clear the position information of the coding units with prediction distortion not less than the preset threshold in the current frame, and at the same time clear the prediction information of these coding units stored in the position prediction filter, so as to re-determine at least one coding unit and its position information in the image area where the coding units with prediction distortion greater than the preset threshold in the current frame are located based on the original frame of the current frame.

[0051] In one embodiment, the coding units in the current frame that have a corresponding relationship with the coding units in the reference frame further include the coding units in the current frame that correspond to the second coding unit. Determine the coding units in the current frame that correspond to the second coding unit, and input their position information into the position prediction filter. Among them, the second coding unit includes a third coding unit and / or a fourth coding unit. The third coding unit is a coding unit in the reference frame with a prediction distortion greater than a preset threshold, and the fourth coding unit is a coding unit in the reference frame that inputs position information into the filter for initialization.

[0052] If there are coding units in the current frame that correspond to the third coding unit, determine the coding units in the current frame that correspond to the third coding unit, and input their position information into the position prediction filter for initialization.

[0053] If there are coding units in the current frame that correspond to the fourth coding unit, determine the coding units in the current frame that correspond to the fourth coding unit, and input their position information into the position prediction filter, so that the position prediction filter uses the position information of the fourth coding unit and the position information of the coding units in the current frame that correspond to the fourth coding unit for prediction.

[0054] Before calculating the degree of difference between each coding unit in the original frame of the current frame and the corresponding coding unit in the predicted frame, it is necessary to first obtain the original frame of the current frame and the coding unit division method in the original frame. There may be some regions with relatively large changes in the original frame compared to the reference frame, and it is necessary to re-determine the coding unit division method for these regions to obtain the true coding unit division result of the current frame.

[0055] In one embodiment, deep learning methods such as object detection algorithms can be used to determine the background region and target regions in the original frame. If there are multiple target regions, the target regions can be numbered in advance. Match the coding units in the background region and target regions in the original frame and the corresponding regions in the reference frame respectively to determine the regions to be updated in the original frame of the current frame, update the coding unit division method in the regions to be updated, and obtain at least one coding unit in the regions to be updated.

[0056] In one example, a target detection algorithm is run, and the image region where each target is located is divided into a target region, and the remaining image regions where no target is detected are classified as background regions. If there is a target region corresponding to a newly emerged target in the original frame, a pre-set image division method can be used to divide the target region to determine at least one coding unit in the target region. If there is a target region in the original frame with the same target number as in the reference frame, further determine the degree of difference between each coding unit in the target region in the reference frame and the original frame. If the degree of difference is greater than the second difference threshold, the region where the coding unit is located is designated as a region to be updated; if the degree of difference is less than the second difference threshold, the coding unit division method corresponding to the region of the coding unit in the reference frame is followed.

[0057] Exemplarily, the frame difference method is used to determine the degree of texture change of each coding unit in the original frame and the reference frame, and the region to be updated in the original frame is determined based on the degree of texture change of each coding unit. A texture change degree threshold is preset. For all regions where the coding units with texture change degrees not exceeding the texture change degree threshold are located, the division method of the corresponding image region in the reference frame is followed; for all regions in the original frame where the coding units with texture change degrees exceeding the texture change degree threshold are located, the region to be updated is divided into multiple image blocks according to a preset fixed size, the gradient texture values of each image block are calculated respectively, and by using image similarity calculation methods such as SSIM (Structural Similarity), the similarity between image blocks is compared, and the image blocks with similar similarities are divided into the same image block group, and each image block group is determined as a coding unit, and finally the coding unit division method of the region to be updated is obtained.

[0058] In some embodiments, if the prediction distortion of the third coding unit in the reference frame is greater than the preset threshold or the position information of the coding unit in the current frame with prediction distortion not less than the preset threshold is cleared, then there is no image division result for the image region corresponding to the region where the third coding unit is located in the reference frame in the current frame and the region where the coding unit in the current frame with prediction distortion not less than the preset threshold is located. In this case, if there is a region without an image division result in the predicted frame of the current frame, then determine this region as the region to be updated image region, and re-perform coding unit division on the region to be updated image region to determine at least one coding unit in the region to be updated image region.

[0059] In one embodiment, a first image region to be updated in the current frame is determined based on the object detection result of the current frame. The first image region to be updated is divided into coding units to determine at least one coding unit in the first image region to be updated. For a first region in the current frame other than the first image region to be updated, the coding unit division method of the corresponding region in the reference frame is followed to determine at least one coding unit in the first region, where the coding units of the corresponding region in the reference frame include at least one of a third coding unit and a fourth coding unit, and the at least one coding unit in the first region corresponds one-to-one with the coding units of the corresponding region in the reference frame.

[0060] In another embodiment, after determining a prediction block with a prediction distortion less than a preset threshold, a second image region to be updated in a second region other than the prediction block with a prediction distortion less than the preset threshold is determined based on the object detection result of the current frame. The second image region to be updated is divided into coding units to determine at least one coding unit in the second image region to be updated. For a third region in the second region other than the second image region to be updated, the coding unit division method of the corresponding region in the reference frame is followed to determine at least one coding unit in the third region, where the coding units of the corresponding region in the reference frame include at least one of a third coding unit and a fourth coding unit, and the at least one coding unit in the third region corresponds one-to-one with the coding units of the corresponding region in the reference frame.

[0061] In other embodiments, a first image region to be updated in the current frame is determined based on the object detection result of the current frame, and the first image region to be updated is divided into coding units to determine at least one coding unit in the first image region to be updated. For a first region in the current frame other than the image region to be updated, the coding unit division method of the corresponding region in the reference frame is followed to determine at least one coding unit in the first region, and the at least one coding unit in the first region corresponds one-to-one with the coding units of the corresponding region in the reference frame. The position information in at least one coding unit in the first region is input into a position prediction filter, so that the position prediction filter predicts the position information of at least a part of the coding units in the first region in a backward frame by using the position information in at least a part of the coding units in the first region and the position information of the corresponding coding units in the reference frame.

[0062] S230: Encode the backward frame by using the position information of at least a part of the coding units in the backward frame.

[0063] The predicted frame of the backward frame is obtained by using the position information of the coding units in the backward frame corresponding to the coding units with a corresponding relationship in the current frame and the reference frame. The predicted frame includes prediction blocks of the coding units corresponding to at least a part of the coding units in the current frame in the backward frame.

[0064] Match the original frame and the predicted frame of the backward frame to determine the degree of difference between at least some of the coding units in the backward frame and those in the current frame, so as to predict subsequent frames using the coding units with a degree of difference not greater than the first difference threshold.

[0065] In one embodiment, match the original frame and the predicted frame of the backward frame to determine the prediction distortion of each coding unit corresponding to at least some of the coding units in the backward frame and those in the current frame. Use the prediction blocks in the backward frame with a prediction distortion less than a preset threshold to code the coding units corresponding to the prediction blocks.

[0066] Assume that the coding units in the backward frame with a prediction distortion less than the preset threshold are the fifth coding units. If there are fifth coding units in the backward frame, use the current frame as the reference frame and the backward frame as the current frame. At least some of the coding units in the current frame include the fifth coding units, and return to execute the step of inputting the position information of the coding units with a corresponding relationship between the coding units in the current frame and the reference frame into the position prediction filter.

[0067] In one embodiment, if there are fifth coding units in the backward frame, use the current frame as the reference frame and the backward frame as the current frame. Assume that the sixth coding unit is the coding unit in the current frame with a prediction distortion greater than the preset threshold. For the sixth coding unit in the current frame, determine the coding unit division method of the area where the sixth coding unit is located based on the object detection result of the current frame, so as to determine at least one coding unit in the area where the sixth coding unit is located, and code at least one coding unit in the area where the sixth coding unit is located.

[0068] In one embodiment, if there are fifth coding units in the backward frame, use the current frame as the reference frame and the backward frame as the current frame. Determine the coding unit division method of the area in the current frame except for the areas where the fifth coding unit and the sixth coding unit are located based on the object detection result of the current frame, so as to determine at least one coding unit in the area in the current frame except for the areas where the fifth coding unit and the sixth coding unit are located, and code at least one coding unit in the area in the current frame except for the areas where the fifth coding unit and the sixth coding unit are located.

[0069] In one embodiment, if at least one coding unit in the area in the current frame except for the areas where the fifth coding unit and the sixth coding unit are located includes the coding unit in the seventh coding unit with a corresponding relationship with the reference frame and the position information has been input into the position prediction filter once, at least some of the coding units in the current frame include the coding unit in the seventh coding unit with the position information input into the position prediction filter once.

[0070] In one embodiment, if at least one coding unit in the region of the current frame other than the regions where the fifth coding unit and the sixth coding unit are located includes a coding unit in the seventh coding unit that has a corresponding relationship with the reference frame and does not input the position information into the position prediction filter, the position information of the coding unit in the seventh coding unit that does not input the position information into the position prediction filter is input into the position prediction filter for initialization.

[0071] To more clearly illustrate the prediction coding method provided by this application, in this embodiment, taking the natural frame order in the video sequence processed frame by frame as an example, the following specific embodiments of the prediction coding method are provided for exemplary illustration:

[0072] 1. Processing of the first frame:

[0073] Use the CNN algorithm to perform foreground object detection on the first frame. Taking Figure 4 as an example, Figure 4 is a schematic flowchart of the coding unit division method when no target is detected provided by this application. The specific process is as follows:

[0074] If no target is detected, the first frame is a background frame, and the image area of the current frame is all background areas. Take this background frame as the starting frame of the encoder, and process its YUV data. Divide the background area into multiple image blocks according to a preset image block size (such as 16×16 pixels), and calculate the gradient texture of each image block respectively. Set a first image texture similarity threshold, and calculate the texture similarity between image blocks through the image similarity calculation method. Group the image blocks with texture similarity greater than the first image texture similarity threshold into one group, and determine that the image blocks in the same group are the same coding unit (CU), and finally determine all the coding units in the background area.

[0075] It should be noted that in this embodiment, taking the background frame as the starting frame of the encoder can establish a stable background reference model for subsequent video processing. In the absence of target interference, the texture features of the background frame are relatively single, and it is easier to group the smallest macroblocks through the image similarity calculation method, reducing the complexity of subsequent processing. Moreover, the coding units determined based on the background frame can continue to be used in subsequent frames as long as the background does not change significantly, and only the target area and the changing background area need to be adjusted, thereby improving the overall efficiency of coding.

[0076] 2. Processing of the second frame:

[0077] Any background frame that has not detected a target before the first detection of a target can be regarded as the first frame, and the image frame where the target is first detected is used as the second frame. Please refer to Figure 5 , Figure 5This is a flow chart of the method for dividing the coding unit when a target is detected provided by the present application. In the processing of the second frame in which the target is detected, the target detection algorithm is run, foreground target detection is performed on the second frame, and the image area where each target is located is divided into a target area, and an ID identifier is set for each target area. The remaining image area where no target is detected is divided into a background area, and the division method of the coding units of each target area and background area in the second frame is determined in the following way.

[0078] A. In each target area, the target area is divided according to a preset image block size in the target detection frame to obtain multiple image blocks in the target area, a second image texture similarity threshold is set, and the texture similarity between the image blocks is calculated by an image similarity calculation method. The image blocks with texture similarity greater than the second image texture similarity threshold are grouped together, and the image blocks in the same group are determined to be the same coding unit, and finally all the coding units in each target area are determined.

[0079] B. In the background area, based on the degree of texture change of each coding unit in the first frame and the second frame, first determine whether the image area where each coding unit is located needs to be updated, determine the area to be updated in the background area, and re-divide the coding units in the image area that needs to be updated.

[0080] A texture change degree threshold is preset, and for the regions where all coding units whose texture change degrees do not exceed the texture change degree threshold are located, the coding unit division method of the corresponding image region in the first frame (i.e., the reference frame of the second frame) is used;

[0081] For the area where all coding units whose texture change degree exceeds the texture change degree threshold in the second frame are located, the area to be updated is divided into multiple image blocks according to a preset fixed size, and the gradient texture value of each image block is calculated respectively. The similarity between image blocks is compared by image similarity calculation methods such as SSIM (Structural Similarity), and image blocks with similar similarity are divided into the same image block group. Each image block group is determined as a coding unit, and finally the coding unit division method of the area to be updated is obtained.

[0082] At this point, after obtaining the coding units of all image areas in the second frame, the position information of each coding unit, such as the diagonal vertex coordinates of the coding unit, is obtained. The diagonal vertex coordinates of each coding unit are input into the Kalman filter for initialization. The Kalman filter is used to predict the position change of the coding unit in the future. The initialization operation is to enable the filter to perform subsequent prediction calculations based on the position information of each coding unit in the second frame, laying the foundation for subsequent dynamic tracking and prediction.

[0083] 3. Processing of the third frame:

[0084] Run the object detection algorithm to perform foreground object detection on the third frame. Update the coding unit partitioning method in this image frame based on the object detection results of the third frame to obtain the coding unit partitioning method in the third frame and the correspondence between each coding unit in the third frame and the coding units in the second frame. The method for updating the coding unit partitioning method is as follows:

[0085] (1) For the target regions corresponding to the newly emerged targets in the third frame: Re-group the image blocks within each target region according to the texture similarity. The image blocks with similar texture similarities are grouped together, and the image blocks in the same group are determined to be the same coding unit, thereby determining the coding unit partitioning method for the newly emerged target regions in the third frame.

[0086] (2) For the target regions in the third frame corresponding to the targets that appeared in the second frame (with the same target ID as in the second frame) and the background regions in the third frame: Based on the degree of texture change of each coding unit between the second frame and the third frame, first determine whether the image region where each coding unit is located needs to be updated, identify the regions to be updated in the target regions and / or background regions, and re-partition the coding units in the image regions that need to be updated.

[0087] Preset a threshold for the degree of texture change. For the regions where the coding units with all degrees of texture change not exceeding the threshold are located, use the coding unit partitioning method of the corresponding image region in the second frame (i.e., the reference frame of the third frame);

[0088] For the regions where the coding units with all degrees of texture change exceeding the threshold in the second frame are located, divide the regions to be updated into multiple image blocks according to a preset fixed size, calculate the gradient texture values of each image block respectively, compare the similarities between the image blocks through image similarity calculation methods such as SSIM (Structural Similarity), group the image blocks with similar similarities into the same image block group, and determine each image block group as a coding unit, finally obtaining the coding unit partitioning method for the regions to be updated.

[0089] So far, obtain the coding unit partitioning method for all image regions in the third frame and the correspondence between each coding unit in the third frame and the second frame. Input the diagonal vertex coordinates of each coding unit into the initialized Kalman filter, and the Kalman filter predicts the position information of each coding unit in the fourth frame based on the position information of each coding unit in the second frame and the third frame.

[0090] 4. Processing of the fourth frame:

[0091] Encode the fourth frame based on the position information of each coding unit in the fourth frame obtained by prediction to obtain the predicted frame of the fourth frame. Run the object detection algorithm, and obtain the coding unit partitioning method of the original frame of the fourth frame based on the object detection result and the coding unit partitioning method of the third frame, so as to evaluate the accuracy of the prediction based on the degree of difference between the predicted frame and the original frame of the fourth frame. The update of the coding unit partitioning method of the original frame of the fourth frame is consistent with the processing logic of the update method of the coding unit partitioning method in the above-mentioned third frame, and will not be elaborated here.

[0092] After determining the coding unit partitioning method in the original frame of the fourth frame, calculate the prediction distortion of the coding units at the same positions in the original frame and the predicted frame of the fourth frame to evaluate the accuracy of the coding unit prediction in the predicted frame. For the area where the coding units with prediction distortion less than the preset threshold are located, it indicates that the prediction of the predicted frame in this area (background or object) is relatively accurate; for the area where the coding units with prediction distortion not less than the preset threshold are located, it means that there is a large error in the prediction of the predicted frame in this area, and it may be necessary to adjust the coding process, such as re-partitioning the CU, adjusting the motion vector or changing the quantization parameter, etc., to improve the quality of video coding.

[0093] Exemplarily, for the first image area where the coding units with prediction distortion less than the preset threshold are located, retain the coding unit partitioning method of the first image area in the predicted frame, and input the position information of these coding units into the Kalman filter to predict the position information corresponding to each coding unit in the next frame based on the coding unit position information in the first image area.

[0094] For the second image area where the coding units with prediction distortion less than the preset threshold are located, clear the coding unit partitioning method in the second image area of the predicted frame, and the image data in the second image area of the predicted frame can be updated based on the image area corresponding to the second image area of the predicted frame in the original frame.

[0095] 5. Processing of the fifth frame:

[0096] Encode the fifth frame based on the position information of at least some coding units obtained by prediction in the fifth frame to obtain the predicted frame of the fifth frame.

[0097] (1) The processing method of the coding units in the image area (the image area corresponding to the first image area) where the coding unit partitioning method is not re-determined in the fifth frame is consistent with the processing method of the fourth frame, and the processing method of the fourth frame is repeatedly executed, which will not be elaborated here.

[0098] (2) Initialize the position information of each coding unit in the image area (the image area corresponding to the second image area) where the coding unit partitioning method is re-determined by inputting it into the Kalman filter.

[0099] 6. Processing of the sixth frame:

[0100] Encode the sixth frame based on the position information of at least some of the coding units predicted in the sixth frame to obtain the predicted frame of the sixth frame.

[0101] (1) Repeatedly execute the processing method of the fourth frame for the coding units with prediction information in the image area where the coding unit division method is not re-determined (the image area corresponding to the first image area in the predicted frame of the fifth frame); repeatedly execute the processing method of the third frame for the coding units without prediction information in this image area, which will not be elaborated here.

[0102] (2) Repeatedly execute the processing method of the fifth frame for the coding units in the image area where the coding unit division method is re-determined (the image area corresponding to the second image area in the predicted frame of the fifth frame), which will not be elaborated here.

[0103] 7. Processing of subsequent frames:

[0104] The processing methods of subsequent frames after the sixth frame are all consistent with the processing logic of the sixth frame. Thus, classification processing is performed on coding units or image areas in different situations to ensure that the processing of each frame can make full use of previous information and appropriate methods, thereby improving the coding efficiency, which will not be elaborated here.

[0105] In this way, by continuously looping the above prediction coding process in this embodiment, the inter-frame information can be fully utilized, thereby reducing data redundancy and improving the coding efficiency and video quality.

[0106] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of an electronic device provided by the present application. The electronic device 60 includes a memory 61 and a processor 62 connected to each other. The memory 61 is used to store a computer program, and when the computer program is executed by the processor 62, it is used to implement the prediction coding method in the above embodiment.

[0107] For the method of the above embodiment, it can exist in the form of a computer program. Therefore, the present application proposes a computer-readable storage medium. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 80 is used to store a computer program 81, which can be executed to implement the prediction coding method in the above embodiment.

[0108] The computer-readable storage medium 80 may be various media that can store program codes, such as a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0109] The above are only embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A predictive coding method, characterized in that: The method comprises: Inputting position information of a plurality of coding units in a reference frame into a position prediction filter; Inputting the position information of the coding units in the current frame and the coding units in the reference frame that have a corresponding relationship into the position prediction filter, so that the position prediction filter predicts the position information of each coding unit in the backward frame by using the position information of each coding unit in the current frame and the corresponding position information of the reference frame; wherein the reference frame is the forward frame of the current frame; Obtaining a prediction frame of the backward frame using the position information of the coding units in the current frame and the reference frame that have a corresponding relationship with each other, and the position information of the backward frame, wherein the prediction frame includes prediction blocks of the coding units in the backward frame that correspond to at least part of the coding units in the current frame; matching an original frame of the backward frame with the predicted frame to determine prediction distortions of respective coding units in the backward frame corresponding to at least some of the coding units in the current frame; Encoding a coding unit corresponding to a prediction block using a prediction block in the backward frame whose prediction distortion is less than a preset threshold; If there is a fifth coding unit in the backward frame, the fifth coding unit is a coding unit in the backward frame whose prediction distortion is less than a preset threshold, the current frame is used as a reference frame, the backward frame is used as the current frame, and at least some of the coding units of the current frame include the fifth coding unit, return to the step of inputting the position information of the coding units in the current frame that correspond to the coding units in the reference frame into the position prediction filter.

2. The predictive coding method according to claim 1, characterized in that: The step of inputting the position information of the coding units in the current frame and the reference frame corresponding to each other into the position prediction filter comprises: If the position prediction filter predicts the position information of the first coding unit in the current frame based on the position information of the first coding unit, the current frame is predicted using the reference frame and the position information of the first coding unit in the current frame to obtain a predicted frame of the current frame, wherein the first coding unit is at least part of the coding units in the reference frame, and the predicted frame includes a predicted block of the coding unit corresponding to the first coding unit in the current frame; Matching an original frame of the current frame with the predicted frame to determine prediction distortions of respective coding units in the current frame corresponding to the first coding unit; The step of inputting the position information of the coding units in the current frame and the reference frame that have a corresponding relationship to the position prediction filter comprises: The position information of the coding unit in the current frame whose prediction distortion is less than a preset threshold is input into the position prediction filter, so that the position prediction filter uses the position information of the coding unit in the current frame whose prediction distortion is less than the preset distortion threshold and the position information of the corresponding coding unit in the reference frame to predict the position information of the coding unit in the backward frame whose prediction distortion is less than the preset distortion threshold.

3. The predictive coding method according to claim 2, characterized in that: The step of inputting the position information of the coding units in the current frame and the reference frame that have a corresponding relationship to the position prediction filter further includes: A coding unit in the current frame corresponding to a second coding unit in the reference frame other than the first coding unit is determined, and position information thereof is input into the position prediction filter.

4. The predictive coding method according to claim 3, characterized in that: The determining a coding unit in the current frame corresponding to a second coding unit other than the first coding unit in the reference frame, and inputting position information thereof into the position prediction filter, comprises: The second coding unit includes a third coding unit, the third coding unit is a coding unit in the reference frame whose prediction distortion is greater than a preset threshold; determining a coding unit in the current frame corresponding to the third coding unit, and inputting its position information into the position prediction filter for initialization; and / or, The second encoding unit includes a fourth encoding unit, which is an encoding unit that inputs position information in the reference frame into the position prediction filter for initialization; determines the encoding unit corresponding to the fourth encoding unit in the current frame, and inputs its position information into the position prediction filter, so that the position prediction filter uses the position information of the fourth encoding unit and the position information of the encoding unit corresponding to the fourth encoding unit in the current frame for prediction.

5. The predictive coding method according to claim 4, characterized in that: The determining a coding unit in the current frame corresponding to a second coding unit other than the first coding unit in the reference frame, and inputting position information thereof into the position prediction filter, comprises: Determine a first image region to be updated in the current frame based on the target detection result of the current frame; divide the first image region to be updated into coding units to determine at least one coding unit in the first image region to be updated; For a first area other than the first image area to be updated in the current frame, the coding unit division method of the corresponding area in the reference frame is used to determine at least one coding unit in the first area, wherein the coding unit of the corresponding area in the reference frame includes at least one of the third coding unit and the fourth coding unit, and the at least one coding unit in the first area corresponds one-to-one to the coding unit of the corresponding area in the reference frame; or, After determining the prediction block whose prediction distortion is less than a preset threshold, determine, based on the target detection result of the current frame, a second image area to be updated in a second area other than the prediction block whose prediction distortion is less than the preset threshold; divide the second image area to be updated into coding units to determine at least one coding unit in the second image area to be updated; For a third area other than the second image area to be updated in the second area, the coding unit division method of the corresponding area in the reference frame is used to determine at least one coding unit in the third area, wherein the coding unit of the corresponding area in the reference frame includes at least one of the third coding unit and the fourth coding unit, and at least one coding unit in the third area corresponds one-to-one to the coding unit of the corresponding area in the reference frame.

6. The predictive coding method according to claim 1, characterized in that: The step of inputting position information of coding units in the current frame and the reference frame that have a corresponding relationship to each other into the position prediction filter comprises: determining a first image region to be updated in the current frame based on the target detection result of the current frame; dividing the first image region to be updated into coding units to determine at least one coding unit in the first image region to be updated; for a first region other than the image region to be updated in the current frame, dividing the coding units of the corresponding region in the reference frame to determine at least one coding unit in the first region, wherein the at least one coding unit in the first region corresponds one-to-one to the coding unit of the corresponding region in the reference frame; The position information of at least one coding unit in the first area is input into the position prediction filter, so that the position prediction filter uses the position information of at least part of the coding units in the first area and the position information of the corresponding coding units in the reference frame to predict the position information of at least part of the coding units except the first area in the backward frame.

7. The predictive coding method according to claim 1, characterized in that: If there is a fifth coding unit in the backward frame, the fifth coding unit is a coding unit in the backward frame whose prediction distortion is less than a preset threshold, the current frame is used as a reference frame, the backward frame is used as a current frame, and at least some coding units of the current frame include the fifth coding unit, returning to the step of inputting the position information of the coding units in the current frame and the reference frame that have a corresponding relationship into the position prediction filter, including: If the fifth coding unit exists in the backward frame, the current frame is used as a reference frame, and the backward frame is used as the current frame. For the sixth coding unit in the current frame, the sixth coding unit is a coding unit in the current frame whose prediction distortion is greater than a preset threshold. Based on the target detection result of the current frame, a coding unit division method for the area where the sixth coding unit is located is determined to determine at least one coding unit in the area where the sixth coding unit is located, and encode at least one coding unit in the area where the sixth coding unit is located.

8. The predictive coding method according to claim 1, characterized in that: If there is a fifth coding unit in the backward frame, the fifth coding unit is a coding unit in the backward frame whose prediction distortion is less than a preset threshold, the current frame is used as a reference frame, the backward frame is used as a current frame, and at least some coding units of the current frame include the fifth coding unit, returning to the step of inputting the position information of the coding units in the current frame and the reference frame that have a corresponding relationship into the position prediction filter, including: If the fifth coding unit exists in the backward frame, the current frame is used as a reference frame, and the backward frame is used as a current frame, a coding unit division method of an area other than an area where the fifth coding unit and the sixth coding unit are located in the current frame is determined based on a target detection result of the current frame, so as to determine at least one coding unit of an area other than an area where the fifth coding unit and the sixth coding unit are located in the current frame, and at least one coding unit of an area other than an area where the fifth coding unit and the sixth coding unit are located in the current frame is encoded, wherein the sixth coding unit is a coding unit in the current frame whose prediction distortion is greater than a preset threshold; If at least one coding unit in an area other than the area where the fifth coding unit and the sixth coding unit are located in the current frame includes a coding unit in the seventh coding unit that has a corresponding relationship with the reference frame and whose position information has been input once into the position prediction filter, at least some of the coding units of the current frame include coding units in the seventh coding unit whose position information has been input once into the position prediction filter; and / or, if at least one coding unit in an area other than the area where the fifth coding unit and the sixth coding unit are located in the current frame includes a coding unit in the seventh coding unit that has a corresponding relationship with the reference frame and whose position information has not been input into the position prediction filter, the position information of the coding unit in the seventh coding unit that has not been input into the position prediction filter is input into the position prediction filter for initialization.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the processor is coupled to the memory, and the processor is configured to execute one or more steps of the predictive encoding method according to any one of claims 1 to 8 based on instructions stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the predictive encoding method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image predictive coding method and image encoder

    CN104125463A

  • Sub-pixel target tracking method applied to precision guide star measurement system

    CN111193496A