A Motion Estimation Method and Device Applicable to Camera Movement Scenarios
By using the motion estimation information of the previous video frame to calculate the initial search center of the current video frame and perform coarse and fine-grained search, the problem of difficulty in motion estimation in the camera moving scene is solved, and the video encoding efficiency is significantly improved.
Patent Information
- Application Number
- CN202111323428.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-08
AI Technical Summary
In camera movement scenarios, the prior art is difficult to effectively process the additional motion vector introduced by camera movement, resulting in motion estimation that the optimal matching results of the foreground object and the background object cannot be searched, reducing coding efficiency.
By using the motion estimation information of the previous video frame, the initial search center of the motion estimation of each image block on the current video frame is calculated, and coarse-grained and fine-grained searches are performed in the reference frame to obtain the optimal motion vector. This method does not require external sensors, avoiding the risk of generating outlier motion vectors.
It significantly improves the video encoding efficiency of camera mobile scenes, solves the difficult problems of motion estimation in camera mobile scenes, and improves the encoding efficiency and accuracy.
Smart Images

Figure CN113965754B_ABST
Abstract
Description
Technical Field
[0001] This application relates to a digital video coding technology, and particularly to a motion estimation method applicable to a camera movement scenario and suitable for hardware implementation. Background Art
[0002] Video coding is a technology that compresses redundant components in video images and uses as little data as possible to represent video information. Currently, common video coding standards include HEVC (High Efficiency Video Coding, also known as H.265) and AVC (Advanced Video Coding, also known as H.264), etc.
[0003] Video coding technology uses image blocks as the most basic coding units. In HEVC, the basic coding unit is CU (Coding Unit). A CU can be an image block of 64 pixels × 64 pixels, 32 pixels × 32 pixels, 16 pixels × 16 pixels, or 8 pixels × 8 pixels.
[0004] Motion estimation is to search for the optimal matching block in the encoded video frame (referred to as the reference frame) for the current image block in the current coding frame, so as to minimize the Rate Distortion Cost (RD Cost). Motion estimation is one of the core technologies of video coding algorithms, including integer motion estimation (IME) and fractional motion estimation (FME). Its role is to eliminate the temporal information redundancy of video signals, thereby improving the coding efficiency. "Motion estimation" in this application all refers to integer motion estimation. Please refer to Figure 1 that the relative offset between the optimal matching block in the reference frame and the co-located image block (image block at the same position) of the current image block in the reference frame in the current coding frame is the optimal motion vector (MV) of the current image block. During motion estimation, the current image block on the current frame is matched one by one with all image blocks (including or not including the co-located image block, depending on the position of the search center and the search range) within the reference frame search window (i.e., the search range of motion estimation) to find the optimal matching block. Figure 1 Applicable to different scenarios: If it is motion estimation with coarse pixel accuracy (i.e., coarse-grained search), Figure 1 the optimal motion vector in Figure 1The optimal motion vector in [[]] is the optimal motion vector with integer pixel precision.
[0005] To improve the video encoding speed, it has become a common practice in the industry to use application-specific integrated circuits (ASICs) to accelerate the video encoding process in hardware. Among the various algorithm modules of video encoding, the motion estimation module accounts for more than 70% of the total computational workload of the entire video encoder and is one of the most important modules in the hardware video encoder. In a hardware video encoder, the motion estimation module is generally designed in a combination of coarse-grained search and fine-grained search. Coarse-grained search is a large-scale downsampling search centered on a preset point to obtain the optimal motion vector with coarse pixel precision for the current image block. The preset point is usually set at the (0, 0) coordinate point, which is the upper left corner position of the corresponding co-located image block of the current image block on the reference frame. The search center of the coarse-grained search can also be referred to as the "initial search center of motion estimation". Fine-grained search is a small-scale local full search centered on the optimal motion vector with coarse pixel precision to obtain the optimal motion vector with integer pixel precision for the current image block.
[0006] Encoding scenarios where the camera moves often occur in video encoding. For example, when recording sports videos such as football games or equestrian competitions, the camera's perspective needs to move continuously to track the football or horse riders; some intelligent surveillance cameras are installed on rotatable pan-tilt heads, and the camera's perspective can move to cover blind spots; for cameras on drones, the camera's field of view will move continuously as the drone flies.
[0007] In a typical video scene, there are generally moving foreground objects and stationary background objects. For example, in a road surveillance video scene, objects such as roads, street lights, and buildings that do not move on their own can be determined as stationary background objects, while objects such as pedestrians and vehicles that move on their own can be determined as moving foreground objects. Of course, there may also be mutual conversions between moving foreground objects and stationary background objects. For example, during video recording, a vehicle that has been parked by the roadside without moving can be determined as part of the stationary background objects. In this application, a "foreground object" generally refers to an object that moves on its own during video recording, and a "background object" generally refers to an object that does not move on its own during video recording. Please refer to Figure 2 , in the camera stationary scene, the motion vector generated by the foreground object due to its own motion is denoted as MV F , MV F has a non-zero magnitude. The motion vector generated by the background object due to its own motion is denoted as MV B , MV B has a magnitude of 0.
[0008] During the process of video recording, if the camera moves, an additional motion vector (denoted as MV) caused by the camera movement will be superimposed on each object (including foreground and background objects) in the field of view. C Please refer to Figure 3 , in the case of camera movement (e.g., the camera field of view pans to the left), the originally stationary background objects will also "move" in the field of view, and its final motion vector MV' B = MV C ; for the originally moving foreground objects, its final motion vector MV F ' will be superimposed with the MV generated by the camera movement on the basis of its own motion vector MV F , and there is MV' C = MV F + MV F + MV C . It can be seen that when the camera moves, both foreground and background objects will superimpose a motion vector generated by the camera movement on the basis of their own motion vectors. If the displacement of the camera is relatively large, the final motion vectors of the foreground and background objects will exceed the search range of the motion estimation with a preset point as the search center, resulting in the inability of the motion estimation to search for the optimal matching results of the foreground and background objects, greatly reducing the coding efficiency. In a hardware encoder, the size of the search range of the motion estimation is generally preset. Currently, video coding standards are all based on image blocks. Regardless of the shape of the object in the video, it is divided into individual image blocks for coding operations.
[0009] It should be noted that the above example is only a simple abstraction of the camera movement problem. In the actual scenario, for each frame in the video, the additional MVs introduced due to the camera movement for objects at different positions in the field of view C may all be different from each other. This is because there are various situations of the camera's motion. In addition to the relatively simple translational motion, there are also axial motion, rotational motion, etc. When the camera undergoes axial motion or rotation, the MVs of objects at different positions in the field of view C are all different. Even if the camera's movement is translational motion, due to the different optical properties of the camera lens, the MVs of the objects in the field of view C are not necessarily the same. Wide-angle and fish-eye lenses have larger distortions, resulting in smaller MVs for objects located at the center of the field of view C , while larger MVs for objects located at the edge of the field of view C ; telephoto lenses have smaller distortions, and the MVs of objects at all positions in the field of view C are approximately the same.
[0010] To sum up, during the process of video coding, if the camera moves, how to handle the MVs introduced by the camera movement during the motion estimation processC It becomes a difficult point to perform effective compensation.
[0011] In existing hardware video encoders, there are already some solutions to handle the motion estimation problem when the camera moves.
[0012] One solution is to set a frame-level Global Motion Vector (GMV) from outside the encoder to the hardware encoder to compensate for the MV introduced by camera movement. In this solution, the user uses a program command or a configuration file to set a GMV applicable to all image blocks within the current video frame for each video frame, that is, the frame-level GMV. When encoding each video frame, the hardware encoder uses the GMV of the current video frame set by the user as the initial search center for the motion estimation of all image blocks within this video frame. This solution has many problems: First, all image blocks within each video frame share the same GMV, and it cannot solve the problem that the MVs of objects at different positions within the field of view are different; Second, additional sensors or computing units are required to sense the motion state of the camera in real time, and then calculate the GMV based on the motion state of the camera. The cost is high and it is difficult to achieve accurately and effectively, increasing the difficulty of use for users. C to perform compensation. In this solution, the user uses a program command or a configuration file to set a GMV applicable to all image blocks within the current video frame for each video frame, that is, the frame-level GMV. When encoding each video frame, the hardware encoder uses the GMV of the current video frame set by the user as the initial search center for the motion estimation of all image blocks within this video frame. This solution has many problems: First, all image blocks within each video frame share the same GMV, and it cannot solve the problem that the MVs of objects at different positions within the field of view are different; Second, additional sensors or computing units are required to sense the motion state of the camera in real time, and then calculate the GMV based on the motion state of the camera. The cost is high and it is difficult to achieve accurately and effectively, increasing the difficulty of use for users. C are different; Second, additional sensors or computing units are required to sense the motion state of the camera in real time, and then calculate the GMV based on the motion state of the camera. The cost is high and it is difficult to achieve accurately and effectively, increasing the difficulty of use for users.
[0013] Another solution does not require setting the GMV from outside the encoder. Instead, before performing the motion estimation of the current video frame, the encoder internally uses the motion estimation results of each image block in the previous video frame to statistically obtain the frame-level GMV of the current frame. A variant of this solution is to divide the current encoded video frame into multiple regions, each region contains multiple image blocks, and then statistically obtain the region-level GMV within each divided region. After obtaining the frame-level GMV or the region-level GMV, set the GMV as the initial search center for the motion estimation of the image blocks within the current frame or each divided region. The advantage of this solution is that no additional external components are required to obtain the GMV, but it still cannot completely solve the problem that the MVs of objects at different positions within the video frame or the divided regions are different. Moreover, from the previous description, it can be seen that the MV C is only considered in terms of image matching factors. In this solution, the motion estimation results of each image block in the previous video frame used are generally motion vectors after complete rate-distortion optimization. In addition to considering the cost of image matching, the cost of encoding bits is also considered, which is prone to generating outlier motion vectors and reducing the efficiency of motion estimation. C is only considered in terms of image matching factors. In this solution, the motion estimation results of each image block in the previous video frame used are generally motion vectors after complete rate-distortion optimization. In addition to considering the cost of image matching, the cost of encoding bits is also considered, which is prone to generating outlier motion vectors and reducing the efficiency of motion estimation.
[0014] An outlier motion vector refers to a motion vector with a large difference in amplitude or direction from the motion vectors of surrounding image blocks during the motion estimation process, which is caused by the failure to find the correct image matching block. In the encoding scenario where the camera moves, due to the movement of the camera, new background objects continuously enter the field of view from the boundary region of the field of view. There is no suitable matching block for this newly entered part of the background object in the previous video frame, and the amplitude and direction of the motion vector obtained through motion estimation are relatively random. After multiple rounds of motion estimation, outlier motion vectors are likely to be generated in the boundary region of the field of view. Summary of the Invention
[0015] The technical problem to be solved by this application is to provide a motion estimation method applicable to the camera movement scenario, without the need for external sensors, and not easily generating outlier motion vectors.
[0016] To solve the above technical problem, the motion estimation method applicable to the camera movement scenario proposed by this application includes the following steps. Step S10: Calculate the initial search center for the motion estimation of each image block on the current video frame according to the motion estimation information of the previous video frame. Step S20: Taking the initial search center for the motion estimation of each image block on the current video frame as the center, perform a downsampling search within a first range in the reference frame to obtain the optimal coarse pixel accuracy motion vector for each image block on the current video frame and the motion estimation information of the current video frame. The motion estimation information of the current video frame refers to two lists: a list composed of the optimal coarse pixel accuracy motion vectors of all image blocks within the current video frame, simply referred to as the coarse motion vector list; a list composed of the sum of absolute differences (SAD) costs of all image blocks within the current video frame, simply referred to as the SAD cost list; the motion estimation information of the current video frame is used for calculating the initial search center for the motion estimation of the next video frame. Step S30: Taking the optimal coarse pixel accuracy motion vector of each image block on the current video frame as the search center, perform a local full search within a second range in the reference frame to obtain the optimal integer pixel accuracy motion vector for each image block on the current video frame; the first range > the second range. The above method utilizes the motion estimation information of the previous encoded frame to automatically calculate the initial search center for the motion estimation of the next encoded frame, greatly improving the encoding efficiency in the camera movement scenario.
[0017] Further, the step S10 includes the following steps; wherein, the motion estimation information of the previous video frame refers to the coarse motion vector list and SAD cost list of the previous video frame. Step S11: Divide the video frame into one or more divided areas, and the number of coded image blocks in each divided area must be greater than one. Step S12: Count the global motion vectors of the previous video frame in each divided area. Step S13: Determine whether the global motion vectors of each divided area of the previous video frame are credible. For a certain divided area on the previous video frame, if its global motion vector is determined to be credible, it is called a credible divided area, and the process proceeds to step S14. For a certain divided area on the previous video frame, if its global motion vector is determined to be untrustworthy, it is called an untrustworthy divided area, and the process proceeds to step S15. Step S14: Use a combination of motion vector amplitude determination and SAD cost determination to determine whether the optimal coarse pixel precision motion vector of each image block in each credible divided area on the previous video frame belongs to an outlier motion vector. If the optimal coarse pixel precision motion vector of a certain image block in a certain credible divided area on the previous video frame is determined to belong to an outlier motion vector, the global motion vector of the credible divided area to which the image block belongs is used as the initial search center for motion estimation of the same-position image block in the current video frame. If the optimal coarse pixel precision motion vector of a certain image block in a certain credible divided area on the previous video frame is determined not to belong to an outlier motion vector, proceed to step S15. Step S15: For each untrustworthy divided area on the previous video frame, determine whether the optimal coarse pixel precision motion vector of each image block in the divided area exceeds the search range of motion estimation with a preset point as the search center. For each optimal coarse pixel precision motion vector of each image block in each credible divided area on the previous video frame that is determined not to belong to an outlier motion vector, determine whether the optimal coarse pixel precision motion vector of the image block exceeds the search range of motion estimation with a preset point as the search center. If exceeded, the optimal coarse pixel precision motion vector of the image block on the previous video frame is used as the initial search center for motion estimation of the same-position image block on the current video frame. If not, the same image block of the image block on the previous video frame on the current video frame uses the preset point as the initial search center for motion estimation. The calculation step of the initial search center for motion estimation is one of the core technical innovations of this application, without the need to introduce a large number of additional calculations and external sensors.
[0018] Furthermore, in the step S11, the divided regions are divided in the same way for each video frame in the entire video stream, no matter it is an encoded video frame or a video frame to be encoded.
[0019] Further, in the step S12, a motion vector MV histogram is used for statistics; an MV horizontal component histogram and an MV vertical component histogram are set, and based on the rough motion vector list of the previous video frame, the MV distribution probability of the image blocks in each divided area of the previous video frame is statistically analyzed to obtain the MV horizontal component and vertical component with the highest occurrence probability in each divided area, and the global motion vector GMV of each divided area is calculated accordingly.
[0020] Further, in the step S13, the judgment method is as follows: for each image block in each divided area, calculate the deviation value between the optimal rough pixel precision motion vector corresponding to the image block in the rough motion vector list of the previous video frame and the GMV of the divided area; if the deviation value is within a preset range, it is determined that the image block belongs to the background part of the divided area, and the image block is marked as a background image block; after processing all the image blocks in the divided area, calculate the ratio of the sum of the areas of all the background image blocks to the area of the divided area, if the ratio is less than the preset threshold, it is determined that the GMV of the divided area is not credible, otherwise it is determined that the GMV of the divided area is credible.
[0021] Furthermore, in step S14, the judgment method includes the following three parts. The first part: calculating the amplitude of the GMV of a certain credible divided area on the previous video frame; calculating the amplitude of the optimal coarse pixel precision motion vector of each image block in the divided area; calculating the reasonable deviation threshold between the amplitude of the optimal coarse pixel precision motion vector of the image block in the divided area and the amplitude of the GMV, which is called the first reasonable deviation threshold; calculating the difference between the amplitude of the optimal coarse pixel precision motion vector of each image block in the divided area and the amplitude of the GMV, which is called the first difference; if the first difference of a certain image block in the divided area is ≤ the first reasonable deviation threshold, the optimal coarse pixel precision motion vector of the image block is not an outlier motion vector; otherwise, proceed to the third part. Part 2: Calculate the average SAD cost of all background image blocks in a credible partition area on the previous video frame; Calculate the reasonable deviation threshold between the SAD cost of the image block in the partition area and the average SAD cost of all background image blocks, called the second reasonable deviation threshold; Calculate the difference between the SAD cost of each image block in the partition area and the average SAD cost of all background image blocks, called the second difference; If the second difference of an image block in the partition area ≤ the second reasonable deviation threshold, the optimal coarse pixel precision motion vector of the image block is not an outlier motion vector; otherwise, enter the third part. Part 3: In a credible partition area on the previous video frame, if an image block satisfies at the same time: the first difference > the first reasonable deviation threshold, and the second difference > the second reasonable deviation threshold, then the optimal coarse pixel precision motion vector of the image block is determined to be an outlier motion vector; otherwise, the optimal coarse pixel precision motion vector of the image block is determined to be not an outlier motion vector. The above-mentioned outlier MV determination method adopts a combination of MV amplitude determination and SAD cost determination, which greatly improves the accuracy of outlier MV determination and improves coding efficiency.
[0022] Furthermore, in step S14, if the optimal coarse pixel precision motion vector of an image block in a credible partition area on the previous video frame is determined to be an outlier motion vector, the optimal coarse pixel precision motion vector of the image block is also corrected to the global motion vector of the credible partition area to which it belongs.
[0023] Furthermore, in step S15, the preset point refers to the upper left corner position of the co-located image block corresponding to the current image block on the reference frame.
[0024] Furthermore, in step S20, when performing downsampling search in the first range in the reference frame, the rate-distortion cost is calculated by only considering the SAD cost, and the bit cost of the encoded image block is not introduced.
[0025] The present application also provides a motion estimation device applicable to a camera movement scenario, including a motion estimation initial search center calculation module, a coarse-grained search module, and a fine-grained search module. The motion estimation initial search center calculation module is configured to calculate the initial search center of motion estimation for each image block on the current video frame according to the motion estimation information of the previous video frame. The coarse-grained search module takes the initial search center of motion estimation for each image block on the current video frame as the center, performs downsampling search within a first range, and obtains the optimal coarse pixel accuracy motion vector for each image block on the current video frame; meanwhile, obtains the motion estimation information of the current video frame; the motion estimation information of the current video frame refers to two lists: a list composed of the optimal coarse pixel accuracy motion vectors of all image blocks within the current video frame, simply referred to as the coarse motion vector list; a list composed of the sum of absolute differences (SAD) costs of all image blocks within the current video frame, simply referred to as the SAD cost list; the motion estimation information of the current video frame is used for calculating the initial search center of motion estimation for the next video frame. The fine-grained search module takes the optimal coarse pixel accuracy motion vector of each image block on the current video frame as the search center, performs local full search within a second range, and obtains the optimal integer pixel accuracy motion vector for each image block on the current video frame. The first range > the second range. The above device utilizes the motion estimation information of the previous encoded frame to automatically calculate the initial search center of motion estimation for the next encoded frame, greatly improving the encoding efficiency in the camera movement scenario.
[0026] The technical effects achieved by the present application are as follows: A motion estimation method applicable to a camera movement scenario and suitable for hardware implementation is proposed. This method combines a coarse-grained search module and a fine-grained search module, which is suitable for hardware circuit implementation. At the same time, the present application innovatively designs the calculation steps for the initial search center of motion estimation. Without introducing a large amount of additional operations and external sensors, it utilizes the motion estimation information of the previous encoded frame to automatically calculate the initial search center of motion estimation for the next encoded frame, greatly improving the encoding efficiency in the camera movement scenario. In addition, the present application also proposes a new method for determining outlier motion vectors (MVs), which combines MV amplitude determination and SAD cost determination, greatly improving the accuracy of outlier MV determination and enhancing the encoding efficiency. Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the optimal motion vector of an image block.
[0028] Figure 2 It is a schematic diagram of a camera static scenario.
[0029] Figure 3 It is a schematic diagram of a camera movement scenario.
[0030] Figure 4It is a schematic flowchart of the motion estimation method applicable to the camera movement scenario proposed in this application.
[0031] Figure 5 It is a schematic sub-step flowchart of step S10.
[0032] Figure 6 It is a schematic diagram of dividing a video frame into one or more divided regions.
[0033] Figure 7 It is a schematic block diagram of the motion estimation device applicable to the camera movement scenario proposed in this application.
[0034] Explanation of reference numerals in the figure: 10 is the initial search center calculation module for motion estimation, 20 is the coarse-grained search module, and 30 is the fine-grained search module. Detailed implementation manner
[0035] Please refer to Figure 4 , the motion estimation method applicable to the camera movement scenario and suitable for hardware implementation proposed in this application includes the following steps.
[0036] Step S10: Calculate the initial search center for the motion estimation of each image block on the current video frame according to the motion estimation information of the previous video frame.
[0037] Step S20: Taking the initial search center for the motion estimation of each image block on the current video frame as the center, perform a downsampling search in the first range (i.e., coarse-grained search) in the reference frame to obtain the optimal coarse pixel accuracy motion vector for each image block on the current video frame and the motion estimation information of the current video frame.
[0038] The motion estimation information of the current video frame refers to two lists: a list composed of the optimal coarse pixel accuracy motion vectors of all image blocks within the current video frame, simply referred to as the coarse motion vector (Coarse MV) list; a list composed of the minimum SAD (Sum of Absolute Difference) costs of all image blocks within the current video frame, simply referred to as the SAD cost list. The motion estimation information of the current video frame is used for calculating the initial search center for the motion estimation of the next video frame.
[0039] Step S30: Taking the optimal coarse pixel accuracy motion vector of each image block on the current video frame as the search center, perform a local full search in the second range (i.e., fine-grained search) in the reference frame to obtain the optimal integer pixel accuracy motion vector for each image block on the current video frame; the first range > the second range.
[0040] Please refer to Figure 5, the step S10 further includes the following steps. Among them, the motion estimation information of the previous video frame refers to the coarse motion vector list and the SAD cost list obtained after the previous video frame passes through the step S20.
[0041] Step S11: Divide the video frame (also called the image frame) into one or more divided regions, and the number of coded image blocks in each divided region must be greater than one. Here, the divided regions are actually divided in the same way for each video frame in the entire video stream, whether it is a coded video frame or a video frame to be coded.
[0042] There are many ways to divide the video frame into one or more divided regions. The specific division method can be selected according to the characteristics of the video scene. For example, it can be selected according to the movement mode of the camera. For example, for the case where the camera makes a translational movement, the MVs C at different positions in the field of view are approximately the same, and the entire video frame can be used as one divided region for processing. For the case where the camera makes an axial movement, the MVs C at different positions in the field of view vary greatly, and the entire video frame needs to be divided into multiple divided regions for separate processing. Figure 6 Examples of dividing the video frame into one or more divided regions are given. Generally speaking, for the case where the image block motions at different positions in the video frame are inconsistent (the motion speeds or motion directions are inconsistent), such as when the camera is making an axial movement, rotation, or the lens distortion is large, it is inclined to divide the video frame into more divided regions; for the case where the image block motions at different positions in the video frame are relatively consistent, such as when the camera is making a translational movement and the lens distortion is small, the image frame can be divided into fewer divided regions.
[0043] Step S12: Statistically calculate the global motion vector (regional-level GMV) of the previous video frame in each divided region.
[0044] There are various ways to statistically calculate the GMV of each divided region of the previous video frame. A relatively common method is to use the MV histogram for statistics. For example, set the MV horizontal component histogram and the MV vertical component histogram, and statistically calculate the MV distribution probability of the image blocks in each divided region of the previous video frame according to the coarse motion vector list of the previous video frame, obtain the MV horizontal component and vertical component with the highest occurrence probability in each divided region, and calculate the GMV of each divided region accordingly.
[0045] Step S13: Determine whether the global motion vectors of each divided region in the previous video frame are credible. The method of determination is as follows: For each image block in each divided region, calculate the deviation value between the optimal coarse pixel precision motion vector corresponding to the image block in the coarse motion vector list of the previous video frame and the GMV in the divided region. If the deviation value is within a preset range, it is determined that the image block belongs to the background part of the divided region, and the image block is marked as a background image block. After processing all the image blocks in the divided region, calculate the ratio of the sum of the areas of all the background image blocks to the area of the divided region. If the ratio is less than a preset threshold (such as 0.3), it is determined that the GMV of the divided region is not credible; otherwise, it is determined that the GMV of the divided region is credible.
[0046] For a certain divided region on the previous video frame, if its global motion vector is determined to be credible, it is called a credible divided region, indicating that there is an obvious background region inside the divided region. In this case, go to step S14.
[0047] For a certain divided region on the previous video frame, if its global motion vector is determined to be not credible, it is called a non-credible divided region, indicating that there is no obvious background region inside the divided region. In this case, go to step S15.
[0048] Step S14: Adopt a combination of motion vector amplitude determination and SAD cost determination to determine whether the optimal coarse pixel precision motion vectors of each image block in each credible divided region on the previous video frame are outlier motion vectors. The method of determination includes the following three parts.
[0049] The first part: (1) Calculate the amplitude of the GMV of a certain credible divided region on the previous video frame. (2) Calculate the amplitude of the optimal coarse pixel precision motion vector of each image block in the divided region. The optimal coarse pixel precision motion vector of each image block can be obtained by looking up the coarse motion vector list of the previous video frame. (3) Calculate the reasonable deviation threshold between the amplitude of the optimal coarse pixel precision motion vector of the image blocks in the divided region and the amplitude of the GMV, which is called the first reasonable deviation threshold. (4) Calculate the difference between the amplitude of the optimal coarse pixel precision motion vector of each image block in the divided region and the amplitude of the GMV, which is called the first difference. If the first difference of a certain image block in the divided region ≤ the first reasonable deviation threshold, then the optimal coarse pixel precision motion vector of the image block is not an outlier motion vector; otherwise, go to the third part.
[0050] Part Two: (1) Calculate the average SAD cost of all background image blocks within a certain reliable division area on the previous video frame. In step S13, it has been marked which image blocks are background image blocks, and by looking up the SAD cost list of the previous video frame, the average SAD cost of all background image blocks within a division area can be calculated. (2) Calculate the reasonable deviation threshold between the SAD cost of the image blocks within this division area and the average SAD cost of all background image blocks, which is called the second reasonable deviation threshold. The SAD cost of each image block within this division area can be obtained by looking up the SAD cost list of the previous video frame. (3) Calculate the difference between the SAD cost of each image block within this division area and the average SAD cost of all background image blocks, which is called the second difference. If the second difference of a certain image block within this division area ≤ the second reasonable deviation threshold, then the optimal coarse pixel precision motion vector of this image block is not an outlier motion vector; otherwise, enter Part Three.
[0051] Part Three: Within a certain reliable division area on the previous video frame, if a certain image block simultaneously satisfies: the first difference > the first reasonable deviation threshold and the second difference > the second reasonable deviation threshold, then it is determined that the optimal coarse pixel precision motion vector of this image block is an outlier motion vector. Otherwise, it is determined that the optimal coarse pixel precision motion vector of this image block is not an outlier motion vector.
[0052] If the optimal coarse pixel precision motion vector of a certain image block within a certain reliable division area on the previous video frame is determined to be an outlier motion vector, then correct the optimal coarse pixel precision motion vector of this image block to the global motion vector of its corresponding reliable division area, and use it as the initial search center for the motion estimation of the in-position (same position) image block within the current video frame. The rate-distortion cost corresponding to the outlier MV is very large and needs to be corrected before it can be used as the initial search center for the motion estimation of the corresponding image block, otherwise the coding efficiency will be greatly reduced. Therefore, the determination and correction of outlier MVs are very important for improving the coding efficiency of the camera motion scene.
[0053] If the optimal coarse pixel precision motion vector of a certain image block within a certain reliable division area on the previous video frame is determined not to be an outlier motion vector, enter step S15.
[0054] Step S15: For each unreliable division area on the previous video frame, determine whether the optimal coarse pixel precision motion vector of each image block within this division area exceeds the search range of the motion estimation with a preset point as the search center.
[0055] For the optimal coarse pixel precision motion vector of each image block determined not to belong to the outlier motion vector within each reliable partition region on the previous video frame, determine whether the optimal coarse pixel precision motion vector of the image block exceeds the search range of motion estimation with a preset point as the search center.
[0056] If it exceeds, the optimal coarse pixel precision motion vector of the image block on the previous video frame is used as the initial search center for motion estimation of the corresponding image block on the current video frame.
[0057] If it does not exceed, the corresponding image block of the image block on the previous video frame on the current video frame uses the preset point as the initial search center for motion estimation. The preset point is usually set to the coordinate point (0, 0), that is, the upper left corner position of the corresponding image block of the current image block on the reference frame.
[0058] Preferably, in step S20, during the coarse-grained search process, the rate-distortion cost calculation only considers the SAD cost and does not introduce the bit cost of the encoded image block. This is not likely to generate outlier motion vectors and is beneficial to improving the coding efficiency.
[0059] In the motion estimation method proposed in this application, the calculation of the initial search center of motion estimation is most different from other motion estimation methods. For the video coding problem in the camera movement scenario, when other motion estimation methods calculate the initial search center of motion estimation, generally, after calculating the GMV of the entire frame or each partition region within the frame through external input or internal statistics of the encoder, the position pointed to by the GMV is used as the initial search center for motion estimation of all image blocks within the current coding frame or all image blocks within each partition region. In the solution proposed in this application, the initial search centers of motion estimation for each image block within the current coding frame are divided into three cases. Although this application also calculates the GMV of the entire frame or each partition region within the frame, the GMV is only used for the determination of outlier MVs and the correction of the initial search center of motion estimation for the image blocks to which the outlier MVs belong. For the image blocks to which the non-outlier MVs belong, the initial search center of their motion estimation comes from the optimal coarse pixel precision motion vector of the corresponding image block on the previous video frame or the preset point, and the GMV will not have any impact on the initial search center of their motion estimation.
[0060] Please refer to Figure 7 The motion estimation device proposed in this application, which is applicable to the camera movement scenario and suitable for hardware implementation, includes an initial search center calculation module 10 for motion estimation, a coarse-grained search module 20, and a fine-grained search module 30.
[0061] The initial search center calculation module 10 for motion estimation is used to calculate the initial search center for motion estimation of each image block on the current video frame according to the motion estimation information of the previous video frame.
[0062] The coarse-grained search module 20 performs downsampling search within a first range centered on the initial search center for motion estimation of each image block on the current video frame, to obtain the optimal coarse pixel accuracy motion vector of each image block on the current video frame; meanwhile, the motion estimation information of the current video frame is obtained. The motion estimation information of the current video frame refers to two lists: a list composed of the optimal coarse pixel accuracy motion vectors of all image blocks within the current video frame, simply referred to as the coarse motion vector list; a list composed of the sum of absolute differences (SAD) costs of all image blocks within the current video frame, simply referred to as the SAD cost list; the motion estimation information of the current video frame is used for the calculation of the initial search center for motion estimation of the next video frame.
[0063] The fine-grained search module 30 performs local full search within a second range centered on the optimal coarse pixel accuracy motion vector of each image block on the current video frame, to obtain the optimal integer pixel accuracy motion vector of each image block on the current video frame. The first range > the second range.
[0064] Compared with the prior art, the present application has the following beneficial technical effects.
[0065] First, when calculating the initial search center for motion estimation, the present application performs operations such as statistics of the global motion vectors (GMVs) of each partition region of the previous video frame, determination of the GMV credibility, determination and correction of outlier motion vectors (MVs), etc., based on the motion estimation information obtained after coarse-grained search of the previous video frame, and finally obtains the initial search center for motion estimation of each image block on the current video frame. This does not introduce a large amount of additional operations, nor does it require the introduction of additional sensors, saving computing resources and hardware costs. At the same time, since the initial search centers for motion estimation of each image block do not completely depend on the GMVs of each partition region of the previous video frame, the initial search centers for motion estimation of each image block within the same partition region can be different from each other, solving the problem that the MVs C of different positions of objects within each partition region of the video frame are different, and improving the coding efficiency.
[0066] Second, the present application adopts a combination of MV amplitude determination and SAD cost determination when determining outlier MVs, greatly improving the accuracy of outlier MV determination and the coding efficiency.
[0067] Third, when performing coarse-grained search, the present application only considers the SAD cost in the rate-distortion cost calculation, which is not likely to generate outlier MVs and is beneficial to improving the coding efficiency.
[0068] Compared with existing motion estimation methods, the motion estimation method proposed in this application can significantly improve the video coding efficiency in the camera movement scenario. To verify its beneficial effects, this application selects three typical camera movement video scenarios, namely Jokey (equestrian competition scenario 1), ReadySetGo (equestrian competition scenario 2), and TouchDownPass (football game scenario), where the camera perspectives all move significantly, for encoding, and conducts a comparative experiment on the coding performance of the motion estimation method in this application and existing motion estimation methods. In the experiment, we perform temporal downsampling on the three YUV sources of Jokey, ReadySetGo, and TouchDownPass to obtain bitstreams of the same YUV source with different frame rates. The lower the frame rate, the more significant the simulated camera movement. The existing motion estimation methods used for the comparative experiment include the scheme of setting frame-level GMV externally to the encoder (i.e., "Existing Scheme 1") and the scheme of uniformly dividing regions and calculating GMV internally to the encoder (i.e., "Existing Scheme 2"). The basis for measuring the experimental results is the decrease in BD-rate ( delta bit rate, delta bit rate. The smaller the BD-rate, the fewer bits are used for encoding when achieving the same video quality, and the higher the coding efficiency) compared to "Existing Scheme 1". The greater the decrease, the higher the coding efficiency. The experimental results are as follows
[0069] shown in Table 1.
[0070]
[0071] Table 1: Comparison table of coding efficiency between the motion estimation method of this application and two existing motion estimation methods
[0072] It can be seen that for video coding in the camera movement scenario, when using the motion estimation method proposed in this application, compared with using the motion estimation methods in "Existing Scheme 1" and "Existing Scheme 2", the coding BD-rate is reduced by 8.64% and 3.00% on average, and the coding efficiency is significantly improved.
[0073] The above is only the preferred embodiment of this application and is not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A motion estimation method applicable to a camera movement scenario, characterized in that, The method comprises the following steps: Step S10: calculating the initial search center of motion estimation for each image block on the current video frame according to the motion estimation information of the previous video frame; the motion estimation information of the previous video frame refers to the coarse motion vector list and SAD cost list of the previous video frame; Step S20: taking the initial search center of the motion estimation of each image block on the current video frame as the center, performing a downsampling search in a first range in the reference frame to obtain the optimal coarse pixel precision motion vector of each image block on the current video frame and the motion estimation information of the current video frame; The motion estimation information of the current video frame refers to two lists: a list consisting of the best coarse pixel precision motion vectors of all image blocks in the current video frame, referred to as the coarse motion vector list; a list consisting of the minimum absolute error and SAD cost of all image blocks in the current video frame, referred to as the SAD cost list; the motion estimation information of the current video frame is used to calculate the initial search center of the motion estimation of the next video frame; Step S30: taking the optimal coarse pixel precision motion vector of each image block on the current video frame as the search center, performing a local full search of the second range in the reference frame to obtain the optimal integer pixel precision motion vector of each image block on the current video frame; the first range>the second range; The step S10 further includes the following steps: Step S11: Divide the video frame into one or more divided areas, and the number of coded image blocks in each divided area must be greater than one; Step S12: Counting the global motion vectors of the previous video frame in each divided area; Step S13: determining whether the global motion vectors of each divided area of the previous video frame are credible; For a certain segmented area on the previous video frame, if its global motion vector is determined to be credible, it is called a credible segmented area, and the process goes to step S14; For a certain segmented area on the previous video frame, if its global motion vector is determined to be unreliable, it is called an unreliable segmented area, and the process goes to step S15; Step S14: using a combination of motion vector amplitude determination and SAD cost determination to determine whether the optimal coarse pixel precision motion vector of each image block in each credible divided area on the previous video frame belongs to an outlier motion vector; If the optimal coarse pixel precision motion vector of a certain image block in a certain credible partition area on the previous video frame is determined to belong to an outlier motion vector, the global motion vector of the credible partition area to which the image block belongs is used as the initial search center for motion estimation of the co-located image block in the current video frame; If the optimal coarse pixel precision motion vector of a certain image block in a certain credible divided area on the previous video frame is determined to be not an outlier motion vector, proceed to step S15; Step S15: for each untrustworthy divided area on the previous video frame, determining whether the optimal coarse pixel precision motion vector of each image block in the divided area exceeds the search range of motion estimation with the preset point as the search center; For each optimal coarse pixel accuracy motion vector of an image block determined not to belong to an outlier motion vector within each reliable partition region on the previous video frame, determine whether the optimal coarse pixel accuracy motion vector of the image block exceeds the search range of motion estimation with a preset point as the search center; If it exceeds, then the optimal coarse pixel accuracy motion vector of the image block on the previous video frame is used as the initial search center for motion estimation of the corresponding image block on the current video frame; If it does not exceed, then the preset point is used as the initial search center for motion estimation of the corresponding image block on the current video frame of the image block on the previous video frame.
2. The motion estimation method applicable to the camera movement scenario according to claim 1, characterized in that In the step S11, the partition region is partitioned in the same way for each video frame in the entire video stream, whether it is an encoded video frame or a video frame to be encoded.
3. The motion estimation method applicable to the camera movement scenario according to claim 1, wherein In the step S12, a motion vector MV histogram is used for statistics; an MV horizontal component histogram and an MV vertical component histogram are set up, and based on the coarse motion vector list of the previous video frame, the MV distribution probabilities of the image blocks within each partition region of the previous video frame are statistically analyzed to obtain the MV horizontal component and vertical component with the highest occurrence probability within each partition region, and the global motion vector GMV of each partition region is calculated accordingly.
4. The motion estimation method applicable to the camera movement scenario according to claim 1, wherein In the step S13, the judgment method is as follows: for each image block within each partition region, calculate the deviation value between the optimal coarse pixel accuracy motion vector corresponding to the image block in the coarse motion vector list of the previous video frame and the GMV within the partition region; if the deviation value is within a preset range, then it is determined that the image block belongs to the background part of the partition region, and the image block is marked as a background image block; after processing all the image blocks within the partition region, calculate the ratio of the sum of the areas of all the background image blocks to the area of the partition region, and if this ratio is less than a preset threshold, then it is determined that the GMV of the partition region is not reliable, otherwise it is determined that the GMV of the partition region is reliable.
5. The motion estimation method applicable to the camera movement scenario according to claim 1, characterized in that, In the step S14, the judgment method includes the following three parts; The first part: calculate the magnitude of the GMV of a certain reliable partition region on the previous video frame; calculate the magnitude of the optimal coarse pixel accuracy motion vector of each image block within the partition region; calculate the reasonable deviation threshold between the magnitude of the optimal coarse pixel accuracy motion vector of the image block within the partition region and the magnitude of the GMV, which is called the first reasonable deviation threshold; calculate the difference between the magnitude of the optimal coarse pixel accuracy motion vector of each image block within the partition region and the magnitude of the GMV, which is called the first difference; if the first difference of a certain image block within the partition region ≤ the first reasonable deviation threshold, then the optimal coarse pixel accuracy motion vector of the image block is not an outlier motion vector; otherwise, enter the third part; Part 2: Calculate the average SAD cost of all background image blocks in a credible partition area on the previous video frame; Calculate the reasonable deviation threshold between the SAD cost of the image blocks in the partition area and the average SAD cost of all background image blocks, called the second reasonable deviation threshold; Calculate the difference between the SAD cost of each image block in the partition area and the average SAD cost of all background image blocks, called the second difference; If the second difference of a certain image block in the partition area is ≤ the second reasonable deviation threshold, the optimal coarse pixel precision motion vector of the image block is not an outlier motion vector; Otherwise, proceed to Part 3; Part 3: In a certain credible divided area on the previous video frame, if a certain image block simultaneously satisfies: the first difference > the first reasonable deviation threshold, and the second difference > the second reasonable deviation threshold, then the optimal coarse pixel precision motion vector of the image block is determined to be an outlier motion vector; otherwise, the optimal coarse pixel precision motion vector of the image block is determined not to be an outlier motion vector.
6. The motion estimation method applicable to the camera movement scenario according to claim 1, characterized in that In step S14, if the optimal coarse pixel precision motion vector of an image block in a credible partition area on the previous video frame is determined to be an outlier motion vector, the optimal coarse pixel precision motion vector of the image block is corrected to the global motion vector of the credible partition area to which it belongs.
7. The motion estimation method applicable to the camera movement scenario according to claim 1, characterized in that, In the step S15, the preset point refers to the upper left corner position of the co-located image block corresponding to the current image block on the reference frame.
8. The motion estimation method applicable to the camera movement scenario according to claim 1, characterized in that, In the step S20, when performing downsampling search in the first range in the reference frame, the rate-distortion cost is calculated by only considering the SAD cost, and the bit cost of the coded image block is not introduced.
9. A motion estimation device applicable to a camera movement scenario, characterized in that, It includes a motion estimation initial search center calculation module, a coarse-grained search module and a fine-grained search module; The motion estimation initial search center calculation module is used to calculate the initial search center of motion estimation of each image block on the current video frame according to the motion estimation information of the previous video frame; the motion estimation information of the previous video frame refers to the coarse motion vector list and SAD cost list of the previous video frame; The coarse-grained search module performs a down-sampling search in a first range with the initial search center of the motion estimation of each image block on the current video frame as the center, and obtains the optimal coarse pixel precision motion vector of each image block on the current video frame; and obtains the motion estimation information of the current video frame at the same time; the motion estimation information of the current video frame refers to two lists: a list consisting of the optimal coarse pixel precision motion vectors of all image blocks in the current video frame, referred to as the coarse motion vector list; and a list consisting of the minimum absolute error and SAD cost of all image blocks in the current video frame, referred to as the SAD cost list; the motion estimation information of the current video frame is used for calculating the initial search center of the motion estimation of the next video frame; The fine-grained search module uses the optimal coarse pixel precision motion vector of each image block on the current video frame as the search center, performs a local full search in the second range, and obtains the optimal integer pixel precision motion vector of each image block on the current video frame; the first range>the second range; The initial search center calculation module for motion estimation specifically performs the following: dividing a video frame into one or more divided regions, where the number of coded image blocks within each divided region must be greater than one; counting the global motion vectors of the previous video frame in each divided region; and determining whether the global motion vectors of each divided region of the previous video frame are reliable. Processing method 1: For a certain divided region on the previous video frame, if its global motion vector is determined to be reliable, it is called a reliable divided region; a combined method of motion vector magnitude determination and SAD cost determination is used to determine whether the optimal coarse pixel accuracy motion vectors of each image block within each reliable divided region on the previous video frame belong to outlier motion vectors; if the optimal coarse pixel accuracy motion vector of a certain image block within a certain reliable divided region on the previous video frame is determined to belong to an outlier motion vector, then the global motion vector of the reliable divided region to which the image block belongs is used as the initial search center for motion estimation of the corresponding image block within the current video frame; if the optimal coarse pixel accuracy motion vector of a certain image block within a certain reliable divided region on the previous video frame is determined not to belong to an outlier motion vector, proceed to processing method 2. Processing method 2: For a certain divided region on the previous video frame, if its global motion vector is determined to be unreliable, it is called an unreliable divided region; for each unreliable divided region on the previous video frame, determine whether the optimal coarse pixel accuracy motion vector of each image block within the divided region exceeds the search range of motion estimation with a preset point as the search center; for the optimal coarse pixel accuracy motion vector of each image block within each reliable divided region on the previous video frame that is determined not to belong to an outlier motion vector, determine whether the optimal coarse pixel accuracy motion vector of the image block exceeds the search range of motion estimation with a preset point as the search center; if it exceeds, then the optimal coarse pixel accuracy motion vector of the image block on the previous video frame is used as the initial search center for motion estimation of the corresponding image block on the current video frame; if it does not exceed, then the corresponding image block of the image block on the previous video frame on the current video frame uses the preset point as the initial search center for motion estimation.
Citation Information
Patent Citations
Compression method of light field image
CN107135393A
Video motion estimation method and device, equipment and computer readable storage medium
CN112203095A