Video Motion Estimation Method, Apparatus, Device, and Computer Program

By classifying scenes, extracting features, and determining a search range based on these features, the method enhances the efficiency and accuracy of video motion estimation, addressing the computational and accuracy challenges in existing technologies.

JP7684391B2Active Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023518852
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-04
Filing Date
2021-12-03
Publication Date
2025-05-27
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Current video motion estimation methods are computationally intensive, taking up to 70% of the total computational resources in video encoding, and struggle with accuracy and efficiency, especially in high-resolution videos and scenarios with camera shake or low contrast.

Method used

The proposed method improves video motion estimation by performing scene classification, extracting contour and color features of foreground objects, and determining a search range based on these features, thereby narrowing the search area and enhancing the accuracy of motion estimation.

Benefits of technology

This approach reduces the search time and improves the accuracy of motion estimation by limiting the search range based on contour features, leading to more efficient video compression and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684391000001
    Figure 0007684391000001
  • Figure 0007684391000002
    Figure 0007684391000002
  • Figure 0007684391000003
    Figure 0007684391000003
Patent Text Reader

Abstract

A video motion estimation method, apparatus, device, computer-readable storage medium, and computer program product, the method including the steps of acquiring a plurality of image frames in a video to be processed, performing a scene segmentation process on the plurality of image frames to obtain a plurality of image frame sets, each image frame set including at least one image frame; extracting contour features and color features of a foreground object in each image frame in each image frame set; determining a search range corresponding to each image frame set based on the contour features of the foreground object in each image frame set; determining a starting search point for each predicted frame in each image frame set; and performing a motion estimation process within a search area corresponding to the search range in each predicted frame based on the starting search point of each predicted frame, a target block in a reference frame, and the color feature of the foreground object, to obtain a motion vector corresponding to the target block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video motion estimation, and relates to, but is not limited to, a video motion estimation method, apparatus, device, computer-readable storage medium, and computer program product.

[0002] The embodiments of this application are proposed based on a Chinese patent application with an application number of No. 202011401743.1 and an application date of December 4, 2020, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is incorporated herein by reference into the embodiments of this application.

Background Art

[0003] With the rapid development of Internet technology and the wide spread of digital devices, video has gradually become an important medium for people to obtain and exchange information. While users' needs for video services are increasing day by day, the requirements for video quality are also increasing day by day. Therefore, improving video encoding and transmission efficiency has become a hot issue attracting attention in the industry.

[0004] High-quality video data has a high redundancy and a large amount of information. Therefore, in order to meet the requirements of transmission and storage in related network fields, it is necessary to compress video data. Video compression can adopt inter-frame prediction to eliminate the temporal redundancy in the sequence of frames. Motion estimation is a key technology widely applied in inter-frame prediction in video encoding. However, it takes a very long time and occupies 70% of the total computational amount of the entire video encoding. Therefore, in the case of higher-resolution videos, this ratio may be even higher. Therefore, the motion estimation algorithm is the main factor for determining video compression efficiency. However, reducing the computational cost of motion estimation, improving the accuracy of motion estimation, and making the search process of motion estimation more robust, faster, and more efficient are the key goals for accelerating the video compression process.

Summary of the Invention

Means for Solving the Problem

[0005] Embodiments of the present application provide a video motion estimation method, apparatus, device, computer-readable storage medium, and computer program product, which can improve the search efficiency and the accuracy of motion estimation.

[0006] The technical solution of the embodiments of the present application is realized as follows.

[0007] Embodiments of the present application provide a video motion estimation method, which is executed by a video motion estimation device, acquiring a plurality of image frames in a video to be processed, performing scene classification processing on the plurality of image frames to obtain a plurality of image frame sets, each of the image frame sets including at least one image frame; extracting contour features and color features of foreground objects in each image frame in each of the image frame sets; determining a search range corresponding to each of the image frame sets based on the contour features of the foreground objects in each of the image frame sets; determining a start search point of each prediction frame in each of the image frame sets; performing motion estimation processing within a search area corresponding to the search range in each of the prediction frames based on the start search point of each of the prediction frames, a target block in a reference frame, and the color features of the foreground objects, to obtain a motion vector corresponding to the target block.

[0008] Embodiments of the present application provide a video motion estimation apparatus, a first acquisition module configured to acquire a plurality of image frames in a video to be processed and perform scene classification processing on the plurality of image frames to obtain a plurality of image frame sets, each of the image frame sets including at least one image frame; A feature extraction module configured to extract contour features and color features of foreground objects in each image frame in each of the image frame sets; A first determination module configured to determine a search range corresponding to each of the image frame sets based on the contour features of the foreground objects in each of the image frame sets; A second determination module configured to determine a start search point for each prediction frame in each of the image frame sets; A motion estimation module configured to perform motion estimation processing within a search area corresponding to the search range in each of the prediction frames based on the start search point of each of the prediction frames, a target block in a reference frame, and the color features of the foreground object, and obtain a motion vector corresponding to the target block.

[0009] Embodiments of the present application provide a device, including a memory used to store executable instructions, and a processor used to implement the method when executing the executable instructions stored in the memory.

[0010] Embodiments of the present application provide a computer-readable storage medium on which executable instructions are stored and which are used to implement the method when executed by a processor.

[0011] Embodiments of the present application provide a computer program product that includes a computer program or instructions, and causes a computer to execute the method by the computer program or instructions.

Advantages of the Invention

[0012] Embodiments of the present application have the following beneficial effects. Based on the contour features of foreground objects in each set of image frames, determine the search range corresponding to each set of image frames, and perform motion estimation within the search area corresponding to the search range in each predicted frame based on each starting search point, the target block in the reference frame, and the color features of the foreground objects, thereby searching within a predetermined range. Therefore, the search range can be narrowed, thereby reducing the search time, and since the search range is limited based on the contour features of foreground objects in each scene, the accuracy of motion estimation can be further improved.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0014] To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be described in more detail below in combination with the drawings. The described embodiments should not be regarded as limitations on the present application, and all other embodiments obtained on the premise that those skilled in the art have not performed creative labor all belong to the protection scope of the present application.

[0015] In the following description, regarding "some embodiments", although it describes a subset of all possible embodiments, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and can be combined with each other under non-contradictory situations. Unless otherwise defined, all technologies and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the technical field of the embodiments of the present application. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended for the purpose of limiting the present application.

[0016] Before providing a more detailed description of the embodiments of the present application, nouns and terms related to the embodiments of the present application will be explained. The nouns and terms related to the embodiments of the present application are applicable to the following interpretations.

[0017] 1) Motion estimation: A kind of technology used in video coding, which is a process of calculating the motion vector between the current frame and the reference frame during the compression coding process.

[0018] 2) Motion vector: A vector representing the relative displacement between the current coded block and the optimal matching block in the reference image.

[0019] 3) Optical flow: When a person's eyes observe a moving object, a series of continuously changing images of the object are formed on the retina of the person's eyes. This series of continuously changing information "flows" continuously on the retina (i.e., the image plane), similar to a kind of light "flow". Therefore, it is called optical flow.

[0020] 4) Optical flow method: A method that utilizes the change of pixels in the time domain in an image sequence and the correlation between adjacent frames to determine the correspondence relationship existing between the previous frame and the current frame, thereby calculating the motion information of the object between adjacent frames.

[0021] To better understand the video motion estimation method provided in the embodiments of the present application, first, an explanation will be given regarding video motion estimation and the video motion estimation methods in related technologies.

[0022] In the inter-frame predictive coding method for video compression, the video content in consecutive frames has a predetermined correlation in time. Therefore, the motion estimation method in related technologies first divides each image frame of the video sequence into a plurality of blocks or macroblocks that are of the same size and do not overlap with each other, and assumes that the displacement amounts of all pixels within the macroblock are equal. Next, according to a predetermined matching means, the target matching block that is most similar to each block or macroblock in the neighboring reference frame is searched. Finally, the relative offset amount of the spatial position between the macroblock and the target matching block, that is, the motion vector, is calculated, and the process of obtaining the motion vector is the motion estimation.

[0023] The core idea of motion estimation is to obtain the motion vectors between frames of a video sequence as accurately as possible, which is mainly used for inter-frame motion compensation. The compensation residue needs to go through operations such as transformation, quantization, and encoding, and then entropy encoding is performed together with the motion vectors and transmitted to the decoding side in the form of a bit stream. At the decoding side, the current block or the current macroblock can be restored through these two pieces of data (i.e., the compensation residue and the motion vector). By applying the motion estimation method in video transmission, the inter-frame data redundancy can be effectively removed, thereby reducing the amount of data to be transmitted. The accuracy of the motion vector determines the quality of the predicted and compensated video frame. The higher the quality, the smaller the compensation residue, the fewer the number of bits required for compensation encoding, and the lower the requirement for the transmission bit rate.

[0024] The motion estimation methods of related technologies include spatial domain motion estimation and frequency domain motion estimation. Here, spatial domain motion estimation includes motion estimation based on global, pixel point, macroblock, region, and grid, etc., and frequency domain motion estimation includes methods such as the phase method, discrete cosine transform method, and wavelet domain method. The spatial domain motion estimation method has characteristics such as relatively fast calculation speed, low complexity, and being easy to be realized on many hardware platforms, so it has become a method favored by many researchers in recent years. The spatial domain motion estimation method can be classified into global search and fast search according to the matching search range. The global search method mainly performs exhaustive search on all regions within the search range, which has the highest accuracy, but also has relatively high computational complexity and is difficult to achieve real-time processing. On the other hand, fast search performs search on the macroblocks of some search regions in the search area according to the set rules. Therefore, compared with the global search, its search speed is fast, but there is a possibility that the optimal block is not searched. For example, the diamond search method (DS, Diamond Search), three-step search method (TSS, Three step Search), and four-step search method (FSS, Four Step Search) are all fast motion estimation methods based on local search, mainly by limiting the number of search steps or the number of search points, and accelerating the search speed through the appropriate use of search templates. The official test model of High Efficiency Video Coding (HEVC) provides two search methods, namely the basic full search algorithm and TZSearch. Here, the TZSearch algorithm is a fast search method based on a hybrid search model (renga search and grid search).In related technologies, most of the research on block search methods aims to improve the search speed and accuracy of blocks based on the TZSearch algorithm. Here, most of the work is optimized from aspects such as reducing the number of search blocks or macro-blocks, introducing thresholds, changing search strategies, and data reuse. However, for a large number of videos in the real situation, problems such as camera shake during shooting, low contrast of image frames, continuous change of moving scenes, and complexity exist, which easily cause incorrect block matching, and there are relatively obvious blurs or block effects in the obtained compensation frames. For these problems, the proposed fast motion estimation methods need to make an effective trade-off between computing resources and computing accuracy, which poses a challenge to even more efficient motion estimation methods.

[0025] The fast motion estimation method provided in related technologies is superior to the full search method in terms of speed. However, for many fast search computing methods, there is irregularity in their data access, and there is still room for improvement in search efficiency. In addition, when the motion estimation method in related technologies processes special videos with problems such as camera shake during shooting, low contrast of image frames, and continuous change of moving scenes, the optimal motion vector of the current block it obtains may have the phenomenon of incorrect sub-block matching, which easily causes obvious blurs and block effects in the obtained interpolation frames.

[0026] Based on this, in the video motion estimation method provided by the embodiments of the present application, by taking the continuous image frames of the video as the entire computing object and adding the foreground-background processing of the video to the restriction of the search range, the search time can be effectively reduced and the accuracy of motion estimation can be improved through the constraints on video content features.

[0027] Hereinafter, exemplary applications of the video motion estimation device provided by the embodiments of the present application will be described. The video motion estimation device provided by the embodiments of the present application may be implemented as any terminal having a screen display function, such as a notebook computer, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), a smart robot, etc., or may be implemented as a server.

[0028] Hereinafter, exemplary applications when the video motion estimation device is implemented as a terminal will be described.

[0029] As shown in FIG. 1, FIG. 1 is a schematic architecture diagram of a video motion estimation system 10 provided according to an embodiment of the present application. As shown in FIG. 1, the video motion estimation system 10 includes a terminal 400, a network 200, and a server 100. Here, an application program is running on the terminal 400. For example, this may be an image collection application program, or may further be an instant messaging application program, etc. When implementing the video motion estimation method of the embodiment of the present application, the terminal 400 acquires the video to be processed. Here, the video may be acquired through an image collection device built in the terminal 400. For example, it may be a video recorded by a camera in real time, or may further be a video locally stored in the terminal. After the terminal 400 acquires the video to be processed, it classifies a plurality of image frames included in the video based on the scene, and extracts the contour features of the foreground object in the plurality of image frames included in each scene, and determines the search range based on the contour features. Each scene corresponds to one search range. In a plurality of image frames of the same scene, a target block corresponding to the reference frame is searched in the search area determined by the search range, and further the motion vector is determined to complete the motion estimation process. Further, the terminal 400 sends the reference frame and the motion vector obtained through motion estimation to the server, and the server can perform motion compensation based on the motion vector, thereby obtaining a complete video file.

[0030] Hereinafter, an exemplary application when the video motion estimation device is implemented as a server will be described.

[0031] As shown in FIG. 1, here, an application program is running on the terminal 400. For example, it may be an image collection application program, or it may be an instant messaging application program or the like. When implementing the video motion estimation method of the embodiments of the present application, the terminal 400 acquires the video to be processed and sends the video to be processed to the server 100. After the server 100 acquires the video to be processed, it classifies a plurality of image frames included in the video based on the scene, extracts the contour features of the foreground object in the plurality of image frames included in each scene, and determines the search range based on the contour features. Each scene corresponds to one search range. In a plurality of image frames of the same scene, a target block corresponding to the reference frame is searched in the search area determined by the search range, and further the motion vector is determined to complete the motion estimation process, and motion compensation is performed based on the motion vector, thereby obtaining a complete video file.

[0032] As shown in FIG. 2, FIG. 2 is a structural schematic diagram of a video motion estimation device provided according to an embodiment of the present application. For example, the terminal 400 shown in FIG. 1 and the terminal 400 shown in FIG. 2 include at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is integrally coupled via a bus system 440. As can be understood, the bus system 440 is used to realize connection communication between these components. In addition to the data bus, the bus system 440 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, all various buses in FIG. 2 are marked as the bus system 440.

[0033] Processor 410 may be a kind of integrated circuit chip, having signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gates, or transistor logic devices, discrete hardware components, etc. Here, the general-purpose processor may be a microprocessor or any general processor, etc.

[0034] User interface 430 includes one or more output devices 431 that enable the presence of media content, which includes one or more speakers and / or one or more visual display screens. User interface 430 further includes one or more input devices 432, which include user interface members that are advantageous for user input, such as a keyboard, a mouse, a microphone, a touch panel display screen, a camera, other input buttons, and control members.

[0035] Memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, and optical disk drives, etc. Memory 450 includes one or more storage devices that are selectively physically located away from processor 410.

[0036] Memory 450 includes volatile memory, non-volatile memory, or may further include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). Memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0037] In some embodiments, the memory 450 can store data, thereby supporting various operations. Examples of these data include programs, modules, and data structures, or subsets or supersets thereof. This will be illustrated below.

[0038] The operating system 451 includes system programs used for processing various basic system services and executing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc. These are used for realizing various basic services and processing tasks based on hardware. The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth (registered trademark), Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc. The presence module 453 is used to present information via one or more output devices 431 (such as a display screen, speaker, etc.) associated with the user interface 430 (for example, used to operate peripheral devices, and the user interface for displaying content and information). The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432, and translate the detected input or interaction.

[0039] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented by adopting a software method. FIG. 2 shows a video motion estimation device 455 stored in a memory 450, which may be software in the form of a program, a plug-in, etc., and includes software modules called a first acquisition module 4551, a feature extraction module 4552, a first determination module 4553, a second determination module 4554, and a motion estimation module 4555. These modules are logical, and thus, any combination or further partitioning can be performed according to the realized functions.

[0040] The functions of each module will be described below.

[0041] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented by adopting a hardware method. For example, the apparatus provided by the embodiments of the present application may be a processor adopting a hardware decode processor form, which is programmed to execute the video motion estimation method provided by the embodiments of the present application. For example, a processor in the form of a hardware decode processor may adopt one or a plurality of application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0042] Next, an exemplary application of the terminal 400 provided by the embodiments of the present application and implementation will be combined to describe the video motion estimation method provided by the embodiments of the present application. As shown in FIG. 3, FIG. 3 is a schematic flowchart of one implementation of the video motion estimation method provided by the embodiments of the present application, and the steps shown in FIG. 3 will be described in combination.

[0043] Step S101: Obtain a plurality of image frames in the video to be processed, perform scene classification processing on the plurality of image frames, and obtain a plurality of image frame sets.

[0044] Here, the video to be processed may be a video recorded by the terminal in real time, or may be a video locally stored in the terminal, or may be a video downloaded by the terminal from the server.

[0045] When realizing step S101, scene classification may be performed based on the background image of the image frame. When the background images of the plurality of image frames are similar, they can be considered to be located in the same scene. For example, when the video is a holiday performance, it can be divided into a plurality of different programs, and the backgrounds of different programs can be different. Therefore, each program can be classified into one scene. Each scene corresponds to one set of image frames, and each set of image frames includes at least one image frame.

[0046] In some embodiments, one set of image frames can be understood as one 3D image block. The three dimensions in the 3D image block are the number of frames, the frame width, and the frame height, respectively. Here, the number of frames is the number of image frames included in the set of image frames, the frame width is the width of the image frame, and in actual implementation, it can be represented by the number of pixel points in the width direction. The frame height is the height of the image frame, and it can be represented by the number of pixel points in the height direction of the image frame.

[0047] Step S102: Extract the contour features and color features of the foreground object in each image frame in each set of image frames.

[0048] Here, calculate and collect the contour features of the foreground object through the optical flow vector gradient of the optical flow field, and extract the color features of the foreground object according to the image area where the foreground object exists.

[0049] In some embodiments, the prior knowledge of the background structure of consecutive sample image frames and the foreground prior area in the sample image frames can be further used to train the image segmentation model. Use the trained image segmentation model to realize foreground object segmentation and scene classification estimation, and extract the color features of the image area of the segmented foreground object. The color features of each image frame in a set of image frames further constitute a color feature sequence.

[0050] In some embodiments, the scene classification process in step S101 and the process of extracting the contour features and color features of the foreground object in step S102 can input the video to be processed into the trained image segmentation model, thereby completing the scene classification and feature extraction for multiple image frames in the video to be processed.

[0051] Step S103: Determine the search range corresponding to each set of image frames based on the contour features of the foreground object in each set of image frames.

[0052] Here, the contour of the foreground object can be represented through a rectangle, a square, an irregular figure, etc., and the contour features of the foreground object can include the coordinates of the contour vertices. When realizing step S103, the search range including the foreground object in all image frames can be determined through the coordinates of the contour vertices of the foreground object in each image frame in the set of image frames.

[0053] Step S104: Determine the start search point of each predicted frame in each set of image frames.

[0054] Here, in each set of image frames, sorting is performed according to the time sequence of the multiple image frames included in the set of image frames. It is common to use the first image frame as the reference frame and the other image frames as predicted frames. Further, it is also possible to use the i-th image frame as the reference frame and the (i + 1)-th image frame as the predicted frame, where i is a positive integer that increases incrementally. What needs to be explained is that the reference frame is the image frame to be referred to when calculating the motion vector in the set of image frames, and the predicted frame is the image frame used for calculating the motion vector in the set of image frames.

[0055] When Step S104 is realized, based on the relevance in the spatial region and the temporal region of the video sequence frames, median prediction, upper layer prediction, and origin prediction in the spatial region are adopted in sequence to perform prediction on the motion vector of the target block in each predicted frame, thereby determining the position of the start search point in each predicted frame.

[0056] Step S105: Based on the start search point of each predicted frame, the target block in the reference frame, and the color feature of the foreground object, perform motion estimation within the search region corresponding to the search range in each predicted frame to obtain the motion vector corresponding to the target block.

[0057] Here, when searching for a target block that matches the reference target block in the reference frame (i.e., a certain target block in the reference frame) from each prediction frame based on the reference target block, a two-way motion estimation calculation can be performed with the search area corresponding to the search range centered on each start search point in each prediction frame and the color feature of the foreground object as a constraint condition. It is maintained to complete the search for the target block of the prediction frame within the search area centered on the start search point. For example, an asymmetric cross search template where the w-axis search point is twice the h-axis search point is adopted, and according to the magnitudes of the components of the prediction vector on the w-axis and h-axis, it can be determined whether the foreground motion in the slider P is horizontal motion or vertical motion. If it is horizontal motion, the horizontal-direction asymmetric cross shape of the UMHexagonS original template is adopted. If it is determined to be vertical motion, a template where the h-axis is twice the w-axis search point is adopted for the search, thereby determining the target block in the prediction frame. Next, based on the position information of the target block and the reference target block, the motion vector corresponding to the target block is determined.

[0058] In the video motion estimation method provided by the embodiments of the present application, after obtaining a plurality of image frames in the video to be processed, first perform scene classification on the plurality of image frames to obtain a plurality of image frame sets. That is, each scene corresponds to one image frame set, and each image frame set contains one or more image frames, and the backgrounds of the image frames belonging to the same scene are similar. Further, extract the contour features and color features of the foreground objects in each image frame in each image frame set, and determine the search range corresponding to each image frame set based on the contour features of the foreground objects in each image frame set. Next, determine each start search point of each prediction frame in each image frame set, and further perform motion estimation within the search area corresponding to the search range of each prediction frame based on each start search point, the target block in the reference frame, and the color features of the foreground object, to obtain the motion vector corresponding to the target block. In the embodiments of the present application, by searching within a predetermined range, the search range can be narrowed, thereby reducing the search time, and since the search range is limited based on the contour features of the foreground objects in each scene, the accuracy of motion estimation can be further improved.

[0059] In some embodiments, the step of performing scene classification on a plurality of image frames and obtaining a plurality of image frame sets in step S101 can be realized through the following method. For any image frame in any image frame set, perform the process of determining the background image regions in the plurality of image frames and determining the image similarity between the plurality of background image regions, and perform the scene classification process on the plurality of image frames and obtain a plurality of image frame sets based on the image similarity between the plurality of background image regions.

[0060] Here, when determining the background image regions in multiple image frames and determining the image similarity between multiple background image regions, first target detection is performed, and first the foreground objects in multiple image frames are identified. For example, the detection and extraction of foreground objects can be performed by using the difference between two consecutive frames or several frames of images in a video sequence, and by using time information, the tone difference value of corresponding pixel points is obtained by comparing some consecutive frames in the image. If all are greater than a predetermined threshold, it can be determined that a foreground object exists at this position, and at this time, the other regions other than this position are background image regions.

[0061] In some embodiments, the optical flow field method can be further used to detect foreground objects. When realizing, the change of the two-dimensional image is evaluated by using the tone preservation principle of corresponding pixels in two adjacent frames, and related foreground objects can be detected relatively well from the image frames. The optical flow field method is applied to the detection of foreground targets that are relatively moving in the movement process of a video camera.

[0062] After obtaining the background image regions, the image similarity between each background image region can be calculated. When realizing, the histogram matching algorithm can be used to calculate the image similarity between each image background region. For example, there are background image regions A and B, and the histograms of the two images, namely HistA and HistB, are calculated respectively. Next, the normalization correlation coefficients of the two histograms (such as the Bhattacharyya distance, histogram intersection distance, etc.) are calculated, and thereby the similarity between the two is determined. In some embodiments, the image similarity calculation can be further performed based on feature points. When realizing, the feature points in each background image region are extracted respectively, the Hamming distance between the feature points is calculated, and thereby the similarity value between the background image regions can be determined.

[0063] Here, when performing scene classification on a plurality of image frames after calculating the image similarity between a plurality of background image regions, based on the time information of each image frame, the image frames corresponding to the background image regions with continuous time and high similarity are classified into one set of image frames.

[0064] In the above embodiment, through the background image, scene classification is performed on a plurality of image frames in the video. Since the movement range of the foreground object in the same scene is relatively small, the search range of the foreground object can be determined based on the scene, while minimizing the search range and ensuring accuracy at the same time.

[0065] In some embodiments, the step of extracting the contour features and color features of the foreground object in each image frame in each set of image frames in step S102 shown in FIG. 3 can be realized through the method described below. For any image frame in any set of image frames, perform the process of determining the foreground image region where the foreground object exists in the image frame, the process of using the position information of the foreground image region as the contour feature of the foreground object in the image frame, and the process of performing color extraction processing based on the foreground image region to obtain the color feature of the foreground object in the image frame.

[0066] Here, when obtaining the background image region, the foreground and background of the image frame are divided, whereby the background image region, the foreground image region where the foreground object exists, and the position information of the foreground image region can be determined.

[0067] In the embodiments of the present application, the contour of the foreground image region does not necessarily perfectly fit the foreground object. For example, when the foreground object is a person, the contour of the foreground image region may be a rectangle or a square that can include the person, and does not necessarily have to be the contour of a human. Therefore, the position information of the foreground image region can be represented using the vertex coordinates on the contour of the foreground object, that is, the contour features of the foreground object include the coordinates of each vertex of the foreground image region.

[0068] Here, the color feature of the foreground object may be further understood as the color feature of the foreground object. The color feature is a visual feature applied in image retrieval, and the color is often highly related to the objects or scenes included in the image. Also, compared with other visual features, the color feature has relatively low dependence on the size, direction, and viewpoint of the image itself, and thus has relatively high robustness.

[0069] In actual implementation, the color feature can be represented using multiple methods such as color histograms, color moments, color sets, color coherence vectors, and color correlograms.

[0070] In the above embodiments, the contour features and color features of the foreground object in each image frame in each set of image frames can be extracted, thereby providing a data basis for setting the search range and setting the constraint conditions for motion estimation when determining the target block that matches the reference target block in frame prediction.

[0071] As shown in FIG. 4A, based on the contour features of the foreground object in each set of image frames in step S103 shown in FIG. 3, the step of determining the search range for each set of image frames can be realized through the following steps.

[0072] Step S1031: Execute the following process for an arbitrary set of image frames. Based on the position information of each foreground image region in the set of image frames, determine the vertex coordinates in each foreground image region.

[0073] Here, the position information of the foreground image region can be represented using the vertex coordinates of the foreground image region. For example, when the foreground image region is a rectangular region, it is necessary to determine the coordinates of the four vertices of the rectangular region. Assume that the coordinates of the four vertices A, B, C, and D of the foreground image region in a certain prediction frame are (100, 100), (100, 500), (300, 500), and (300, 100) respectively.

[0074] Step S1032: Determine the first maximum value and the first minimum value corresponding to the first dimension, and determine the second maximum value and the second minimum value corresponding to the second dimension from the respective vertex coordinates.

[0075] Here, the first dimension and the second dimension are different. For example, the first dimension may be the width, and the second dimension may be the height. When realizing Step S1032, determine the first maximum value and the first minimum value corresponding to the first dimension, and determine the second maximum value and the second minimum value corresponding to the second dimension from the respective vertex coordinates of each foreground image region belonging to the same set of image frames.

[0076] For example, if a set of image frames contains 100 image frames and the foreground image region in each image frame is a rectangular region, when realizing Step S1032, determine the first maximum value and the first minimum value of the first dimension and the second maximum value and the second minimum value of the second dimension from 400 vertex coordinates.

[0077] Step S1033: Determine the search range corresponding to the set of image frames based on the first minimum value, the first maximum value, the second minimum value, and the second maximum value.

[0078] Here, after determining the first minimum value, the first maximum value, the second minimum value, and the second maximum value, the search range corresponding to the set of image frames can be determined. That is, the search range in the first dimension is greater than or equal to the first minimum value and less than or equal to the first maximum value, and the search range in the second dimension is greater than or equal to the second minimum value and less than or equal to the second maximum value. For example, using these four values, four vertex coordinates, namely (the first minimum value, the second minimum value), (the first minimum value, the second maximum value), (the first maximum value, the second minimum value), and (the first maximum value, the second maximum value), can be determined, and based on these four vertex coordinates, the search range corresponding to the set of image frames is determined.

[0079] Taking an example for explanation, when the first minimum value is 100, the first maximum value is 600, the second minimum value is 100, and the second maximum value is 800, the four vertex coordinates determined based on these four values are (100, 100), (100, 800), (600, 100), and (600, 800) respectively. Therefore, the search range is the rectangular area determined by these four vertices.

[0080] Through the above steps S1031 to S1033, based on the contour features of the foreground objects in multiple image frames belonging to the same scene, the search area is determined. Since the movement range of the foreground objects in the same scene is generally relatively small, it is possible to ensure that the foreground objects in all the image frames in the scene are included in the search range determined by the maximum and minimum coordinates in two dimensions in multiple foreground image areas, thereby ensuring the calculation accuracy of motion estimation.

[0081] In some embodiments, the step of determining the start search point of each predicted frame in each set of image frames in step S104 shown in FIG. 3 can be realized through the following method. Determine the position information of the reference target block in each reference frame in each set of image frames. Here, the reference target block is any target block in the reference frame. Predict the motion vector of each predicted frame through the set prediction mode to obtain the predicted motion vector of each predicted frame. The prediction mode includes at least one of the median prediction mode, the upper layer block prediction mode, and the origin prediction mode. Based on the position information of the reference target block and the predicted motion vector, determine the start search point of each predicted frame.

[0082] Here, the reference frame is also divided into foreground and background, and after determining the foreground image area in the reference frame, the foreground image area can be divided to obtain a plurality of reference target blocks. The size of the reference target block may be 4*4, 8*8, etc. The position information of the reference target block is represented by using the coordinates of one vertex of the reference target block. For example, it can be represented by using the coordinates of the upper left corner vertex in the reference target block.

[0083] Here, the prediction mode includes at least one of median prediction, upper layer block prediction, and origin prediction.

[0084] Due to the integrity of the moving object and the continuity of the video motion, there must be temporal and spatial correlations in the video motion, and there is a correlation between adjacent blocks. Therefore, the motion vector of the current block can be predicted through the motion vectors of adjacent blocks. When realizing, the initial motion vector of the current block can be predicted according to the motion vector of the block adjacent to the current block in the spatial position (median prediction) or the block at the same position in the previous frame image in time (origin prediction), thereby determining the initial search point.

[0085] Here, when realizing means for determining the start search point of each predicted frame based on the position information of the reference target block and the predicted motion vector, the position information of the reference target block can be moved according to the predicted motion vector, thereby determining each start search point of each predicted frame.

[0086] A high-precision start search point can promote the search point to approach the target block in the predicted frame as much as possible, thereby improving the search speed. In the above embodiment, at least one of median prediction, upper layer prediction, and origin prediction in the spatial region is adopted based on the relevance in the spatial region and the temporal region of the video sequence frame to perform prediction on the current motion vector, thereby determining the position of the optimal start search point and ensuring the accuracy of the start search point.

[0087] In some embodiments, based on the start search point of each predicted frame, the target block in the reference frame, and the color feature of the foreground object in step S105 shown in FIG. 3, the motion estimation process is performed within the search region corresponding to the search range of each predicted frame to obtain the motion vector corresponding to the target block. The steps can be realized through steps S1051 to S1058 shown in FIG. 4B. Hereinafter, each step will be described in combination with FIG. 4B.

[0088] Step S1051: Determine the first search template corresponding to each predicted frame.

[0089] Here, the first search template may be an asymmetric cross template, or may be a hexagonal template, a rhombic template, etc. The first search template may be determined according to the predicted motion direction of the foreground object.

[0090] In some embodiments, the above step S1051 can be implemented through the following method. Based on the predicted motion vector of the predicted target block of each predicted frame, determine the first motion direction of the foreground object in the predicted frame, and based on the first motion direction of the foreground object, determine the search template corresponding to each predicted frame.

[0091] Here, the first motion direction may be the horizontal direction, the vertical direction, or even the oblique direction.

[0092] In some embodiments, when predicting the first motion direction of the foreground object, it may also be determined based on the motion direction of the reference frame of the frame before the predicted frame. For example, the motion direction of the reference frame of the frame before the predicted frame can be determined as the first motion direction of the foreground object.

[0093] Step S1052: Centering on each starting search point, perform a search process within the search area corresponding to the search range in the predicted frame through the first search template to obtain the reference target block and the corresponding predicted target block in the predicted frame.

[0094] Here, when implementing step S1052, by performing a search process within the search area corresponding to the search range in the predicted frame through the first search template centered on the starting search point, each candidate target block can be determined. Next, the candidate target block is matched with the reference target block, thereby determining the predicted target block corresponding to the reference target block.

[0095] In the embodiments of the present application, in order to fully utilize the constraints on the motion estimation of color features, color feature constraints are added on the basis of the motion estimation target function. That is, assuming that SADcolor represents the bidirectional motion estimation function of color features, SADobj represents the bidirectional motion estimation function of the target sequence of foreground objects, and λ1 and λ2 are the weight coefficients of color features and the target sequence of foreground objects respectively, the weight coefficients can be dynamically adjusted according to the ratio of the binary sequence features in the preprocessing stage. Therefore, the motion estimation target function SAD in the embodiments of the present application can be expressed as SAD = λ1SADcolor + λ2SADobj.

[0096] Step S1053: Determine the degree of texture difference between the reference target block and the predicted target block.

[0097] Here, when realizing step S1053, the texture features between the reference target block and the predicted target block can be extracted, and then the texture difference degree value between the two can be determined through the texture features between the reference target block and the predicted target block.

[0098] Step S1054: Determine whether the degree of texture difference is less than the difference threshold.

[0099] Here, when the degree of texture difference is less than the preset difference threshold, it is explained that the texture difference between the reference target block and the predicted target block is relatively small, so that the predicted target block is considered to be the correct target block. At this time, step S1055 is entered. When the degree of texture difference is greater than or equal to the difference threshold, it is explained that the texture difference between the reference target block and the predicted target block is relatively large, so that the predicted target block is considered to be an incorrect target block. At this time, step S1056 is entered.

[0100] Step S1055: Determine the motion vector corresponding to the predicted target block based on the position information of the reference target block and the position information of the predicted target block.

[0101] Here, after obtaining the position information of the reference target block and the position information of the predicted target block, the motion vector corresponding to the predicted target block can be determined. When implementing, the two vertex coordinates used to characterize the position information can be subtracted. That is, by subtracting the vertex coordinates of the predicted target block from the vertex coordinates of the reference target block, the motion vector corresponding to the predicted target block can be obtained.

[0102] Step S1056: Determine the degree of color difference and the degree of texture difference between each predicted block in the search area corresponding to the search range in the predicted frame and the reference target block.

[0103] Here, when the texture difference between the predicted target block and the reference target block is relatively large, the degree of color difference and the degree of texture difference between each predicted block in the search area and the reference target block can be determined in sequence.

[0104] Step S1057: Determine the predicted target block corresponding to the reference target block from each predicted block based on the degree of color difference and the degree of texture difference between each predicted block and the reference target block.

[0105] Here, when implementing step S1057, those with a color difference degree between each predicted block and the reference target block less than the color threshold and a texture difference degree less than the difference threshold may be selected. If there are those with a color difference degree between the predicted block and the reference target block less than the color threshold and a texture difference degree less than the difference threshold, the predicted block with the minimum difference is determined as the predicted target block.

[0106] Step S1058: Determine a motion vector corresponding to the predicted target block based on the position information of the reference target block and the position information of the predicted target block.

[0107] In the above steps S1051 to S1058, after determining each search template, centering on the start search point, using the search template to search for a predicted target block that matches the reference target block in the reference frame within the search area of the predicted frame, and further comparing the degree of texture difference between the predicted target block and the reference target block is necessary. When the degree of texture difference is less than the difference threshold, it is considered to match the accurate predicted target block, and when the degree of texture difference is greater than or equal to the difference threshold, it is considered not to match the accurate predicted target block. At this time, each predicted block in the search area of the predicted frame can be traversed, and an accurate predicted target block can be determined from them. In this way, the accuracy of the predicted target block can be ensured, and the accuracy of motion estimation can be further improved.

[0108] In the actual implementation process, as shown in FIG. 4C, step S1052 can be realized through the following steps.

[0109] Step S10521: Determine a plurality of first candidate blocks in the search area based on the first search template centering on each start search point.

[0110] Here, taking the first search template as an asymmetric cross template in the horizontal direction, with 6 blocks in the horizontal direction and 3 blocks in the vertical direction as an example for explanation. When realizing step S10521, centering on the start search point, 3 predicted blocks above and below the start search point and 6 adjacent predicted blocks on the left and right can be determined as the first candidate blocks.

[0111] Step S10522: Determine the matching order of a plurality of first candidate blocks based on the predicted motion vector.

[0112] Here, when realizing step S1052, the matching order of a plurality of candidate blocks can be determined based on the region corresponding to the predicted motion vector, or the matching order of a plurality of candidate blocks can be determined according to the distance between the predicted motion vector and each candidate block. Taking the above example, for instance, if the predicted motion vector is in the horizontal leftward direction, it preferentially matches with six candidate blocks on the left side of the starting search point.

[0113] Step S10523: Perform matching processing on each first candidate block and the reference target block based on the matching order, and determine whether there is a first candidate target block that matches the reference target block among the plurality of first candidate blocks.

[0114] Here, when there is a first candidate target block that matches the reference target block among the plurality of first candidate blocks, it enters step S10524; when there is no candidate target block that matches the reference target block among the plurality of first candidate blocks, it enters step S10525.

[0115] In some embodiments, when there is no first candidate target block that matches the reference target block among the plurality of first candidate blocks, traversal can be directly performed on each prediction block in the search region in the prediction frame, thereby determining the prediction target block.

[0116] Step S10524: Determine the candidate target block as the prediction target block.

[0117] Step S10525: Determine a second search template corresponding to each prediction frame based on the second motion direction.

[0118] Here, the second movement direction is different from the first movement direction. If the prediction target block is not matched through the first search template determined based on the first movement direction, it can be considered that the prediction of the first movement direction is incorrect. At this time, the second search template can be determined according to the second movement direction, and the search for the prediction target block can be performed again.

[0119] Step S10526: Centering on the start search point, determine a plurality of second candidate blocks in the search area based on the second search template.

[0120] Step S10527: Determine the matching order of the plurality of second candidate blocks based on the predicted motion vector.

[0121] Step S10528: Perform matching processing on each second candidate block and the reference target block based on the matching order, and determine whether there is a second candidate target block that matches the reference target block among the plurality of second candidate blocks.

[0122] Step S10529: When there is a second candidate target block that matches the reference target block among the plurality of second candidate blocks, determine the second candidate target block as the prediction target block.

[0123] Here, the realization processes of steps S10526 to S10529 are similar to those of steps S10521 to S10524. According to the above steps S10521 to S10529, a search template can be determined through the predicted movement direction of the foreground object, and a candidate block preferentially matched through the predicted movement vector can be determined. Thereby, when the candidate block and the reference target block are matched and the predicted target block is not matched through the search template, the search template can be determined again based on a movement direction different from the predicted movement direction of the foreground object to search for the predicted target block. In this way, the matching speed can be improved, and thereby the processing efficiency of motion estimation can be improved.

[0124] Based on the above-described embodiments, the embodiments of the present application further provide a video motion estimation method applied to the network architecture shown in FIG. 1. FIG. 5 is a schematic diagram of a realization flow of a video motion estimation method provided by an embodiment of the present application. As shown in FIG. 5, the video motion estimation method includes the following flow.

[0125] Step S501: The terminal activates an image collection device based on the received image collection instruction.

[0126] Here, the image collection instruction may be an operation instruction for instructing video collection. The image collection instruction is triggered through an instant messaging application. Of course, it may also be triggered through an office application, or further triggered through a short video application.

[0127] Step S502: The terminal obtains a plurality of image frames collected by the image collection device.

[0128] Here, after the image collection device is activated, image collection is performed, thereby obtaining a plurality of image frames.

[0129] Step S503: The terminal performs scene classification processing on a plurality of image frames to obtain a plurality of sets of image frames.

[0130] Here, the terminal can perform scene segmentation on a plurality of image frames by combining the optical flow field and the geometric scene classification method estimated based on the scene structure, thereby obtaining a plurality of sets of image frames, and each set of image frames includes at least one image frame.

[0131] Step S504: The terminal extracts the contour features and color features of the foreground objects in each image frame in each set of image frames.

[0132] Step S505: The terminal determines the vertex coordinates in each foreground image region based on the position information of each foreground image region in each set of image frames.

[0133] Step S506: The terminal determines a first maximum value and a first minimum value corresponding to the first dimension from each vertex coordinate, and determines a second maximum value and a second minimum value corresponding to the second dimension.

[0134] Step S507: The terminal determines a search range corresponding to the set of image frames based on the first minimum value, the first maximum value, the second minimum value, and the second maximum value.

[0135] Step S508: The terminal determines the start search point of each prediction frame in each set of image frames.

[0136] Step S509: The terminal performs motion estimation within the search region corresponding to the search range in each prediction frame based on each start search point, the target block in the reference frame, and the color features of the foreground object, and obtains a motion vector corresponding to the target block.

[0137] Step S510: The terminal performs video encoding based on the motion vector and a plurality of image frames to obtain an encoded video.

[0138] Step S511: The terminal sends the encoded video to the server.

[0139] Here, the server may be a service server corresponding to the trigger of the image collection instruction application, for example, an instant messaging server, an office application server, or a short video server.

[0140] Step S512: The server performs motion compensation on the encoded video based on the motion vector to obtain each decoded image frame.

[0141] In the video motion estimation method provided by the embodiments of the present application, after the terminal collects a plurality of image frames of the video, first, scene classification is performed on the plurality of image frames to obtain a plurality of image frame sets. That is, each scene corresponds to one image frame set, and each image frame set includes one or more image frames, and the backgrounds of the image frames belonging to the same scene are similar. Further, the contour features and color features of the foreground objects in each image frame in each image frame set are extracted, and based on the vertex coordinates of the contours of the foreground objects in each image frame set, the maximum value and the minimum value of the coordinates are determined. Further, a search range corresponding to each image frame set is determined, and then each start search point of each prediction frame in each image frame set is determined. Further, based on each start search point, the target block in the reference frame, and the color features of the foreground object, motion estimation is performed within the search area corresponding to the search range in each prediction frame to obtain the motion vector corresponding to the target block. Since the search range is determined based on the vertex coordinates of the contour, it is possible to ensure that the search range is as small as possible on the premise of including the foreground object. Thereby, the search time can be reduced, and further the accuracy of motion estimation can be ensured. Thereafter, the terminal sends the reference frame and the motion vector to the server, which can reduce the requirement for the data bandwidth, further reduce the transmission delay, and improve the transmission efficiency.

[0142] The following describes an exemplary application in one actual application scenario of an embodiment of the present application.

[0143] The embodiments of the present application can be applied to video applications such as video storage applications, instant messaging applications, video playback applications, video call applications, and live applications. Taking the instant messaging application as an example, the instant messaging application is executed on the call sending side. The video sender obtains the video to be processed (for example, a recorded video), and searches for the target block corresponding to the reference frame in the search area corresponding to the search range to determine the motion vector, and performs video encoding based on the motion vector, and sends the encoded video to the server. The server sends the encoded video to the video receiving side, and the video receiving side decodes the received encoded video to play the video, thereby improving the video transmission efficiency. Taking the video storage application as an example, the video storage application is executed on the terminal. The terminal obtains the video to be processed (for example, a video recorded in real time), and searches for the target block corresponding to the reference frame in the search area corresponding to the search range to determine the motion vector, and performs video encoding based on the motion vector, and sends the encoded video to the server to realize cloud storage means, thereby saving storage capacity. The following describes the video motion estimation method provided by the embodiments of the present application in combination with video scenes.

[0144] FIG. 6 is a schematic flowchart of the realization of a video motion estimation method based on a 3D image block provided by an embodiment of the present application. As shown in FIG. 6, the flow includes the following.

[0145] Step S601: Perform an initialization definition on the video to be processed.

[0146] When realizing, the video sequence data Vo(f1, f2, …, fn) can be defined as a three-dimensional spatial rectangular parallelepiped of F*W*H. Here, F, W, and H are the number of frames in the time domain of Vo, the frame width in the spatial domain, and the frame height, respectively. In Vo, a three-dimensional rectangular parallelepiped slider P (P ∈ Vo) with a length, width, and height of f, w, and h respectively is set, and it is set that the initial position point of P in Vo is O(0, 0, 0). Here, the initial position point O is the boundary initial position of Vo.

[0147] Step S602: Extract the motion characteristics of the foreground object based on the optical flow field to obtain the foreground motion contour characteristics.

[0148] Considering that the image segmentation method based on video object segmentation is easily affected by situations such as complex environments, lens movement, and unstable light irradiation, in the embodiments of the present application, global optical flow (such as Horn-Schunck optical flow) and image segmentation are combined to process the foreground of the video. When realizing step S602, it is possible to calculate through the optical flow vector gradient of the optical flow field and collect the motion contour characteristics of the foreground object.

[0149] Step S603: Establish a video scene motion model and extract the foreground area of the sequence frame.

[0150] Here, scene segmentation is performed on the video Vo by combining the optical flow field and the geometric scene segmentation method estimated based on the scene structure, and the foreground object and its corresponding color block color sequence characteristics are extracted, and the extraction, segmentation results, and color information are used for the constraints in the matching process of each macroblock in the image.

[0151] For example, based on the motion information of the video foreground object in the Horn-Schunck optical flow field and the prior knowledge of the background structure of consecutive frames, the prior foreground regions are combined to establish a motion model of the video scene, and the prior knowledge of the foreground regions in consecutive image frames is extracted. Further, the extreme value is obtained by iterating the probability density function of the pixel points in consecutive frames, the pixel points of the same type of regions are classified, the scene segmentation is realized, and at the same time, the color sequence information of the divided image blocks is extracted, and the segmentation result is improved based on scene structure estimation and classification.

[0152] In this way, the foreground object segmentation and the estimation of scene classification are completed, and the color sequence information of the divided image is extracted.

[0153] Step S604: Obtain the video foreground motion object sequence and the background segmentation sequence.

[0154] In some embodiments, in the video preprocessing stage where the above steps S601 to S604 exist, the extraction of the video foreground motion information and the segmentation of the video scene can be realized through a method based on a neural network.

[0155] Step S605: Perform motion estimation calculation based on the 3D sequence image frames.

[0156] Here, the side length value range of the slider P is set by combining the video foreground motion information and the features of the video sequence image frames in the three directions of F, W, and H of Vo. The current motion vector is predicted and the position of the starting search point is determined, and the initial position point O of Vo of the slider P is initialized according to the starting search point. When actually realized, based on the relevance in the spatial and temporal regions of the video sequence frames, median prediction, upper layer prediction, and origin prediction in the spatial region are adopted in sequence to predict the current motion vector, and thereby the position of the optimal starting search point can be determined. The initial position point O of Vo of the slider P is set to the center position of the foreground motion region belonging to the three directions of f, w, and h where the position of the determined starting search point exists.

[0157] Centering around the starting prediction point O, under the constraint range of the spatial, temporal regions, and color sequence features limited by P, based on the bidirectional motion estimation idea, the estimation of the motion vector is realized through an improved UMHexagonS search template.

[0158] For example, bidirectional motion estimation calculation is performed with the edge position of the rectangular parallelepiped slider P, the divided scene, and the color sequence characteristics of the video as constraints. Centering on the start prediction point, the completion of the search for the target macroblock of the current frame within P is maintained, and an asymmetric cross search template where the w-axis search point is twice the h-axis search point is adopted. According to the magnitudes of the components of the prediction vector of the previous step on the w-axis and h-axis, it is determined whether the foreground motion in the slider P is horizontal motion or vertical motion. If the foreground motion in the slider P is horizontal motion, the horizontal asymmetric cross shape of the UMHexagonS original template is adopted. If the foreground motion in the slider P is vertical motion, a template where the h-axis is twice the w-axis search point is adopted. Different sub-regions are preferentially searched according to the region corresponding to the prediction motion vector. As shown in FIG. 7, when the prediction motion vector corresponds to the first quadrant, the sub-region shown in 701 is preferentially searched. When the prediction motion vector corresponds to the second quadrant, the sub-region shown in 702 is preferentially searched. When the prediction motion vector corresponds to the third quadrant, the sub-region shown in 703 is preferentially searched. When the prediction motion vector corresponds to the fourth quadrant, the sub-region shown in 704 is preferentially searched. Thereby, in addition to being able to reduce the search time cost, since it is restricted by the scene sequence characteristics of the video, the matching rate of the target macroblock can be improved.

[0159] Step S606: Perform motion estimation optimization based on the energy function.

[0160] Here, considering that the color information of macroblocks at different positions in the image frame may be similar, incorrect block matching is very likely to occur in the macroblock search of consecutive frames. However, since the consistency of the scene segmentation image represents the specific texture information of the video image frame, the texture difference between two similar macroblocks can be effectively discriminated, the motion information of each macroblock in the image frame can be accurately tracked, and the motion vector field of the image can be corrected.

[0161] When implementing, constraints are imposed on the calculation and estimation of the motion vector in step S605 through a consistency energy function, and a consistency division means is adopted to determine whether the color sequence information of each divided image matches, thereby detecting and correcting the situation of macroblock mismatching, and improving the accuracy of the motion vector field.

[0162] The higher the similarity between macroblocks, the smaller the value of the consistency energy function; conversely, the lower the similarity between macroblocks, the larger the value of the consistency energy function. Therefore, when obtaining the optimal motion vector of macroblocks corresponding to consecutive frames, the minimum value of the consistency energy function can be obtained to optimize the incorrect motion vector and improve the accuracy of motion estimation. When actually implementing, to improve the search efficiency, when determining the motion vector of macroblocks corresponding to consecutive frames, if it is further determined that the function value of the consistency energy function is less than a preset threshold, it is considered that the macroblock corresponding to the reference frame has been searched. At this time, the motion vector can be determined based on the position information of the macroblock in the reference frame and the macroblock searched from the current frame.

[0163] The motion estimation method based on macroblocks, considering that it mainly determines the optimal motion vector of a macroblock by obtaining the minimum absolute error between the reference frame and the macroblock corresponding to the current frame, has a long calculation time and high complexity. Especially for videos with particularly complex scenes, the accuracy rate of motion estimation is unstable. However, by combining with a high-speed motion estimation method based on foreground-background preprocessing of video content, the motion estimation time can be reduced, the complexity of motion estimation can be decreased, and the accuracy of the motion vector field can be improved. Therefore, in the embodiments of the present application, by combining the unique structural features of video sequence frames and the advantages and disadvantages of the motion estimation method based on macroblocks, preprocessing is performed on the video sequence image frames as one 3D overall calculation object and its motion vector information is calculated to achieve more efficient motion estimation, and the video encoding time is reduced under the condition of ensuring a predetermined coding rate-distortion performance.

[0164] FIG. 8 is a realization schematic diagram of the motion estimation of a 3D image block provided by an embodiment of the present application. As shown in FIG. 8, a 3D image set of a series of frames included in a video is regarded as one three-dimensional calculation object Vo(f 1 , f 2 , …, f n ). After preprocessing, each scene S 1 (f 1 , f …, f i-1 ), S2(f i , f i+1 , f …), …, S N} of Vo is listed as one search group. Vo is divided into N cuboid blocks according to the number of scenes N, and the slider P sequentially moves from the first cuboid scene S i to the Nth cuboid S 1 until the motion estimation of Vo is completed. NTraverse to. The values of f, w, and h in each scene of the slider P are determined by the motion ranges in three directions of consecutive frames of the foreground motion target. Starting from F = 0, bidirectional motion estimation is performed. Through the prediction start search point and according to the search template provided in the above step S605, the target block search for the current macroblock of the current frame is completed, and the corresponding motion vector is extracted. Each time a search step is executed, the slider P slides along with the search direction of the step of the search template to ensure that the search range is restricted within the three-dimensional slider P, so that the motion characteristics of the foreground object can constrain the search range of the target matching macroblock and achieve the purpose of reducing the search points.

[0165] Also, in the embodiment of the present application, when performing bidirectional motion estimation calculation, in order to fully utilize the color sequence to constrain the motion estimation, a constraint on the color sequence feature is added on the basis of the original bidirectional motion estimation target function. That is, SAD color represents the bidirectional motion estimation function of the color sequence frame, and SAD obj represents the bidirectional motion estimation function of the target sequence of the foreground object. Let λ 1 , and λ 2 be the weight coefficients of the color sequence and the target sequence of the foreground object respectively. The weight coefficients can be dynamically adjusted according to the ratio of the binary sequence features in the preprocessing stage. Therefore, in the embodiment of the present application, the bidirectional motion estimation target function SAD of Vo is SAD = λ 1 SAD color + λ 2 SAD obj and can be expressed as such.

[0166] Also, considering that the color information of macroblocks at different positions in the actual video image frame is similar, when the macroblock search between the reference frame and the current frame matches, the situation of incorrect matching is likely to occur. By using the consistency energy function, the difference in the underlying texture of similar blocks can be distinguished, the motion information of the macroblocks in the image frame can be accurately tracked, the incorrect vector information can be corrected, and the accuracy of the extraction of the motion vector field can be improved.

[0167] Regarding problems such as the execution of the video motion estimation method taking a long time, having a high computational complexity resulting in an overly long video encoding time, blurring in the obtained interpolated frame, and the existence of block effects, the embodiments of the present application propose an efficient high-speed video motion estimation method based on a kind of 3D image block. The difference of this method is that the preprocessing of video content is applied to motion estimation, the foreground motion information of consecutive frames and the scene structure features of the background are fully utilized, the search range in the search process is effectively restricted, the number of search points is reduced, and thereby the time cost of motion estimation is reduced.

[0168] In the embodiments of the present application, the continuous sequence images of the video to be encoded are regarded as one 3D overall calculation object, the content of the three-dimensional image block composed of consecutive frames is used in the constrained motion vector calculation process to realize high-speed motion estimation, and the accuracy of the vector field is improved. Compared with the motion estimation method based on the macroblocks of the reference frame and the current frame in the related art, this method can effectively reduce the complexity of motion estimation and save 10% - 30% of the motion estimation time on the basis of guaranteeing the encoding rate distortion performance.

[0169] The motion estimation method provided by the embodiments of the present application mainly eliminates the temporal redundancy in video sequence frames through inter-frame prediction, is used for the compression encoding of video data, can improve the video transmission efficiency, can be applied to video conferencing, videophone, etc., and can realize the real-time transmission of high-compression-ratio video data under extremely low bitrate transmission conditions. Moreover, the method is further applicable to 2D video and stereoscopic video, and can still maintain good coding rate distortion performance especially for various complex videos, such as situations where camera shake occurs during shooting, the contrast of image frames is low, and the motion scenes change continuously and complexly.

[0170] Hereinafter, an exemplary structure in which the video motion estimation device 455 provided by the embodiments of the present application is implemented as a software module will be continuously described. In some embodiments, as shown in FIG. 2, the software module in the video motion estimation device 455 stored in the memory 450 may be the video motion estimation device in the terminal 400. A first acquisition module 4551 configured to acquire a plurality of image frames in the video to be processed and perform scene classification processing on the plurality of image frames to obtain a plurality of image frame sets, where each of the image frame sets includes at least one image frame; a feature extraction module 4552 configured to extract the contour features and color features of the foreground object in each image frame in each of the image frame sets; a first determination module 4553 configured to determine a search range corresponding to each of the image frame sets based on the contour features of the foreground object in each of the image frame sets; a second determination module 4554 configured to determine the start search point of each prediction frame in each of the image frame sets; and a motion estimation module 4555 configured to perform motion estimation processing within a search area corresponding to the search range in each of the prediction frames based on the start search point of each of the prediction frames, the target block in the reference frame, and the color features of the foreground object, and obtain a motion vector corresponding to the target block.

[0171] In some embodiments, the first acquisition module is further configured to determine a background image region in the plurality of image frames, determine an image similarity between the plurality of background image regions, perform scene classification processing on the plurality of image frames based on the image similarity between the plurality of background image regions, and obtain a plurality of image frame sets.

[0172] In some embodiments, the feature extraction module is further configured to perform, on any image frame in any of the image frame sets, a process of determining a foreground image region where a foreground object exists in the image frame, a process of using the position information of the foreground image region as the contour feature of the foreground object in the image frame, and a process of performing color extraction processing based on the foreground image region to obtain the color feature of the foreground object in the image frame.

[0173] In some embodiments, the first determination module is further configured to perform, on any of the image frame sets, a process of determining vertex coordinates in each foreground image region based on the position information of each foreground image region in the image frame set, a process of determining a first maximum value and a first minimum value corresponding to a first dimension, and a second maximum value and a second minimum value corresponding to a second dimension from each of the vertex coordinates, and a process of determining a search range corresponding to the image frame set based on the first minimum value, the first maximum value, the second minimum value, and the second maximum value.

[0174] In some embodiments, the second determination module further determines position information of a reference target block in each reference frame in each of the image frame sets, where the reference target block is any one of the target blocks in the reference frame, makes a prediction for the motion vector of each prediction frame via a set prediction mode, obtains the predicted motion vector of each prediction frame, the prediction mode includes at least one of a median prediction mode, an upper block prediction mode, and an origin prediction mode, and is configured to determine a start search point of each prediction frame based on the position information of the reference target block and the predicted motion vector.

[0175] In some embodiments, the motion estimation module further performs a process of determining a first search template corresponding to the prediction frame for any one of the prediction frames, performs a search process within a search area corresponding to the search range in the prediction frame via the first search template with the start search point of the prediction frame as the center, obtains a prediction target block corresponding to the reference target block in the prediction frame, determines a degree of texture difference between the reference target block and the prediction target block, and when the degree of texture difference is less than a difference threshold, determines a motion vector corresponding to the prediction target block based on the position information of the reference target block and the position information of the prediction target block.

[0176] In some embodiments, when the degree of texture difference is greater than or equal to the difference threshold, the motion estimation module further determines the degree of color difference and the degree of texture difference between each prediction block in the prediction frame and the reference target block, where the prediction block is a target block within a search region corresponding to the search range in the prediction frame. Based on the degree of color difference and the degree of texture difference between each prediction block and the reference target block, a prediction target block corresponding to the reference target block is determined from each prediction block. Based on the position information of the reference target block and the position information of the prediction target block, a motion vector corresponding to the prediction target block is determined.

[0177] In some embodiments, the motion estimation module further determines a first motion direction of a foreground object in the prediction frame based on a predicted motion vector of a predicted target block in the prediction frame, and determines a first search template corresponding to the prediction frame based on the first motion direction of the foreground object.

[0178] In some embodiments, the motion estimation module further determines a plurality of first candidate blocks in the search region based on the first search template, determines a matching order of the plurality of first candidate blocks based on the predicted motion vector, performs a matching process on the plurality of first candidate blocks and the reference target block based on the matching order, and when there is a first candidate target block that successfully matches the reference target block among the plurality of first candidate blocks, configures the first candidate target block as the prediction target block corresponding to the reference target block in the prediction frame.

[0179] In some embodiments, when there is no first candidate target block that successfully matches the reference target block among the plurality of first candidate blocks, the motion estimation module further determines a second search template corresponding to the prediction frame based on a second motion direction, where the second motion direction is different from the first motion direction. Centering on the starting search point, a plurality of second candidate blocks within the search region are determined based on the second search template. Based on the predicted motion vector, the matching order of the plurality of second candidate blocks is determined. Based on the matching order, a matching process is performed on the plurality of second candidate blocks and the reference target block. When there is a second candidate target block that successfully matches the reference target block among the plurality of second candidate blocks, the second candidate target block is configured to be determined as the predicted target block.

[0180] It should be noted that the description of the apparatus in the embodiments of the present application is similar to the description of the above method embodiments and has beneficial effects similar to those of the method embodiments. Therefore, detailed descriptions are omitted. For the technical details not disclosed in the embodiments of the present apparatus, reference should be made to the description of the embodiments of the method of the present application for understanding.

[0181] The embodiments of the present application provide a storage medium storing executable instructions. The executable instructions are stored therein, and when the executable instructions are executed by a processor, the processor is caused to execute the method provided by the embodiments of the present application, for example, the method shown in FIG. 4.

[0182] In some embodiments, the memory medium may be a computer-readable memory medium, for example, a ferromagnetic random access memory (FRAM (registered trademark), Ferromagnetic Random Access Memory), a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read Only Memory), a flash memory, a magnetic surface memory, an optical disk, or a memory such as a compact disk read-only memory (CD-ROM, Compact Disk-Read Only Memory), and further, it may be various devices including one of the above memories or any combination thereof.

[0183] In some embodiments, the executable instructions may take the form of a program, software, a software module, a script, or code, and can be created in any form of programming language (including a compiled or interpreted language, or a declarative or procedural language), and it can be deployed in any form, including being deployed as an independent program or as other units suitable for use as a module, component, subroutine, or other unit in a computing environment.

[0184] As an example, the executable instructions may or may not correspond to files in a file system, and may be stored in a part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program under consideration, or stored in a plurality of shared files (for example, files that store one or more modules, subroutines, or code portions). As an example, the executable instructions can be executed on one computing device, or on a plurality of computing devices located in one place, or alternatively, deployed to be executed on a plurality of computing devices distributed in a plurality of places and connected to each other via a communication network.

[0185] The above are only examples of the embodiments of the present application and are not used to limit the protection scope of the present application. Any revisions, equivalent substitutions, improvements, etc. made within the spirit and scope of the present application are all included within the protection scope of the present application.

Description of Reference Signs

[0186] 10 Estimation System 100 Server 200 Network 400 Terminal 410 Processor 420 Network Interface 430 User Interface 431 Output Device 432 Input Device 440 Bus System 450 Memory 451 Operating System 452 Network Communication Module 453 Presence Module 454 Input Processing Module 455 Estimation Device 4551 First Acquisition Module 4552 Feature extraction module 4553 First decision module 4554 Second decision module 4555 Estimation module

Claims

1. A video motion estimation method executed by a video motion estimation device, comprising: acquiring a plurality of image frames in a video to be processed, determining a background image region in the plurality of image frames, and determining an image similarity between the plurality of background image regions; performing scene classification processing on the plurality of image frames based on the image similarity between the plurality of background image regions to obtain a plurality of sets of image frames, wherein each set of image frames includes at least one image frame, and each set of image frames corresponds to the image similarity; extracting contour features and color features of foreground objects in each image frame in each set of image frames; determining vertex coordinates in each foreground image region based on the position information of each foreground image region in the set of image frames; determining a first maximum value and a first minimum value corresponding to a first dimension, and a second maximum value and a second minimum value corresponding to a second dimension from each vertex coordinate; determining a region that is greater than or equal to the first minimum value and less than or equal to the first maximum value in the first dimension, and greater than or equal to the second minimum value and less than or equal to the second maximum value in the second dimension as a search range corresponding to the set of image frames; determining a start search point of each prediction frame in each set of image frames; performing motion estimation processing within a search region corresponding to the search range in each prediction frame based on the start search point of each prediction frame, a target block in a reference frame, and the color features of the foreground object, and obtaining a motion vector corresponding to the target block.

2. The step of extracting contour features and color features of foreground objects in each image frame in each set of image frames includes: determining a foreground image region in which a foreground object exists in one of the at least two image frames based on the image difference and time information of at least two image frames included in each set of image frames; using the position information of the foreground image region as the contour features of the foreground object in the one image frame. performing color extraction processing based on the foreground image region to obtain the color feature of the foreground object in the one image frame, the method according to claim 1, comprising:

3. the step of determining a start search point for each prediction frame in each set of image frames, a step of determining position information of a reference target block in each reference frame in each set of image frames, wherein the reference target block is one target block in the reference frame, performing prediction on the motion vectors of each prediction frame via a set prediction mode to obtain the predicted motion vectors of each prediction frame, wherein the prediction mode includes at least one of a median prediction mode, an upper layer block prediction mode, and an origin prediction mode, the method according to claim 1, comprising: determining a start search point for each prediction frame based on the position information of the reference target block and the predicted motion vector.

4. the step of performing motion estimation processing within a search region corresponding to the search range in each prediction frame based on the start search point of each prediction frame, the target block in the reference frame, and the color feature of the foreground object to obtain a motion vector corresponding to the target block, including the step of performing the following processing for each prediction frame, the processing is a step of determining a first search template corresponding to the prediction frame, performing a search process within a search region corresponding to the search range in the prediction frame via the first search template centered on the start search point of the prediction frame to obtain a predicted target block corresponding to the reference target block in the prediction frame, a step of determining the degree of texture difference between the reference target block and the predicted target block, when the degree of texture difference is less than a difference threshold, determining a motion vector corresponding to the predicted target block based on the position information of the reference target block and the position information of the predicted target block, the method according to claim 3, comprising:

5. the method is When the degree of texture difference is greater than or equal to the difference threshold, a step of determining the degree of color difference and the degree of texture difference between each prediction block in the prediction frame and the reference target block, wherein the prediction block is a target block within a search region corresponding to the search range in the prediction frame, the step; A step of determining a prediction target block corresponding to the reference target block from each prediction block based on the degree of color difference and the degree of texture difference between each prediction block and the reference target block; The method according to claim 4, further comprising: a step of determining a motion vector corresponding to the prediction target block based on the position information of the reference target block and the position information of the prediction target block.

6. The step of determining the first search template corresponding to the prediction frame is: A step of determining a first motion direction of a foreground object in the prediction frame based on a predicted motion vector of a prediction target block of the prediction frame; The method according to claim 4, comprising: a step of determining a first search template corresponding to the prediction frame based on the first motion direction of the foreground object.

7. Centering on the start search point of the prediction frame, the step of performing a search process within a search region corresponding to the search range in the prediction frame through the first search template to obtain a prediction target block corresponding to the reference target block in the prediction frame is: A step of determining a plurality of first candidate blocks within the search region based on the first search template; A step of determining a matching order of the plurality of first candidate blocks based on the predicted motion vector; A step of performing a matching process on the plurality of first candidate blocks and the reference target block based on the matching order; When there is a first candidate target block that has successfully matched the reference target block among the plurality of first candidate blocks, the step of setting the first candidate target block as the prediction target block corresponding to the reference target block in the prediction frame. The method according to claim 6.

8. The method is: When there is no first candidate target block that has successfully matched with the reference target block among the plurality of first candidate blocks, determining a second search template corresponding to the prediction frame based on a second motion direction, where the second motion direction is different from the first motion direction; Determining a plurality of second candidate blocks within the search region based on the second search template, with the starting search point as the center; Determining a matching order of the plurality of second candidate blocks based on the predicted motion vector; Performing a matching process on the plurality of second candidate blocks and the reference target block based on the matching order; When there is a second candidate target block that has successfully matched with the reference target block among the plurality of second candidate blocks, determining the second candidate target block as the prediction target block. The method according to claim 7 further includes this step.

9. A video motion estimation device, comprising: A first acquisition module configured to acquire a plurality of image frames in a video to be processed, determine background image regions in the plurality of image frames, determine an image similarity between the plurality of background image regions, and perform scene classification processing on the plurality of image frames based on the image similarity between the plurality of background image regions to obtain a plurality of image frame sets, where each image frame set includes at least one image frame and each image frame set corresponds to the image similarity; A feature extraction module configured to extract contour features and color features of a foreground object in each image frame in each image frame set; Determining vertex coordinates in each foreground image region based on position information of each foreground image region in the image frame set; Determining a first maximum value and a first minimum value corresponding to a first dimension, and a second maximum value and a second minimum value corresponding to a second dimension from each vertex coordinate; Determining a region that is greater than or equal to the first minimum value and less than or equal to the first maximum value in the first dimension, and greater than or equal to the second minimum value and less than or equal to the second maximum value in the second dimension as a search range corresponding to the image frame set; A first determination module configured as such. A second determination module configured to determine a start search point of each predicted frame in each of the image frame sets; A motion estimation module configured to perform motion estimation processing within a search area corresponding to the search range in each of the predicted frames based on the start search point of each of the predicted frames, a target block in the reference frame, and the color characteristics of the foreground object, and obtain a motion vector corresponding to the target block. A video motion estimation device including: **Claim 10** A video motion estimation device, A memory used for storing executable instructions; A processor used for realizing the video motion estimation method according to any one of claims 1 to 8 when executing the executable instructions stored in the memory. A video motion estimation device including: **Claim 11** A computer program for causing a computer to execute the video motion estimation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Motion compensation device and dynamic image corder and its method

    JP1999243551A

  • Motion vector generation method, picture encoder, motion compensation method, motion compensation device and provision medium

    JP2000023190A

  • Image coder, image coding method, image decoder, image decoding method, medium and image processor

    JP2001086507A

  • System and method for compressing 3d computer graphics

    JP2005523541A

  • Dynamic image encoding device, dynamic image encoding method, dynamic image encoding program, dynamic image decoding device, dynamic image decoding method, and dynamic image decoding program

    JP2007043651A