Video encoder, video encoding method, and video decoder

The video encoder addresses the challenge of motion estimation in high-resolution videos by using a DP module to determine an optimal initial search position and an ME module to perform motion estimation, resulting in efficient and accurate video encoding.

JP2025079322APending Publication Date: 2025-05-21SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024186415
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-10-23
Publication Date
2025-05-21

AI Technical Summary

Technical Problem

Existing video encoding technologies face challenges in efficiently performing motion estimation, particularly in mobile devices with limited search ranges, as the resolution of videos increases, making it difficult to find significant motion.

Method used

A video encoder that includes a DP module for performing a full search within a predetermined region to output an optimal initial search position for motion estimation, and an ME module that moves the search position based on the optimal initial search position to perform motion estimation.

Benefits of technology

This approach allows for fast and accurate motion estimation, even in high-resolution videos, by optimizing the initial search position, thereby improving video encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079322000001_ABST
    Figure 2025079322000001_ABST
Patent Text Reader

Abstract

To provide a video encoder, a video encoding method, and a video decoder.SOLUTION: A video encoder according to an embodiment may include a DP module that receives a pair of video frames as input, performs a full search in a predetermined region, and outputs an optimal initial search position for motion estimation, and an ME module that moves a search position in a direction toward the optimal initial search position output from the DP module to perform motion estimation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a video encoder, a video encoding method using the video encoder, and a video decoder. [Background technology]

[0002] With the development of information and communication technology, the shooting, storage, and sharing of video has become more diverse and active. In particular, the number of videos shot or stored through mobile and portable devices has increased, which has led to the need for image signal processing to process the shot video and eliminate physical deterioration, or codec technology for efficient storage and transmission. Image signal processing or codec estimates the correlation between frames in a sequence, which is an image stream for video processing, to improve the quality of the video, or compresses the correlation to store and transmit it in a small capacity. The correlation between frames is based on motion estimation (ME) between images according to the unit in the video to be processed, for example, patch or block. However, when the maximum search range in a system on chip (SoC) is set as in mobile devices, or the search range is set for software optimization, it is difficult to find a certain amount of motion that increases as the video resolution increases. Summary of the Invention [Problem to be solved by the invention]

[0003] SUMMARY OF THE PRESENT EMBODIMENT An object of the present invention is to provide a video encoder, a video encoding method, and a video decoder. [Means for solving the problem]

[0004] According to one aspect, a video encoder may include a DP (differentiable prediction) module that receives a pair of video frames as input, performs a full search of a predetermined region, and outputs an optimal initial search position for motion estimation; and an ME (motion estimation) module that moves the search position in the direction of the optimal initial search position output from the DP module to perform motion estimation.

[0005] The DP module can generate predicted images for a number of initial search positions within a predetermined region, and output an initial search position that minimizes a residual between the generated predicted image and a first frame of the frame pair as the optimal initial search position.

[0006] The DP module may include an affine transformation module that affinely transforms a second frame of the frame pair for each of a plurality of initial search positions; and a motion estimation motion compensation (MEMC) module that performs motion estimation on the second frame affine transformed for each of the plurality of initial search positions and outputs it in the form of a kernel, and performs motion compensation based on the kernel-type motion to generate a predicted image.

[0007] The MEMC module can perform unfolding on each of the multiple blocks of the affine transformed second frame, divide each block into multiple patches, calculate a Sum of Absolute Difference (SAD) between the patch of each block and the corresponding block of the first frame, and generate a motion kernel through softmax based on the calculated SAD.

[0008] The DP module can process multiple frame pairs of the video in parallel using one or more external processors.

[0009] The one or more processors may include a Graphic Processing Unit (GPU).

[0010] The video encoder may further include a scaler that scales the size of video frame pairs to a size for processing by the DP module.

[0011] In this case, at least one of the number and size of the regions may be preset based on at least one of computing performance, a target processing speed, and accuracy of motion estimation.

[0012] According to one aspect, a video encoder includes an ISPO (initial search position optimization) module including a neural network trained to input a pair of video frames and output an optimal initial search position for motion estimation; and an ME module that performs motion estimation by moving a search position in the direction of the optimal initial search position output from the ISPO module, wherein the neural network is trained to output the optimal initial search position using a GT (Ground Truth) initial search position generated by an external DP module through a full search for a specified region.

[0013] In this case, the neural network may include a convolutional neural network (CNN).

[0014] The neural network is trained to output the optimal initial search positions as affine matrix values.

[0015] A neural network is trained using the output GT initial search position by generating predicted images for multiple initial search positions within the region and determining the initial search position that minimizes the residual between the generated predicted image and the first frame of the frame pair.

[0016] According to one aspect, the video encoding method may include a step of: by a DP module, fully searching a predetermined area using a frame pair of the video as input, and outputting an optimal initial search position for motion estimation; and by an ME module, moving the search position in the direction of the optimal initial search position output from the DP module to perform motion estimation.

[0017] The step of outputting the initial search position may include the steps of: generating predicted images for a plurality of initial search positions within a predetermined region; and outputting an initial search position that minimizes a residual between the generated predicted image and a first frame of the frame pair as the optimal initial search position.

[0018] The step of generating a predicted image may include a step of affine transforming a second frame of the frame pair for each of a plurality of initial search positions; a step of performing motion estimation on the affine transformed second frame for each of the plurality of initial search positions and outputting it in the form of a kernel; and a step of performing motion compensation based on the kernel-type motion to generate a predicted image.

[0019] The step of outputting in the form of a kernel may include unfolding each of the plurality of blocks of the affine transformed second frame, dividing each block into a plurality of patches, calculating the SAD between the patch of each block and the corresponding block of the first frame, and generating a motion kernel through softmax based on the calculated SAD.

[0020] The step of outputting initial search positions may involve parallel processing of multiple frame pairs of the video using one or more external processors.

[0021] The video encoding method may further include scaling a size of a frame pair of the video to a size for processing by the DP module.

[0022] According to one aspect, a video decoder may include: an MC (motion compensation) module that performs motion compensation based on an optimal initial search position obtained by a video encoder by fully searching a frame pair of a video, and a motion vector extracted by performing motion estimation of the frame pair; and a decoding module that decodes video based on the results of the motion compensation.

[0023] According to one aspect, the electronic device includes a position estimation device that inputs a pair of frames of a video and outputs an optimal initial search position; an image processing device that performs motion estimation based on the output optimal initial search position and processes an image based on a result of the motion estimation; and one or more processors that process control and requests of the image processing device. The position estimation device may include a DP module that fully searches a predetermined area and outputs the optimal initial search position for motion estimation; or an ISPO module including a neural network trained to output the optimal initial search position using a GT initial search position generated by the DP module through a full search for the area. [Brief description of the drawings]

[0024] [Figure 1] FIG. 2 is a block diagram of a video encoder according to one embodiment. [Figure 2A] 2 is a diagram illustrating an embodiment of the DP module of FIG. 1. [Figure 2B] 2 is a diagram illustrating an embodiment of the DP module of FIG. 1. [Figure 2C] 2 is a diagram illustrating an embodiment of the DP module of FIG. 1. [Diagram 3] 1 is a diagram for explaining parallel processing of a DP module. [Figure 4] FIG. 2 is a block diagram of a video encoder according to another embodiment. [Diagram 5] FIG. 2 is a block diagram of a video encoder according to another embodiment. [Figure 6] 6 is a diagram illustrating the neural network of the ISPO module of FIG. 5. [Figure 7]FIG. 2 is a block diagram of a video encoder according to another embodiment. [Figure 8] 1 is a flowchart of a video encoding method according to one embodiment. [Figure 9] 4 is a flow chart of a method for outputting an optimal initial search position according to one embodiment. [Figure 10] 4 is a flow chart of a method for outputting an optimal initial search position according to one embodiment. [Figure 11] 4 is a flowchart of a video encoding method according to another embodiment. [Figure 12] FIG. 2 is a block diagram of a video decoder. [Figure 13] FIG. 1 is a block diagram of an electronic device according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Other specific details of the embodiments are included in the detailed description and drawings. The advantages and features of the described technology and the manner in which they are accomplished will become apparent with reference to the following detailed description of the embodiments together with the drawings. The same reference numerals refer to the same components throughout the specification.

[0026] Terms such as "first" and "second" are used to describe various components, but the components are not limited by the terms. Terms are used only to distinguish one component from another. A singular expression includes a plural expression unless otherwise specified in the context. In addition, when a part "includes" a certain component, this does not exclude other components, but means that the part may further include other components, unless otherwise specified. In addition, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, and can be embodied as hardware or software, or a combination of hardware and software.

[0027] Fig. 1 is a block diagram of a video encoder according to an embodiment. Figs. 2A to 2C are diagrams illustrating an embodiment of a DP module in Fig. 1. Fig. 3 is a diagram illustrating parallel processing of the DP module.

[0028] 1, the video encoder 100 may include a DP module 110 and an ME module 120. The video encoder 100 may use processed results from the DP module 110 and the ME module 120 to encode video.

[0029] The DP module 110 may be implemented as a neural network that replicates a standard codec. The DP module 110 may include a module that is implemented to be differentiable so that backpropagation can flow. The DP module 110 may input a pair of frames (first frame, second frame) of a video, perform a full search for an initial search position within a predetermined region, and output an optimal initial search position for motion estimation. The first frame is a frame at the current time (t) and a frame at a time (t-1) before the second frame. In this case, the predetermined region is, for example, the entire region of the input frame, and the size and number of the region are adjusted based on at least one of computing performance, target processing speed, and accuracy of motion estimation as necessary. For example, the size of the region may be adjusted to be small for fast processing, and one or more regions may be sampled from the entire frame region and used as the region for the full search. In addition, the size of the input image itself may be downscaled for fast processing, and the entire downscaled image may be set as the region for the full search.

[0030] FIG. 2A is a block diagram illustrating one embodiment of a DP module 110.

[0031] Referring to FIG. 2A, the DP module 110 may include an affine transformation module 210 and a MEMC module 220 that are implemented in a differentiable manner.

[0032] The affine transformation module 210 receives the second frame (FR21) of the video and an initial search position. The input initial search position is each pixel position within an area set for a full search of the second frame (FR21). The affine transformation module 210 can affine transform the second frame (FR21) with respect to the input initial search position (ISP). The affine transformation module 210 transforms the second frame (FR21) into an image similar to the first frame (FR1) so that the search range does not go out when motion is estimated in block units.

[0033] The MEMC module 220 may include a motion kernel estimation module 221 and a motion compensation module 222. The motion kernel estimation module 221 may divide a first frame (FR1) and a second frame (FR22) after affine transformation into a plurality (N) of blocks (BL1, BL2) of a predetermined size, and perform motion estimation for each block. The first frame (FR1) is divided into blocks of a size determined by the image, and the second frame (FR22) is divided so as to overlap with surrounding blocks to a size including a search range. The motion kernel estimation module 221 may output motion in the form of a kernel, which is to enable a differentiable convolution operation during prediction through motion compensation in the future.

[0034] Figure 2B is an example of a differentiable made motion kernel estimation module 221. Figure 2C is an example of the unfolding 2211 process.

[0035] 2B and 2C, each block (BL2) of the affine transformed second frame (FR22) is unfolded into (2S+1) through an unfolding process 2211. 22211, the image may be divided into 49 patches (PA), where S indicates a search range. The patch (PA) size is the block size (B), which is the same as the block size of the first frame (FR1). For example, if the block size (B) is 4 and the search range (S) is 3, the image may be divided into 49 patches (PA) of size 4 through unfolding 2211.

[0036] Next, the SAD between each patch (PA) of a specific block (BL2) in the affine transformed second frame (FR22) and the block (BL1) in the first frame (FR1) corresponding to that specific block (BL2) is calculated (2212), and the patch position with the least difference is selected using a softmax function to generate a motion kernel (MK) (2213).

[0037] The motion compensation module 222 performs motion compensation and / or convolution using the motion kernel (MK) output from the motion kernel estimation module 221, and can align the motion-compensated block to its original position to generate a predicted image (FR3) for the first frame (FR1).

[0038] The DP module 110 can calculate a residual, which is the difference (FR1-FR3) between the first frame (FR1) and a predicted image (FR3) generated for each initial search position, and output the initial search position that minimizes the residual as the optimal initial search position.

[0039] 3, the DP module 110 can process all frame pairs of the input video 300 in parallel by dividing them using a plurality of processors (XPU1, ..., XPUn). In this case, the processors (XPU1, ..., XPUn) can include a GPU, a central processing unit (CPU), a neural processing unit (NPU), a tensor processing unit (TPU), etc. A single processor (XPU1) can process a plurality of frame pairs 3011, ..., 301t in parallel, or can process a plurality of frame pairs 3011, ..., 301t one by one sequentially. For example, the processor (XPU1) can execute the affine transformation module 210 and the MEMC module 220 of the DP module 110 for the initial search position (ISP(-128, -128) to ISP(127, 127)) of the frame pair 310t to output an optimal initial search position (Opt_ISP) 320t. In this manner, one or more processors can be used to process frame pairs in parallel to quickly and accurately estimate optimal initial search positions (Opt_ISP) 3201, . . . , 320t in a complete search manner.

[0040] 1 again, the ME module 120 receives the first and second frames and the optimal initial search position output from the DP module 110 and can perform motion estimation based on the optimal initial search position. This allows for fast and accurate video encoding for input of images of various sizes. The ME module 120 can perform motion estimation in block units between the first and second frames, and can perform motion estimation by moving the center of a search range toward the optimal initial search position.

[0041] FIG. 4 is a block diagram of a video encoder according to another embodiment.

[0042] Referring to FIG. 4, the video encoder 400 may include a DP module 110, an ME module 120, and a scaler 410.

[0043] The scaler 410 may scale the first and second frames of the input video to a size to be processed by the DP module 110 and then input the frames to the DP module 110. For example, since the resolution of an image captured through a camera is various, such as FHD, 4K, and 8K, the scaler 410 may downscale the frames to a smaller size (e.g., 448×256) so that the optimal initial search position can be estimated quickly with a relatively small amount of calculation in the DP module 110 and input the frames to the DP module 110. However, the present invention is not limited thereto, and the scaler 410 may upscale the frames depending on the size of the input video, or may input the frames to the DP module 110 in the original size without scaling.

[0044] When the image size is up- or down-scaled and input by the scaler 410, the DP module 110 can convert the optimal initial search position to correspond to the size of the original image as necessary and provide it to the ME module 120.

[0045] The ME module 120 may perform motion estimation by moving the search position based on the initial search position. For example, the ME module 120 may perform motion estimation after moving the center of the search range to the initial search position. This makes it possible to effectively handle input of images of various sizes.

[0046] 5 is a block diagram of a video encoder according to another embodiment of the present invention, and FIG 6 is a diagram illustrating a neural network of the ISPO module of FIG 5.

[0047] 5, a video encoder 500 may include an ISPO module 510 and an ME module 520. The video encoder 500 may use processed results from the ISPO module 510 and the ME module 520 to encode video.

[0048] The ISPO module 510 includes a neural network 511 trained to receive video frame pairs and output optimal initial search positions for motion estimation, and can input the input frame pairs to the neural network to output optimal initial search positions, where the neural network is trained to output the optimal initial search positions as affine matrix values.

[0049] The neural network 511 can output information about an optimal initial search position by inputting the first and second frames. The neural network 511 is, for example, a convolutional neural network (CNN). However, the neural network 511 is not limited to this. The neural network 511 is trained in advance through an external training device 600.

[0050] The learning device 600 may include a DP module 610 and one or more processors (e.g., GPU). The DP module 610 may include an affine transformation module and a MEMC module, as described above. As described in FIG. 3, the DP module 610 is executed by one or more processors, and can output an optimal initial search position (Opt_ISP) by performing parallel processing (or sequential processing) on ​​a pair of video frames input therethrough in a full search manner. The optimal initial search position (Opt_ISP) output by the DP module 610 may be generated as a GT initial search position for training a neural network of the ISPO module 510.

[0051] The learning device 600 can train the neural network 511 by a supervised learning method using the generated GT initial search position. Referring to FIG. 6, the loss of a loss function between the GT initial search position (GT ISP) generated through the DP module 610 and the initial search position (ES ISP) output by the neural network 511 can be calculated and training can be performed so that the loss is minimized. However, without being limited thereto, the neural network 511 can be trained by an unsupervised learning method using a back propagation technique using the DP module 610 formed to be differentiable. The loss function can include, but is not limited to, Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), Cross-Entropy Loss, Binary Cross-Entropy Loss, Log Likelihood Loss, Frequency Domain Loss, etc.

[0052] The ME module 520 can perform motion estimation by moving the center of the search range in the direction of the optimal initial search position output from the ISPO module 510 .

[0053] The video encoder 500 can be lightweight because it uses the neural network 511 trained by the learning device 600. As a result, the video encoder 500 can perform fast video encoding by being installed in an electronic device having a relatively low computing performance compared to the learning device 600. However, without being limited thereto, the video encoder 500 can be included in the learning device 600.

[0054] FIG. 7 is a block diagram of a video encoder according to another embodiment.

[0055] Referring to FIG. 7, a video encoder 700 may include an ISPO module 510, an ME module 520, and a scaler 710.

[0056] The scaler 710 may scale the first and second frames of the input video to a size to be processed by the ISPO module 510, and then input the frames to the ISPO module 510. The scaler 710 may not scale, or may up- or down-scale the frames, taking into consideration the size of the input video, the desired processing speed, the desired encoding accuracy, and the like.

[0057] When the size of the image is up- or down-scaled by the scaler 710, the ISPO module 510 can convert the output optimal initial search position to correspond to the original size of the image as necessary and provide it to the ME module 520.

[0058] The ME module 520 can move the center of the search range based on the initial search position to perform motion estimation.

[0059] FIG 8 is a flowchart of a video encoding method according to an embodiment. FIG 9 and FIG 10 are flowcharts of a method for outputting an optimal initial search position according to an embodiment. FIG 8 to FIG 10 are examples of a video encoding method performed by the video encoder 100, 400 of FIG 1 or FIG 4, and will be described briefly below.

[0060] First, a DP module may input a pair of video frames and output an optimal initial search position (810). The DP module may be executed by one or more processors (e.g., GPU) to perform a full search on initial search positions within a predetermined region for a full search and output an optimal initial search position for motion estimation. The size and number of regions for the full search are set in consideration of computing performance, target processing speed, and accuracy of motion estimation, and the size of the input image itself may be downscaled for fast processing as necessary.

[0061] Referring to Fig. 9, an embodiment of the step of outputting an optimal initial search position (810) will be described. For a plurality of initial search positions within a predetermined region, a reference frame from among a pair of video frames is affine-transformed (910), and a motion kernel can be estimated based on the affine-transformed reference frame (920). Referring to Fig. 10, the step of estimating a motion kernel (920) may include dividing each block of the affine-transformed reference frame into a plurality of patches through an unfolding process (1010), calculating the SAD between each patch of a specific block of the affine-transformed reference frame and a block of the current frame corresponding to the specific block (1020), and selecting a patch position with the smallest difference using a softmax function (1030) to output a motion kernel (MK) (1040).

[0062] 9 again, a predicted image for the current frame is generated by performing motion compensation using a motion kernel (930), and an initial search position at which the residual between the generated predicted image and the current frame of the video frame pair is smallest can be output as an optimal initial search position (940). A predicted image for the current frame can be generated by performing motion compensation and / or convolution using the motion kernel output from step (920) and matching the motion-compensated block to its original position. Residuals between the generated predicted images and the current frame for all initial search positions can be calculated, and the initial search position at which the residual is smallest can be output as an optimal initial search position.

[0063] 8, the ME module can perform motion estimation (820) based on the frame pair and the optimal initial search position output from step 810. The center of the search range can be moved toward the optimal initial search position to perform block-based motion estimation for the frame pair.

[0064] 11 is a flow chart of a video encoding method according to another embodiment of the present invention, which will be briefly described as an embodiment of a video encoding method performed by the video encoding apparatuses 500 and 700 of FIGS.

[0065] A neural network of the ISPO module can be trained based on the GT initial search positions generated by the external DP module (1110). The neural network is a convolutional neural network (CNN). The DP module can be executed by one or more processors and can process pairs of frames of the video in a parallel or sequential manner in a full search manner to output optimal initial search positions. The optimal initial search positions output by the DP module can be generated as GT initial search positions used to train the neural network of the ISPO module. The neural network is trained in a supervised learning manner using the generated GT initial search positions.

[0066] The ISPO module can then input the video frame pairs to a neural network trained through the DP module to output optimal initial search positions (1120), with the size of the video frame pairs scaled as necessary.

[0067] The ME module can then move the center of the search range toward the optimal initial search position output from the ISPO module to perform motion estimation (1130).

[0068] FIG. 12 is a block diagram of a video decoder according to one embodiment.

[0069] Referring to FIG. 12, a video decoder 1200 may include an MC module 1210 and a decoding module 1220.

[0070] The MC module can perform motion compensation based on the motion vector extracted by the ME module of the video encoding device. As described above, the DP module or ISPO module of the video encoder performs a full search on a pair of video frames to obtain an optimal initial search position, and the ME module of the video encoder performs motion estimation based on the optimal initial search position to obtain a motion vector.

[0071] The decoding module 1220 may decode moving images based on the motion compensation result of the MC module 1210. The decoding module 1220 may perform various moving image decoding processes, such as decoding frames compressed using motion compensation and reconstructing the frames in a format that can be displayed on a screen by a video compression codec.

[0072] FIG. 13 is a block diagram of an electronic device according to one embodiment.

[0073] The electronic device may include various embodiments of the video encoder described above. The electronic device may include a device that requires an application to convert a low-resolution image with a low frame rate into a high-resolution image with a high frame rate in an environment with limited computing resources, such as an edge device, a TV, a monitor, an Internet of Things (IoT) device, a radar device, a smartphone, a wearable device, a tablet computer, a netbook, a laptop, a desktop, a head mounted display (HMD), an autonomous vehicle and a smart vehicle, a virtual reality (VR), an augmented reality (AR), an eXtended Reality (XR) device, an automobile, a mobile robot, and various devices that transmit / receive images, a cloud computing device, and the like.

[0074] 13, an electronic device 1300 may include a location estimation device 1310, a video processing device 1320, a processor 1330, a storage device 1340, an output device 1350, and a communication device 1360. The location estimation device 1310 is included within the video processing device 1320.

[0075] The position estimation device 1310 may receive a video frame as an input and output an optimal initial search position for motion estimation. The position estimation device 1310 may include a DP module as described above and may output an initial search position using the DP module. Alternatively, the position estimation device 1310 may include an ISPO module including a neural network trained by the DP module as described above. In this case, the DP module may be included in another device inside or outside the electronic device 1300.

[0076] The image processing device 1320 may include a video codec device that performs encoding and / or decoding of a video. The video codec may include the above-mentioned video encoder and / or video decoder. The image processing device 1320 may also include a device that processes physical degradation of an image, such as a video stabilizer, noise reduction (NR), high dynamic range (HDR), de-blur, frame rate up conversion (FRUC), etc.

[0077] The processor 1330 may include one or more main processors such as a CPU and an application processor, an IP core (Intellectual Property Core), and auxiliary processors that can operate independently or together with the main processors, such as a GPU, an image signal processor, a sensor hub processor, a communication processor, etc. The processor 1330 may process control and requests of the configuration of the electronic device 1300. For example, one or more GPUs may support parallel processing of a DP module in response to a request from the position estimation device 1310.

[0078] The storage device 1340 may store data required for the operation of the components of the electronic device 1300 (e.g., images (video and still images) captured by the imaging device, data processed by the processor 1330, neural networks used in the position estimation device 1310 and the image processing device 1320, etc.) and instructions for executing functions. The storage device 1340 may include random access memories (RAM), dynamic random access memories (DRAM), static random access memories (SRAM), a magnetic hard disk, an optical disk, a flash memory, an electrically programmable read-only memory (EPROM), or other forms of computer-readable recording media known in the art.

[0079] The output device 1350 may output in a visual / non-visual manner the image captured by the image capture device, the position estimation device 1310, the image processing device 1320, and the data generated or processed by the processor 1330. The output device 1250 may include an audio output device, a display device, an audio module, and / or a haptic module.

[0080] The communication device 1360 may support the establishment of a direct (wired) communication channel and / or a wireless communication channel between the electronic device and other electronic devices, servers, or sensor devices in a network environment using various communication technologies, and the performance of communication through the established communication channel. The communication device 1360 may transmit images captured by the imaging device, data generated or processed by the location estimation device 1310, the image processing device 1320, and the processor 1330 to other electronic devices. The communication device 1360 may also receive images to be processed from the cloud or other electronic devices, store the received images in the storage device 1340, and transmit the images to the processor 1330 so that the images are processed by the processor 1330.

[0081] In addition, the electronic device 1300 may further include a sensor device (e.g., an acceleration sensor, a gyroscope, a magnetic field sensor, a proximity sensor, an illuminance sensor, a fingerprint sensor, a GPS sensor, etc.) that detects various data, a photographing device (camera) that captures images, and an input device (e.g., a microphone, a mouse, a keyboard, and / or a digital pen (such as a stylus pen)) that receives commands and / or data from a user, etc.

[0082] Meanwhile, the present invention may be embodied as computer readable codes on a computer readable recording medium, which may include any type of recording device in which data that can be read by a computer system is stored.

[0083] Examples of the computer readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc., and also include those embodied in the form of carrier waves (e.g., transmission via the Internet). The computer readable recording medium may be distributed in computer systems connected through a network, and may be stored and executed as computer readable code in a distributed manner. Functional programs, codes, and code segments for implementing the present embodiment may be easily construed by a programmer skilled in the art.

[0084] It will be understood by those skilled in the art that the present invention can be embodied in other specific forms without changing the technical ideas and essential features disclosed herein. Therefore, it should be understood that the above-described embodiments are illustrative in all respects and are not limiting. [Explanation of symbols]

[0085] 100 Video Encoders 110 DP module 120 ME Module

Claims

1. a DP module that receives a pair of video frames as input, performs a full search of a predetermined region, and outputs an optimal initial search position for motion estimation; an ME module for performing motion estimation by moving a search position in the direction of the optimal initial search position output from the DP module; A video encoder, including:

2. The DP module includes:

2. The video encoder of claim 1, further comprising: a prediction image generating unit configured to generate a prediction image for each of a plurality of initial search positions within the predetermined region; and an initial search position that minimizes a residual between the generated prediction image and a first frame of the pair of frames is output as the optimal initial search position.

3. The DP module includes: an affine transformation module for affine transforming a second frame of the pair of frames to each of the plurality of initial search positions; a MEMC module for performing motion estimation on a second frame affine-transformed with respect to each of the initial search positions, outputting the motion in a kernel form, and performing motion compensation based on the kernel form to generate a predicted image; The video encoder of claim 2 , comprising:

4. The MEMC module includes: The video encoder of claim 3, further comprising: unfolding each of the affine transformed blocks of the second frame; dividing each block into a plurality of patches; calculating an SAD between each patch of the block and a corresponding block of the first frame; and generating a motion kernel through softmax based on the calculated SAD.

5. The DP module includes: The video encoder of claim 1 , further comprising: one or more external processors for processing multiple frame pairs of the video in parallel.

6. The one or more processors: The video encoder of claim 5 including a GPU.

7. The video encoder of claim 1 , further comprising a scaler that scales a size of a pair of frames of the video to a size for processing by the DP module.

8. The video encoder of claim 1 , wherein at least one of the number and size of the regions is preset based on at least one of computing performance, target processing speed, and accuracy of motion estimation.

9. an ISPO module including a neural network trained to take video frame pairs as input and output optimal initial search positions for motion estimation; an ME module for moving a search position toward the optimal initial search position output from the ISPO module and performing motion estimation; The neural network comprises: A video encoder that is trained to use GT initial search positions generated through a full search for a given region by an external DP module and output the optimal initial search positions.

10. The neural network comprises: The video encoder of claim 9 , comprising a convolutional neural network (CNN).

11. The neural network comprises: The video encoder of claim 9 , trained to output the optimal initial search positions in affine matrix values.

12. The neural network comprises:

10. The video encoder of claim 9, further comprising: generating predicted images for a plurality of initial search positions within the region; and determining an initial search position that minimizes a residual between the generated predicted image and a first frame of the pair of frames, the initial search position being used for learning.

13. a step of performing a full search of a predetermined region using a pair of frames of a moving image as input by a DP module, and outputting an optimal initial search position for motion estimation; moving a search position toward the optimal initial search position output from the DP module by an ME module and performing motion estimation; A video encoding method, including:

14. The step of outputting the optimal initial search position comprises: generating predicted images for a plurality of initial search positions within the predetermined region; outputting an initial search position at which a residual between the generated predicted image and a first frame of the pair of frames is minimized as the optimal initial search position; The video encoding method of claim 13, comprising:

15. The step of generating a predicted image comprises: affinely transforming a second frame of the pair of frames to each of the plurality of initial search locations; performing motion estimation on a second frame affine-transformed with respect to each of the initial search positions, and outputting the motion estimation result in a kernel form; generating a predicted image by performing motion compensation based on the kernel type motion; 15. The video encoding method of claim 14, comprising:

16. The step of outputting in the form of a kernel includes:

16. The video encoding method of claim 15, further comprising: unfolding each of the affine transformed blocks of the second frame; dividing each block into a plurality of patches; calculating a SAD between each patch of the block and a corresponding block of the first frame; and generating a motion kernel through softmax based on the calculated SAD.

17. The step of outputting the optimal initial search position comprises:

14. The method of claim 13, further comprising processing multiple frame pairs of the video in parallel using one or more external processors.

18. The video encoding method of claim 13 , further comprising: scaling a size of the frame pairs of the video to a size for processing by the DP module.

19. an MC module that performs motion compensation based on a motion vector extracted by performing motion estimation on a pair of frames of a video based on an optimal initial search position obtained by a video encoder performing a full search on the pair of frames of the video; a decoding module for decoding moving images based on a result of the motion compensation; a video decoder.

20. a position estimation device that receives a pair of video frames as input and outputs an optimal initial search position; a video processing device that performs motion estimation based on the output optimal initial search position and processes a video based on a result of the motion estimation; one or more processors for processing control and requests of the image processing device; The position estimation device includes: a DP module for fully searching a given region and outputting the optimal initial search position for motion estimation; or and an ISPO module including a neural network trained to output the optimal initial search positions using GT initial search positions generated by the DP module through a full search for the region.