Video encoder, video encoding method, and video decoder
By combining differentiable prediction module and motion estimation module in video encoding, the neural network is used to optimize the initial search position, and the problem of low motion estimation efficiency when video resolution increases in the prior art is solved, achieving more efficient video processing.
Patent Information
- Application Number
- CN202410498228.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-09
- Filing Date
- 2024-04-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively find motion beyond a certain level in video encoding, especially when the video resolution increases, resulting in inefficient video processing.
Using a combination scheme of differentiable prediction (DP) module and motion estimation (ME) module, the DP module outputs the optimal initial search position in the predetermined area through full search, and the ME module performs motion estimation based on this position. At the same time, the Initial Search Position Optimization (ISPO) module was introduced, and neural network training was used to output the optimal initial search position.
Improve the accuracy and efficiency of motion estimation in video encoding, and can quickly find the optimal initial search position in videos of different resolutions, thereby improving the performance of video processing.
Smart Images

Figure CN119996708A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from Korean Patent Application No. 10-2023-0154594 filed on November 9, 2023 in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference for all purposes. Technical Field
[0003] The following description relates to a video encoder, a video encoding method using the video encoder, and a video decoder. Background Art
[0004] With the development of information and communication technology, videos are captured, stored and shared in active and diverse ways. In particular, more and more videos are captured and stored using mobile devices and portable devices, and image signal processing (ISP) is required to process the captured videos to address physical degradation, or codec technology is required to efficiently store and transmit videos. In ISP or codec technology, video processing is performed by estimating the correlation between frames in a sequence as an image stream to improve video quality, or the correlation is compressed so that low-volume videos can be stored and transmitted. The correlation between frames is based on motion estimation (ME) between images in the video unit to be processed (such as a patch or block). However, if the maximum search range is set within the system on chip (SoC) as in a mobile device, or if the search range is set for software optimization, it may not be possible to find motion above a certain level (which increases as the video resolution increases). Summary of the invention
[0005] According to an aspect of an example embodiment, a video encoder is provided, comprising: a differentiable prediction (DP) module configured to output an optimal initial search position by performing a full search in a predetermined area using a pair of frames of a video as input; and a motion estimation (ME) module configured to perform motion estimation by moving the search position toward the optimal initial search position output by the DP module.
[0006] The DP module can be configured to generate a predicted image for each of multiple initial search positions in a predetermined area, and output the following initial search position as an optimal initial search position: at this initial search position, the residual between the generated predicted image and the first frame of the pair of frames is minimized.
[0007] The DP module may include: an affine transformation module configured to perform an affine transformation on the second frame of the pair of frames for each of a plurality of initial search positions; and a motion estimation motion compensation (MEMC) module configured to perform motion estimation on the affine transformed second frame for each of a plurality of initial search positions, output motion in the form of a kernel, and perform motion compensation based on the motion in the form of the kernel to generate a predicted image.
[0008] The MEMC module can be configured to: for each block of the multiple blocks of the second frame that have been affine transformed, expand the block to divide the block into multiple patches, and calculate the sum of absolute differences (SAD) between the multiple patches of the block and the corresponding block of the first frame; and generate a motion in the form of a kernel based on the calculated SAD of the multiple blocks by using softmax.
[0009] The DP module may be configured to process pairs of frames of a video in parallel using one or more processors.
[0010] The one or more processors may include a graphics processing unit (GPU).
[0011] The video encoder may further include: a scaler configured to scale the pair of frames of the video to a size to be processed by the DP module.
[0012] At least one of the number and size of the predetermined areas is preset based on at least one of computing capability, target processing speed, and accuracy of motion estimation.
[0013] According to an aspect of an example embodiment, a video encoder is provided, comprising: an initial search position optimization (ISPO) module, comprising a neural network trained to output an optimal initial search position by using a pair of frames of a video as input; and a motion estimation (ME) module configured to perform motion estimation by moving a search position toward the optimal initial search position output by the ISPO module, wherein the neural network is trained to output the optimal initial search position by using a reference truth (GT) initial search position generated by performing a full search in a predetermined area by a differentiable prediction (DP) module.
[0014] The neural network may include a convolutional neural network (CNN).
[0015] The neural network can be trained to output the optimal initial search position as an affine matrix value.
[0016] A neural network may be trained using the GT initial search position, wherein the GT initial search position is output by generating predicted images for a plurality of initial search positions in a predetermined area and determining an initial search position where a residual between the predicted image and a first frame in a pair of frames is minimized.
[0017] According to an aspect of an exemplary embodiment, a video encoding method is provided, comprising: outputting, by a differentiable prediction (DP) module, an optimal initial search position by performing a full search in a predetermined area using a pair of frames of a video as input; and performing motion estimation, by a motion estimation (ME) module, by moving the search position toward the optimal initial search position output by the DP module.
[0018] Outputting the optimal initial search position may include generating predicted images for multiple initial search positions in a predetermined area, and outputting the following initial search position as the optimal initial search position: at which the residual between the generated predicted image and the first frame of the pair of frames is minimized.
[0019] Generating a predicted image may include: performing an affine transformation on the second frame of the pair of frames for each of the multiple initial search positions; performing motion estimation on the affine transformed second frame for each of the multiple initial search positions to output motion in kernel form; and performing motion compensation based on the kernel form of motion to generate a predicted image.
[0020] Outputting the motion in kernel form may include: for each block of a plurality of blocks of the second frame that has been affine transformed, expanding the block to divide the block into a plurality of patches, and calculating the sum of absolute differences (SAD) between the plurality of patches of the block and a corresponding block of the first frame; and generating the motion in kernel form based on the calculated SAD of the plurality of blocks by using softmax.
[0021] Outputting the optimal initial search position may include performing parallel processing on multiple pairs of frames of the video using one or more processors.
[0022] The video encoding method also includes scaling a pair of frames of the video to a size to be processed by the DP module.
[0023] According to an aspect of an example embodiment, a video decoder is provided, comprising: a motion compensation (MC) module configured to perform motion compensation based on a motion vector, wherein the motion vector is extracted by motion estimating a pair of frames of a video based on an optimal initial search position, wherein the optimal initial search position is obtained by a video encoder performing a full search on the pair of frames of the video; and a decoding module configured to decode the video based on a result of the motion compensation.
[0024] According to an aspect of an example embodiment, an electronic device is provided, comprising: a position estimation device configured to output an optimal initial search position by using a pair of frames of a video as input; an image processing device configured to perform motion estimation based on the output optimal initial search position, and configured to perform image processing based on a result of the motion estimation; and one or more processors configured to control the image processing device and process a request thereof, wherein the position estimation device comprises: a differentiable prediction (DP) module configured to output an optimal initial search position by performing a full search in a predetermined area; or an initial search position optimization (ISPO) module comprising a neural network trained to output an optimal initial search position using a reference truth (GT) initial search position generated by the DP module performing a full search in a predetermined area. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Exemplary embodiments will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.
[0026] Figure 1 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0027] FIG. 2A to FIG. 2C is the explanation Figure 1 Schematic diagram of an example of a differentiable prediction (DP) module.
[0028] Figure 3 is a schematic diagram for explaining parallel processing of DP modules according to example embodiments.
[0029] Figure 4 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0030] Figure 5 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0031] Figure 6 is the explanation Figure 5 Schematic diagram of the neural network of the Initial Search Position Optimization (ISPO) module.
[0032] Figure 7 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0033] Figure 8 is a flowchart illustrating a video encoding method according to an embodiment of the present disclosure.
[0034] Fig. 9 and Fig.10 is a flowchart illustrating a method of outputting an optimal initial search position according to an exemplary embodiment of the present disclosure.
[0035] Fig.11The flowchart of the video encoding method is shown according to an embodiment of the present disclosure.
[0036] Fig.12 is a block diagram of a video decoder according to an example embodiment of the present disclosure.
[0037] Fig.13 is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] Details of the example embodiments are included in the following detailed description and drawings. Advantages and features of the present disclosure and methods of implementing the same will be more clearly understood according to the embodiments described in detail below with reference to the drawings. Throughout the drawings and detailed description, unless otherwise described, the same reference numerals will be understood to refer to the same elements, features and structures.
[0039] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish elements from each other. Unless otherwise explicitly stated, any reference to a singular number may include a plural number. In addition, unless explicitly described to the contrary, expressions such as "including" or "comprising" will be understood to mean including the elements described, but not excluding any other elements. In addition, terms such as "unit" or "module" should be understood as units that perform at least one function or operation and can be embodied as hardware, software, or a combination thereof.
[0040] Figure 1 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure. FIG. 2A to FIG. 2C is the explanation Figure 1 Schematic diagram of an example of a differentiable prediction (DP) module. Figure 3 is a schematic diagram explaining parallel processing of DP modules.
[0041] Reference Figure 1 , the video encoder 100 includes a differentiable prediction (DP) module 110 and a motion estimation (ME) module 120. The video encoder 100 may encode a video by using processing results of the DP module 110 and the ME module 120.
[0042] The DP module 110 may be implemented with a neural network that mimics a standard codec. The DP module 110 may include a differentiable module to back propagate the flow. The DP module 110 may perform a full search for an initial search position in a predetermined area using a pair of frames (e.g., a first frame and a second frame) of a video as input to output an optimal initial search position for motion estimation. The first frame may be a frame at the current time t, and the second frame may be a frame at the previous time (t-1). In this case, the predetermined area may be, for example, the entire input frame area, or the size and number of the area may be adjusted as needed based on at least one of the computing power, the target processing speed, and the accuracy of the motion estimation. For example, for fast processing, the area may be reduced, and one or more areas in all frame areas may be sampled to serve as the area for full search. In addition, the size of the input video itself may be reduced for fast processing, and the entire reduced video may be set as the area for full search.
[0043] Figure 2A is a block diagram illustrating an example of the DP module 110 .
[0044] Reference Figure 2A , the DP module 110 may include an affine transformation module 210 implemented as a differentiable module, and a motion estimation motion compensation (MEMC) module 220.
[0045] The affine transformation module 210 receives the second frame FR21 of the video and its initial search position. The input initial search position may be each pixel position in an area set for a full search of the second frame FR21. The affine transformation module 210 may perform an affine transformation on the second frame FR21 for the input initial search position ISP. Through the affine transformation module 210, the second frame FR21 may be converted into an image similar to the first frame FR1, so that motion estimation of each block may be performed within the search range.
[0046] The MEMC module 220 may include a motion kernel estimation module 221 and a motion compensation module 222. The motion kernel estimation module 221 may divide each of the first frame FR1 and the affine transformed second frame FR22 into a plurality of blocks BL1 of a predetermined size and a plurality of blocks BL2 of a predetermined size, and may perform motion estimation for each block. In the first frame FR1, each block may be divided into an image of a fixed size, and in the second frame FR22, each block may be divided into a size covering a search range to overlap with surrounding blocks. The motion kernel estimation module 221 may output motion in the form of a kernel, which will be used for subsequent differentiable convolution operations during prediction using motion compensation.
[0047] Figure 2Bis a diagram showing an example of the motion kernel estimation module 221 implemented as a differentiable module, and Figure 2C is a diagram illustrating an example of an expansion process 2211 of the motion kernel estimation module 221 .
[0048] Reference Figure 2B and Figure 2C , each block BL2 of the affine transformed second frame FR22 can be divided into (2S+1) by the expansion process 2211 2 Here, S represents the search range. The size of the patch PA is the block size B and is equal to the block size of the first frame FR1. For example, if the block size B is 4 and the search range S is 3, the block can be divided into 49 patches PA with a size of 4 through the expansion process 2211.
[0049] Then, the sum of absolute differences (SAD) (2212) between each patch PA of the specific block BL2 of the second frame FR2 after affine transformation and the block BL1 of the first frame FR1 corresponding to the specific block BL2 is calculated, and the motion kernel (MK) (2213) can be generated by selecting the patch position with the minimum SAD using a softmax function.
[0050] Return to reference Figure 2A The motion compensation module 222 may perform motion compensation and / or convolution by using the motion kernel MK output by the motion kernel estimation module 221, and may generate a predicted image FR3 of the first frame FR1 by merging the motion compensation blocks at their original positions.
[0051] The DP module 110 may obtain a residual which is a difference (FR1-FR3) between the first frame FR1 and the predicted image FR3 generated at each initial search position, and may output an initial search position in which the residual is minimized as an optimal initial search position.
[0052] Reference Figure 3, the DP module 110 may use a plurality of processors XPU1, ..., and XPUn to perform parallel processing by segmenting all frame pairs of the input video 300. In this case, the processors XPU1, ..., XPUn may include a graphics processing unit (GPU), a central processing unit (CPU), a neural processing unit (NPU), a tensor processing unit (TPU), etc. In one processor XPU1, a plurality of frame pairs 3011, ..., 301t may be processed in parallel, or a plurality of frame pairs 3011, ..., 301t may be processed sequentially one by one. For example, the processor XPU1 may output an optimal initial search position Opt_ISP 320t by executing the affine transformation module 210 and the MEMC module 220 for the initial search position ISP (-128, -128) to ISP (127, 127) of the frame pair 310t. In this way, by processing frame pairs in parallel using one or more processors, the optimal initial search position Opt_ISP 3201, ..., 320t may be quickly and accurately estimated by a full search method.
[0053] Return to reference Figure 1 , the ME module 120 may receive the first frame and the second frame and the optimal initial search position output by the DP module 110, and may perform motion estimation based on the optimal initial search position. In this way, video encoding may be performed quickly and accurately on input videos of various sizes. The ME module 120 may perform motion estimation in units of blocks between the first frame and the second frame. In this case, the ME module 120 may perform motion estimation by moving the center of the search range toward the optimal initial search position.
[0054] Figure 4 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0055] Reference Figure 4 , the video encoder 400 includes a DP module 110 , an ME module 120 , and a scaler 410 .
[0056] The scaler 410 may scale the first and second frames of the input video to a size to be processed by the DP module 110, and then the scaled frames may be input to the DP module 110. For example, since the video captured by the camera may have various resolutions such as full high definition (FHD), 4K, 8K, etc., the video may be reduced to a small size (e.g., 448×256) so that the DP module 110 can quickly and accurately estimate the initial search position with a relatively small amount of calculation. However, the size of the input video is not limited thereto, and the scaler 410 may enlarge or reduce the input video according to the size of the video, or may input the video to the DP module 110 at its original size without scaling.
[0057] If a video enlarged or reduced by the scaler 410 is input, the DP module 110 may convert the optimal initial search position into a size matching the original size of the video as needed, and may provide the converted optimal initial position to the ME module 120 .
[0058] The ME module 120 may perform motion estimation by moving the search position based on the initial search position. For example, the ME module 120 may perform motion estimation after moving the center of the search range toward the initial search position. In this way, operations may be efficiently performed in response to input of videos of various sizes.
[0059] Figure 5 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure. Figure 6 is the explanation Figure 5 Schematic diagram of the neural network of the Initial Search Position Optimization (ISPO) module.
[0060] Reference Figure 5 , the video encoder 500 includes an initial search position optimization (ISPO) module 510 and an ME module 520. The video encoder 500 may encode a video by using processing results of the ISPO module 510 and the ME module 520.
[0061] The ISPO module 510 may include a neural network 511, which is trained to output an optimal initial search position for motion estimation by using a frame pair of a video as an input, and the ISPO module 510 may output the optimal initial search position by inputting the input frame pair to the neural network. In an embodiment, the neural network is trained to output the optimal initial search position as an affine matrix value.
[0062] The neural network 511 can output information about the optimal initial search position by using the first frame and the second frame as input. The neural network 511 can be, for example, a convolutional neural network (CNN), but is not limited thereto. The neural network 511 can be pre-trained by an external training device 600.
[0063] The training device 600 may include a DP module 610 and one or more processors (eg, GPUs). As described above, the DP module 610 may include an affine transformation module and a MEMC module. Figure 3 As shown, the DP module 610 can be executed by one or more processors to perform parallel processing (or sequential processing) on the input frame pairs of the video through a full search method to output the optimal initial search position Opt_ISP. The optimal initial search position Opt_ISP output by the DP module 610 can be generated as a reference truth (GT) initial search position for training the neural network of the ISPO module 510.
[0064] The training device 600 can use the generated GT initial search position to train the neural network 511 through supervised learning. Figure 6 , the training device 600 can calculate the loss of the loss function between the GT initial search position GT ISP generated by the DP module 610 and the initial search position ES ISP output by the neural network 511, and can train the neural network 511 to minimize the loss. However, the neural network is not limited to this, and the neural network can be trained by unsupervised learning and back propagation methods using the DP module 610 implemented as a differentiable module. The loss function includes peak signal-to-noise ratio (PSNR), mean square error (MSE), cross entropy loss, binary cross entropy loss, log likelihood loss, frequency domain loss, etc., but is not limited to this.
[0065] The ME module 520 may perform motion estimation by moving the center of the search range toward the optimal initial search position output by the ISPO module 510 .
[0066] By using the neural network 511 trained by the training device 600, the video encoder 500 can be lightweight. Therefore, the video encoder 500 can be installed in an electronic device having relatively low computing power compared with the training device 600, and thus can quickly perform video encoding. However, the video encoder 500 is not limited thereto and may be included in the training device 600.
[0067] Figure 7 is a block diagram illustrating a video encoder according to an exemplary embodiment of the present disclosure.
[0068] Reference Figure 7 , the video encoder 700 includes an ISPO module 510, an ME module 520 and a scaler 710.
[0069] The scaler 710 may scale the first frame and the second frame of the input video to a size to be processed by the ISPO module 510, and may input the scaled frames to the ISPO module 510. The scaler 710 may enlarge or reduce the frames or may not perform scaling by considering the size of the input video, the desired processing speed, the desired encoding accuracy, etc.
[0070] If the video is enlarged or reduced by the scaler 710, the ISPO module 510 may convert the output optimal initial search position into a size matching the original size of the video as needed, and may provide the converted optimal initial position to the ME module 520.
[0071] The ME module 520 may perform motion estimation by moving the center of the search range based on the initial search position.
[0072] Figure 8 is a flowchart illustrating a video encoding method according to an embodiment of the present disclosure. Fig. 9 and Fig.10 is a flowchart illustrating a method of outputting an optimal initial search position according to an exemplary embodiment of the present disclosure. Figures 8 to 10 Is Figure 1 The video encoder 100 or Figure 4 This is an example of a video encoding method performed by the video encoder 400, and thus will be briefly described below.
[0073] First, at 810, the DP module may output an optimal initial search position by using a frame pair of a video. The DP module may be executed by one or more processors (e.g., a GPU) to perform a full search for an initial search position in a search area set for a full search to output an optimal initial search position for motion estimation. The size and number of areas for full search may be set in consideration of computing power, target processing speed, accuracy of motion estimation, etc., and in an embodiment, the size of the input video itself may be reduced for fast processing.
[0074] Reference Fig. 9 In the example of operation 810 of outputting the optimal initial search position, an affine transformation may be performed on a reference frame among the frame pairs of the video for a plurality of initial search positions in a predetermined area in 910, and motion kernel estimation may be performed based on the affine transformed reference frame in 920. Fig.10 , the motion kernel estimation in 920 may include: in 1010, performing expansion for each block of the reference frame after affine transformation to divide each block of the reference frame into multiple patches; in 1020, calculating the sum of absolute differences (SAD) between each patch of a specific block of the reference frame after affine transformation and the block BL1 of the first frame FR1 corresponding to the specific block; and outputting the motion kernel (MK) in 1040 by selecting the patch position with the minimum SAD using a softmax function in 1030.
[0075] Return to reference Fig. 9 , in 930, motion compensation may be performed by using a motion kernel to generate a predicted image of the current frame, and in 940, an initial search position as an optimal initial search position is outputted at which a residual between the predicted image and the current frame in the frame pair of the video is minimized. Motion compensation and / or convolution may be performed by using the motion kernel outputted in 920, and a predicted image of the current frame may be generated by merging motion compensation blocks at their original positions. A residual between the current frame and the predicted image generated at each of all initial search positions may be obtained, and the initial search position where the residual is minimized may be outputted as the optimal initial search position.
[0076] Return to reference Figure 8 The ME module may perform motion estimation for the frame pair in 820 based on the optimal initial search position output in 810. The ME module may perform motion estimation for each block of the frame pair by moving the center of the search range toward the optimal initial search position.
[0077] Fig.11 is a flowchart illustrating a video encoding method according to an embodiment of the present disclosure. Fig.11 Is Figure 5 The video encoder 500 or Figure 7 This is an example of a video encoding method performed by the video encoder 700, and thus will be briefly described below.
[0078] In 1110, a neural network of an ISPO module may be trained based on a GT initial search position generated by an external DP module. The neural network may be a convolutional neural network (CNN). The DP module may be executed by one or more processors to perform processing (e.g., in parallel or sequentially) on a frame pair of a video by a full search method to output an optimal initial search position. The optimal initial search position output by the DP module may be generated as a reference truth (GT) initial search position for training the neural network of the ISPO module. The generated GT initial search position may be used to train the neural network by supervised learning.
[0079] Then, in 1120, the ISPO module can output the optimal initial search position by inputting the frame pair of the video to the neural network trained by the DP module. In this case, the size of the frame pair of the video can be scaled as needed.
[0080] Subsequently, in 1130, the ME module may perform motion estimation by moving the center of the search range toward the optimal initial search position output by the ISPO module.
[0081] Fig.12 is a block diagram of a video decoder according to an example embodiment of the present disclosure.
[0082] Reference Fig.12 , the video decoder 1200 includes a motion compensation (MC) module 1210 and a decoding module 1220.
[0083] The MC module 1210 may perform motion compensation based on the motion vector extracted by the ME module. As described above, the DP module or ISPO module of the video encoder performs a full search for a frame pair of the video to obtain an optimal initial search position, and the ME module of the video encoder performs motion estimation based on the optimal initial search position to obtain a motion vector.
[0084] The decoding module 1220 may decode the video based on the result of the motion compensation performed by the MC module 1210. The decoding module 1220 may perform various video decoding operations, such as decoding compressed frames by using motion compensation and reconstructing frames into a format suitable for display on a screen according to a video compression codec.
[0085] Fig.13 is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure.
[0086] The electronic device 1300 according to an example embodiment includes one or more of the various examples of the above-mentioned video encoder. The electronic device 1300 may include, for example, the following devices: edge devices, which need to convert low-resolution and low-frame-rate videos into high-resolution and high-frame-rate videos in an environment with limited computing resources; various image transmission / reception devices, such as televisions, displays, Internet of Things (IoT) devices, radar devices, smart phones, wearable devices, tablet computers, netbooks, laptops, desktop computers, head-mounted displays (HMDs), autonomous driving and smart vehicles, virtual reality (VR) devices, augmented reality (AR) devices, extended reality (XR) devices, vehicles, mobile robots, etc.; and cloud computing devices, etc.
[0087] Reference Fig.13 , the electronic device 1300 may include a position estimating device 1310, an image processing device 1320, a processor 1330, a storage device 1340, an output device 1350, and a communication device 1360. The position estimating device 1310 may be included in the image processing device 1320.
[0088] The position estimation device 1310 can use the video frame as an input to output the optimal initial search position for motion estimation. The position estimation device 1310 may include a DP module as described above, and may output the initial search position by using the DP module. Alternatively, as described above, the position estimation device 1310 may include an ISPO module, and the ISPO module includes a neural network trained by the DP module. In this case, the DP module may be included in other devices installed inside or outside the electronic device 1300.
[0089] The image processing device 1320 may include a video codec device for performing video encoding and / or video decoding. The video codec device may include the aforementioned video encoder and / or video decoder. In addition, the image processing device 1320 may include a device for addressing physical degradation, such as a stabilizer or noise reduction (NR), high dynamic range (HDR), deblurring, frame rate up conversion (FRUC), etc.
[0090] The processor 1330 may include a main processor (e.g., a central processing unit (CPU) or an application processor (AP), etc.), an intellectual property (IP) core, and an auxiliary processor (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor, or a communication processor (CP)), which may operate independently of the main processor or in conjunction with the main processor, etc. The processor 1330 may control components of the electronic device 1300 and process requests thereof. For example, one or more GPUs may support parallel processing in response to a request from the position estimation device 1301.
[0091] The storage device 1340 may store data required to operate the components of the electronic device 1300 (e.g., images (still images or moving images captured by the image capture device), data processed by the processor 1330, neural networks used by the position estimation device 1310 and the image processing device 1320, etc.), and instructions for performing functions. The storage device 1340 may include a computer-readable storage medium, such as a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a magnetic hard disk, an optical disk, a flash memory, an electrically programmable read-only memory (EPROM), or other types of computer-readable storage media known in the art.
[0092] The output device 1350 may visually and / or non-visually output an image captured by the image capture device, and data generated or processed by the image processing device 1320 and the processor 1330. The output device 1350 may include a sound output device, a display device (e.g., a display), an audio module, and / or a haptic module.
[0093] The communication device 1360 may support establishing a direct (e.g., wired) communication channel and / or a wireless communication channel between the electronic device and other electronic devices, servers, or sensor devices within the network environment, and performing communication via the established communication channel by using various communication technologies. The communication device 1360 may send the image captured by the image capture device and the data generated or processed by the image processing device 1320 and the processor 1330 to other electronic devices. In addition, the communication device 1360 may receive an image to be processed from a cloud device or other electronic device, and may store the received image in the storage device 1340, and may send the image to the processor 1330 so that the image may be processed by the processor 1330.
[0094] In addition, the electronic device 1300 may also include a sensor device for detecting various data (for example, an accelerometer, a gyroscope, a magnetic field sensor, a proximity sensor, an illumination sensor, a fingerprint sensor, a GPS sensor, etc.), an image capture device for acquiring images (for example, a camera), an input device for receiving instructions and / or data, etc. (for example, a microphone, a mouse, a keyboard and / or a digital pen (for example, a stylus, etc.), etc.).
[0095] The present disclosure may be implemented as computer-readable codes written on a computer-readable recording medium. The computer-readable recording medium may be any type of recording device that stores data in a computer-readable manner.
[0096] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage, and carrier wave (e.g., data transmission via the Internet). Computer-readable recording media can be distributed on multiple computer systems connected to a network so that computer-readable code is written therein and executed therefrom in a decentralized manner. A programmer of ordinary skill in the technical field to which the present disclosure belongs can easily infer the functional programs, codes, and code segments required to implement the present disclosure.
[0097] According to example embodiments, at least one of the components, elements, modules or units described herein may be embodied as various numbers of hardware, software and / or firmware structures that perform the above-mentioned various functions. For example, at least one of these components, elements or units may use a direct circuit structure, such as a memory, a processor, a logic circuit, a lookup table, etc., which may perform various functions by the control of one or more microprocessors or other control devices. In addition, at least one of these components, elements or units may be specifically implemented by a module, a program or a portion of code that contains one or more executable instructions for performing a specified logical function and is performed by one or more microprocessors or other control devices. In addition, at least one of these components, elements or units may also include a processor, a microprocessor, etc., such as a central processing unit (CPU) that performs various functions, or is implemented by it. Two or more of these components, elements or units may be combined into a single component, element or unit that performs all operations or functions of the two or more components, elements or units combined. In addition, at least part of the functions of at least one of these components, elements or units may be performed by another of these components, elements or units. In addition, although the bus is not shown in the block diagram, the communication between components, elements or units may be performed by the bus. The functional scheme of the above exemplary embodiments can be implemented as an algorithm executed on one or more processors. In addition, the components, elements or units or processing operations represented by the blocks can use any number of related technologies for electronic configuration, signal processing and / or control, data processing, etc.
[0098] The present disclosure has been described above with reference to preferred embodiments. However, it is obvious to those skilled in the art that various changes and modifications can be made without changing the technical concept and basic features of the present disclosure. Therefore, it is clear that the above embodiments are illustrative in all aspects and are not intended to limit the present disclosure.
Claims
1. A video encoder, comprising: a differentiable prediction DP module configured to output an optimal initial search position by performing a full search in a predetermined area using a pair of frames of a video as input; as well as The motion estimation ME module is configured to perform motion estimation by moving the search position toward the optimal initial search position output by the DP module.
2. The video encoder according to claim 1, wherein the DP module is configured to generate a predicted image for each of a plurality of initial search positions in the predetermined area, and is configured to output the following initial search position as an optimal initial search position: at this initial search position, the residual between the generated predicted image and the first frame in the pair of frames is minimized.
3. The video encoder according to claim 2, wherein the DP module comprises: an affine transformation module configured to perform an affine transformation on a second frame of the pair of frames for each of the plurality of initial search positions; as well as The motion estimation and motion compensation MEMC module is configured as follows: performing motion estimation on the affine transformed second frame for each of the plurality of initial search positions to output motion in the form of a kernel; as well as Motion compensation is performed based on the motion of the kernel form to generate a predicted image.
4. The video encoder according to claim 3, wherein the MEMC module is configured as follows: For each block of the plurality of blocks of the second frame that have been affine transformed, performing expansion of the block to divide the block into a plurality of patches, and calculating a sum of absolute differences (SADs) between the plurality of patches of the block and a corresponding block of the first frame; and The motion in the kernel form is generated based on the calculated SAD of the plurality of blocks by using softmax. 5 . The video encoder of claim 1 , wherein the DP module is configured to perform parallel processing on multiple pairs of frames of the video using one or more processors.
6. The video encoder of claim 5, wherein the one or more processors include a graphics processing unit (GPU).
7. The video encoder according to claim 1, further comprising: A scaler is configured to scale the pair of frames of the video to a size to be processed by the DP module.
8. The video encoder according to claim 1, wherein: At least one of the number and size of the predetermined areas is preset based on at least one of computing capability, target processing speed, and accuracy of motion estimation.
9. A video encoder comprising: An initial search position optimization ISPO module, comprising a neural network trained to output an optimal initial search position by using a pair of frames of a video as input; as well as a motion estimation ME module configured to perform motion estimation by moving a search position toward an optimal initial search position output by the ISPO module, Among them, the neural network is trained to output the optimal initial search position by using the benchmark truth GT initial search position generated by the differentiable prediction DP module performing a full search in a predetermined area.
10. The video encoder of claim 9, wherein the neural network comprises a convolutional neural network (CNN).
11. The video encoder of claim 9, wherein the neural network is trained to output an optimal initial search position as an affine matrix value.
12. The video encoder of claim 9, wherein the neural network is trained by using the GT initial search position, wherein the GT initial search position is output by generating predicted images of a plurality of initial search positions in a predetermined area and determining an initial search position where a residual between the predicted image and a first frame in the pair of frames is minimized.
13. A video encoding method, comprising: The differentiable prediction DP module uses a pair of frames of the video as input to perform a full search in a predetermined area to output an optimal initial search position; Motion estimation is performed by the motion estimation ME module by moving the search position towards the optimal initial search position output by the DP module.
14. The video encoding method according to claim 13, wherein: Output of the optimal initial search position includes: generating predicted images for a plurality of initial search positions in the predetermined area; and An initial search position at which a residual between the generated predicted image and a first frame of the pair of frames is minimized is output as an optimal initial search position.
15. The video encoding method according to claim 14, wherein: Generating a predicted image includes: performing an affine transform on a second frame of the pair of frames for each of the plurality of initial search positions; performing motion estimation on the affine transformed second frame for each of the plurality of initial search positions to output motion in the form of a kernel; and Motion compensation is performed based on the motion of the kernel form to generate a predicted image.
16. The video encoding method according to claim 15, wherein: Output kernel forms of motion include: For each block of the plurality of blocks of the second frame that have been affine transformed, performing expansion of the block to divide the block into a plurality of patches, and calculating a sum of absolute differences (SADs) between the plurality of patches of the block and a corresponding block of the first frame; and The motion in the kernel form is generated based on the calculated SAD of the plurality of blocks by using softmax.
17. The video encoding method according to claim 13, wherein: Outputting the optimal initial search position includes performing parallel processing on multiple pairs of frames of the video using one or more processors.
18. The video encoding method according to claim 13, further comprising: The pair of frames of the video are scaled to a size to be processed by the DP module.
19. A video decoder comprising: a motion compensation MC module configured to perform motion compensation based on a motion vector extracted by performing motion estimation on a pair of frames of a video based on an optimal initial search position, wherein the optimal initial search position is obtained by a video encoder performing a full search on the pair of frames of the video; as well as The decoding module is configured to decode the video based on the result of motion compensation.
20. An electronic device comprising: A position estimation device configured to output an optimal initial search position by using a pair of frames of a video as input; an image processing device configured to perform motion estimation based on the output optimal initial search position, and configured to perform image processing based on a result of the motion estimation; as well as one or more processors configured to control the image processing device and process requests of the image processing device, The position estimation device comprises: a differentiable prediction DP module configured to output an optimal initial search position by performing a full search in a predetermined area; or an initial search position optimization ISPO module comprising a neural network trained by using a reference truth GT initial search position to output an optimal initial search position, wherein the GT initial search position is generated by the DP module performing a full search in a predetermined area.
Citation Information
Patent Citations
Well ceiling assembly for easy installation
KR1020230154594A