Method and apparatus for converting interlaced video into higher-resolution progressive video
By generating movement maps and using advanced interpolation and attention-based blocks, the method effectively converts interlaced HDTV to progressive UHD videos, addressing existing challenges in image quality and stability.
Patent Information
- Application Number
- PCT/KR2024/016492
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing technologies face challenges in converting interlaced HDTV (1080i FHD) to progressive UHD (3840x2160, 60p) videos, due to limitations in predicting accurate field line pixels, low-resolution training datasets, and instability during dial racing with large movements, leading to poor image quality and performance.
The method involves generating movement maps between fields, extracting feature maps, and using bicubic interpolation and directional attention-based residual dense blocks to perform dial racing and upscaling simultaneously, thereby converting FHD 60i videos to UHD 60p videos.
This approach simplifies the learning and execution process, improves image quality by correcting distortions in the CG area, and enhances the stability of the dial racing process, resulting in higher quality UHD progressive videos.
Smart Images

Figure KR2024016492_08052025_PF_FP_ABST
Abstract
Description
Method and device for converting interlaced video to higher resolution progressive video
[0001] The present invention relates to a method and apparatus for converting interlaced video into higher resolution progressive video.
[0002] In Korea, HDTV broadcasts are standardized at 1080i FHD (Full High Definition). 1080i FHD uses interlaced scanning at a resolution of 1920x1080 (1920x1080, 60i). In contrast, the UHD (Ultra High Definition) broadcast standard has a resolution four times higher than FHD and uses progressive scanning (3840x2160, 60p).
[0003] Therefore, in order to transmit existing FHD video as UHD broadcast, it is essential to convert FHD 60i video to UHD 60p.
[0004] In the past, in an image of an interlaced scanning method, odd field images and even field images acquired at different times were combined without being separated to form a single frame, and then this was used as input to a deep neural network to predict and output a target pixel line within the input field image, thereby forming a frame image.
[0005] When using a method like this, it is difficult to learn to predict accurate field line pixels because the input uses a frame image that combines two field images acquired at different times due to the characteristics of the CNN (Convolutional Neural Network) that operates in a sliding window manner.
[0006] In addition, the dataset used for training existing deep learning-based deinterlacing technology has low image resolution and small image size compared to the patch size used for training, which makes it difficult to utilize various areas within the image, resulting in low randomness when extracting patches.
[0007] The part where the image quality deterioration is noticeable in the deinterlacing result image is when the object movement in the image is large. However, the dataset contains many images with small movements, so if training is performed using only the dataset, image quality deterioration is observed for images with large motion.
[0008] Additionally, stable deinterlacing of large-motion field image inputs is difficult due to the size limitation of the perceptive field using the basic CNN structure.
[0009] Additionally, there is a disadvantage in that the deinterlacing performance is low due to the absence of an attention function that can accurately learn the pixel correspondence between field images.
[0010] In addition, conventional deinterlacing technologies have a structure that converts FHD 60i video into FHD 30p video. Therefore, deinterlacing must be performed twice to create a 60p video according to the UHD standard. In addition, there is no way to compensate for the results of each deinterlacing performed twice. In addition, when converting FHD 60i video into FHD 30p video, the first line of the middle field, which is the reference in the top field first and bottom field first methods, is different, so a phase difference may occur between the two methods when generating at 60p.
[0011] Additionally, the conventional method has the disadvantage of requiring a separate step of upscaling FHD to UHD after deinterlacing.
[0012] In addition, conventional deinterlacing technologies do not take into account the phenomenon of CG (Computer Graphics) such as subtitles or speech bubbles being inserted on top of interlaced images, so when performing deinterlacing, the CG area is not properly restored, causing image shaking.
[0013] The present invention aims to solve the above-mentioned problems of the prior art that occur when converting an image of a non-interlaced scan method (e.g., an FHD 60i image) into an image of a higher resolution of a progressive scan method (e.g., an UHD 60p image).
[0014] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0015] According to one embodiment of the present invention, an inter-field motion map generation unit for separating odd fields and even fields of Hx2W resolution from an input image frame of an interlaced scanning method of a 2Hx2W resolution, calculating the degree of motion of the odd fields and the even fields, and generating an inter-field motion map for each field, a first deinterlacing unit for obtaining a first intermediate result of deinterlacing of a 2Hx2W resolution based on first feature maps extracted for each of the odd fields and the even fields, and bicubic row interpolation results obtained for the odd fields and the even fields, a second deinterlacing unit for calculating a second feature map for the odd fields and the even fields from the first feature maps and inter-field motion maps for the odd fields and the even fields, and outputting a second deinterlacing result of a progressive scanning method of a 2Hx2W resolution based on at least the second feature map and the inter-field motion map, and from the second deinterlacing result, An image conversion device including a super-resolution unit that generates a sequential scan output image with a resolution of 4Hx4W is provided.
[0016] An image conversion device for converting an interlaced image into a progressive scan image of higher resolution according to a preferred embodiment of the present invention comprises: an inter-field motion map generation unit for separating an odd field and an even field of Hx2W resolution from an input image frame of an interlaced image of 2Hx2W resolution, calculating the degree of motion of the odd field and the even field, and generating an inter-field motion map for each field; a first deinterlacing unit for obtaining a first intermediate result of deinterlacing of 2Hx2W resolution based on first feature maps extracted for the odd field and the even field, and a bicubic row interpolation result obtained for the odd field and the even field; and a second feature map for calculating a first feature map for the odd field and the even field from the first feature map and the inter-field motion map for the odd field and the even field, and generating a second deinterlacing result of the progressive scan of 2Hx2W resolution based on at least the second feature map and the inter-field motion map. It comprises a second deinterlacing unit for outputting, and a super-resolution unit for generating a sequentially scanned output image with a resolution of 4Hx4W from the second deinterlacing result.
[0017] In one embodiment, the first deinterlacing unit increases the resolution of an image generated from the first feature map to 2Hx2W, and adds the bicubic row interpolation results obtained for the odd fields and the even fields to obtain a first intermediate deinterlacing result with a resolution of 2Hx2W.
[0018] In one embodiment, the second deinterlacing unit receives a feature map and an inter-field motion map for each of the odd fields and the even fields, calculates a feature value through a DARDB (Directional Attention-based Residual Dense Block) block, converts the second feature map generated through upscaling into a 2Hx2W image and outputs it, adds the output result to the first intermediate result of deinterlacing, and then outputs a second deinterlacing result through a combination with the inter-field motion map and the input image frame.
[0019] In one embodiment, the super-resolution unit passes the second feature map through a DARDB (Directional Attention-based Residual Dense Block) for each of the odd fields and the even fields, then combines the features in SRNet, increases the resolution to 4Hx4W, interpolates the second deinterlacing result in a bicubic interpolation module, and then adds it to the output of SRNet to generate a sequentially scanned output image with a resolution of 4Hx4W.
[0020] In one embodiment, the degree of motion is calculated based on the difference in pixel values between the previous and next frames, and an area with motion and an area without motion are distinguished and displayed based on whether the degree of motion is greater than a predetermined threshold.
[0021] A preferred embodiment of the present invention provides an image conversion method for converting an interlaced image into a progressive image of higher resolution, comprising: an inter-field motion map generation step of separating an odd field and an even field of Hx2W resolution from an input image frame of an interlaced image of 2Hx2W resolution, calculating the degree of motion of the odd field and the even field to generate a motion map between each field; a first deinterlacing and spatial interpolation step of obtaining a first intermediate result of deinterlacing of 2Hx2W resolution based on first feature maps extracted for each of the odd field and the even field and bicubic row interpolation results obtained for each of the odd field and the even field; A second deinterlacing and spatiotemporal interpolation step for calculating a second feature map for the odd field and the even field from the first feature map and the inter-field motion map for the odd field and the even field, and outputting a second deinterlacing result in a 2Hx2W resolution sequential scanning manner based on at least the second feature map and the inter-field motion map; and a 2x upscaling step for generating an output image in a 4Hx4W resolution sequential scanning manner from the second deinterlacing result.
[0022] In one embodiment, the first deinterlacing and spatial interpolation step may increase the resolution of an image generated from the first feature map to 2Hx2W, and add the bicubic row interpolation results obtained for the odd fields and the even fields to obtain a first intermediate deinterlacing result with a resolution of 2Hx2W.
[0023] In one embodiment, the second deinterlacing and spatiotemporal interpolation step includes a step of receiving a feature map and an inter-field motion map for each of the odd fields and the even fields, calculating a feature value through a DARDB (Directional Attention-based Residual Dense Block) block, converting the second feature map generated through upscaling into a 2Hx2W image and outputting the converted image, and a step of adding the output results for each of the odd fields and the even fields to a first intermediate result of deinterlacing, and then outputting a second deinterlacing result through a combination with the inter-field motion map and an input image frame.
[0024] In one embodiment, the 2x upscaling step includes a step of passing the second feature map through a Directional Attention-based Residual Dense Block (DARDB) for each of the odd fields and the even fields, combining features in an SRNet, and then increasing the resolution to 4Hx4W, and a step of interpolating the second deinterlacing result for each of the odd fields and the even fields in a bicubic interpolation module, and then adding it to the output of the SRNet to generate a sequentially scanned output image having a resolution of 4Hx4W.
[0025] In one embodiment, the degree of motion is calculated based on the difference in pixel values between the previous and next frames, and an area with motion and an area without motion are distinguished and displayed based on whether the degree of motion is greater than a predetermined threshold.
[0026] According to the present invention, since deinterlacing and upscaling are performed at the same time, the learning and execution process is simpler than in the prior art, stable deinterlacing can be performed, and improved image quality of the result can be expected.
[0027] According to the present invention, by inputting a motion map together during learning, distortion in a CG area without motion can be improved through motion information during the learning process.
[0028] In addition, according to the present invention, by learning deinterlacing and super-resolution (SR) at the same time, distortion occurring in deinterlacing can be additionally corrected in the super-resolution process, thereby obtaining improved results.
[0029] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0030] FIG. 1 is an explanatory diagram for explaining an image conversion concept according to one embodiment of the present disclosure.
[0031] FIG. 2 is a block diagram showing the configuration of an image conversion device according to an embodiment of the present disclosure.
[0032] FIG. 3 is a drawing for explaining the operation of a field-to-field motion map generation unit according to an embodiment of the present disclosure.
[0033] FIG. 4 is a drawing for explaining the operation of a first deinterlacing unit according to an embodiment of the present disclosure.
[0034] Figure 5 is a structural diagram of a feature extraction block according to an embodiment of the present disclosure.
[0035] Figure 6 is a structural diagram of a first image generation block according to an embodiment of the present disclosure.
[0036] FIG. 7 is a drawing for explaining the operation of a second deinterlacing unit according to an embodiment of the present disclosure.
[0037] Figure 8 is a diagram for explaining a directional attention-based residual dense block.
[0038] Figure 9 is a diagram for explaining directional attention.
[0039] FIG. 10 is a diagram illustrating a feature fusion and upscaling block according to an embodiment of the present disclosure.
[0040] FIG. 11 is a structural diagram of a second image generation block according to an embodiment of the present disclosure.
[0041] FIG. 12 is a drawing for explaining the operation of a super-resolution unit according to an embodiment of the present disclosure.
[0042] FIG. 13 is a diagram illustrating a process for preparing a training image dataset according to an embodiment of the present disclosure.
[0043] FIG. 14 is a conceptual diagram illustrating a method for training an artificial intelligence network according to an embodiment of the present disclosure.
[0044] Hereinafter, some embodiments of the present disclosure will be described in detail using exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, when describing the present disclosure, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present disclosure.
[0045] In describing components of embodiments according to the present disclosure, symbols such as first, second, i), ii), a), b) may be used. These symbols are only for distinguishing the components from other components, and the nature, order, or sequence of the components are not limited by the symbols. When a part in the specification is said to 'include' or 'have' a component, this does not mean that other components are excluded, but rather that other components may be further included, unless explicitly stated to the contrary. In addition, terms such as 'part' and 'module' described in the specification mean a unit that processes at least one function or operation, and this may be implemented by hardware, software, or a combination of hardware and software.
[0046] The following description of the invention, together with the accompanying drawings, is intended to illustrate exemplary embodiments of the invention and is not intended to represent the only embodiments in which the invention may be practiced.
[0047] Figure 1 is a diagram illustrating an image conversion concept according to one embodiment of the present disclosure. The present invention performs bidirectional deinterlacing and upscaling simultaneously using a single artificial intelligence network to convert, for example, an FHD 60i image into a UHD 60p image. For convenience of explanation below, the resolution of the input image is set to 2Hx2W.
[0048] The image conversion device (100) of the present invention is a device that converts three consecutive 2Hx2W resolution interlaced scanning image frames (I t-2,t-1 , I t,t+1 , I t+2, t+3 ) is input and sequentially scanned video frames (O) of 2 4Hx4W resolution t ,O t+1 ) is output. Input video frame I t,t+1 is an odd field X at time t + t and the even field X at time t+1- t+1 It consists of.
[0049] In the image conversion device (100) of the present invention, three consecutive 2Hx2W resolution interlaced scanning image frames (I) t-2,t-1 , I t,t+1 , I t+2, t+3 ) output image frames (O) in sequential scanning mode with a resolution of 4Hx4W on two sheets t ,O t+1 ) during the process of converting two sequentially scanned video frames (O) of the same resolution (2Hx2W) t ,O" t+1 ) is generated. If necessary, an intermediate output image (O") of 2Hx2W resolution sequential scanning method is generated in the middle. t ,O" t+1 ) can also be used as the final output. It is also possible to configure the output image resolution to be 4Hx4W, which is twice the input image resolution of 2Hx2W, or 2Hx2W, which is the same as the input image resolution.
[0050] FIG. 2 is a block diagram showing the configuration of an image conversion device according to an embodiment of the present disclosure. The image conversion device (100) of the present invention includes an inter-field motion map generation unit (110), a first deinterlacing unit (120), a second deinterlacing unit (130), and a super-resolution unit (140), and converts an input image (I) of an interlaced scanning method with a resolution of 2Hx2W into an output image (O) of a progressive scanning method with a resolution of 4Hx4W. When resolution upscaling is not required, an image (O") of a progressive scanning method with a resolution of 2Hx2W output from the second deinterlacing unit (130) may be used.
[0051] The present invention can be used to convert an input image of FHD 60i (1920x1080, 60i) into an output image of UHD 60p (3840x2160, 60p), but the resolution to which the present invention can be applied is not limited thereto.
[0052] 1. Field-to-field movement map generation unit
[0053] A motion map is intended to provide inter-image motion information to deinterlaced images using a deep neural network. For CG regions inserted over interlaced images, it is essential to maintain the original image without distortion in both odd and even fields. To achieve this, the present invention utilizes inter-image motion information to improve deinterlacing performance.
[0054] That is, in the present invention, the degree of motion is calculated separately for odd fields and even fields to identify motion between images, and these are combined and used to increase the accuracy of the motion map.
[0055] In one embodiment of the present invention, an inter-field motion map is used to combine input frames and deinterlaced frames to output a deinterlacing result that reflects motion information. By using the motion map, areas deemed motionless, such as CG areas, can be trained to maintain the input frame.
[0056] The motion map displays areas where motion is determined to exist and areas where motion is determined to exist.
[0057] In one embodiment, a field-to-field motion map M fd t can be expressed as in mathematical formula 1.
[0058]
[0059] Here, is the input odd field at time t, is a preset threshold value for the difference in pixel values. For example, can be 15. The sum of the absolute value of the difference in pixel values between the previous frame (t-1) and the current frame (t) and the absolute value of the difference in pixel values between the current frame (t) and the immediately following frame (t+1) If it is smaller than , the motion information for the pixel at that location in the motion map is 1, otherwise it is 0. If the motion information is 1, the area where the pixel is located is an area without motion, and if the motion information is 0, the area where the pixel is located is an area with motion.
[0060] Inter-field motion map M for input even fields fd t+1 can be obtained by applying mathematical expression 1 to the input even field.
[0061] Figure 3 is a diagram for explaining the operation of the field-to-field motion map generation unit according to one embodiment of the present disclosure. Input image frames (I) of the interlaced scanning method with a continuous 2Hx2W resolution t-2,t-1 , I t,t+1 , I t+2,t+3 ) is an odd field (X) with a resolution of Hx2W + ) and even fields (X - ) is divided into. And using mathematical expression 1, the motion map (M) of the odd field at time t fd t ) and the motion map of the even field at time t+1 (M fd t+1 ) are created respectively.
[0062] 2. 1st deinterlacing unit
[0063] The first deinterlacing unit (120) deinterlaces the input video frame (I t-2,t-1 , I t,t+1 , I t+2,t+3 ) odd field frames of Hx2W resolution divided into + ) and even field frames (X -) input and the first deinterlacing result (O') of 2Hx2W resolution spatially interpolated t , O' t+1 ) is obtained.
[0064] FIG. 4 is a drawing for explaining the operation of a first deinterlacing unit according to an embodiment of the present disclosure.
[0065] Referring to Figure 4, three consecutive frames (I t-2,t-1 , I t,t+1 , I t+2,t+3 ) has four fields, that is, odd fields (X + t , X + t+2 ) and even fields (X - t-1 , X - t+1 ) is input and features are extracted from each feature extraction block (121a, 121b, 121c, 121d) to generate a first feature map (FM1a, FM1b, FM1c, FM1d).
[0066] The feature extraction blocks (121a, 121b, 121c, 121d) are blocks for extracting features of an input field, and extract features using only spatial information without information of the preceding and following fields. The structure of the feature extraction block is as illustrated in FIG. 5. Referring to FIG. 5, the feature extraction blocks (121a, 121b, 121c, 121d) may be composed of a plurality of convolutional layers (CONV3-48) and activation layers (LeakyReLU) that compress the features of the input field, and a plurality of directional attention-based residual dense blocks (DARDB) that generate feature maps with weights assigned based on directionality, but may be modified in various implementation examples. The directional attention-based residual dense block (DARDB) will be described below with reference to FIGS. 8 and 9.
[0067] The first image generation block (123a, 123b) generates an image from the first feature map (FM1a, FM1b, FM1c, FM1d). The first image generation block (123a, 123b) increases the resolution of the generated image from Hx2W to 2Hx2W through pixel shuffle. The first image generation block (123a, 123b) generates a 3-channel color image. The structure of the first image generation block is as illustrated in Fig. 6. Referring to FIG. 6, the first image generation block (123a, 123b) may be composed of a number of convolution layers (CONV3-192, CONV3-48) for feature extraction, an activation layer (LeakyReLU), a pixel shuffle (PixelShuffle) for increasing resolution, and a convolution layer (CONV3-3) for generating a 3-channel color image, but may be modified in various implementation examples.
[0068] Also odd field X + t and even field X - t+1 Bicubic row interpolation (122a, 122b), that is, vertical bicubic interpolation is performed to output an image with a resolution of 2Hx2W.
[0069] The first deinterlacing result (O') of 2Hx2W resolution is obtained from the sum of the output of the first image generation block (123a, 123b) and the output of the bicubic row interpolation (122a, 122b). t , O' t+1 ) is obtained.
[0070] 3. Second deinterlacing unit
[0071] The second deinterlacing unit (130) receives three consecutive first feature maps (FM1a, FM1b, FM1c, FM1d) and an inter-field motion map and generates a second deinterlacing result (O") that is subjected to temporal-spatial interpolation. t , O"t+1 ) is obtained.
[0072] The second deinterlacing unit (130) calculates feature values through a directional attention-based residual dense block (DARDB), converts the second feature map (FM2a, FM2b) generated through upscaling into a 2Hx2W image, and outputs the output result as the first deinterlacing result (O'). t , O' t+1 ) and then combined with the inter-field motion map and input video frame to obtain the second deinterlacing result (O") of 2Hx2W resolution. t , O" t+1 ) is printed.
[0073] FIG. 7 is a drawing for explaining the operation of a second deinterlacing unit according to an embodiment of the present disclosure.
[0074] The second deinterlacing unit (130) is X - t-1 , X + t , X - t+1 Feature maps (FM1a, FM1b, FM1c) and motion map at time t (M fd t ) is connected (131a) and input to DARDB (132a). The second deinterlacing unit (130) is X + t , X - t+1 , X + t+2 Feature maps (FM1b, FM1c, FM1d) and motion map at time t+1 (M fd t+1 ) is connected (131b) and input into DARDB (132b).
[0075] Referring to FIGS. 8 and 9 together, the directional attention-based residual dense block (DARDB) is described.
[0076] Referring to Fig. 8, DARDB may include M dense blocks, one compression block, and a directional attention block. Here, M may be 8.
[0077] DARDB generates a feature map weighted based on directionality from the input frame. It combines the generated feature map with the input frame. DARDB weights the regions containing information necessary for predicting the deinterlaced image using directional attention.
[0078] Referring to Fig. 9, the directional attention block projects the input feature maps into the 0-degree attention direction (0°Avg Pool), the 45-degree attention direction (45°Avg Pool), the 90-degree attention direction (90°Avg Pool), and the 135-degree attention direction (135°Avg Pool), respectively. Accordingly, a map projected in the 0-degree direction, a map projected in the 45-degree direction, a map projected in the 90-degree direction, and a map projected in the 135-degree direction are generated.
[0079] The directional attention block merges maps projected in the 0-degree direction, 45-degree direction, 90-degree direction, and 135-degree direction and performs a convolution operation (Conv2d). The directional attention block performs batch normalization (BatchNorm) on the merged maps. Batch normalization is an artificial neural network technique that prevents the input values entering the layer from being skewed, spread out, or narrowed. Batch normalization can perform a transformation using gamma or beta. Batch normalization can be performed while maintaining non-linear properties.
[0080] After batch normalization is performed on the merged map, the merged map can be split into four maps again. The directional attention block performs a convolution operation (Conv2d) on each of these four maps. Afterwards, the sigmoid function is applied to each of the four maps on which the convolution operation was performed. The sigmoid function is a function that has an S-shaped curve or sigmoid curve and is used as an activation function. After the sigmoid function is applied to each of the four maps, each map can be rearranged. The directional attention block re-weights the four rearranged maps. After the weights are assigned, the directional attention block generates a feature map with weights assigned based on directionality.
[0081] Returning to FIG. 7, the feature values of three consecutive fields calculated in DARDB (132a, 132b) are combined in the feature fusion and upscaling module (133a, 133b) and upscaled to a size of 2Hx2W.
[0082] The feature fusion and upscaling module (133a, 133b) combines the features of three consecutive input fields and then increases the resolution to output a second feature map (FM2a, FM2b). The structure of the feature fusion and upscaling block is as illustrated in Fig. 10. Referring to Fig. 10, the feature fusion and upscaling module (133a, 133b) may be composed of a plurality of convolution layers (CONV3-48) and an activation layer (LeakyReLU) for feature extraction, and a pixel shuffle (PixelShuffle) for increasing the resolution of the feature maps, but may be modified in various implementation examples.
[0083] The second image generation block (134a, 134b) generates an image from the second feature map (FM2a, FM2b). The generated image is the first deinterlacing result (O' t , O' t+1 ) and is added (135a, 135b), and again the motion map (M fdt , M fd t+1 ), input frame (I t,t+1 ) and the second deinterlacing result (O") is obtained through the sum (136a, 136b) t , O" t+1 ) is output. The structure of the second image generation block is as illustrated in Fig. 11. Referring to Fig. 11, the second image generation block (134a, 134b) may be composed of a convolution layer (CONV3-48) for feature extraction, an activation layer (LeakyReLU), and a convolution layer (CONV3-3) for image generation of a 3-channel color, but may be modified in various implementation examples.
[0084] The output of the second image generation block (134a, 134b) is each FeatIMG t and FeatIMG t+1 When this is the case, the second deinterlacing result (O" t , O" t+1 ) can be expressed as mathematical expressions 2 and 3, respectively.
[0085]
[0086]
[0087] 4. Super-resolution department
[0088] The super-resolution unit (140) is a second feature map (FM2a, FM2b) with a resolution of 2Hx2W and a second deinterlacing result (O" t , O" t+1 ) with a final output image resolution of 4Hx4W, upscaled by 2x (O t, O t+1 ) to get.
[0089] FIG. 12 is a drawing for explaining the operation of a super-resolution unit according to an embodiment of the present disclosure.
[0090] Referring to Fig. 12, the second feature maps (FM2a, FM2b) output in the middle from the second deinterlacing unit (130) pass through DARDB (141a, 141b) and are then input to SRNet (142a, 142b). SRNet (142a, 142b) combines the features and then outputs an image with a resolution of 4Hx4W. Here, SRNet refers to a super-resolution network and may include, but is not limited to, VDSR (Very Deep Super Resolution), EDSR (Enhanced Deep Super Resolution), SRGAN (Super Resolution Generative Adversarial Network), etc.
[0091] Second deinterlacing result (O") t , O" t+1 ) are interpolated through the bicubic interpolation module (143a, 143b) and then added to the output of SRNet (142a, 142b) (144a, 144b).
[0092] That is, the final output image O with a resolution of 4Hx4W t Wow O t+1 can be expressed as in mathematical expressions 4 and 5. In the mathematical expressions below, Interpolation represents bicubic interpolation, and SRNet t is the output of SRNet(142a), SRNet t+1 represents the output of SRNet (142b).
[0093]
[0094]
[0095] 5. Learning
[0096] First, a training image dataset used for learning the artificial intelligence network of the image conversion device (100) of the present invention will be described. FIG. 13 is a diagram illustrating a process for preparing a training image dataset according to an embodiment of the present disclosure.
[0097] Referring to Fig. 13, odd field images and even field images of FHD 60i are extracted from the ground truth UHD 60p image. For example, a training image dataset can be prepared by obtaining an FHD 60p image (Fig. 13 (b)) by performing 1 / 2 bicubic downsampling on a UHD 60p image (Fig. 13 (a)), and then obtaining an FHD 60i image (Fig. 13 (c)) from which odd and even fields are extracted.
[0098] FIG. 14 is a conceptual diagram illustrating a method for training an artificial intelligence network according to an embodiment of the present disclosure.
[0099] Referring to Figure 14, the training image dataset (X) of the interlaced scanning method with a resolution of 2Hx2W from which odd and even fields are extracted - t-1 , X + t , X - t+1 , X + t+2 ) from the correct UHD 60p video (O t , O t+1 ) is trained to generate an artificial intelligence network of an image conversion device (100).
[0100] For training artificial intelligence networks, loss functions such as Equation 6 are used to minimize loss. In Equation 6, Y is the original image, and Y is the 1 / 2 is an image of Y that has been bicubically downsampled.
[0101]
[0102] Each component of the device or method according to the present invention may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.
[0103] Various implementations of the systems and techniques described herein can be implemented as digital electronic circuits, integrated circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations of one or more computer programs executable on a programmable system. The programmable system includes at least one programmable processor (which may be a special purpose processor or a general purpose processor) coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device. Computer programs (also known as programs, software, software applications, or code) include instructions for the programmable processor and are stored on a "computer-readable recording medium."
[0104] A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system. Such a computer-readable recording medium may be a non-volatile or non-transitory medium such as a ROM, CD-ROM, magnetic tape, floppy disk, memory card, hard disk, magneto-optical disk, storage device, and may further include a transitory medium such as a data transmission medium. Furthermore, the computer-readable recording medium may be distributed across network-connected computer systems, so that computer-readable code can be stored and executed in a distributed manner.
[0105] Although the flowchart of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart of this specification without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart of this specification is not limited to a chronological order.
[0106] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art to which the present embodiment pertains may make various modifications and variations without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The protection scope of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present embodiment.
[0107] (Explanation of symbols)
[0108] 110 Field-to-field movement map generation unit,
[0109] 120 1st deinterlacing unit,
[0110] 130 Second deinterlacing unit,
[0111] 140 Super Resolution Department.
[0112] CROSS-REFERENCE TO RELATED APPLICATION
[0113] This patent application claims priority to Korean Patent Application No. 10-2023-0148357, filed on October 31, 2023, the entire contents of which are incorporated herein by reference.
Claims
An inter-field motion map generation unit that separates odd and even fields of Hx2W resolution from an input image frame of an interlaced scanning method of 1.2Hx2W resolution, calculates the degree of motion of the odd field and the even field, and generates a motion map between each field; A first deinterlacing unit that obtains a first intermediate result of deinterlacing with a resolution of 2Hx2W based on the first feature map extracted for each of the odd field and the even field and the bicubic row interpolation result obtained for each of the odd field and the even field, A second deinterlacing unit that calculates a second feature map for the odd field and the even field from the first feature map and the inter-field motion map for the odd field and the even field, and outputs a second deinterlacing result of a sequential scanning method with a resolution of 2Hx2W based on at least the second feature map and the inter-field motion map, and A super-resolution unit that generates a sequentially scanned output image with a resolution of 4Hx4W from the above secondary deinterlacing results. An image conversion device including:
2. In paragraph 1, The above first deinterlacing unit, An image conversion device that increases the resolution of an image generated from the first feature map to 2Hx2W and adds the bicubic row interpolation results obtained for the odd fields and the even fields to obtain a first intermediate result of deinterlacing with a resolution of 2Hx2W.
3. In paragraph 2, The above second deinterlacing unit, For the above odd fields and even fields, respectively, An image conversion device that receives a feature map and an inter-field motion map as input, calculates feature values through a DARDB (Directional Attention-based Residual Dense Block), converts a second feature map generated through upscaling into a 2Hx2W image and outputs it, adds the output result to a first intermediate result of deinterlacing, and then outputs a second deinterlacing result through combination with the inter-field motion map and the input image frame.
4. In paragraph 3, The above super-resolution part is, For the above odd fields and even fields, respectively, After passing the above second feature map through DARDB, the features are combined in SRNet and the resolution is increased to 4Hx4W. An image conversion device that interpolates the secondary deinterlacing result in a bicubic interpolation module and then adds it to the output of SRNet to generate a sequential scan output image with a resolution of 4Hx4W.
5. In any one of paragraphs 1 to 4, The above motion degree is calculated based on the difference in pixel values between the previous and next frames. An image conversion device that distinguishes and displays areas with movement and areas without movement based on whether the degree of movement is greater than a predetermined threshold. A step of generating an inter-field motion map by separating odd fields and even fields of Hx2W resolution from an input image frame of an interlaced scanning method of 6.2Hx2W resolution, calculating the degree of motion of the odd fields and the even fields, and generating an inter-field motion map for each field; A first deinterlacing and spatial interpolation step for obtaining a first intermediate result of deinterlacing with a resolution of 2Hx2W based on the first feature maps extracted for the odd field and the even field, and the bicubic row interpolation results obtained for the odd field and the even field; A second deinterlacing and spatiotemporal interpolation step of calculating a second feature map for the odd field and the even field from the first feature map and the inter-field motion map for the odd field and the even field, and outputting a second deinterlacing result of a sequential scanning method with a resolution of 2Hx2W based on at least the second feature map and the inter-field motion map; and A 2x upscaling step that generates a sequential scan output image with a resolution of 4Hx4W from the above secondary deinterlacing result. An image conversion method including:
7. In paragraph 6, The above first deinterlacing and spatial interpolation step is, An image conversion method for increasing the resolution of an image generated from the first feature map to 2Hx2W and obtaining a first intermediate result of deinterlacing with a resolution of 2Hx2W by adding the bicubic row interpolation results obtained for the odd fields and the even fields.
8. In paragraph 7, The above second deinterlacing and spatiotemporal interpolation steps are: A step of receiving a feature map and a motion map between fields for each of the odd fields and the even fields, calculating feature values through DARDB (Directional Attention-based Residual Dense Block), and converting the second feature map generated through upscaling into a 2Hx2W image and outputting it; An image conversion method comprising a step of adding the output results for each of the odd fields and the even fields to the first intermediate result of deinterlacing, and then outputting the second deinterlacing result through combination with the inter-field motion map and the input image frame.
9. In paragraph 8, The above 2x upscaling step is A step of increasing the resolution to 4Hx4W after combining the features in SRNet after passing the second feature map through DARDB for each of the odd fields and the even fields, An image conversion method comprising a step of interpolating the second deinterlacing result for each of the odd fields and the even fields in a bicubic interpolation module and then adding it to the output of SRNet to generate an output image in a sequential scanning method with a resolution of 4Hx4W.
10. In any one of paragraphs 6 to 9, The above motion degree is calculated based on the difference in pixel values between the previous and next frames. An image conversion method that distinguishes and displays areas with movement and areas without movement based on whether the degree of movement is greater than a predetermined threshold.
Citation Information
Patent Citations
Digital video signal scaling method
KR100655040B1
Apparatus and method for video signal reconstitution
KR1020020044714A
Molding Methods for Recycling of scrap Waste for Steelmaking
KR1020210036768A
Method and de-interlacing apparatus that employs recursively generated motion history maps
US20080036908A1
Video image de-interlacing method and video image de-interlacing device
US20230156147A1