Electronic device and operating method thereof
Patent Information
- Application Number
- PCT/KR2025/002919
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-06
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies lack an efficient method to convert still images into high-quality videos while maintaining structural and detailed texture information without artifacts.
An electronic device employs an image processing module comprising a feature extraction unit, noise generation unit, flow generation unit, residual generation unit, and residual synthesis unit, utilizing neural networks with encoder-decoder structures to generate a series of frame images from a still image, incorporating noise and flow information to enhance image quality and detail.
The solution effectively generates high-quality video frames from still images, preserving structural and texture details without degradation, achieving a seamless conversion process.
Smart Images

Figure KR2025002919_02102025_PF_FP_ABST
Abstract
Description
Electronic device and method of operation thereof
[0001] The present disclosure relates to an electronic device for converting a still image into a moving image and a method of operating the same.
[0002] With the advancement of computer technology and the exponential growth of data traffic, artificial intelligence has become a key trend driving future innovation. Because AI mimics human thought processes, it has virtually limitless applications across all industries. Representative AI technologies include pattern recognition, machine learning, expert systems, neural networks, and natural language processing.
[0003] Neural networks mathematically model the characteristics of human biological neurons, mimicking the human capacity for learning. Neural networks can create mappings between input and output data. This ability to create mappings can be described as the neural network's learning ability. Furthermore, neural networks possess the ability to generalize, generating correct output data for input data not used in learning, based on learned results.
[0004] Neural networks can be used for image processing, for example, a deep neural network (DNN) can be used to generate images, remove noise or artifacts from images, or perform image processing to increase the resolution of images.
[0005] An electronic device according to one embodiment can convert a still image into a video.
[0006] An electronic device according to one embodiment may include at least one processor including a memory storing one or more instructions and a processing circuit.
[0007] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain flow information for the first image.
[0008] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain a plurality of modified images that have modified the first image based on the flow information.
[0009] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain residual information about the plurality of transformed images based on the plurality of transformed images and the first image.
[0010] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate a plurality of frame images based on the plurality of transformed images and the residual information.
[0011] A method of operating an electronic device according to one embodiment may include a step of obtaining flow information for a first image.
[0012] A method of operating an electronic device according to one embodiment may include a step of obtaining a plurality of transformed images obtained by transforming the first image based on the flow information.
[0013] A method of operating a display device according to one embodiment may include a step of obtaining residual information for the plurality of transformed images based on the plurality of transformed images and the first image.
[0014] A method of operating an electronic device according to one embodiment may include a step of generating a plurality of frame images based on the plurality of transformed images and the residual information.
[0015] The above and other aspects, features and advantages relating to specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0016] FIG. 1 is a diagram illustrating an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0017] Figure 2 is a block diagram showing the configuration of an image processing module according to one embodiment.
[0018] FIG. 3 is a diagram showing an encoder included in a feature extraction unit according to one embodiment.
[0019] FIG. 4 is a diagram illustrating a residual block according to one embodiment.
[0020] Fig. 5 is a diagram showing a flow generation unit according to one embodiment.
[0021] FIG. 6 is a diagram illustrating an attention block according to one embodiment.
[0022] Figure 7 is a diagram showing a down block according to one embodiment.
[0023] FIG. 8 is a diagram illustrating a spatial attention block according to one embodiment.
[0024] FIG. 9 is a diagram showing an up block according to one embodiment.
[0025] FIG. 10 is a diagram showing a feature transformation unit and a feature decoding unit according to one embodiment.
[0026] Fig. 11 is a drawing showing a residual generation unit according to one embodiment.
[0027] Fig. 12 is a drawing showing a residual synthesis unit according to one embodiment.
[0028] Fig. 13 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0029] FIG. 14 is a diagram for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0030] FIG. 15 is a drawing for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video.
[0031] FIG. 16 is a diagram for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0032] FIG. 17 is a diagram for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0033] Fig. 18 is a block diagram showing the configuration of an electronic device according to one embodiment.
[0034] Fig. 19 is a block diagram showing the configuration of an electronic device according to one embodiment.
[0035] The terms used in this specification will be briefly explained, and the present invention will be described in detail.
[0036] The terms used in this invention have been selected from widely used, current terms, taking into account their functions. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this invention should not be defined simply as names, but rather based on their inherent meanings and the overall content of the invention.
[0037] When a part throughout the specification is said to "include" a certain component, this does not mean that other components are excluded, but rather that other components may be included, unless otherwise specifically stated. In addition, terms such as "unit" and "module" described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware or software, or a combination of hardware and software. As used herein, "unit" and "module" refer to a hardware component such as a processor or a circuit, and / or a software component executed by a hardware component such as a processor. A "unit" and a "module" may be implemented by a program stored on an addressable storage medium and executed by a processor. For example, a "unit" and a "module" may be implemented as a component such as a software component, an object-oriented software component, a class component, a task component, a process, a function, an attribute, a procedure, a subroutine, a portion of program code, a driver, firmware, microcode, a circuit, data, a database, a data structure, a table, an array, and a parameter.
[0038] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement the present invention. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.
[0039] In the embodiments of this specification, the term "user" means a person who controls a system, function or operation, and may include a developer, administrator or installer.
[0040] Additionally, in the embodiments of the present specification, 'image' or 'picture' may represent a still image, a moving image composed of a plurality of consecutive still images (or frames), or a video.
[0041] When the phrase "at least one of" is used with a list of items, it means that various combinations including one or more of the listed items can be used, and that only one item from the list is required. For example, "at least one of A, B, and C" includes the following combinations: A, B, C, A and B, A and C, B and C, A, B, and C, and variations thereof. As a further example, the phrase "at least one of a, b, or c" can include a alone, b alone, c alone, a and b, a and c, b and c, or all of a, b, and c, and variations thereof. Similarly, the term "set" means one or more items. Thus, a set of items can be a single item or a collection of two or more items.
[0042] FIG. 1 is a diagram illustrating an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0043] An electronic device according to one embodiment may be implemented as various electronic devices such as a mobile phone, a tablet PC, a digital camera, a camcorder, a laptop computer, a desktop, an e-book reader, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a navigation device, an MP3 player, a camcorder, an IPTV (Internet Protocol Television), a DTV (Digital Television), a wearable device, etc.
[0044] Referring to FIG. 1, an electronic device (100) according to one embodiment may generate a video (20) including a plurality of frame images by processing a first image (10) using an image processing module (or an image processing network). At this time, the plurality of frame images included in the video (20) may be images in which details included in the first image (10) are maintained (e.g., structural information such as edges included in the first image, detailed texture information, etc.). In addition, the plurality of frame images may be high-quality images that do not include artifacts due to warping and in which the quality of the plurality of frame images is not degraded compared to the first image (10).
[0045] An image processing module according to one embodiment may include suitable logic, circuitry, interfaces, and / or code that operate to convert a still image into a video. The image processing module may include one or more neural networks. For example, the image processing module may include an encoder that converts an input image into a compressed form or extracts features of the input image, a decoder that restores the compressed form to the original resolution, a U-Net having an encoder-decoder structure, etc. However, the present invention is not limited thereto. The image processing module may include various neural networks.
[0046] Hereinafter, an image processing module according to one embodiment will be described in detail with reference to the drawings.
[0047] Figure 2 is a block diagram showing the configuration of an image processing module according to one embodiment.
[0048] Referring to FIG. 2, an image processing module (200) according to one embodiment may include a feature extraction unit (210), a feature transformation unit (220), a feature decoding unit (230), a noise generation unit (240), a flow generation unit (250), a residual generation unit (260), and a residual synthesis unit (270). However, the present invention is not limited thereto.
[0049] The feature extraction unit (210) can extract feature information from the first image (201) using an encoder. The encoder is a network that extracts feature information from the image and may include one or more convolutional neural networks (CNNs). The feature extraction unit (210) will be described in detail with reference to FIGS. 3 and 4.
[0050] The noise generation unit (240) can randomly generate noise images. For example, the noise generation unit (240) can generate noise images following a Gaussian distribution (normal distribution). However, the present invention is not limited thereto.
[0051] The flow generation unit (250) can generate flow information (flow map) based on the feature information extracted from the feature extraction unit (210) and the noise images generated from the noise generation unit (240) using a flow generation network. The flow generation network may include a U-Net having an encoder-decoder structure. The flow generation network will be described in detail with reference to FIGS. 5 to 9.
[0052] The feature transformation unit (220) can generate transformed feature information by applying the flow information generated by the flow generation unit (250) to the feature information extracted by the feature extraction unit (210) and warping it. The feature decoding unit (230) can generate transformed images by processing the transformed feature information using a decoder. The decoder can restore compressed information (e.g., feature information) to its original resolution. The feature transformation unit (220) and the feature decoding unit (230) will be described in detail with reference to FIG. 10.
[0053] The residual generation unit (260) can generate residual information (residual map) based on the transformed images generated by the feature transformation unit (220), the first image (201), and the noise images generated by the noise generation unit (240) using a residual generation network. The residual generation network can include a U-Net having an encoder-decoder structure. The residual generation network will be described in detail with reference to FIG. 11.
[0054] The residual synthesis unit (270) can generate a plurality of frame images (202) by synthesizing the residual information generated by the residual generation unit (260) and the transformed images generated by the feature decoding unit (230). The residual synthesis unit (270) will be described in detail with reference to FIG. 12.
[0055] Hereinafter, with reference to the drawings, the feature extraction unit (210), feature transformation unit (220), feature decoding unit (230), noise generation unit (240), flow generation unit (250), residual generation unit (260), and residual synthesis unit (270) will be described in detail.
[0056] FIG. 3 is a diagram showing an encoder included in a feature extraction unit according to one embodiment.
[0057] Referring to FIG. 3, a feature extraction unit (210) according to one embodiment may include an encoder (300). The encoder (300) according to one embodiment may extract feature information (302) of a first image (301). The encoder (300) may include one or more convolution layers (310), one or more residual blocks (320, “ResBlk” 320 of FIG. 3), a normalization layer (330, “Norm” 330), and an activation layer (340, “activation function” 340).
[0058] In a convolution layer (“Conv” 310) according to one embodiment, a convolution operation can be performed between input data (or input image) input to the convolution layer (310) and a kernel included in the convolution layer (310).
[0059] According to one embodiment, a residual block (320) may include a skip connection that skips one or more layers included in the residual block (320). The residual block will be described in detail with reference to FIG. 4.
[0060] FIG. 4 is a diagram illustrating a residual block according to one embodiment.
[0061] Referring to FIG. 4, a residual block (320, “ResBlk” 320 of FIG. 4) may include one or more normalization layers (410, “Norm” 410), activation layers (420, “Activation Function” 420), convolution layers (430, “Conv” 430), and a summation layer (440). In addition, the residual block (320) may include a skip connection (450) that performs a convolution operation on input data input to the residual block (320) and passes the convolution operation to the summation layer (440).
[0062] In the normalization layer (410), the range of values of data input to the normalization layer (410) can be adjusted. In the normalization layer (410), batch normalization, layer normalization, instance normalization, group normalization, etc. can be performed.
[0063] In the activation layer (420), an activation function operation may be performed to apply an activation function to input data input to the activation layer (420). The activation function operation provides non-linear characteristics, and the activation function may include a sigmoid function, a Tanh function, a ReLU (Rectified Linear Unit) function, a leaky ReLu function, an ELU (Exponential Linear Unit) function, a Swish function (SiLU (Sigmoid Linear Unit) function), etc. However, the present invention is not limited thereto.
[0064] In the convolution layer (430), a convolution operation can be performed between the input data input to the convolution layer (430) and the kernel included in the convolution layer.
[0065] In the summation layer (440), an element-by-element summation operation of data input to the summation layer (440) can be performed.
[0066] Again, referring to FIG. 3, in the normalization layer (330), the range of values of data input to the normalization layer (330) can be adjusted. In the normalization layer (330), batch normalization, layer normalization, instance normalization, group normalization, etc. can be performed.
[0067] In the activation layer (340), an activation function operation may be performed to apply an activation function to input data input to the activation layer (340). The activation function operation may impart non-linear characteristics. For example, the activation function may include a sigmoid function, a Tanh function, a ReLU (Rectified Linear Unit) function, a leaky ReLu function, an ELU (Exponential Linear Unit) function, a Swish function (SiLU (Sigmoid Linear Unit) function), etc. However, the present invention is not limited thereto.
[0068] An encoder (300) according to one embodiment can output feature information (302) for a first image (301).
[0069] Fig. 5 is a diagram showing a flow generation unit according to one embodiment.
[0070] Referring to FIG. 5, the flow generation unit (250) may include a flow generation network (500). The flow generation network (500) may include a U-Net having an encoder-decoder structure.
[0071] The flow generation network (500) may include one or more concatenation layers, one or more convolution layers, a temporal attention block, a spatial attention block, a normalization layer, one or more residual blocks, down blocks, and up blocks.
[0072] The flow generation network (500) may be input with noise information (501) generated by the noise generation unit (240) and feature information (302) for the first image. The noise information (501) and feature information (302) for the first image may be input to the first connection layer (511).
[0073] In the first connection layer (511), data input to the first connection layer (511) can be concatenated. For example, in the first connection layer (511), data that concatenates noise information (501, first input data) and feature information (302, second input data) for the first image in the channel direction can be output.
[0074] In the first convolution layer (“Conv” 521), a convolution operation can be performed between the input data and the kernel included in the first convolution layer (“521”).
[0075] Data output from the first convolution layer (521) can be input to the first time attention block (531). In addition, data output from the first convolution layer (521) can be input to the second connection layer (512) of FIG. 5, which will be described later.
[0076] In the first time attention block (531, “Temporal attn” 531 of FIG. 5), attention can be applied within the time axis for each location of pixels included in the input data. Here, attention can mean finding related features using similarity between features, integrating related features into one (aggregation), and extracting integrated feature information.
[0077] In the first time attention block (531), an attention operation can be performed. The attention operation means an operation that obtains correlation information (e.g., similarity information) between query data (“q”) and key data (“k”), obtains a weight based on the correlation information, reflects the weight to value data (“v”) mapped to the key data (“k), and performs a weighted sum on the value data (“v”) to which the weight is reflected.
[0078] At this time, the attention operation performed based on query data (“q”), key data (“k”), and value data (“v”) obtained from the same input data may be referred to as a self-attention operation.
[0079] The attention operation performed in the first time attention block (531) will be described in detail with reference to FIG. 6.
[0080] FIG. 6 is a diagram illustrating an attention block according to one embodiment.
[0081] Referring to FIG. 6, in the first time attention block (531), query data (“q”), key data (“k”), and value data (“v”) can be obtained based on input data input to the first time attention block (531).
[0082] For example, in the linear layer (“Linear” 611), a linear transformation of data input to the linear layer (611) can be performed. For example, a multiplication operation between data input to the linear layer (611) and a weight matrix included in the linear layer (611) can be performed.
[0083] Data on which a multiplication operation with a weight matrix is performed in the linear layer (611) can be input to the split layer (“Split” 620).
[0084] Data input to the split layer (620) can be split into a preset number of pieces. For example, data input to the split layer (620) can be divided into query data (“q”), key data (“k”), and value data (“v”), respectively.
[0085] Query data (“q”), key data (“k”), and value data (“v”) can be input into transformation layers (“reshape” 631, “reshape” 632, “reshape” 633), respectively. In the transformation layers (“631,” “632,” “633”), the input data can be rearranged into specific dimensions.
[0086] First association data (“e”) can be obtained through an element-wise multiplication operation of the rearranged query data (“q”) and key data (“k”) in each of the transformation layers (631, 632).
[0087] By adding position bias (“pos bias”) to the first association data, the second association data can be obtained.
[0088] The first time attention block (531) can obtain weight data (A) by applying a softmax function (“softmax” in FIG. 6) to the second correlation data, and can perform an element-wise multiplication operation of the weight data (A) and the value data rearranged in the transformation layer (633).
[0089] In the transformation layer (“reshape” 634), the first output data (a) can be obtained by rearranging the data on which the element-wise multiplication operation is performed.
[0090] The first time attention block (531) can obtain second output data by performing a multiplication operation between the first output data (a) and the weight matrix included in the linear layer (612, “Linear” 612).
[0091] Referring again to FIG. 5, data output from the first time attention block (531) can be sequentially image-processed in the normalization layer (“Norm” 540) and the first residual block (“ResBlk” 551).
[0092] Data output from the first time attention block (531) can be input to the normalization layer (540).
[0093] In the normalization layer (540), the range of values of data input to the normalization layer (540) can be adjusted. In the normalization layer (540), batch normalization, layer normalization, instance normalization, group normalization, etc. can be performed.
[0094] Data output from the normalization layer (540) can be input to the first residual block (551).
[0095] In the first residual block (551), the operations illustrated and described in FIG. 4 can be performed, and data output from the first residual block (551) can be input to the first down block (561, “Down Block” 561 of FIG. 5).
[0096] The first down block (561) will be described in detail with reference to Fig. 7.
[0097] Figure 7 is a diagram showing a down block according to one embodiment.
[0098] Referring to FIG. 7, a first down block (561) according to one embodiment may include one or more residual blocks, a spatial attention block, a temporal attention block, and a down sampling layer.
[0099] According to one embodiment, data input to the first down block (561) may be sequentially image-processed in two residual blocks (711, 712, “ResBlk” 711, “ResBlk” 712), a spatial attention block (720, “Spatial Attn” 720), a temporal attention block (730, “Temporal Attn” 730), and a down sampling layer (740, “Down sample” 740).
[0100] In the two residual blocks (711, 712), the operations illustrated and described in FIG. 4 can be performed. In addition, in the time attention block (730), the operations illustrated and described in FIG. 6 can be performed.
[0101] In the spatial attention block (720), attention can be applied within a spatial axis for each input frame. Here, attention can mean finding related features using similarity between features, integrating related features into one (aggregation), and extracting integrated feature information.
[0102] An attention operation can be performed in the spatial attention block (720). The attention operation means an operation that obtains association information (e.g., similarity information) between query data (“q”) and key data (“k”), obtains a weight based on the association information, reflects the weight to value data (“v”) mapped to the key data (“k”), and performs a weighted sum on the value data (“v”) to which the weight is reflected.
[0103] At this time, the attention operation performed based on query data (“q”), key data (“k”), and value data (“v”) obtained from the same input data may be referred to as a self-attention operation.
[0104] The attention operation performed in the spatial attention block (720) will be described in detail with reference to FIG. 8.
[0105] FIG. 8 is a diagram illustrating a spatial attention block according to one embodiment.
[0106] Referring to FIG. 8, in the spatial attention block (720), query data (“q”), key data (“k”), and value data (“v”) can be obtained based on input data input to the spatial attention block (720).
[0107] For example, input data may be input to a convolution layer (811). A convolution operation may be performed between the input data input to the convolution layer (“Conv” 811) and the kernel included in the convolution layer (811).
[0108] Data on which a convolution operation with a kernel is performed in the convolution layer (811) can be input to the split layer (“Split” 820).
[0109] The data input to the split layer (820) can be split into a preset number of pieces. For example, the data input to the split layer (820) can be split into three pieces of data. The three pieces of data split in the split layer (820) can be input to the transformation layers (831, 832, 833, “reshape” 831, 832, 833), respectively. In the transformation layers (831, 832, 833), the input data can be rearranged into a specific dimension.
[0110] The rearranged data in each of the transformation layers (831, 832, 833) can be query data (“q”), key data (“k”), and value data (“v”), respectively.
[0111] A weight matrix (w) can be obtained through an element-wise multiplication operation of query data (“q”) and key data (“k”).
[0112] The spatial attention block (720) can obtain weight data (A) by applying a softmax function to a weight matrix (w), and the spatial attention block (720) can perform an element-wise multiplication operation of the weight data (A) and the value data (v). The data on which the element-wise multiplication operation has been performed can be input to a transformation layer (“reshape” 834).
[0113] In the transformation layer (834), the first output data (a) can be obtained by rearranging the data on which the element-wise multiplication operation is performed.
[0114] The first output data can be input to a convolution layer (“Conv” 812). The spatial attention block (720) can obtain the second output data by performing a convolution operation between the input first output data and the kernel included in the convolution layer (812) in the convolution layer (812).
[0115] Referring again to FIG. 7, data output from the spatial attention block (720) can be input to the temporal attention block (730).
[0116] In the time attention block (730), the operations illustrated and described in FIG. 6 can be performed. Data processed in the time attention block (730) can be input to the down sampling layer (740) and the first up block (“Up Block” 581) of FIG. 5. The first up block (581) of FIG. 5 will be described in detail later.
[0117] The downsampling layer (740) may include a convolution layer. In the downsampling layer, the size (dimension) or resolution of input data may be reduced through a convolution operation.
[0118] Again, referring to FIG. 5, data output from the first down block (561) can be input to the second down block (“Down Block” 562) located next to the first down block (561). In the second down block (562), the operations illustrated and described in FIG. 7 can be performed. In addition, the operations illustrated and described in FIG. 7 can also be performed in the third down block (“Down Block” 563) and the fourth down block (“Down Block” 564).
[0119] Data output from the fourth down block (564) can be input to the second residual block (552). In the second residual block (“ResBlk” 552), the operations illustrated and described in FIG. 4 can be performed.
[0120] Data output from the second residual block (552) can be input to a spatial attention block (“Spatial Attn” 570). The spatial attention block (570) can perform operations illustrated and described in FIG. 8.
[0121] Data output from the spatial attention block (570) can be input to the third residual block (553). The operations illustrated and described in FIG. 4 can be performed in the third residual block (553).
[0122] Data output from the third residual block (“ResBlk” 553) can be input to the second temporal attention block (“Temporal Attn” 532). In the second temporal attention block (532), the operations illustrated and described in FIG. 6 can be performed.
[0123] Data output from the second time attention block (532) can be input to the first up block (581). The first up block (581) will be described in detail with reference to FIG. 9.
[0124] FIG. 9 is a diagram showing an up block according to one embodiment.
[0125] Referring to FIG. 9, a first up block (581) according to one embodiment may include an up-sampling layer, a summation layer, one or more residual blocks, a spatial attention block, and a temporal attention block.
[0126] According to one embodiment, data input to the first up block (581) may be sequentially image-processed in an upsampling layer (910, “Upsample” 910), a summation layer (920), two residual blocks (a first residual block (931, “ResBlk” 931), a second residual block (932, “ResBlk” 932)), a spatial attention block (940, “Spatial Attn” 940), and a temporal attention block (950, “Temporal Attn” 950).
[0127] Data input to the first up block (581) can be input to the up sampling layer (910).
[0128] The upsampling layer (910) can increase the size (dimension) or resolution of the input data. For example, the upsampling layer (910) can restore the resolution reduced through downsampling to the original size or a larger size. In one embodiment, the upsampling layer (910) can restore the details of the image, thereby improving the quality of the image. The upsampling layer (910) can increase the size (dimension) or resolution of the input data through deconvolution operations, nearest neighbor upsampling, bilinear, bicubic interpolation, pixel shuffling, etc., but is not limited thereto.
[0129] Data output from the upsampling layer (910) can be input to the summation layer (920). In addition, data output from the temporal attention block (730) included in the first down block (561) of FIG. 7 can also be input to the summation layer (920).
[0130] In the summation layer (920), an element-by-element summation operation of data input to the summation layer (920) can be performed.
[0131] Data output from the summation layer (920) can be input to the first residual block (931). Operations illustrated and described in FIG. 4 can be performed in the first residual block (931), and data output from the first residual block (931) can be input to the second residual block (932). Operations illustrated and described in FIG. 4 can also be performed in the second residual block (932), and data output from the second residual block (932) can be input to the spatial attention block (940).
[0132] In the spatial attention block (940), the operations illustrated and described in FIG. 8 can be performed, and data output from the spatial attention block (940) can be input to the temporal attention block (950).
[0133] In the time attention block (950), the operations illustrated and described in FIG. 6 can be performed, and the final output data of the first up block (581) can be obtained.
[0134] Again, referring to FIG. 5, data output from the first up block (581) can be input to the second up block (“Up Block” 582), and the operations illustrated and described in FIG. 9 can be performed in the second up block (582).
[0135] Additionally, the operations illustrated and described in FIG. 9 can also be performed in the third up block (“Up Block” 583) and the fourth up block (“Up Block” 584).
[0136] Data output from the 4th up block (584) can be input to the connection layer (512).
[0137] In the connection layer (512), data output from the convolution layer (521) of FIG. 5 and data output from the fourth up block (584) may be connected. For example, data may be output by connecting data output from the convolution layer (521) of FIG. 5 and data output from the fourth up block (584) in the channel direction.
[0138] Data output from the connection layer (512) can be input to a residual block (“ResBlk 554”). Operations illustrated and described in FIG. 4 can be performed in the residual block (554), and data output from the residual block (554) can be input to a convolution layer (“Conv” 522).
[0139] In the convolution layer (522), a convolution operation between the input data and the kernel can be performed. In the convolution layer (522), flow information can be obtained.
[0140] FIG. 10 is a diagram showing a feature transformation unit and a feature decoding unit according to one embodiment.
[0141] Referring to FIG. 10, according to one embodiment, flow information (502) and feature information (302) extracted from an encoder (300) may be input to a feature transformation unit (220). The flow information (502) may be generated through a flow generation network (500) illustrated and described in FIG. 5. In one embodiment, the flow information (502) may also be generated based on a user input. For example, the flow information may be generated based on flow direction information set by a user. However, the present invention is not limited thereto.
[0142] A feature transformation unit (220) according to one embodiment may include a warping module (1001). The warping module (1001) may apply flow information (502) to feature information (302) to warp the feature information. The warping module (1001) may obtain the transformed feature information.
[0143] The transformed feature information can be input to the feature decoding unit (230).
[0144] A feature decoding unit (230) according to one embodiment may include a decoder (1002) capable of restoring transformed feature information to its original resolution. Accordingly, the decoder (1002) may obtain transformed frame images (1003).
[0145] Referring to FIG. 10, the decoder (1002) may include one or more convolution layers, residual blocks, upsampling layers, and normalization layers.
[0146] The transformed feature information output from the warping module (1001) can be input to the first convolution layer (“Conv” 1011) of the decoder (1002). In the first convolution layer (1011), a convolution operation between the input data and the kernel can be performed.
[0147] Data output from the first convolution layer (1011) can be input to the first residual block (“ResBlk” 1021). Operations illustrated and described in FIG. 4 can be performed in the first residual block (1021), and data output from the first residual block (1021) can be input to the second residual block (“ResBlk” 1022). The operations illustrated and described in FIG. 4 can also be performed in the second residual block (1022), the third residual block (1023, “ResBlk” 1023), the fourth residual block (1024, “ResBlk” 1024), and the fifth residual block (1025, “ResBlk” 1025), and the data output from the fifth residual block (1025) can be input to the first upsampling layer (1031, “Upsample” 1031).
[0148] In the first up-sampling layer (1031), an operation may be performed to increase the size (dimension) or resolution of the input data. For example, in the first up-sampling layer (1031), a deconvolution operation, nearest neighbor up-sampling, bilinear, bicubic interpolation, pixel shuffling, etc. may be performed.
[0149] Data output from the first up-sampling layer (1031) can be input to the second convolution layer (“Conv” 1012). In the second convolution layer (1012), a convolution operation can be performed between the input data and the kernel included in the second convolution layer (1012).
[0150] Data output from the second convolution layer (1012) can be input to the sixth residual block (“ResBlk” 1026).
[0151] The operations illustrated and described in FIG. 4 can also be performed in the sixth to eighth residual blocks (1026, 1027, 1028, “ResBlk” 1026, 1027, 1028), and the data output from the eighth residual block (1028) can be input to the second up-sampling layer (1032, “Upsample” 1032).
[0152] In the second up-sampling layer (1032), an operation may be performed to increase the size (dimension) or resolution of the input data. For example, in the second up-sampling layer (1032), a deconvolution operation, nearest neighbor up-sampling, bilinear, bicubic interpolation, pixel shuffling, etc. may be performed.
[0153] Data output from the second up-sampling layer (1032) can be input to the third convolution layer (“Conv” 1013). In the third convolution layer (1013), a convolution operation can be performed between the input data and the kernel included in the third convolution layer (1013).
[0154] Data output from the third convolution layer (1013) can be input to the ninth residual block (“ResBlk” 1029).
[0155] The operations illustrated and described in FIG. 4 can also be performed in the 9th to 11th residual blocks (1029, 1051, 1052, “ResBlk” 1029, 1051, 1052), and the data output from the 11th residual block (1052) can be input to the normalization layer (“Norm” 1040).
[0156] In the normalization layer (1040), the range of values of data input to the normalization layer (1040) can be adjusted. In the normalization layer (1040), batch normalization, layer normalization, instance normalization, group normalization, etc. can be performed.
[0157] Data output from the normalization layer (1040) can be input to the fourth convolution layer (“Conv” 1014). In the fourth convolution layer (1014), a convolution operation can be performed between the input data and the kernel included in the fourth convolution layer (1014).
[0158] In the fourth convolution layer (1014), transformed frame images (1003) can be obtained.
[0159] Fig. 11 is a drawing showing a residual generation unit according to one embodiment.
[0160] Referring to FIG. 11, the residual generation unit (260) may include a residual generation network (1100). The residual generation network (1100) may include a U-Net having an encoder-decoder structure.
[0161] The residual generation network (1100) may include one or more concatenation layers, one or more convolution layers, a temporal attention block, a spatial attention block, a normalization layer, a residual block, a down block, and an upsample block.
[0162] Noise information (1101), a first image (301), and transformed frame images (1003) generated by a noise generation unit (240) can be input to a residual generation network (1100). The noise information (1101), the first image (301), and transformed frame images (1003) can be input to a first connection layer (1111).
[0163] In the first connection layer (1111), data input to the first connection layer (1111) can be concatenated. For example, when noise information (1101), a first image (301), and transformed frame images (1003) are input to the first connection layer (1111), input data that concatenates the noise information (1101), the first image (301), and the transformed frame images (1003) in the channel direction can be output from the first connection layer (1111).
[0164] In the first convolution layer (“Conv” 1121), a convolution operation can be performed between the input data and the kernel included in the first convolution layer (“Conv”).
[0165] Data output from the first convolution layer (1121) can be input to the first temporal attention block (“Temporal Attn” 1131). In addition, data output from the first convolution layer (1121) can be input to the second connection layer (1112) of FIG. 11.
[0166] In the first time attention block (1131), attention can be applied within the time axis for each location of pixels included in the input data. Here, attention can mean finding related features using similarity between features, integrating related features into one (aggregation), and extracting integrated feature information.
[0167] In the first time attention block (1131), the operations illustrated and described in Fig. 6 can be performed. Data output from the first time attention block (1131) can be input to a normalization layer (“Norm” 1140).
[0168] In the normalization layer (1140), the range of values of data input to the normalization layer (1140) can be adjusted. In the normalization layer (1140), batch normalization, layer normalization, instance normalization, group normalization, etc. can be performed.
[0169] Data output from the normalization layer (1140) can be input to the first residual block (“ResBlk” 1151).
[0170] In the first residual block (1151), the operations illustrated and described in FIG. 4 can be performed, and data output from the first residual block (1151) can be input to the first down block (“Down Block” 1161).
[0171] In the first down block (1161), the operations illustrated and described in FIG. 7 can be performed, and data output from the first down block (1161) can be input to the second down block (1162).
[0172] The operations illustrated and described in FIG. 7 can also be performed in the second to fourth down blocks (1162, 1163, 1164, “Down Block” 1162, 1163, 1164).
[0173] Data output from the fourth down block (1164) can be input to the second residual block (“ResBlk” 1152). In the second residual block (1152), the operations illustrated and described in FIG. 4 can be performed.
[0174] Data output from the second residual block (1152) can be input to a spatial attention block (“Spatial Attn” 1170). The spatial attention block (1170) can perform operations illustrated and described in FIG. 8.
[0175] Data output from the spatial attention block (1170) can be input to the third residual block (“ResBlk” 1153). The operations illustrated and described in FIG. 4 can be performed in the third residual block (1153).
[0176] Data output from the third residual block (1153) can be input to the second temporal attention block (“Temporal Attn” 1132). In the second temporal attention block (1132), the operations illustrated and described in FIG. 6 can be performed.
[0177] Data output from the second time attention block (1132) can be input to the first up block (1181). In the first up block (“Up Block” 1181), operations illustrated and described in FIG. 9 can be performed, and data output from the first up block (1181) can be input to the second up block (“Up Block” 1182).
[0178] The operations illustrated and described in FIG. 9 can also be performed in the second to fourth up blocks (“Up Block” 1182, 1183, 1184), and data output from the fourth up block (1184) can be input to the second connection layer (1112).
[0179] In the second connection layer (1112), the data output from the fourth up block (1184) and the data output from the first convolution layer (1121) can be connected in the channel direction to be output to the fourth residual block (“ResBlk” 1154).
[0180] In the fourth residual block (1154), the operations illustrated and described in FIG. 4 can be performed, and the data output from the fourth residual block (1154) can be input to the second convolution layer (“Conv” 1122).
[0181] In the second convolution layer (1122), a convolution operation can be performed between the input data and the kernel included in the second convolution layer (1122). In the second convolution layer (1122), residual information (1103) can be obtained.
[0182] Fig. 12 is a drawing showing a residual synthesis unit according to one embodiment.
[0183] Referring to FIG. 12, a residual synthesis unit (270) according to one embodiment may include a summation layer. In the summation layer (1210), an element-wise summation operation of data input to the summation layer (1210) may be performed. In the summation layer (1210), an element-wise summation operation of residual information (1103) generated by the residual generation unit (260) and transformed frame images (1003) output from the feature decoding unit (230) may be performed. The summation layer (1210) may obtain a plurality of frame images (1220).
[0184] The plurality of frame images (1220) may be images that maintain details included in the first image (301) (e.g., structural information such as edges included in the first image, detailed texture information, etc.). In addition, the plurality of frame images may be high-quality images that do not include artifacts due to warping and whose image quality is not degraded compared to the first image (301).
[0185] Fig. 13 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0186] Referring to FIG. 13, an electronic device (100) according to one embodiment can obtain flow information for a first image (S1310).
[0187] For example, the electronic device (100) can obtain flow information for a first image using the flow generation network (500) illustrated and described in FIG. 5. At this time, noise information generated by the noise generation unit and feature information for the first image can be input to the flow generation network (500).
[0188] In one embodiment, the electronic device (100) may generate flow information based on user input. For example, if a user input is received that sets the direction of motion appearing in a video, the electronic device (100) may generate flow information so that motion occurs in that direction.
[0189] In one embodiment, the electronic device (100) may segment the selected object based on a user input selecting a specific object from the first image so that motion occurs only for the selected object, thereby generating flow information only for the selected object. However, the present invention is not limited thereto.
[0190] An electronic device (100) according to one embodiment can obtain a plurality of transformed images that have transformed a first image based on flow information (S1320).
[0191] The electronic device (100) can extract feature information of the first image. For example, the electronic device (100) can extract feature information of the first image using the encoder (300) illustrated and described in FIG. 3.
[0192] The electronic device (100) can generate transformed feature information by applying flow information to the extracted feature information and warping it.
[0193] The electronic device (100) can obtain a plurality of transformed images by decoding transformed feature information using the decoder (1002) illustrated and described in FIG. 10.
[0194] An electronic device (100) according to one embodiment can obtain residual information for a plurality of transformed images based on a plurality of transformed images and a first image (S1330).
[0195] The electronic device (100) can generate residual information for a plurality of transformed images using the residual generation network (1100) illustrated and described in FIG. 11. At this time, the noise information generated by the noise generation unit, the first image, and the plurality of transformed images can be input to the residual generation network (1100).
[0196] An electronic device (100) according to one embodiment can generate a plurality of frame images based on a plurality of transformed images and residual information (S1340).
[0197] The electronic device (100) can generate multiple frame images by adding multiple transformed images and residual information.
[0198] FIG. 14 is a diagram for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0199] Referring to FIG. 14, an electronic device (100) according to one embodiment can display multiple images on a display. For example, the electronic device (100) can execute a photo application based on a user input requesting execution of the photo application. When the photo application is executed, the electronic device (100) can display images previously stored in the electronic device (100).
[0200] When the electronic device (100) receives a user input for selecting a first image (1410) from among a plurality of images, the electronic device (100) can enlarge and display the selected first image (1410). Based on the user input, the electronic device (100) can select a portion of the first image (1410) or edit the properties of the first image (1410). In addition, the electronic device (100) can convert the first image (1410) into a video (1430) based on a user input for selecting a video generation menu (1420). For example, the electronic device (100) can generate a plurality of frame images included in the video (1430) by performing image processing on the first image (1410) using the image processing module (200) illustrated and described in FIGS. 2 to 12. The electronic device (100) can extract feature information from the first image (1410) using the feature extraction unit (210). The electronic device (100) can generate first noise images using the noise generation unit (240). The electronic device (100) can generate flow information (flow map) based on the feature information extracted by the feature extraction unit (210) and the first noise images using the flow generation unit (250). The electronic device (100) can generate transformed feature information by applying the flow information to the extracted feature information and warping it using the feature transformation unit (220). The electronic device (100) can generate transformed images by decoding the transformed feature information using the feature decoding unit (230).
[0201] The electronic device (100) can generate second noise images using the noise generation unit (240). The electronic device (100) can generate residual information (residual map) based on the transformed images, the first image (1410), and the second noise images using the residual generation unit (260). The electronic device (100) can generate a plurality of frame images by synthesizing the residual information and the transformed images using the residual synthesis unit (270).
[0202] The electronic device (100) can display a video (1430) including a plurality of generated frame images on a display.
[0203] FIG. 15 is a drawing for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video.
[0204] Referring to FIG. 15, an electronic device (100) according to one embodiment can display multiple images on a display. For example, the electronic device (100) can execute a photo application based on a user input requesting execution of the photo application. When the photo application is executed, the electronic device (100) can display images previously stored in the electronic device (100).
[0205] When the electronic device (100) receives a user input for selecting a first image (1510) from among a plurality of images, the electronic device (100) can enlarge and display the selected first image (1510). Based on the user input, the electronic device (100) can select a portion of the first image (1510) or edit the properties of the first image (1510).
[0206] In addition, the electronic device (100) can set a motion for the first image (1510) based on a user input that selects a motion setting menu (1520). The user can set the direction of the motion with a drag input. For example, the electronic device (100) can receive a touch input that drags from left to right and set the direction of the motion to the first direction (1550). Setting the direction of the motion to the first direction (1550) may mean setting a motion in which at least one object included in the first image (1510) moves in the first direction (1550), or moving a viewpoint looking at at least one object included in the first image (1510) in the first direction (1550). However, the present invention is not limited thereto.
[0207] The electronic device (100) can obtain flow information based on the direction of the set motion. For example, the electronic device (100) can generate flow information (1560) indicating a first direction.
[0208] In one embodiment, the electronic device (100) may convert the first image (1510) into a video (1540) based on a user input of selecting a video generation menu (1530). For example, the electronic device (100) may generate a plurality of frame images included in the video (1540) by performing image processing on the first image (1510) using the flow information (1560) and the image processing module (200) illustrated and described in FIGS. 2 to 12. The electronic device (100) may extract feature information from the first image (1510) using the feature extraction unit (210). The electronic device (100) may apply the flow information (1560) generated based on a motion input by the user to the extracted feature information using the feature transformation unit (220), thereby generating transformed feature information. For example, the electronic device (100) may utilize flow information (1560) indicating the first direction (1550) without generating flow information through a flow generation network. The electronic device (100) may obtain transformed feature information by warping the flow information (1560) indicating the first direction (1550) by applying the flow information (1560) indicating the first direction (1550) to feature information extracted from the first image (1510). The electronic device (100) may generate transformed images by decoding the transformed feature information using the feature decoding unit (230).
[0209] The electronic device (100) can generate noise images using the noise generation unit (240). The electronic device (100) can generate residual information (residual map) based on the transformed images, the first image (1510), and the noise images using the residual generation unit (260). The electronic device (100) can generate a plurality of frame images by synthesizing the residual information and the transformed images using the residual synthesis unit (270).
[0210] The electronic device (100) may display a video (1540) including a plurality of generated frame images on a display. The generated video (1540) may include a motion in which at least one object included in the first image (1310) moves in a first direction (1550). In one embodiment, the generated video (1540) may include a motion in which a viewpoint looking at at least one object included in the first image (1510) moves in the first direction (1550). However, the present invention is not limited thereto.
[0211] FIG. 16 is a diagram for explaining an operation of an electronic device according to one embodiment of the present invention to convert a still image into a video image.
[0212] Referring to FIG. 16, an electronic device (100) according to one embodiment can generate a moving image in which only a specific object among objects included in a still image moves.
[0213] For example, the electronic device (100) may display a first image (1610) on the display. The first image (1610) may be an image selected based on a user input. The first image (1610) may include at least one object. The electronic device (100) may receive a user input for selecting a first object (1620) from among the objects included in the first image (1610).
[0214] The electronic device (100) can segment a first object (1620) through object segmentation. Object segmentation may refer to identifying and separating a specific object or a specific region in an image. Object segmentation can be used to identify the boundary of an object in an image at the pixel level, assign different colors to each object, or separate the object from the background. The electronic device (100) can generate a mask (1630) representing a first object region in the first image (1610) through object segmentation.
[0215] The electronic device (100) can convert the first image (1610) into a video (1640) based on a user input requesting video generation. The electronic device (100) can generate a video (1640) in which only the first object (1620) moves in the first image (1610). For example, the electronic device (100) can generate a plurality of frame images included in the video (1640) by performing image processing on the first image (1610) using the image processing module (200) illustrated and described in FIGS. 2 to 12.
[0216] The electronic device (100) can extract feature information from the first image (1610) using the feature extraction unit (210). The electronic device (100) can generate first noise images using the noise generation unit (240). The electronic device (100) can generate flow information (flow map) based on the feature information extracted by the feature extraction unit (210) and the first noise images using the flow generation unit (250). At this time, the electronic device (100) can apply a mask (1630) indicating a first object area to the flow information generated through the flow generation network. The flow information to which the mask (1630) is applied can only indicate flow information of the first object area. The electronic device (100) can generate transformed feature information by applying the flow information to which the mask (1630) is applied to the extracted feature information using the feature transformation unit (220) and warping it. The electronic device (100) can generate transformed images by decoding transformed feature information using the feature decoding unit (230).
[0217] The electronic device (100) can generate second noise images using the noise generation unit (240). The electronic device (100) can generate residual information (e.g., a residual map) based on the transformed images, the first image (1410), and the second noise images using the residual generation unit (260). At this time, the electronic device (100) can apply a mask (1630) indicating a first object area to the residual information generated through the residual generation network. The electronic device (100) can generate a video (1640) including a plurality of frame images by synthesizing the transformed images with the residual information to which the mask (1630) is applied using the residual synthesis unit (270).
[0218] Accordingly, the electronic device (100) can generate a video (1640) in which only the first object moves. The electronic device (100) can display the generated video (1640) on a display.
[0219] FIG. 17 is a drawing for explaining an operation of an electronic device according to an embodiment to convert a still image into a video.
[0220] Referring to FIG. 17, an electronic device (100) according to one embodiment can generate a moving image in which only a specific object among objects included in a still image moves.
[0221] For example, the electronic device (100) may display a first image (1710) on the display. The first image (1710) may be an image selected based on a user input. The first image (1710) may include at least one object. The electronic device (100) may receive a user input for selecting a first object (1720) from among the objects included in the first image (1710).
[0222] The electronic device (100) can segment a first object (1720) through object segmentation. The electronic device (100) can generate a mask (1730) representing a first object area in a first image (1710) through object segmentation.
[0223] Additionally, the electronic device (100) can set a motion for the first object (1720) based on a user input selecting a motion setting menu (1740). The user can set the direction of the motion with a drag input. For example, the electronic device (100) can receive a touch input of dragging from left to right and set the direction of the motion to the first direction (1750). Setting the direction of the motion to the first direction (1750) may mean setting a motion in which the first object (1720) moves in the first direction (1750). However, the present invention is not limited thereto.
[0224] The electronic device (100) can obtain flow information based on the direction of the set motion. For example, the electronic device (100) can generate flow information (1760) indicating a first direction.
[0225] The electronic device (100) can convert the first image (1710) into a video (1780) based on a user input requesting video generation. For example, the electronic device (100) can perform an image processing operation on the first image (1710) using flow information (1560) and the image processing module (200) illustrated and described in FIGS. 2 to 12, thereby generating a plurality of frame images included in the video (1780).
[0226] The electronic device (100) can extract feature information from the first image (1710) using the feature extraction unit (210). The electronic device (100) can apply flow information (1760) generated based on a motion input by the user to the extracted feature information using the feature transformation unit (220), thereby generating transformed feature information. For example, the electronic device (100) can apply a mask (1730) indicating a first object area to flow information (1760) indicating a first direction (1750) without generating flow information through a flow generation network. Flow information (1770) to which the mask (1730) is applied can only indicate flow information of the first object area.
[0227] The electronic device (100) can obtain transformed feature information by applying flow information (1770) to which a mask (1730) is applied to feature information extracted from a first image (1710) using a feature transformation unit (220) and warping the same. The electronic device (100) can generate transformed images by decoding the transformed feature information using a feature decoding unit (230).
[0228] The electronic device (100) can generate noise images using the noise generation unit (240). The electronic device (100) can generate residual information (residual map) based on the transformed images, the first image (1710), and the noise images using the residual generation unit (260). At this time, the electronic device (100) can apply a mask (1730) indicating a first object area to the residual information generated through the residual generation network. The electronic device (100) can generate a video (1780) including a plurality of frame images by synthesizing the transformed images with the residual information to which the mask (1730) is applied using the residual synthesis unit (270).
[0229] Accordingly, the electronic device (100) can generate a video (1780) in which only the first object moves in the first direction. The electronic device (100) can display the generated video (1780) on the display.
[0230] Fig. 18 is a block diagram showing the configuration of an electronic device according to one embodiment.
[0231] The electronic device (100) of FIG. 18 may be a device that performs an image processing operation using an image processing module (200). The image processing module (200) according to one embodiment may include one or more neural networks. For example, the image processing module (200) may include an encoder that converts an input image into a compressed form or extracts features of the input image, a decoder that restores the compressed form to the original resolution, a U-Net having an encoder-decoder structure, etc. However, the present invention is not limited thereto. The image processing module (200) may include various neural networks.
[0232] Referring to FIG. 18, an electronic device (100) according to one embodiment may include a processor (110), a memory (120), and a display (130).
[0233] A processor (110) according to one embodiment can control the overall electronic device (100). A processor (110) according to one embodiment can execute one or more programs stored in a memory (120).
[0234] According to one embodiment, the memory (120) may store various data, programs, or applications for driving and controlling the electronic device (100). The program stored in the memory (120) may include one or more instructions. The program (one or more instructions) or application stored in the memory (120) may be executed by the processor (110).
[0235] A processor (110) according to one embodiment may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a VPU (Video Processing Unit). Alternatively, according to one embodiment, the processor (110) may be implemented in the form of a SoC (System On Chip) that integrates at least one of a CPU, a GPU, and a VPU. The processor (110) according to one embodiment may further include an NPU (Neural Processing Unit).
[0236] According to one embodiment, a processor (110) may generate a plurality of frame images (video) by processing a first image using an image processing module (200) including one or more neural networks. At this time, the plurality of frame images included in the video may be images that maintain details included in the first image (e.g., structural information such as edges and detailed texture information included in the first image). In addition, the plurality of frame images may be high-quality images that do not include artifacts due to warping and whose image quality is not degraded compared to the first image.
[0237] According to one embodiment, the processor (110) can obtain feature information of a first image. For example, the feature information of the first image can be extracted using the encoder (300) illustrated and described in FIG. 3. The structure and operation of the encoder (300) have been described in detail in FIGS. 3 and 4, and therefore, a detailed description thereof will be omitted.
[0238] According to one embodiment, a processor (110) can obtain flow information for a first image. For example, the processor (110) can generate flow information for the first image using a flow generation network (500) illustrated and described in FIG. 5. The structure and operation of the flow generation network (500) have been described in detail in FIGS. 5 to 9, and thus a detailed description thereof will be omitted.
[0239] According to one embodiment, the processor (110) may generate flow information based on user input. For example, upon receiving a user input setting the direction of motion appearing in a video, the processor (110) may generate flow information based on the direction of motion so that the generated video moves in the direction set by the user.
[0240] According to one embodiment, a processor (110) may segment a selected object based on a user input selecting a specific object from a first image so that motion occurs only for the selected object, thereby generating flow information only for the selected object. However, the present invention is not limited thereto.
[0241] According to one embodiment, the processor (110) can obtain a plurality of transformed images by transforming a first image based on flow information. For example, the processor (110) can generate transformed feature information by warping the extracted feature information by applying the flow information. In addition, the processor (110) can obtain a plurality of transformed images by decoding the transformed feature information using the decoder (1002) illustrated and described in FIG. 10. Since the structure and operation of the decoder (1002) have been described in detail in FIG. 10, a detailed description thereof will be omitted.
[0242] According to one embodiment, the processor (110) may obtain residual information for a plurality of transformed images based on a plurality of transformed images and a first image. For example, the processor (110) may generate residual information for a plurality of transformed images using the residual generation network (1100) illustrated and described in FIG. 11. At this time, noise information generated by the noise generation unit, the first image, and the plurality of transformed images may be input to the residual generation network (1100). Since the structure and operation of the residual generation network (1100) have been described in detail in FIG. 11, a detailed description thereof will be omitted.
[0243] According to one embodiment, the processor (110) can generate a plurality of frame images based on a plurality of transformed images and residual information. For example, the processor (110) can generate a plurality of frame images by summing a plurality of transformed images and residual information.
[0244] Meanwhile, the encoder (300), flow generation network (500), decoder (1002), and residual generation network (1100) according to one embodiment may be networks trained by a server or an external device. The external device may train the encoder (300), flow generation network (500), decoder (1002), and residual generation network (1100) based on training data.
[0245] A server or external device can determine, through training, parameter values used in each of the plurality of layers and the plurality of blocks included in the encoder (300), the flow generation network (500), the decoder (1002), and the residual generation network (1100).
[0246] An electronic device (100) according to one embodiment may receive an encoder (300), a flow generation network (500), a decoder (1002), and a residual generation network (1100), which have been trained, from a server or an external device, and store them in a memory (120). For example, the memory (120) may store structures and parameter values of the encoder (300), the flow generation network (500), the decoder (1002), and the residual generation network (1100) according to one embodiment, and the processor (110) may use the parameter values stored in the memory (120) to generate a plurality of frame images (videos) from a first image according to one embodiment.
[0247] A display (130) according to one embodiment converts image signals, data signals, OSD signals, control signals, etc. processed by a processor (110) to generate a driving signal. The display (130) may be implemented as a PDP, LCD, OLED, flexible display, etc., and may also be implemented as a 3D display. In addition, the display (130) may be configured as a touch screen and may be used as an input device in addition to an output device.
[0248] A display (130) according to one embodiment can display a plurality of generated frame images (videos) using an image processing module (200).
[0249] Fig. 19 is a block diagram showing the configuration of an electronic device according to one embodiment.
[0250] The electronic device (1900) of FIG. 19 may be an embodiment of the electronic device (100) illustrated and described in FIGS. 1 to 18.
[0251] Referring to FIG. 19, an electronic device (1900) according to one embodiment may include a sensing unit (1910), a communication unit (1920), a processor (1930), an A / V input unit (1940), an output unit (1950), a memory (1960), and a user input unit (1970).
[0252] The processor (1930) of FIG. 19 corresponds to the processor (110) of FIG. 18, the memory (1960) of FIG. 19 corresponds to the memory (120) of FIG. 18, and the display (1951) of FIG. 19 corresponds to the display (130) of FIG. 18, so the same description will be omitted.
[0253] The sensing unit (1910) may include a sensor that detects the state of the electronic device (1900) or the state around the electronic device (1900). In addition, the sensing unit (1910) may transmit information detected by the sensor to the processor (1930).
[0254] The communication unit (1920) may include, but is not limited to, a short-range wireless communication unit, a mobile communication unit, etc., depending on the performance and structure of the electronic device (1900).
[0255] The short-range wireless communication unit may include, but is not limited to, a Bluetooth communication unit, a BLE (Bluetooth Low Energy) communication unit, a near field communication unit, a WLAN (Wi-Fi) communication unit, a Zigbee communication unit, an infrared (IrDA, infrared Data Association) communication unit, a WFD (Wi-Fi Direct) communication unit, an UWB (ultra wideband) communication unit, an ANT+ communication unit, a microwave (uWave) communication unit, etc.
[0256] The mobile communication unit transmits and receives wireless signals with at least one of a base station, an external terminal, or a server on a mobile communication network. Here, the wireless signals may include various forms of data, such as voice call signals, video call signals, or text / multimedia message transmission and reception.
[0257] According to one embodiment, the communication unit (1920) can receive an image from an external device or transmit an image.
[0258] A processor (1930) according to one embodiment can convert a first image into a plurality of frame images (videos) using an image processing module (200) according to one embodiment.
[0259] A processor (1930) according to one embodiment may include a single core, a dual core, a triple core, a quad core, and a multiple thereof. Additionally, the processor (1930) may include multiple processors.
[0260] The memory (1960) according to one embodiment may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk.
[0261] The A / V (Audio / Video) input unit (1940) is for inputting audio signals or video signals, and may include a camera (1941) and a microphone (1942). The camera (1941) can obtain image frames, such as still images or moving images, through an image sensor in video call mode or shooting mode. Images captured through the image sensor can be processed through a processor (1930) or a separate image processing unit.
[0262] The image frame processed by the camera (1941) can be stored in the memory (1960) or transmitted externally through the communication unit (1920). Two or more cameras (1941) may be provided depending on the configuration of the electronic device (1900).
[0263] The microphone (1942) receives external acoustic signals and processes them into electrical voice data. For example, the microphone (1942) can receive acoustic signals from an external device or a speaker. The microphone (1942) can utilize various noise removal algorithms to remove noise generated during the process of receiving external acoustic signals.
[0264] The output unit (1950) is for outputting an audio signal, a video signal, or a vibration signal, and may include a display (1951), an audio output unit (1952), a vibration motor (1953), etc.
[0265] A display (1951) according to one embodiment can display a plurality of frame images (videos) generated through an image processing module (200).
[0266] The audio output unit (1952) outputs audio data received from the communication unit (1920) or stored in the memory (1960). In addition, the audio output unit (1952) outputs audio signals related to functions performed in the electronic device (1900) (e.g., call signal reception sound, message reception sound, notification sound). The audio output unit (1952) may include a speaker, a buzzer, or the like.
[0267] The vibration unit (1953) can output a vibration signal. For example, the vibration unit (1953) can output a vibration signal corresponding to the output of audio data or video data (e.g., a call signal reception sound, a message reception sound, etc.). In addition, the vibration unit (1953) can also output a vibration signal when a touch is input to the touch screen.
[0268] The user input unit (1970) refers to a means for a user to input data for controlling the electronic device (1900). For example, the user input unit (1970) may include, but is not limited to, a key pad, a dome switch, a touch pad (contact electrostatic capacitance type, pressure resistive film type, infrared detection type, surface ultrasonic conduction type, integral tension measurement type, piezo effect type, etc.), a jog wheel, a jog switch, etc.
[0269] Meanwhile, the block diagrams of the electronic devices (100, 1900) illustrated in FIGS. 18 and 19 are block diagrams for one embodiment. Each component of the block diagram may be integrated, added, or omitted depending on the specifications of the electronic devices (100, 1900) actually implemented. That is, two or more components may be combined into one component, or one component may be subdivided into two or more components, as needed. In addition, the functions performed by each block are for explaining embodiments, and the specific operations or devices thereof do not limit the scope of the present invention.
[0270] An electronic device according to one embodiment may include at least one processor including a memory storing one or more instructions and a processing circuit.
[0271] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain flow information for the first image.
[0272] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain a plurality of modified images that have modified the first image based on the flow information.
[0273] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain residual information about the plurality of transformed images based on the plurality of transformed images and the first image.
[0274] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate a plurality of frame images based on the plurality of transformed images and the residual information.
[0275] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can extract feature information of the first image.
[0276] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate the flow information based on the feature information.
[0277] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate first noise information.
[0278] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can perform a first image processing operation on the first noise information and the feature information to generate the flow information.
[0279] The above first image processing operation can be performed by a U-Net network including an encoder and a decoder.
[0280] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can extract feature information of the first image.
[0281] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain the plurality of transformed images by warping the feature information based on the flow information.
[0282] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate second noise information.
[0283] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can perform a second image processing operation on the second noise information, the plurality of transformed images, and the first image to generate the residual information.
[0284] The above second image processing operation can be performed by a U-Net network including an encoder and a decoder.
[0285] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can output a moving image including the plurality of frame images.
[0286] The electronic device may further include a display.
[0287] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can control the display to display a plurality of images.
[0288] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device may receive a first user input for selecting the first image from among the plurality of images and a second user input for requesting generation of a video for the first image.
[0289] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate the plurality of frame images based on the second user input.
[0290] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device may receive a third user input that sets motion information for the first image.
[0291] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate the flow information based on the third user input.
[0292] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device may receive a fourth user input selecting a first object from among at least one object included in the first image.
[0293] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can generate the plurality of frame images in which only the first object among the first images moves, based on the fourth user input.
[0294] When the one or more instructions are individually or collectively executed by at least one processor according to one embodiment, the electronic device can obtain the plurality of transformed images by applying the flow information only to the first object.
[0295] A method of operating an electronic device according to one embodiment may include a step of obtaining flow information for a first image.
[0296] A method of operating an electronic device according to one embodiment may include a step of obtaining a plurality of transformed images obtained by transforming the first image based on the flow information.
[0297] A method of operating an electronic device according to one embodiment may include a step of obtaining residual information for the plurality of transformed images based on the plurality of transformed images and the first image.
[0298] A method of operating an electronic device according to one embodiment may include a step of generating a plurality of frame images based on the plurality of transformed images and the residual information.
[0299] The step of obtaining the above flow information may include a step of extracting feature information of the first image.
[0300] The step of obtaining the above flow information may include a step of generating the flow information based on the feature information.
[0301] Based on the above characteristic information, the step of generating the flow information may include the step of generating first noise information.
[0302] Based on the above characteristic information, the step of generating the flow information may include the step of generating the flow information by performing a first image processing operation on the first noise information and the characteristic information.
[0303] The above first image processing operation can be performed by a U-Net network including an encoder and a decoder.
[0304] The step of obtaining the plurality of transformed images may include a step of extracting feature information of the first image.
[0305] The step of obtaining the plurality of transformed images may include a step of obtaining the plurality of transformed images by warping the feature information based on the flow information.
[0306] The step of obtaining the residual information may include a step of generating second noise information.
[0307] The step of obtaining the residual information may include a step of generating the residual information by performing a second image processing operation on the second noise information, the plurality of transformed images, and the first image.
[0308] The above second image processing operation can be performed by a U-Net network including an encoder and a decoder.
[0309] The above operating method may further include a step of outputting a video including the plurality of frame images.
[0310] The above method of operation may include a step of displaying a plurality of images.
[0311] The above operating method may further include a step of receiving a first user input for selecting the first image from among the plurality of images and a second user input for requesting generation of a video for the first image.
[0312] The step of generating the plurality of frame images may include the step of generating the plurality of frame images based on the second user input.
[0313] The step of obtaining the above flow information may include the step of receiving a third user input that sets motion information for the first image.
[0314] The step of obtaining the above flow information may include a step of generating the flow information based on the third user input.
[0315] The above operating method may further include a step of receiving a fourth user input selecting a first object from among at least one object included in the first image.
[0316] The step of obtaining the above flow information may include a step of obtaining flow information for the first object.
[0317] Based on the flow information, the step of obtaining a plurality of transformed images by transforming the first image may include a step of obtaining the plurality of transformed images by applying the flow information for the first object only to the first object among the first images.
[0318] The above plurality of frame images may include a plurality of frame images in which only the first object among the first images moves.
[0319] An operating method of an electronic device according to one embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the medium may be those specially designed and configured for the present invention or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0320] Additionally, the image processing device and the operating method thereof according to the disclosed embodiments may be provided as a computer program product. The computer program product may be traded as a commodity between sellers and buyers.
[0321] A computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, a computer program product may include a product in the form of a software program (e.g., a downloadable app) distributed electronically by an electronic device manufacturer or through an electronic marketplace (e.g., Google Play Store, App Store). For electronic distribution, at least a portion of the software program may be stored on a storage medium or temporarily created. In this case, the storage medium may be a storage medium of a manufacturer's server, an electronic marketplace server, or a relay server that temporarily stores the software program.
[0322] In a system comprising a server and a client device, the computer program product may include a storage medium of the server or a storage medium of the client device. In one embodiment, if there is a third device (e.g., a smartphone) that is communicatively connected to the server or the client device, the computer program product may include a storage medium of the third device. In one embodiment, the computer program product may include a software program itself that is transmitted from the server to the client device or the third device, or from the third device to the client device.
[0323] In this case, one of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments. In one embodiment, two or more of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments in a distributed manner.
[0324] For example, a server (e.g., a cloud server or an artificial intelligence server, etc.) may execute a computer program product stored on the server, thereby controlling a client device in communication with the server to perform a method according to the disclosed embodiments.
[0325] The scope of the present invention is not limited to the embodiments described above. Various modifications and improvements made by those skilled in the art utilizing the basic concepts of the present invention defined in the following claims also fall within the scope of the present invention.
Claims
1. In an electronic device (100) that converts a still image into a video, A memory (130) storing one or more instructions; and comprising at least one processor (120), When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Obtain flow information for the first image, Based on the flow information for the first image, a plurality of transformed images are obtained by transforming the first image, Based on the plurality of transformed images and the first image, residual information for the plurality of transformed images is obtained, An electronic device that generates a plurality of frame images based on the plurality of transformed images and the residual information for the plurality of transformed images.
2. In paragraph 1, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Extracting feature information of the first image, An electronic device that generates the flow information for the first image based on the above characteristic information.
3. In paragraph 2, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Generate the first noise information, By performing a first image processing operation on the first noise information and the feature information, the flow information for the first image is generated, An electronic device wherein the first image processing operation is performed by a U-Net network including an encoder and a decoder.
4. In any one of paragraphs 1 to 3, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Extracting feature information of the first image, An electronic device that obtains the plurality of deformed images by warping the feature information based on the flow information for the first image.
5. In any one of paragraphs 1 to 4, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Generate second noise information, Performing a second image processing operation on the second noise information, the plurality of transformed images, and the first image to generate the residual information for the plurality of transformed images, An electronic device wherein the second image processing operation is performed by a U-Net network including an encoder and a decoder.
6. In any one of paragraphs 1 to 5, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), An electronic device that outputs a video including the plurality of frame images.
7. In any one of paragraphs 1 to 6, The above electronic device, Including more displays, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Controlling the display to display multiple images, Receiving a first user input for selecting the first image among the plurality of images, and receiving a second user input for requesting video generation for the first image, An electronic device that generates the plurality of frame images based on the second user input.
8. In any one of paragraphs 1 to 7, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Receive a third user input that sets motion information for the first image, An electronic device that generates the flow information for the first image based on the third user input.
9. In any one of paragraphs 1 to 8, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), Receiving a fourth user input selecting a first object from among at least one object included in the first image; An electronic device that generates a plurality of frame images in which only the first object moves based on the fourth user input.
10. In paragraph 9, When the one or more instructions are individually or collectively executed by the at least one processor (120), the electronic device (100), An electronic device that obtains the plurality of transformed images by applying the flow information for the first image only to the first object.
11. In a method of operation performed by an electronic device (100) for converting a still image into a video, Step of obtaining flow information for the first image (S1310); A step (S1320) of obtaining a plurality of transformed images of the first image based on the flow information for the first image; A step (S1330) of obtaining residual information for the plurality of transformed images based on the plurality of transformed images and the first image; and An operating method of an electronic device, comprising a step (S1340) of generating a plurality of frame images based on the plurality of transformed images and the residual information for the plurality of transformed images.
12. In paragraph 11, The step of obtaining the flow information for the first image comprises: A step of extracting feature information of the first image; and An operating method of an electronic device, comprising the step of generating the flow information for the first image based on the feature information.
13. In paragraph 12, Based on the above feature information, the step of generating the flow information for the first image is: a step of generating first noise information; and A step of generating the flow information for the first image by performing a first image processing operation on the first noise information and the feature information, An operating method of an electronic device, wherein the first image processing operation is performed by a U-Net network including an encoder and a decoder.
14. In any one of paragraphs 11 to 13, The step of obtaining the above multiple transformed images is: A step of extracting feature information of the first image; and An operating method of an electronic device, comprising a step of obtaining the plurality of deformed images by warping the feature information based on the flow information for the first image.
15. One or more computer-readable recording media storing a program for performing any one of the methods of Articles 11 to 14.