Image restoration method and apparatus based on continuous shot images

The method and apparatus enhance image quality in continuous-shot videos by using an anchor image within a neural network model to emphasize anchor information, addressing the computational challenges of center line alignment and improving image restoration efficiency.

JP7806382B2Active Publication Date: 2026-01-27SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022000881
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-08
Filing Date
2022-01-06
Publication Date
2026-01-27
Estimated Expiration
2042-01-06

AI Technical Summary

Technical Problem

Existing video restoration methods face challenges in efficiently improving image quality of continuous-shot videos without the need for center line alignment, which increases computational complexity and time.

Method used

A method and apparatus that utilize an anchor image or video within a set of continuous-shot images, employing a neural network model to generate a restored image or video by emphasizing anchor information, thereby bypassing the need for center line alignment and reducing computational demands.

Benefits of technology

The approach effectively enhances image quality of continuous-shot videos by combining image information at corresponding positions, suppressing blurring, and reducing calculation time and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007806382000001
    Figure 0007806382000001
  • Figure 0007806382000002
    Figure 0007806382000002
  • Figure 0007806382000003
    Figure 0007806382000003
Patent Text Reader

Abstract

To provide a continuously shot image-based image restoration method and apparatus.SOLUTION: The continuously shot image-based image restoration method and apparatus are disclosed. In accordance with one embodiment, the image restoration method includes the steps of: determining an anchor image based on individual images of a continuously shot image set; executing a feature extraction network based on an anchor image set and based on the continuously shot image set; and generating a restored image based on a feature map corresponding to an output of the feature extraction network.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The following embodiments relate to a method and apparatus for restoring video based on continuous-shot videos. [Background technology]

[0002] Video restoration is a technique for restoring degraded video to video with improved image quality. Deep learning-based neural networks can be used for video restoration. After being trained based on deep learning, neural networks can perform tailored inference by mapping nonlinearly related input and output data to each other. The ability to generate such mappings is similar to the learning ability of neural networks. Furthermore, neural networks trained for specialized purposes such as video restoration may have generalization capabilities, such as generating relatively accurate outputs even for input patterns not included in the training data. Summary of the Invention [Problem to be solved by the invention]

[0003] SUMMARY OF THE INVENTION An object of the present invention is to provide a method and apparatus for restoring an image based on continuous-shot images. [Means for solving the problem]

[0004] According to one embodiment, a video restoration method includes the steps of determining an anchor video based on an individual video of a set of continuous-shot videos, running a feature extraction network based on the set of continuous-shot videos using anchor information of the anchor video, and generating a restored video based on a feature map corresponding to the output of the feature extraction network.

[0005] According to one embodiment, the video restoration device includes a processor and a memory containing instructions executable by the processor, and when the instructions are executed by the processor, the processor determines an anchor image based on an individual image of a set of continuous shot images, executes a feature extraction network based on the set of continuous shot images using anchor information of the anchor image, and generates a restored image based on a feature map corresponding to the output of the feature extraction network.

[0006] According to one embodiment, an electronic device includes a camera that generates a set of continuous shots, and a processor that determines an anchor video based on individual videos of the set of continuous shots, runs a feature extraction network based on the set of continuous shots using anchor information of the anchor video, and generates a restored video based on a feature map corresponding to the output of the feature extraction network. [Effects of the Invention]

[0007] According to the present invention, a method and apparatus for image restoration based on continuous-shot images can be provided. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an operation of a video restoration device according to an embodiment. [Figure 2] 10 illustrates an operation for selecting an anchor video according to various embodiments. [Figure 3] 10 illustrates an operation for selecting an anchor video according to various embodiments. [Figure 4] 10 illustrates an operation for selecting an anchor video according to various embodiments. [Figure 5] 10 illustrates an operation of generating restored video using anchor information according to an embodiment. [Figure 6] 1 illustrates the configuration and operation associated with a neural network model according to one embodiment. [Figure 7] 10 illustrates an operation of using anchor information in an input video input process according to an embodiment. [Figure 8]10 illustrates an operation of using anchor information in the process of outputting an output feature map according to an embodiment. [Figure 9] 10 illustrates an operation of using anchor information in the process of extracting global features according to one embodiment. [Figure 10] An example of the operation shown in FIG. 9 is shown. [Figure 11] 10 illustrates an operation of using anchor information in the process of extracting local features according to one embodiment. [Figure 12] 10 is a flowchart illustrating a video reset operation according to an embodiment. [Figure 13] 1 is a block diagram showing a configuration of a video decompression device according to an embodiment. [Figure 14] 1 is a block diagram showing a configuration of an electronic device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] The specific structural or functional descriptions disclosed in this specification are merely examples for the purpose of describing the embodiments, and the embodiments may be implemented in various different forms. The present invention is not limited to the embodiments described in this specification, and the scope of the present invention includes modifications, equivalents, or alternatives that fall within the technical idea described in the embodiments.

[0010] Although terms such as "first" or "second" may be used to describe multiple components, such terms should be construed only to distinguish one component from the other components. For example, a first component may be designated as a second component, and similarly, a second component may be designated as a first component.

[0011] When a component is referred to as being "coupled" or "connected" to another component, it should be understood that although it is directly coupled or connected to the other component, there may be other components in between.

[0012] The singular expression includes the plural expression unless the context clearly dictates otherwise. In this specification, the words "comprise" or "have" and the like indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0013] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention pertains. Commonly used predefined terms should be interpreted as having a meaning consistent with the meaning they have in the context of the relevant art, and should not be interpreted as having an ideal or overly formal meaning unless expressly defined herein.

[0014] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. In the description with reference to the accompanying drawings, the same reference numerals will be used to designate the same components regardless of the reference numerals, and redundant description thereof will be omitted.

[0015] FIG. 1 schematically illustrates the operation of an image restoration apparatus according to an embodiment. Referring to FIG. 1, an image restoration apparatus 100 receives a burst image set 101, generates a restored image 102 based on the burst image set 101, and outputs the restored image 102. The burst image set 101 may be generated by a camera (not shown). The burst image set 101 may include a plurality of images captured continuously. For example, the burst image set 101 may be video images generated through a video capture function or a series of still images generated through a burst capture function. Each of the images in the burst image set 101 may be referred to as an individual image. In the case of a video, each image frame corresponds to an individual image, and in the case of a burst image, each still image corresponds to an individual image.

[0016] Assuming that a target object is photographed using a camera to generate a continuous-shot video set 101, each individual video in the continuous-shot video set 101 may have different characteristics due to the movement of the camera and / or the target object, and / or changes in ambient light (e.g., illuminance, color, etc.). If the continuous-shot video set 101 is photographed in a poor environment, such as a low-light environment, and / or each individual video has degraded image quality, a restored video 102 with improved image quality can be derived by appropriately combining various characteristics of each individual video. Therefore, a restored video 102 with high image quality can be derived by performing restoration on individual videos with low image quality.

[0017] Because the position of an object changes in each individual image depending on the movement of the camera and / or the target object, a pre-processing step is required to align the centerlines of each individual image and match the objects in each individual image. The centerlines are not actual lines displayed in each individual image, but virtual lines used as a reference for aligning the individual images. If this pre-processing step is not performed, blurring may become severe, and performing this pre-processing step may significantly increase the amount and time of calculations. Because the centerline alignment step requires iterative processing, the increase in the amount and time of calculations increases as the number of individual images increases.

[0018] The video restoration device 100 may determine an anchor image based on individual images of a continuous-shot image set 101 and generate a restored image 102 by executing a neural network model using anchor information of the anchor image. For example, generating the restored image 102 using the anchor information may include repeatedly using (e.g., emphasizing) the anchor information during the image restoration process to generate the restored image 102 centered on the anchor information. The image reset operation of the video restoration device 100 uses the anchor information of the anchor image as a reference instead of a center line, thereby deriving a restored image 102 with improved image quality without the need for center line alignment. Because the center line alignment is not required, the amount and time of calculations for image restoration are reduced, and the tendency for the amount and time of calculations to increase significantly depending on the number of individual images is also eliminated.

[0019] The video restoration device 100 may select an anchor video from among the individual videos in the continuous-shot video set 101, or may generate an anchor video using video information of the individual videos. For example, the video restoration device 100 may select an anchor video from among the individual videos based on quality-based selection, time interval-based selection, or arbitrary selection. Alternatively, the video restoration device 100 may assign a weight to each individual video based on criteria such as video quality and apply the weight to each individual video to generate an anchor video.

[0020] The video restoration device 100 generates a restored image 102 by executing a neural network model based on a continuous-shot image set 101. For example, the neural network model may include a feature extraction network that extracts features from individual images of the continuous-shot image set 101 and an image restoration network that converts the extracted features into a restored image 540. At least a portion of each of the feature extraction network and the image restoration network corresponds to a deep neural network (DNN) including multiple layers. Here, the multiple layers may include an input layer, at least one hidden layer, and an output layer.

[0021] The deep neural network may include at least one of a fully connected network (FCN), a convolutional neural network (CNN), and a recurrent neural network (RNN). For example, at least some of the layers in a neural network are CNNs and others are FCNs. In this case, the CNN may be referred to as a convolutional layer, and the FCN may be referred to as a fully connected network.

[0022] In the case of CNN, the data input to each layer is called an input feature map, and the data output from each layer is called an output feature map. The input feature map and the output feature map may also be called activation data. When a convolutional layer corresponds to the input layer, the input feature map of the input layer is the input image.

[0023] After being trained based on deep learning, a neural network can perform inference suitable for the training purpose by mapping input data and output data that have a nonlinear relationship to each other. Deep learning is a machine learning method for solving problems such as video or voice recognition using big data sets. Deep learning can be understood as an optimization problem process in which a neural network is trained using prepared training data and the point at which energy is minimized is found.

[0024] Through supervised or unsupervised learning of deep learning, weights corresponding to the structure or model of a neural network are obtained, and input data and output data can be mapped to each other through these weights. If the width and depth of a neural network are sufficiently large, it can have the ability to realize any function. If a neural network learns a sufficiently large amount of training data through an appropriate training process, it can achieve optimal performance.

[0025] Hereinafter, a neural network will be referred to as being "pre-trained," where the term "pre-trained" refers to before the neural network is "started." "Starting" a neural network means that the neural network is ready for inference. For example, "starting" a neural network may include loading the neural network into memory, or inputting input data for inference into the neural network after loading the neural network into memory.

[0026] The video restoration apparatus 100 may execute a neural network model using anchor information of an anchor image. For example, the video restoration apparatus 100 may emphasize the anchor information by performing at least one of the following operations: inputting an input image to a neural network model; extracting features from the input image using the neural network model; and outputting the extracted features. For example, the anchor information may include image information of the anchor image and / or feature information extracted from the anchor image. Such anchor information provides a geometric reference for image restoration. Therefore, even if centerlines are not aligned, image information at corresponding positions can be combined with each other, thereby suppressing blurring and improving image quality.

[0027] 2 to 4 illustrate an operation of selecting an anchor picture according to various embodiments. Referring to FIG. 2, the video decompression apparatus may select an anchor picture 220 from a plurality of individual pictures 211 to 216 in a continuous-shot picture set 210. For example, the video decompression apparatus may select the anchor picture 220 based on picture quality. Specifically, the video decompression apparatus may determine the quality of each of the individual pictures 211 to 216 based on at least one of noise, blur, signal-to-noise ratio (SNR), and sharpness, and select the picture with the best quality as the anchor picture 220. The video decompression apparatus may use a deep learning network and / or a calculation module to calculate such quality.

[0028] As another example, the video decompression apparatus may select the anchor video 220 based on the video order. Specifically, the individual videos 211 to 216 may be shot in a sequential order, and the video decompression apparatus may select the first individual video 211 as the anchor video 220. As another example, the video decompression apparatus may select any one of the individual videos 211 to 216 as the anchor video 220. This is because the anchor video 220 provides a reference for video restoration, and even if the quality of the anchor video 220 is not high, the video quality can be improved through the video information of the individual videos 211 to 216.

[0029] 3, the video decompression apparatus may select an anchor picture 320 from among the individual pictures in the determined time interval. For example, a first time interval 331 may cover a certain time period from the start of filming, and the video decompression apparatus may select the anchor picture 320 in the first time interval 331. Alternatively, multiple time intervals may be used. For example, a second time interval 332 and a third time interval 333 may cover different filming times, and the video decompression apparatus may select the anchor picture 320 from the second time interval 332 and the third time interval 333.

[0030] Referring to FIG. 4, the image restoration apparatus determines a weight set 420 for a continuous-shot image set 410 and assigns weights W 41 ~W 46 For example, the image restoration device may apply a weight W based on the image quality of the individual images 411 to 416 to generate the anchor image 430. 41 ~W 46 Determine the weight W 41 ~W 46 Thus, the anchor video 430 can be generated by reflecting the video information of the individual videos 411 to 416. Here, the higher the weight value of a video, the more video information the anchor video 430 can provide.

[0031] FIG. 5 illustrates an operation of generating a restored video using anchor information according to an embodiment. Referring to FIG. 5, the video restoration apparatus may determine an anchor video based on individual videos 521-524 of a continuous video set 520, and generate a restored video 540 by executing a neural network model 510 while emphasizing anchor information 530. The continuous video set 520 may include individual videos 521-524, and the video restoration apparatus may determine an anchor video based on the individual videos 521-524 according to various criteria. FIG. 5 illustrates an example in which individual video 521 is selected as the anchor video. Hereinafter, an example in which the number of individual videos 521-524 is four will be described, but the number of individual videos 521-524 may be more or less than four.

[0032] The video restoration apparatus sequentially inputs individual videos 521-524 to the neural network model 510, and executes the neural network model 510 by emphasizing anchor information 530. For example, the video restoration apparatus can emphasize the anchor information 530 when performing at least one of the following operations: inputting the individual videos 521-524 to the neural network model 510; extracting features from the individual videos 521-524 using the neural network model 510; and outputting the extracted features. The anchor information 530 may include video information of the anchor videos and / or feature information extracted from the anchor videos.

[0033] The neural network model 510 may include a feature extraction network 511 and an image restoration network 512. The feature extraction network 511 receives the individual images 521-524 as input and extracts features from the individual images 521-524. For example, the feature extraction network 511 may extract local features from the individual images 521-524 and extract global features from the local features. The image restoration network 512 converts the extracted features into a restored image 540. The feature extraction network 511 corresponds to an encoder that converts image information into feature information, and the image restoration network 512 corresponds to a decoder that converts feature information into image information.

[0034] Figure 6 illustrates the configuration and operation of a neural network model according to one embodiment. Referring to Figure 6, a feature extraction network 610 may include a local feature extractor 611 and a global feature extractor 612. The local feature extractor 611 extracts local features from each individual image in a continuous image set 620, and the global feature extractor 612 extracts global features from the local features. The image reconstruction network 640 converts the global features into a reconstructed image 650. The feature extraction network 610 and the image reconstruction network 640 may include neural networks and may be pre-trained to perform the extraction and conversion operations.

[0035] The video decompression apparatus may execute the feature extraction network 610 by repeatedly using and / or emphasizing the anchor information 630. For example, the video decompression apparatus may emphasize the anchor information 630 in performing at least one of the following operations: inputting individual videos to the feature extraction network 610; extracting features from the individual videos using the feature extraction network 610; and outputting the extracted features. Operations related to the use of the anchor information 630 will be described in more detail below.

[0036] 7 illustrates an operation of using anchor information during an input process of an input image according to an embodiment. Referring to FIG. 7, the video restoration apparatus can use anchor information of an anchor image during a process of inputting individual images 721 to 724 to a feature extraction network. Here, the feature extraction network corresponds to a local feature extractor. In FIG. 7, it is assumed that the individual image 721 is an anchor image, and the video restoration apparatus can fuse the image information of the individual image 721 with the individual images 721 to 724 as anchor information.

[0037] For example, the fusion may include concatenation and / or addition. Concatenation is combining elements, while addition is summing elements. Thus, concatenation affects dimensions, while addition does not. Concatenation may be performed in the channel direction. For example, if each of the individual images 721-724 has dimensions of W×H×C, the concatenated result may have dimensions of W×H×2C, while the added result has dimensions of W×H×C.

[0038] The video decompression apparatus can input the fusion results as input images to the feature extraction network, thereby performing local feature extraction operations 711-714. For example, the video decompression apparatus can input the fusion results of individual image 721 and anchor video information to the feature extraction network to perform local feature extraction operation 711, thereby obtaining local feature map 731. Similarly, the fusion results of the remaining individual images 722-724 and anchor video information are sequentially input to the feature extraction network to perform local feature extraction operations 712-714, thereby obtaining local feature map 734.

[0039] FIG. 8 illustrates an operation using anchor information in a process of outputting an output feature map according to an embodiment. Referring to FIG. 8, the video restoration apparatus may use anchor information of an anchor image in a process of extracting a global feature map 850 from local feature maps 831 to 834 through a global feature extraction operation 840 and outputting the global feature map 850. The video restoration apparatus may perform the global feature extraction operation 840 using a feature extraction network. Here, the feature extraction network corresponds to a global feature extractor. In FIG. 8, it is assumed that a local feature map 831 is extracted from an anchor image, and the video restoration apparatus may fuse feature information of the local feature map 831 with the global feature map 850 as anchor information. Here, the fusion may include concatenation and / or addition. The fusion result corresponds to an output feature map of the feature extraction network, and the video restoration apparatus may convert the output feature map into a restored image using the video restoration network.

[0040] FIG. 9 illustrates an operation using anchor information in a global feature extraction process according to an embodiment. Referring to FIG. 9, the video restoration apparatus extracts a global feature map 950 from local feature maps 931-934 through a global feature extraction operation 940. In this process, anchor information of the anchor image may be used as guide information. In FIG. 9, it is assumed that the local feature map 931 is extracted from the anchor image. The video restoration apparatus may use feature information of the local feature map 931 as anchor information and as guide information. For example, if the local feature map 931 is an anchor local feature and the local feature maps 932-934 are peripheral local features, the video restoration apparatus may perform the global feature extraction operation 940 by assigning a higher weight to the anchor local feature than to the peripheral local features. Therefore, information on the anchor local feature may have a greater influence on the global feature map 950 than information on the peripheral local features.

[0041] FIG. 10 illustrates an example of the operation shown in FIG. 9. As shown in FIG. 9, various weighting methods may be used to use anchor information as guide information. FIG. 10 illustrates a method for emphasizing anchor information through a weighting operation 1040 and a weighted fusion operation 1060. This method performs a pooling operation such as max pooling or average pooling to emphasize anchor information corresponding to a geometric criterion for image restoration, and is more effective for combining corresponding image information than a method for extracting global features from local features. Referring to FIG. 10, the video restoration apparatus may assign different weights to each of the local feature maps 1031 to 1034. The video restoration apparatus may fuse the local feature maps 1031 to 1034 in consideration of a weight set 1050, and generate a global feature map 1070 as a result of the weighted fusion operation 1060. For example, the video restoration apparatus may assign weights W to the local feature maps 1034 to 1034 through softmax. 101 ~W 104 and weighted by W 101 ~W 104 The local feature maps 1034 to 1034 can be added together to generate the global feature map 1070 taking into account the above.

[0042] It is assumed that a local feature map 1031 is extracted from an anchor video and local feature maps 1032 to 1034 are extracted from other individual videos. In this case, the video restoration device calculates the weights W 102 ~W 104 The weight value W of the local feature map 1031 compared to 101may be set even higher. Therefore, anchor information may be emphasized through feature information of the local feature map 1031. Alternatively, the video restoration apparatus may determine the similarity between the local feature map 1031 and each of the local feature maps 1032 to 1034, and assign a high weight to not only the local feature map 1031 but also other feature maps that are similar to the local feature map 1031. Specifically, if the local feature map 1032 has a high similarity to the local feature map 1031 and the local feature maps 1033 and 1034 have a low similarity to the local feature map 1031, the video restoration apparatus may assign a weight W 103 , W 104 Compared to the weights W of the local feature maps 1031 and 1032, 101 , W 102 Therefore, the anchor information can be emphasized through the feature information of the local feature maps 1031 and 1032.

[0043] 11 illustrates an operation of using anchor information in a process of extracting local features according to an embodiment. Referring to FIG. 11, the video decompression apparatus may extract local feature maps 1131-1134 from individual images 1151-1154 of a continuous-shot image set 1050 through local feature extraction operations 1110-1140. The video decompression apparatus may emphasize anchor information in the process of extracting the local feature maps 1131-1134. For example, the video decompression apparatus may select individual image 1151 as an anchor image and generate anchor information 1101 based on the individual image 1151. The anchor information 1101 may include image information and / or feature information of the individual image 1151.

[0044] A feature extraction network (e.g., a local feature extractor) may include multiple layers, which may be classified into layer groups each including several layers. For example, each layer group may include a convolutional layer and / or a pooling layer. The video decompression apparatus may extract local features using anchor information for each of the multiple layer groups of the feature extraction network. The video decompression apparatus extracts local features for each layer group and combines the extracted local features with anchor information 1101. The video decompression apparatus may generate local feature maps 1131-1134 by repeating this process for all layer groups.

[0045] Specifically, the video decompression apparatus may extract primary local features from an individual image 1151 through a feature extraction operation 1111 using a first layer group, and transform the primary local features by fusing anchor information 1101 with the primary local features. The video decompression apparatus may extract secondary local features from the transformed primary local features through a feature extraction operation 1112 using a second layer group, and transform the secondary local features by fusing anchor information 1101 with the secondary local features. The video decompression apparatus may also extract tertiary local features from the transformed secondary local features through a feature extraction operation 1113 using a third layer group, and transform the tertiary local features by fusing anchor information 1101 with the tertiary local features. When the feature extraction operation 1115 using the last layer group is completed, a local feature map 1131 is generated as a result. The remaining local feature extraction operations 1120-1140 for the remaining individual images 1152-1154 correspond similarly to the local feature operation 1110 for the individual image 1151, and can result in the generation of local feature maps 1132-1134.

[0046] Here, the video restoration apparatus may fuse the same anchor information 1101 to the output of each layer group, or fuse anchor information 1101 specific to each layer group. First, fusion through common anchor information 1101 will be described. The common anchor information 1101 may be video information and / or feature information of the anchor video. Here, to obtain the feature information, an operation of extracting features from the anchor video may be performed in advance. For example, this operation may be performed using a feature extraction network (e.g., feature extraction network 511 shown in FIG. 5 or local feature extractor 611 shown in FIG. 6) used for the local feature extraction operations 1110 to 1140, or a separately provided feature extraction network. Once the common video information and / or feature information is provided, it can be fused into the output of each layer group by the local feature extraction operations 1110 to 1140, in other words, into local features.

[0047] Next, fusion through specialized anchor information 1101 will be described. Unlike the common anchor information 1101, the specialized anchor information 1101 may be information processed to suit each layer group. The specialized anchor information 1101 may include tiered local features of the anchor video extracted by each layer group. For example, if first to third local features of the anchor video are extracted by the first to third layer groups, the first to third local features are used as anchor information 1101 specialized for the first to third layer groups. Therefore, the first local feature may be fused with each of the local features extracted through feature extraction operations 1111 and 1121, the second local feature may be fused with each of the local features extracted through feature extraction operations 1112 and 1122, and the third local feature may be fused with each of the local features extracted through feature extraction operations 1113 and 1123.

[0048] 12 is a flowchart illustrating a video reset operation according to an embodiment. Referring to FIG. 12, in step S1210, the video decompression apparatus determines an anchor video based on individual videos in a continuous-shot video set. The video decompression apparatus may select the anchor video from the individual videos based on the quality of the individual videos. Alternatively, the video decompression apparatus may select any video from the individual videos as the anchor video.

[0049] In step S1220, the video decompression apparatus executes a feature extraction network based on the set of continuous shots using anchor information of the anchor video. The video decompression apparatus may extract primary local features from a first individual video among the individual videos using a first layer group of the feature extraction network, transform the primary local features by fusing the primary local features with the anchor information, and extract secondary local features from the transformed primary local features using a second layer group of the feature extraction network. Alternatively, the video decompression apparatus may transform the secondary local features by fusing the anchor information with the secondary local features, extract tertiary local features from the transformed secondary local features using a third layer group of the feature extraction network, and determine global features based on the tertiary local features.

[0050] The video restoration apparatus may extract anchor local features from the anchor image, extract local features from images other than the anchor image among the individual images, and extract global features from the anchor local features and the local features of the other images using the anchor local features. Here, the video restoration apparatus may assign a higher weight to the anchor local features than to the local features of the other images, and extract global features from the anchor local features and the local features of the other images.

[0051] The video restoration device may extract anchor local features from the anchor video, extract local features from other videos other than the anchor video among the individual videos, extract global features from the anchor local features and the local features of the other videos, and fuse the anchor local features with the global features. The video restoration device may also fuse anchor information with each individual video to generate an input video for the neural network model.

[0052] In step S1230, the video restoration apparatus generates a restored video based on a feature map corresponding to the output of the feature extraction network. Here, the video restoration apparatus may execute the video restoration network based on the feature map. In addition, the description of Figures 1 to 11 may be applied to the video restoration method.

[0053] 13 is a block diagram showing the configuration of a video restoration device according to an embodiment. Referring to FIG. 13, the video restoration device 1300 includes a processor 1310 and a memory 1320. The memory 1320 is connected to the processor 1310 and stores instructions executable by the processor 1310, data operated on by the processor 1310, or data processed by the processor 1310. The memory 1320 may include a non-transitory computer-readable recording medium, such as a high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0054] The processor 1310 executes instructions for performing the operations described with reference to Figures 1 to 12. For example, the processor 1310 determines an anchor video based on individual videos of a continuous-shot video set, executes a feature extraction network based on the continuous-shot video set using anchor information of the anchor video, and generates a restored video based on a feature map corresponding to the output of the feature extraction network. In addition, the details of the video restoration device 1300 may be applied to the descriptions of Figures 1 to 12.

[0055] FIG. 14 is a block diagram showing a configuration of an electronic device according to an embodiment. Referring to FIG. 14, the electronic device 1400 may include a processor 1410, a memory 1420, a camera 1430, a storage device 1440, an input device 1450, an output device 1460, and a network interface 1470, which may communicate with each other via a communication bus 1480. For example, the electronic device 1400 may be implemented as at least a part of a mobile device such as a mobile phone, a smartphone, a PDA, a netbook, a tablet computer, a laptop computer, etc.; a wearable device such as a smart watch, a smart band, smart glasses, etc.; a computing device such as a desktop, a server, etc.; a home appliance such as a television, a smart TV, a refrigerator, etc.; a security device such as a door rack, etc.; or a vehicle such as an autonomous vehicle, a smart vehicle, etc. The electronic device 1400 may structurally and / or functionally include the video restoration device 100 shown in FIG. 1 and / or the video restoration device 1300 shown in FIG. 13.

[0056] The processor 1410 executes functions and instructions to be executed within the electronic device 1400. For example, the processor 1410 processes instructions stored in the memory 1420 or the storage device 1440. The processor 1410 may perform one or more of the operations described with reference to FIGS. 1-13. The memory 1420 may include a computer-readable storage medium or a computer-readable storage device. The memory 1420 stores instructions to be executed by the processor 1410 and stores related information during execution of software and / or applications by the electronic device 1400.

[0057] Camera 1430 takes photos and / or videos. Camera 1430 may continuously take photos or record videos to generate a continuous video set. If the continuous video set is a series of photos, each individual video in the continuous video set corresponds to a photo. If the continuous video set is a video, each individual video in the continuous video set corresponds to a frame of the video. Storage device 1440 includes a computer-readable storage medium or computer-readable storage device. Storage device 1440 can store more information than memory 1420 and can store information for a longer period of time. For example, storage device 1440 may include a magnetic hard disk, an optical disk, a flash memory, a floppy disk, or other forms of non-volatile memory known in the art.

[0058] The input device(s) 1450 can receive input from a user through traditional input methods such as a keyboard and a mouse, as well as newer input methods such as touch input, voice input, and image input. For example, the input device(s) 1450 may include a keyboard, a mouse, a touchscreen, a microphone, or any other device capable of detecting input from a user and communicating the detected input to the electronic device 1400. The output device(s) 1460 can provide output of the electronic device 1400 to a user through a visual, auditory, or tactile channel. The output device(s) 1460 may include, for example, a display, a touchscreen, a speaker, a vibration generator, or any other device capable of providing output to a user. The network interface 1470 can communicate with external devices via a wired or wireless network.

[0059] The above-described embodiments may be implemented using hardware components, software components, or a combination of hardware and software components. For example, the devices and components described herein may be implemented using one or more general-purpose or special-purpose computers, such as a processor, controller, arithmetic logic unit (ALU), digital signal processor, microcomputer, field programmable array (FPA), programmable logic unit (PLU), microprocessor, or other device that executes and responds to instructions. The processing device may execute an operating system (OS) and one or more software applications that run on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of software. For convenience of understanding, a processing device may be described as being a single device, but those skilled in the art will recognize that a processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing device may include multiple processors or one processor and one controller. Other processing configurations, such as parallel processors, are also possible.

[0060] Software includes computer programs, codes, instructions, or a combination of one or more thereof, which can configure a processing device to operate as intended or independently or in combination to instruct the processing device. The software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave to be interpreted by the processing device or to provide instructions or data to the processing device. The software can be distributed across computer systems coupled to a network and stored and executed in a distributed manner. The software and data can be stored on one or more computer-readable recording media.

[0061] The method according to the present invention may be embodied in the form of program instructions that can be executed by various computer means and recorded on a computer-readable recording medium. The recording medium may include program instructions, data files, data structures, and the like, alone or in combination. The recording medium and program instructions may be specially designed and constructed for the purposes of the present invention, or may be well-known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tape, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, flash memory, and the like. Examples of program instructions include not only machine language code, such as that generated by a compiler, but also high-level language code that is executed by a computer using an interpreter, for example.

[0062] The hardware devices described above may be configured to operate as one or more software modules to perform the operations described in the present invention, and vice versa.

[0063] Although the embodiments have been described above with reference to limited drawings, those skilled in the art may apply various technical modifications and variations based on the above description. For example, the described techniques may be performed in a different order than described, and / or the components of the described systems, structures, devices, circuits, etc. may be combined or combined in a different manner than described, or may be replaced or substituted with other components or equivalents, and still achieve suitable results.

[0064] Therefore, the scope of the present invention is not limited to the disclosed embodiments, but is defined by the appended claims and their equivalents.

Claims

1. A video restoration method executed by a video restoration device, comprising: determining an anchor video based on individual videos of the set of continuous-shot videos; running a feature extraction network based on the set of continuous shots using anchor information of the anchor video; generating a reconstructed image based on a feature map corresponding to the output of the feature extraction network; In the feature extraction network, a plurality of layer groups each including one or more layers are connected in order, and the step of executing the feature extraction network includes: extracting primary local features from a first individual image corresponding to the individual image using a first layer group of the plurality of layer groups; fusing the anchor information with the primary local features to transform the primary local features; extracting secondary local features from the transformed primary local features using a second layer group of the plurality of layer groups; and determining a global feature based on the second-order local features.

2. The video restoration method of claim 1 , wherein determining the anchor video comprises selecting the anchor video from the individual videos based on a quality-based selection, a time interval-based selection, or an arbitrary selection.

3. The image restoration method of claim 1 , wherein determining the anchor image comprises applying a weight to the individual images to generate the anchor image.

4. 4. The video restoration method according to claim 1, wherein the anchor information includes at least one of video information of the anchor video and feature information of the anchor video.

5. The video restoration method according to any one of claims 1 to 4, further comprising the step of generating an input video for the feature extraction network by fusing the anchor information with each of the individual videos.

6. Executing the feature extraction network comprises: extracting anchor local features from the anchor video; extracting local features from images other than the anchor image among the individual images; extracting global features from the anchor local features and the local features of the other images; fusing the anchor local features with the global features; 6. The video restoration method according to claim 1, further comprising:

7. Executing the feature extraction network comprises: extracting anchor local features from the anchor video; extracting local features from images other than the anchor image among the individual images; extracting global features from the anchor local features and local features of the other images using the anchor local features; 6. The video restoration method according to claim 1, further comprising:

8. The image restoration method of claim 7, wherein the step of extracting the global features includes a step of assigning a higher weighting value to the anchor local features than to the local features of the other images, and extracting the global features from the anchor local features and the local features of the other images.

9. An image restoration method as described in Claim 1, wherein the anchor information is common to the multiple layer groups.

10. An image restoration method as described in Claim 1, wherein the anchor information is specialized for each of the multiple layer groups.

11. The video restoration method according to any one of claims 1 to 10, wherein the step of generating the restored video comprises the step of running a video restoration network based on the feature map.

12. A computer program stored on a computer-readable recording medium for executing the method according to any one of claims 1 to 11 in combination with hardware.

13. a processor; a memory containing instructions executable by the processor; Including, When the instruction is executed by the processor, the processor: determining an anchor video based on each video of the continuous-shot video set; running a feature extraction network based on the set of continuous shots using anchor information of the anchor video; A reconstructed image is generated based on a feature map corresponding to an output of the feature extraction network, and in the feature extraction network, a plurality of layer groups, each including one or more layers, are connected in order, and when the processor executes the feature extraction network, extracting first-order local features from a first individual image corresponding to the individual image using a first layer group of the plurality of layer groups; Transforming the primary local features by fusing the anchor information with the primary local features; extracting secondary local features from the transformed primary local features using a second layer group of the plurality of layer groups; A video restoration device that determines global features based on the second-order local features.

14. the processor selects the anchor picture from the individual pictures based on a quality-based selection, a time interval-based selection, or an arbitrary selection; The video restoration apparatus of claim 13 , wherein the anchor video is generated by applying a weight to the individual video.

15. The processor: Extracting anchor local features from the anchor video; extracting local features from images other than the anchor image among the individual images; 15. The video restoration device according to claim 13, wherein the anchor local features are used to extract global features from the anchor local features and local features of the other video.

16. 16. The video restoration device according to claim 13, wherein the processor extracts local features by using the anchor information for each of a plurality of layer groups of the feature extraction network.

17. a camera for generating a set of burst images; determining an anchor image based on the individual images of the set of continuous-shot images; running a feature extraction network based on the set of continuous shots using anchor information of the anchor video; a processor for generating a reconstructed image based on a feature map corresponding to the output of the feature extraction network; In the feature extraction network, a plurality of layer groups each including one or more layers are connected in order, and when the processor executes the feature extraction network, extracting first-order local features from a first individual image corresponding to the individual image using a first layer group of the plurality of layer groups; Transforming the primary local features by fusing the anchor information with the primary local features; extracting secondary local features from the transformed primary local features using a second layer group of the plurality of layer groups; An electronic device that determines a global feature based on the second-order local features.

18. the processor selects the anchor picture from the individual pictures based on a quality-based selection, a time interval-based selection, or an arbitrary selection; The electronic device of claim 17 , further comprising: applying weights to the individual images to generate the anchor image.

19. The processor: Extracting anchor local features from the anchor video; extracting local features from images other than the anchor image among the individual images; 19. The electronic device according to claim 17 or 18, wherein the anchor local features are used to extract global features from the anchor local features and local features of the other images.

20. The electronic device according to any one of claims 17 to 19, wherein the processor extracts local features using the anchor information for each of a plurality of layer groups of the feature extraction network.

Citation Information

Patent Citations

  • Timing fixed-point scene anomaly detection method

    CN111767826A

  • Ghost artifact detection and removal methods in HDR image processing using multi-scale normalized cross-correlation

    JP2015011717A

  • Image processing method and apparatus, facial recognition method and apparatus, and computer device

    US20200372243A1

  • Methods to maintain image quality in ultrasound imaging at reduced cost, size, and power

    WO2020139775A1