Method and Apparatus for Improving Image Quality by Using Noise Classification

The method addresses slow computation and noise removal issues in super-resolution by using a noise classifier to enhance image quality through a pre-trained network, achieving improved image quality and efficiency.

KR102994237B1Active Publication Date: 2026-07-21SK TELECOM CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
SK TELECOM CO LTD
Filing Date
2021-10-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing deep neural network-based super-resolution methods struggle with slow computation speed and are ineffective in removing noise from input images due to the use of large parameter counts and lack of noise classification.

Method used

A method involving a noise classifier to generate a noise class map, merge it with the input frame, and input the merged frame into a pre-trained image quality improvement network to enhance image quality.

Benefits of technology

Effectively removes noise by classifying noise intensity, improving image quality and computation efficiency by reducing the number of parameters and optimizing network training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112021124394549-PAT00001_ABST
    Figure 112021124394549-PAT00001_ABST
Patent Text Reader

Abstract

A method and apparatus for improving image quality using noise classification are disclosed. According to one aspect of the present disclosure, a method for improving image quality is provided, comprising: a step of generating a noise class map corresponding to the noise intensity of an input frame; a step of merging the input frame and the noise class map to generate a merged frame; and a step of inputting the merged frame into a pre-trained image quality improvement network to generate an output frame with improved image quality compared to the input frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and apparatus for improving image quality using noise classification. Background Technology

[0002] The content described in this section merely provides background information regarding the present invention and does not constitute prior art.

[0003] Recently, Super Resolution (SR) technology based on deep learning-based deep neural networks has been actively developed, and the performance of SR is significantly improving compared to other previous technologies. However, in existing deep neural network-based SR research, the data used for training the neural networks consists mostly of distortion-free images that differ in quality from actual broadcast footage. Therefore, when applying the results of SR implementation based on existing research to broadcast content, there is a limitation in that it is difficult to obtain high-quality images at the target level.

[0004] A representative technology for super-resolution deep neural networks is EDSR (Enhanced Deep Super Resolution Network), which is a neural network improved upon ResNet, widely used for image transformation. To implement super-resolution technology, EDSR includes 32 Residual Blocks (RBs). Compared to ResNet's RBs, EDSR's RBs remove the Batch Normalization (BN) layer and have a single Rectified Linear Unit (ReLU) layer. When EDSR, utilizing these structural modifications, was announced, it achieved the highest performance in the field of super-resolution implementation. However, compared to the 64 feature map channels typically used in SR implementations, EDSR uses 256 channels, so 43x10 6It has the disadvantage of slow computation speed because it includes a vast number of parameters.

[0005] Another super-resolution deep neural network is the RCAN (Residual Channel Attention Network) model, in which the depth of the neural network is significantly increased to achieve super-resolution. Since it was difficult to train deep networks using previously studied neural network structures to achieve super-resolution, the RCAN model includes a structure for training deep networks. The RCAN model includes RIR (Residual In Residual), which uses long skip connections; RG (Residual Group), which is contained within the RIR and uses short skip connections; and RCAB (Residual Channel Attention Block), which is inside the RG and utilizes CA (Channel Attention) to readjust channel-level features based on inter-channel interdependencies. In the RCAN model, approximately 800 convolutional layers were used to implement the structure characterized by RIR, RG, RCAB, and CA. Based on the results of comparative experiments with existing methods, the RCAN model demonstrated improved SR implementation. However, compared to EDSR, the number of parameters is 16×10 6 Even though the number has been reduced, the RCAN model still has the disadvantage of slow computation speed because it uses about 800 convolution layers. The problem to be solved

[0006] The main purpose of the present disclosure is to provide a super-resolution method and apparatus capable of effectively removing noise from an input image by classifying the noise intensity of the input image to divide it into classes and learning parameters accordingly.

[0007] The problems that the present invention aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0008] According to one aspect of the present disclosure, a method for improving image quality is provided, comprising: a step of generating a noise class map corresponding to the noise intensity of an input frame; a step of merging the input frame and the noise class map to generate a merged frame; and a step of inputting the merged frame into a pre-trained image quality improvement network to generate an output frame with improved image quality compared to the input frame.

[0009] According to another aspect of the present disclosure, a computer program stored on a computer-readable recording medium is provided to execute each process included in the image quality improvement method described above.

[0010] According to another aspect of the present disclosure, an image quality enhancement device is provided comprising: a memory for storing instructions; and at least one processor, wherein the at least one processor generates a noise class map corresponding to the noise intensity of an input frame by executing the instructions, generates a merged frame by merging the input frame and the noise class map, and inputs the merged frame to a pre-trained image quality enhancement network to generate an output frame with improved image quality compared to the input frame. Effects of the invention

[0011] As explained above, according to the embodiment of the present disclosure, by classifying the noise intensity of the input image to divide it into classes and learning parameters accordingly, the noise of the input image can be effectively removed.

[0012] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0013] FIG. 1 is a block diagram schematically showing an image quality improvement model according to one embodiment of the present disclosure. FIG. 2 is a block diagram schematically showing a noise classifier according to one embodiment of the present disclosure. FIG. 3 is an exemplary diagram showing the structure of a noise class classification network according to one embodiment of the present disclosure. FIG. 4 is an exemplary diagram showing the structure of a reconstruction module according to one embodiment of the present disclosure. FIG. 5 is an exemplary diagram showing the structure of an LRB included in a reconstruction module according to one embodiment of the present disclosure. FIG. 6 is an exemplary diagram showing the structure of an SRB included in a reconstruction module according to one embodiment of the present disclosure. FIGS. 7a and 7b are exemplary diagrams for explaining the learning process of an image quality improvement model according to one embodiment of the present disclosure. FIG. 8 is a flowchart illustrating a method for improving image quality according to one embodiment of the present disclosure. FIG. 9 is a flowchart illustrating a learning method of an image quality improvement network according to one embodiment of the present disclosure. FIGS. 10a to 10c are illustrative diagrams for explaining the effects of an image quality improvement model according to one embodiment of the present disclosure. Specific details for implementing the invention

[0014] Some embodiments of the present disclosure are described in detail below with reference to the exemplary drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known components or functions could obscure the essence of the present disclosure, such detailed description is omitted.

[0015] In addition, terms such as first, second, A, B, (a), (b), etc. may be used to describe the components of the present disclosure. These terms are intended only to distinguish the components from other components, and the nature, order, or sequence of the components is not limited by these terms. Throughout the specification, when a part is described as 'comprising' or 'equipped' with a certain component, unless specifically stated otherwise, this means that it does not exclude other components but may include additional components. Furthermore, terms such as '…part' or 'module' described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.

[0016] The detailed description set forth below, together with the accompanying drawings, is intended to describe exemplary embodiments of the present disclosure and is not intended to represent the only embodiment in which the present disclosure can be practiced.

[0017] FIG. 1 is a block diagram schematically showing an image quality improvement model according to one embodiment of the present disclosure.

[0018] As illustrated in FIG. 1, an image quality improvement model (10) according to one embodiment of the present disclosure may include all or part of a noise classifier (100) and an image quality improvement network (110). Not all blocks illustrated in FIG. 1 are essential components, and some blocks included in the image quality improvement model (10) in other embodiments may be added, changed, or deleted. Each component of the image quality improvement model (10) and the image quality improvement model (10) may be implemented in hardware or software, or in a combination of hardware and software. Additionally, the function of each component may be implemented in software, and one or more processors may be implemented to execute the function of the software corresponding to each component.

[0019] The noise classifier (100) can classify the noise intensity of an input frame to generate a noise class map and output the input frame and the noise class map by merging them.

[0020] According to embodiments, the image quality improvement model (10) can input a plurality of consecutive input frames to the noise classifier (100) so that temporal information of the input image can be reflected.

[0021] For example, when the frame to be improved at time t is called the target input frame I(t), the noise classifier (100) receives the previous input frame I(t-1), the target input frame, and the subsequent input frame I(t+1) and can output the previous merged frame, the target merged frame, and the subsequent merged frame.

[0022] A detailed description of the noise classifier (100) will be explained with reference to FIGS. 2 and FIGS. 3.

[0023] The image quality enhancement network (110) generates an output frame with improved image quality and / or resolution compared to an input frame by using a merged frame containing information about noise intensity. To this end, the image quality enhancement network (110) according to one embodiment of the present disclosure may include all or part of a frame alignment module (120), a reconstruction module (130), and an upsampling layer (140).

[0024] The frame alignment module (120) aligns other merge frames based on the target merge frame. For example, the frame alignment module (120) can align the previous merge frame and the subsequent merge frame, respectively, based on the target merge frame. The frame alignment module (120) can generate a feature map by concatenating the aligned features into one and then passing them through a convolution layer.

[0025] According to embodiments, the frame alignment module (120) may be implemented using a deformable convolution network, but is not limited thereto, and the present disclosure does not limit it in a specific way.

[0026] The reconstruction module (130) reconstructs a high-quality frame using a feature map output by the frame alignment module (). The reconstruction module (130) according to one embodiment of the present disclosure may be implemented as a deep neural network based on a plurality of convolution layers. A detailed description of the reconstruction module (130) will be given with reference to FIGS. 4 to 6.

[0027] An upsampling layer (140) may be provided to generate an output frame with improved resolution compared to an input frame. For example, if the resolution of the input frame is FHD (Full High Definition), the upsampling layer (140) may generate an output frame having a resolution such as 4K or 8K. According to embodiments, the image quality enhancement network (110) may generate an output frame having the same resolution as the input frame by omitting the upsampling layer (140).

[0028] FIG. 2 is a block diagram schematically showing a noise classifier according to one embodiment of the present disclosure.

[0029] Referring to FIG. 2, a noise classifier (100) according to one embodiment of the present disclosure may include all or part of a scene change detection module (200), a noise class classification network (210), and a merging module (220).

[0030] The scene change detection module (200) detects scene changes between input frames. Since frames within the same scene have similar noise intensities, the scene change detection module (200) does not input every input frame into the noise class classification network (210), but inputs the input frame into the noise class classification network (210) only when a scene change occurs. Accordingly, the problem of frames within the same scene being classified with different noise intensities can be prevented, while simultaneously improving processing speed.

[0031] According to embodiments, the scene change detection module (200) can determine whether there is a scene change by comparing pixel differences between a plurality of frames. For example, the scene change detection module (200) can calculate a Peak Signal to Noise Ratio (PSNR) based on pixel-to-pixel difference values ​​between a previous input frame and a target input frame, and if the PSNR is less than or equal to a preset threshold (e.g., 15 dB), it can determine that there is a scene change between the previous input frame and the target input frame.

[0032] The noise class classification network (210) classifies the noise intensity of an input frame. For example, the noise class classification network (210) defines the noise intensity into five classes (0, 1, 2, 3 and 4), and can classify the input frame with more noise into a larger value.

[0033] FIG. 3 is an exemplary diagram showing the structure of a noise class classification network according to one embodiment of the present disclosure.

[0034] As illustrated in FIG. 3, a noise class classification network (210) according to one embodiment of the present disclosure may be implemented as a deep neural network based on a plurality of convolution layers. The noise class classification network (210) may include all or part of a plurality of convolution layers and a plurality of linear layers. For example, the noise class classification network (210) may include eight convolution layers and three linear layers, but is not necessarily limited thereto, and an appropriate number may be set by compromising the computational speed and classification accuracy of the noise class classification network (210).

[0035] The specific structure of the noise class classification network (210) according to one embodiment of the present disclosure may be as shown in Table 1, but is not limited to such examples.

[0036] filter Stride Output channel conv1 3×3 1 64 conv2 3×3 2 64 conv3 3×3 1 64 conv4 3×3 2 64×4 conv5 3×3 1 64×4 conv6 3×3 2 64×4 conv7 3×3 1 64×4 conv8 3×3 2 64×4 linear1 - - 1024 linear2 - - 256 linear3 - - 5

[0037] According to embodiments, a downsampling layer (not shown) that downsamples an input frame by 1 / 4 may be further provided at the input end of the noise class classification network (210). When the input frame has a resolution of FHD (1920×1080), the size of the frame input to the first convolution layer (conv1) is 3×480×270, and the size of the feature map input to the first linear layer (linear1) through the first to eighth convolution layers may be 64×4×17×30.

[0038] The number of output channels of the last linear layer (e.g., linear3) of the noise class classification network (210) according to one embodiment of the present disclosure is 5, and it outputs probability values ​​for 5 classes (0, 1, 2, 3 and 4). The noise class classification network (210) can output the class with the highest probability value by applying an argmax function to the output of the last linear layer.

[0039] Referring again to FIG. 2, the merging module (220) can generate a noise class map based on a class (hereinafter, noise class) output by the noise class classification network (210) and merge the input frame and the noise class map. For example, if the size of the input frame is 3×H×W, the merging module (220) can generate a noise class map of size 1×H×W and merge the input frame and the noise class map to generate a frame of size 4×H×W (hereinafter, merged frame).

[0040] Here, the value of each pixel of the noise class map may be a value obtained by normalizing the noise class of the input frame. A merging module (220) according to one embodiment of the present disclosure may divide the noise class by the number of classes (e.g., 5) to use as the value of each pixel of the noise class map in order to normalize the noise class to a value between 0 and 1. For example, when the output of the noise class classification network (210) is '3', the merging module (220) may generate a noise class map with a size of 1×H×W and a value of '0.6' for each pixel.

[0041] According to embodiments, if it is confirmed that no scene change has occurred between input frames, the merging module (220) can merge a noise class map generated at a previous point in time into the input frames.

[0042] According to embodiments, if it is confirmed that a scene change has occurred between a previous input frame and a target input frame, the merging module (220) can merge a noise class map generated based on the target input frame into the target input frame and the subsequent input frame, respectively.

[0043] FIG. 4 is an exemplary diagram showing the structure of a reconstruction module according to one embodiment of the present disclosure.

[0044] As illustrated in FIG. 4, the reconstruction module (130) may include all or part of a plurality of LRBs (Long Residual Blocks, 402), a plurality of convolution layers (406, 408, 410), and a first skip connection (404). The reconstruction module (130) can generate a high-quality frame by processing the input feature map in the order of the first convolution layer (406), a plurality of LRBs, a second convolution layer (408), and a third convolution layer (410).

[0045] A first convolution layer (406) located on the input side of the reconstruction module (130) extracts a plurality of feature maps from the input feature map. An appropriate number of feature map channels extracted by the first convolution layer (406) can be set by compromising the computational speed of the reconstruction module (130) and the quality of the output frame. In FIGS. 4 to 6, C (where C is a natural number) represents the number of feature map channels, for example, C can be 64.

[0046] As illustrated in FIG. 4, the reconstruction module (130) according to the present embodiment may include eight LRBs, but is not necessarily limited thereto, and an appropriate number of LRBs may be set by compromising the computational speed of the reconstruction module (130) and the quality of the output frame. Each LRB included in the reconstruction module (130) is implemented with the same structure and performs the same function. Accordingly, the structure and operation of one LRB (402) will be described below.

[0047] FIG. 5 is an exemplary diagram showing the structure of an LRB included in a reconstruction module according to one embodiment of the present disclosure.

[0048] As illustrated in FIG. 5, an LRB (402) according to one embodiment of the present disclosure may include all or part of two SRBs (Short Residual Blocks, 502), a second skip connection (504), and a fourth convolution layer (506). The LRB (402) processes inputs in the order of the two SRBs and the fourth convolution layer (506).

[0049] Each SRB included in the LRB (402) is implemented with the same structure and performs the same function. Therefore, the structure and operation of one SRB (502) will be described below.

[0050] FIG. 6 is an exemplary diagram showing the structure of an SRB included in a reconstruction module according to one embodiment of the present disclosure.

[0051] As illustrated in FIG. 6, an SRB (502) according to one embodiment of the present disclosure may include all or part of a third skip connection (602), a first convolution group (604), a concatenation layer (606), and a second convolution group (608). Each of the first convolution group (604) and the second convolution group (608) may include two convolution layers and one Rectified Linear Unit (ReLU). Here, the ReLU is an activation function that limits the range of the output.

[0052] The SRB (502) processes the input in the order of the first convolution group (604), the chain layer (606), and the second convolution group (608).

[0053] The first convolution group (604) generates a residual output for the input of the SRB (502). The chaining layer (606) chains the input of the SRB (502) and the residual output of the first convolution group (604) and passes them to the second convolution group (608). The second convolution group (608) generates a residual output from the result of the chaining layer (606). The third skip connection (602) adds the input of the SRB (502) and the residual output of the second convolution group (608).

[0054] Meanwhile, since the Batch Normalization (BN) layer is known to be effective for classifying image classes but has little effect on implementing super-resolution that transforms images at the pixel level, it was not used in the SRB (502) according to one embodiment of the present disclosure. In addition, in order to focus on transforming the high-frequency region of the image compared to existing methods, the SRB (502) according to one embodiment of the present disclosure uses a chain layer (606).

[0055] Referring again to FIG. 5, the fourth convolution layer (506) included in the LRB (402) generates a residual output from the output of the second SRB (the result of adding the input of the SRB and the residual output of the second convolution group). Additionally, the second skip connection (504) adds the input of the LRB (402) and the residual output of the fourth convolution layer (506).

[0056] Referring again to FIG. 4, the second convolution layer (408) included in the reconstruction module (130) generates a residual output from the output inside the last LRB (the result of adding the input of the LRB and the residual output of the fourth convolution layer). The first skip connection (404) can transmit the characteristics extracted from the input to the output side of the reconstruction module (130) by adding the output of the first convolution layer (406) and the residual output of the second convolution layer (408).

[0057] The reconstruction module (130) can generate a high-quality image from the result of adding the output of the first convolution layer (406) (features extracted from the input) and the residual output of the second convolution layer (408) using the third convolution layer (410).

[0058] As illustrated in FIGS. 4 to 6, the reconstruction module (130) is expanded through SRB (502), LRB (402), and reconstruction module (130), increasing the number of layers and thus deepening the neural network. Nevertheless, each component of the reconstruction module (130) includes skip connections (404, 504, and 602) to allow training to proceed effectively.

[0059] FIGS. 7a and 7b are exemplary diagrams for explaining the learning process of an image quality improvement model according to one embodiment of the present disclosure.

[0060] Learning of an image quality improvement model according to one embodiment of the present disclosure is performed by a learning device, and the learning device may be executed on a computing device. The learning device may include a computer-readable storage connected to one or more processors available to the computing device, each performing a function, and having instructions stored internally.

[0061] FIG. 7a illustrates the learning process of a noise class classification network (210) according to one embodiment of the present disclosure.

[0062] The learning device can train a noise class classification network (210) using pairs of input frames and noise class labels as training data.

[0063] According to embodiments, the learning device may have a pair of a target frame and an input frame in advance, and may receive a noise class label based on the difference value between the target frame and the input frame and use it as training data. According to embodiments, the learning device may generate an input frame by using a learning frame generation module (700) to add noise corresponding to the noise class label to the target frame.

[0064] The learning device can input an input frame into a noise class classification network (210) to predict a noise class. The learning device can train the noise class classification network (210) through back-propagation based on the error between the noise class label and the predicted noise class. Training of the noise class classification network (210) may include updating the network parameters of the noise class classification network (210).

[0065] FIG. 7b illustrates the learning process of an image quality improvement network (110) according to one embodiment of the present disclosure.

[0066] The learning device can train an image quality improvement network (110) using a pair of target frames and input frames as training data.

[0067] According to embodiments, the learning device may have a pair of a target frame and an input frame in advance. According to embodiments, the learning device may generate an input frame by using a learning frame generation module (700) to add noise corresponding to a noise class label to the target frame.

[0068] The learning device can input an input frame into a noise classifier (100) to obtain a merged frame in which the input frame and the noise class map are merged. The noise classifier (100) may include a noise class classification network (210) that has been pre-trained to classify the noise intensity of the input frame.

[0069] The learning device can input a merged frame into the image quality improvement network (110) to predict an output frame. The learning device can train the image quality improvement network (110) through backpropagation based on the error between the target frame and the predicted output frame. Training of the image quality improvement network (110) may include updating the network parameters of the image quality improvement network (110).

[0070] FIG. 8 is a flowchart illustrating a method for improving image quality according to one embodiment of the present disclosure.

[0071] The method illustrated in FIG. 8 is executed by the aforementioned image quality enhancement model (10) or an electronic device equipped with the same (hereinafter, image quality enhancement device), and the image quality enhancement device may be executed on a computing device. The image quality enhancement device may include a computer-readable storage having instructions stored internally connected to one or more processors that can be used by the computing device to perform each function.

[0072] The image quality enhancement device generates a noise class map corresponding to the noise intensity of the input frame (S800). The image quality enhancement device can generate a noise class map with the same width and height as the input frame. Here, the value of each pixel of the noise class map may be a normalized value of the noise class, which is the result of classifying the noise intensity of the input frame.

[0073] The image quality enhancement device can obtain the noise class of an input frame by using a noise class classification network (210) that has been pre-trained to classify the noise intensity of an input frame. According to embodiments, the image quality enhancement device may use a target input frame to be improved and one or more adjacent input frames adjacent to the target input frame as input frames. The image quality enhancement device may detect a scene change between the target input frame and the adjacent input frames, and input the target input frame to the noise class classification network (210) only when a scene change is detected. Here, the image quality enhancement device may calculate a Peak Signal to Noise Ratio (PSNR) based on the pixel-to-pixel difference value between the adjacent input frame and the target input frame, and detect a scene change by comparing the PSNR with a preset threshold.

[0074] The image quality enhancement device merges the input frame and the noise class map to generate a merged frame (S800).

[0075] The image quality enhancement device inputs the merged frame into a pre-trained image quality enhancement network (110) to generate an output frame with improved image quality compared to the input frame (S820).

[0076] FIG. 9 is a block diagram schematically showing an image quality improvement device according to one embodiment of the present disclosure.

[0077] Since steps S900 to S920 of FIG. 9 may be identical to or corresponding to steps S800 to S820 of FIG. 8, detailed explanations regarding overlapping content are omitted.

[0078] The learning device generates a noise class map corresponding to the noise intensity of the input frame (S900). Here, the input frame is a frame with degraded image quality compared to the target frame, and the learning device may store a pair of the target frame and the input frame in advance, or generate the input frame from the target frame.

[0079] The learning device merges the input frame and the noise class map to generate a merged frame (S910).

[0080] The learning device inputs the merged frame into the image quality improvement network (110) to predict the output frame (S920).

[0081] The learning device trains the image quality improvement network (110) using the target frame and the output frame (S930).

[0082] As described above, according to one embodiment of the present disclosure, by merging a noise class map containing information on the noise intensity of an input frame into the input frame and inputting it to the image quality improvement network (110), the image quality improvement network (110) can be trained according to the noise intensity. Accordingly, the image quality improvement network (110) can select different network parameters according to the noise intensity of the input frame.

[0083] FIGS. 10a to 10c are illustrative diagrams for explaining the effects of an image quality improvement model according to one embodiment of the present disclosure.

[0084] FIG. 10a illustrates an input frame containing noise, FIG. 10b illustrates an output image of a quality improvement model that does not have a noise classifier according to a comparative example, and FIG. 10bc illustrates an output image of a quality improvement model (10) according to one embodiment of the present disclosure.

[0085] When training is performed all at once without dividing the noise intensity of the input images, the details of the images with low noise are lost due to the influence of the images with high noise. Referring to Fig. 10b, it can be seen that the details of the trees are lost in the output image of the image quality enhancement model that does not have a noise classifier.

[0086] On the other hand, in the case of the image quality improvement model (10) according to one embodiment of the present disclosure, a noise classifier (100) is provided at the input end, and the image quality improvement network (110) is trained according to the noise intensity. Accordingly, as shown in FIG. 10c, it can be confirmed that noise is removed while maintaining detail compared to the comparative embodiment.

[0087] Although FIGS. 8 and 9 describe the processes as being executed sequentially, this is merely an illustrative explanation of the technical concept of one embodiment of the present disclosure. In other words, a person skilled in the art to which one embodiment of the present disclosure belongs may modify and adapt the process in various ways, such as changing the order described in FIGS. 8 and 9 or executing one or more of the processes in parallel, without departing from the essential characteristics of one embodiment of the present disclosure; therefore, FIGS. 8 and 9 are not limited to a chronological order.

[0088] Various embodiments of the systems and techniques described herein may be realized as digital electronic circuits, integrated circuits, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented as one or more computer programs executable on a programmable system. A programmable system comprises a storage system, at least one input device, and at least one programmable processor (which may be a special-purpose processor or a general-purpose processor) coupled to receive data and instructions from at least one output device and to transmit data and instructions to them. Computer programs (which are also known as programs, software, software applications, or code) include instructions for the programmable processor and are stored on a "computer-readable recording medium."

[0089] Computer-readable recording media include all types of recording devices in which data that can be read by a computer system is stored. Such computer-readable recording media may be non-volatile or non-transitory media such as ROM, CD-ROM, magnetic tape, floppy disk, memory card, hard disk, magneto-optical disk, and storage device, and may also include transitory media such as data transmission media. Additionally, computer-readable recording media may be distributed across networked computer systems, and computer-readable code may be stored and executed in a distributed manner.

[0090] Various embodiments of the systems and techniques described herein may be implemented by a programmable computer. Here, the computer includes a programmable processor, a data storage system (including volatile memory, non-volatile memory, or other types of storage systems, or a combination thereof), and at least one communication interface. For example, the programmable computer may be one of a server, a network device, a set-top box, an embedded device, a computer expansion module, a personal computer, a laptop, a PDA (Personal Data Assistant), a cloud computing system, or a mobile device.

[0091] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment. Explanation of the symbols

[0092] 10: Image quality improvement model 100: Noise Classifier 110: Image Quality Improvement Network

Claims

Claim 1 A method for improving image quality, comprising: a process of generating a noise class map corresponding to the noise intensity of an input frame; a process of merging the input frame and the noise class map to generate a merged frame; and a process of inputting the merged frame into a pre-trained image quality improvement network to generate an output frame with improved image quality compared to the input frame, wherein the process of generating the noise class map includes a process of predicting a single noise class representing the noise intensity of the input frame using a pre-trained noise class classification network to classify the noise intensity of the input frame into one of a pre-defined number of noise classes. Claim 2 delete Claim 3 A method for improving image quality according to claim 1, wherein the process of generating the noise class map comprises the process of generating the noise class map having the same width and height as the input frame and having all pixels having the same value corresponding to the noise class predicted from the input frame, wherein the value of each pixel of the noise class map is a value obtained by normalizing the predicted noise class based on the number of noise classes. Claim 4 A method for improving image quality according to claim 1, wherein the input frame comprises a target input frame to be improved and one or more adjacent input frames adjacent to the target input frame, and the process of generating the noise class map comprises: a process of detecting a scene change between the target input frame and the adjacent input frames; and a process of inputting the target input frame into the noise class classification network only when the scene change is detected. Claim 5 In claim 4, the process of detecting scene changes comprises: a process of calculating a Peak Signal to Noise Ratio (PSNR) based on a pixel-to-pixel difference value between the target input frame and the adjacent input frame; and a process of comparing the PSNR with a preset threshold, thereby improving image quality. Claim 6 A computer program stored on a computer-readable recording medium to execute each process included in the image quality improvement method according to any one of paragraphs 1, 3 through 5. Claim 7 A picture quality enhancement device comprising: a memory for storing instructions; and at least one processor, wherein the at least one processor generates a noise class map corresponding to the noise intensity of an input frame by executing the instructions, generates a merged frame by merging the input frame and the noise class map, and generates an output frame with improved picture quality relative to the input frame by inputting the merged frame into a pre-trained picture quality enhancement network, wherein the at least one processor is configured to predict a single noise class representing the noise intensity of the input frame by using a pre-trained noise class classification network to classify the noise intensity of the input frame into one of a pre-defined number of noise classes.