Apparatus and method for upscaling resolution based on slice images

By dividing low-resolution images into slices and using cascaded and upgrade blocks of convolutional neural networks, the problem of increased memory and computational load on convolutional neural networks in on-chip systems is solved, achieving efficient image resolution upgrade.

CN113538326BActive Publication Date: 2025-11-07SILICON WORKS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110332433.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-29
Publication Date
2025-11-07
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Existing single-image super-resolution technologies based on convolutional neural networks suffer from increased memory and computational demands in on-chip system implementations, making them unsuitable for upgrading display device resolutions.

Method used

High-resolution images are generated by dividing low-resolution images into multiple slices and processing them using a convolutional neural network, including cascaded and upgraded blocks, adjusting the size of the convolutional filters, and training the neural network using a pixel similarity loss function.

Benefits of technology

It enables efficient upscaling of low-resolution images to high-resolution images on a system-on-a-chip, reducing memory requirements and computational load, and is suitable for resolution upgrades of display devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113538326B_ABST
    Figure CN113538326B_ABST
Patent Text Reader

Abstract

The present application relates to an apparatus and method for upscaling resolution based on slice images. The present disclosure provides an apparatus for upscaling resolution based on slice images, in which a low-resolution image is divided into a plurality of slice images to generate a high-resolution image. The apparatus includes a convolution operation unit configured to convert a low-resolution input slice image into a high-resolution output slice image using a convolutional neural network. The convolutional neural network includes a cascade block configured to perform a convolution operation using a convolution filter having a predetermined size and a residual operation on an input feature map generated from the low-resolution input slice image to generate an output feature map, and an upscaling block configured to upscale the output feature map to generate the high-resolution output slice image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to image processing, and more particularly to a technique for increasing resolution of an image. BACKGROUND

[0002] Recently, a display apparatus capable of outputting an image of 8K resolution, which is higher than 4K resolution, which is an ultra-high definition (UHD) resolution, has been released. However, conventional broadcast content and video content are produced only in 2K or 4K resolution compared to the resolution of the display apparatus, and thus a technique for converting a low resolution (LR) image into a high resolution (HR) image is being developed.

[0003] As an example of an image conversion technique, a single image super resolution (SISR) technique has been proposed. SISR refers to a technique for generating an HR image corresponding to a single LR image. In particular, recently, with the development of deep learning technology, an SISR technique using a convolutional neural network (CNN) is common.

[0004] However, a conventional CNN-based SISR algorithm has many layers and filters, and thus there is a limitation that the conventional CNN-based SISR algorithm is not suitable for a system on chip (SoC) implementation because the number of memories and the amount of computation inevitably increase. SUMMARY

[0005] The disclosure is designed to solve the above problems and to provide an apparatus and a method for upscaling resolution based on slice images in which an image of low resolution is divided into a plurality of slice images, thereby generating an image of high resolution.

[0006] The disclosure is also designed to provide an apparatus and a method for upscaling resolution based on slice images in which the size of a vertical receptive field is reduced by adjusting the size of a convolution filter.

[0007] The disclosure is also designed to provide an apparatus and a method for upscaling resolution based on slice images in which, when training a neural network, a loss function based on similarity between pixels in an original image and similarity between pixels in an output image is used to adjust parameters of a convolution filter.

[0008] One aspect of the present disclosure provides an apparatus for upgrading resolution based on slice images, the apparatus including a convolution operation unit configured to convert a low-resolution input slice image into a high-resolution output slice image using a convolutional neural network. The convolutional neural network includes a cascading block configured to perform a convolution operation using a convolution filter having a predetermined size and a residual operation on an input feature map generated from the low-resolution input slice image to generate an output feature map, and an upgrading block configured to upgrade the output feature map to generate the high-resolution output slice image.

[0009] Another aspect of the present disclosure provides a method of upgrading resolution based on slice images, the method including dividing a low-resolution input image and acquiring a plurality of low-resolution input slice images, performing a convolution operation using a convolution filter having a predetermined size and a residual operation on an input feature map generated from the low-resolution input slice image by a convolutional neural network and generating an output feature map, upgrading the output feature map and generating a high-resolution output slice image, and sequentially concatenating the high-resolution output slice images corresponding to the low-resolution input slice images and generating a high-resolution output image. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. In the drawings:

[0011] Figure 1 is a schematic block diagram illustrating a configuration of an apparatus for upgrading resolution based on slice images according to an embodiment of the present disclosure;

[0012] Figure 2 is a conceptual diagram illustrating a process of upgrading a low-resolution input slice image to a high-resolution output slice image by a convolutional network according to the present disclosure;

[0013] Figure 3 is a schematic diagram illustrating Figure 1 a configuration of the convolutional network shown;

[0014] Figure 4 is a block diagram illustrating Figure 3 a configuration of the cascading block shown;

[0015] Figure 5A is a block diagram illustrating Figure 4 a configuration of the first residual block shown;

[0016] Figure 5B is a block diagram illustrating Figure 4 a configuration of the second residual block shown;

[0017] Figure 6 is a block diagram illustrating a configuration of an upgrading block shown in Figure 4

[0018] Figure 7 is an exemplary block diagram of a method of upgrading an output feature map by a factor of 4 by connecting two first upgrading blocks in series as shown in Figure 6

[0019] Figure 8A is an example graph illustrating a comparison of an output image converted to have a high resolution according to the present disclosure with a high resolution original image and an output image converted by another algorithm;

[0020] Figure 8B is another example graph illustrating a comparison of an output image converted to have a high resolution according to the present disclosure with a high resolution original image and an output image converted by another algorithm;

[0021] Figure 9 is a flowchart illustrating a method of upgrading a resolution based on a slice image according to an embodiment of the present disclosure; and

[0022] Figure 10 is a flowchart illustrating a method of generating an output feature map using a cascaded block performed by an upgrading apparatus according to the present disclosure. DETAILED DESCRIPTION

[0023] In the specification, it should be noted that, insofar as possible, similar reference numerals are used for elements that have been used to designate similar elements in other drawings. In the following description, detailed descriptions of functions and configurations that are known to those skilled in the art will be omitted when they are not related to the essential configurations of the present disclosure. The terms described in the specification should be understood as follows.

[0024] Advantages and features of the present disclosure and methods of accomplishing the same will be described by embodiments described below with reference to the accompanying drawings. The present disclosure may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that the disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Also, the present disclosure is defined only by the scope of the claims.

[0025] The shapes, sizes, ratios, angles, and numbers disclosed in the accompanying drawings for describing the embodiments of the present disclosure are merely examples and thus the present disclosure is not limited to the illustrated details. Like reference numerals refer to like elements throughout. In the following description, detailed descriptions of functions or configurations that are determined to make the gist of the present disclosure unnecessarily obscure will be omitted.

[0026] ​​In the case of using "include", "have", and "comprise" in the specification, another component can be added unless "only" is used. The singular form can include the plural form unless the contrary is mentioned.

[0027] In interpreting elements, although not explicitly described, the elements are interpreted to include an error range.

[0028] In describing the time relationship, for example, when the time sequence is described as "after", "followed by", "next", and "before", unless "only" or "directly" is used, discontinuous cases can be included.

[0029] It will be understood that, although the terms "first", "second", etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the present disclosure.

[0030] The term "at least one of" should be construed as including any and all combinations of one or more of the associated listed items. For example, the meaning of "at least one of a first item, a second item, and a third item" indicates all of the items from a combination of two or more of the first item, the second item, and the third item and a combination of the first item, the second item, or the third item.

[0031] The features of various embodiments of the present disclosure can be partially or wholly coupled or combined with each other and can interoperate with each other and be technically driven in various ways, as can be sufficiently understood by those skilled in the art. Embodiments of the present disclosure can be executed independently of each other or can be executed together in a mutually dependent relationship.

[0032] Hereinafter, embodiments of the present specification will be described in detail with reference to the accompanying drawings.

[0033] Figure 1 is a block diagram illustrating an apparatus for upgrading resolution based on a sliced image according to an embodiment of the present disclosure. Figure 1 The illustrated sliced image-based upgrading apparatus 100 (hereinafter, referred to as "upgrading apparatus") uses a convolutional neural network (CNN)-based super-resolution (SR) technique to convert an input image of low resolution (LR) into an output image of high resolution (HR).

[0034] Specifically, the upgrading device 100 according to the present disclosure can acquire a plurality of input slice images from an LR input image to implement super-resolution based on the slice images. To this end, the upgrading device 100 according to the present disclosure includes an image division unit 110, a convolution operation unit 120, a CNN 130, an image concatenation unit 140, and a training unit 150, as shown in Figure 1

[0035] The image division unit 110 receives an LR input image from an external device and acquires a plurality of input slice images from the received input image. Specifically, as shown in Figure 2

[0036] In the present disclosure, the reason why the input slice images are composed of a plurality of horizontal lines is as follows. If the input slice images are composed of a single horizontal line, this inevitably is a limitation on the performance of super-resolution, because the performance of super-resolution is proportional to the number of horizontal lines. On the other hand, if the input slice images are composed of a plurality of horizontal lines, it is possible to maintain a performance similar to that of super-resolution using frame memory, and at the same time, it is easier to implement hardware because frame memory is not required.

[0037] In an embodiment, the image division unit 110 can divide the LR input image 210 in units of 15 horizontal lines. Accordingly, the image division unit 110 can generate a plurality of input slice images 210a to 210n each composed of 15 horizontal lines.

[0038] In the present disclosure, the reason why the LR input image is divided into a plurality of input slice images by the image division unit 110 is as follows. In general SR technology using a CNN, many frame memories are required, and thus it is difficult to reduce the weight of the CNN, making it difficult to implement a system on chip (SoC). However, as proposed in the present disclosure, when slice images are used as input images, it is possible to replace frame memories with line memories, thereby making it easier to make the CNN 130 lightweight and implement an SoC.

[0039] The image division unit 110 sequentially inputs the plurality of generated input slice images 210a to 210n as input images of the CNN 130 through the convolution operation unit 120.

[0040] Referring again to Figure 1 ​​The convolution operation unit 120 converts the LR input slice images into HR output slice images using the CNN 130. The convolution operation unit 120 performs a convolution operation and a residual operation performed through the CNN 130 to convert the LR input slice images into the HR output slice images. The convolution operation unit 120 provides the converted HR output slice images to the image concatenation unit 140.

[0041] When the LR input slice images are input from the convolution operation unit 120, the CNN 130 performs a plurality of convolution operations based on the input slice images to generate output feature maps and upscales the output feature maps to generate the HR output slice images.

[0042] Hereinafter, a configuration of the CNN 130 according to the present disclosure will be described in more detail with reference to Figure 3 The CNN 130 according to the embodiment of the present disclosure includes an input convolution layer 310, a plurality of cascade blocks 320a to 320c, an upsampling block 330, and an output convolution layer 340.

[0043] Figure 3 is a schematic diagram illustrating a configuration of a convolutional network according to an embodiment of the present disclosure. As Figure 3 indicated, the CNN 130 according to the embodiment of the present disclosure includes the input convolution layer 310, the plurality of cascade blocks 320a to 320c, the upsampling block 330, and the output convolution layer 340.

[0044] The input convolution layer 310 performs a convolution operation on the input slice images 210a to 210n using a predetermined input convolution filter to generate input feature maps for the input slice images 210a to 210n. In an embodiment, the input convolution filter can be a convolution filter having a square size. For example, the input convolution filter can be a convolution filter having a size of 3x3, as Figure 3 indicated.

[0045] According to the above-described embodiment, the number of channels of the input feature maps can be determined by the number of channels of the input convolution filter. Specifically, when the input slice images 210a to 210n are composed of i channels and the input convolution filter is composed of j channels, the input feature maps are composed of j channels. For example, when the input slice images 210a to 210n are composed of three channels R, G, and B and the size of the input convolution filter is 3x3x64, the input feature maps composed of 64 channels are generated by the input convolution layer 310.

[0046] The plurality of cascade blocks 320a to 320c performs a convolution operation and a residual operation on the input feature maps input to each of the cascade blocks 320a to 320c and generates output feature maps based on the results of the operations. Here, the residual operation refers to an operation in which an input value passes through a number of layers using a skip connection structure and then an output value and an input thereto are calculated to obtain a result value.

[0047] The reason why CNN 130, as disclosed in this publication, includes cascaded blocks 320a to 320c for performing residual operations and generating output feature maps is as follows. Typical CNN architectures use backpropagation for training, and as the network structure becomes deeper, the gradient value of the next layer decreases exponentially based on the magnitude of the gradient of the previous layer.

[0048] Therefore, the gradients near the output layer have values, but the gradients near the input layer have values ​​close to zero, leading to the vanishing gradient problem where training no longer progresses. Therefore, the CNN 130 according to this disclosure includes cascaded blocks 320a to 320c capable of performing residual operations to solve the vanishing gradient problem, thus enabling efficient training even in deep network structures.

[0049] In the implementation method, such as Figure 3 As shown, multiple cascaded blocks 320a to 320c are connected in series, so the output feature map of the previous cascaded block becomes the input feature map of the next cascaded block.

[0050] exist Figure 3 In the original text, CNN 130 is described as comprising three cascaded blocks 320a to 320c, but this is merely exemplary, and in modified embodiments, CNN 130 may include only two cascaded blocks or include four or more cascaded blocks. However, as the number of cascaded blocks decreases, the accuracy of the output feature maps decreases, and as the number of cascaded blocks increases, the computational cost increases, making it difficult to reduce the weight of the network. Therefore, the number of cascaded blocks can be variably set depending on the environment in which CNN 130 is applied.

[0051] In the following text, for ease of description, it is assumed and described that CNN 130 consists of three cascaded blocks 320a to 320c.

[0052] First, a first concatenated block 320a is placed at the back end of the input convolutional layer 310, and performs convolution and residual operations on the input feature map output from the input convolutional layer 310 to generate a first output feature map based on the results of the operations. The first concatenated block 320a inputs the first output feature map into the second concatenated block 320b.

[0053] The second concatenated block 320b is positioned after the first concatenated block 320a and performs convolution and residual operations on the first output feature map output from the first concatenated block 320a to generate a second output feature map based on the results of the operations. The second concatenated block 320b then inputs the second output feature map into the third concatenated block 320c.

[0054] The third cascaded block 320c is disposed at a rear end of the second cascaded block 320b and performs a convolution operation and a residual operation on the second output feature map output from the second cascaded block 320b to generate a third output feature map based on a result of the operations. The third output feature map generated by the third cascaded block 320c becomes a final output feature map and is input to the upgrading block 330.

[0055] Hereinafter, a description will be given with reference to Figure 4 The configuration of the cascaded blocks 320a to 320c will be described in more detail. Since all configurations of the cascaded blocks 320a to 320c are the same, a description will be given based on the first cascaded block 320a in the first stage 310. Figure 4 Hereinafter, for convenience of description, the first cascaded block 320a will be described as a cascaded block 320.

[0056] Figure 4 is a schematic block diagram illustrating a configuration of a cascaded block according to the disclosure. As Figure 4 indicated, the cascaded block 320 according to the disclosure includes a first residual block 400, a first concatenation layer 410, a first dimension reduction layer 420, a second residual block 430, a second concatenation layer 440, and a second dimension reduction layer 450.

[0057] The first residual block 400 sequentially performs a group convolution operation using a first type of convolution filter and a point-wise convolution operation using a second type of convolution filter having a square size on the input feature map IFM, performs a residual operation based on a result of the operations, and generates a first residual feature map RFM_1. In an embodiment, the first type of convolution filter can be a filter in which a size of a vertical receptive field is smaller than a size of a horizontal receptive field, and the second type of convolution filter can be a filter having a square size in which the size of the vertical receptive field is the same as the size of the horizontal receptive field.

[0058] The reason why the first residual block 400 according to the disclosure performs the group convolution using the first type of convolution filter in which the size of the vertical receptive field is smaller than the size of the horizontal receptive field is as follows. In general, in the SR technology, since the size of the receptive field is directly related to performance, it is common to increase the size of the receptive field.

[0059] However, as the size of the receptive field increases, the amount of information to be stored in the CNN increases. Therefore, in the disclosure, by reducing the size of the vertical receptive field of the first type of convolution filter, the size of the feature map can be reduced in the vertical direction when the feature map passes through the first type of convolution filter, and the number of line memories required to store the feature map can be reduced. For example, the first residual block 400 according to the disclosure can use a filter in which the size of the vertical receptive field is one as the first type of convolution filter.

[0060] Hereinafter, a first residual block according to the present disclosure will be described with reference to Figure 5A The first residual block according to the present disclosure will be described in more detail. Figure 5A is a block diagram illustrating a configuration of the first residual block according to an embodiment of the present disclosure. As Figure 5A indicated, the first residual block 400 according to an embodiment of the present disclosure includes a first group convolutional block 510a, a second group convolutional block 520a, a point-wise convolutional block 530a, an operation unit 540a, and an output activation layer 550a.

[0061] In Figure 5A , for convenience of description, the first group convolutional block 510a and the second group convolutional block 520a are illustrated as having a first type convolutional filter having a size of 1x5, but a filter having a size of a vertical receptive field smaller than that of a horizontal receptive field is sufficient as the first type convolutional filter, and the first type convolutional filter can have any size other than the size of 1x5.

[0062] The first group convolutional block 510a performs a group convolution operation on the input feature map IFM and includes a first convolutional layer 512a and a first activation layer 514a.

[0063] The first convolutional layer 512a performs a group convolution on the input feature map IFM using the first type convolutional filter having a size of 1x5 to generate a first feature map FM_1, and the first activation layer 514a applies an activation function to the first feature map FM_1 to nonlinearize the first feature map FM_1.

[0064] In an embodiment, as Figure 5A indicated, the first activation layer 514a applies a rectified linear unit (ReLu) function in which positive pixel values among pixel values of the first feature map FM_1 are output without change and negative pixel values are output as zero to the first feature map FM_1, so that a nonlinear characteristic can be imparted to the first feature map FM_1.

[0065] The second group convolutional block 520a performs a group convolution operation on the first feature map FM_1 generated by the first group convolutional block 510a and includes a second convolutional layer 522a and a second activation layer 524a.

[0066] The second convolution layer 522a performs group convolution on the first feature map FM_1 using first-type convolution filters having a size of 1x5 to generate a second feature map FM_2, and the second activation layer 524a applies an activation function to the second feature map FM_2 to nonlinearize the second feature map FM_2. In an embodiment, the second activation layer 524a applies a ReLu function in which positive pixel values among pixel values of the second feature map FM_2 are outputted without change and negative pixel values are outputted as zero to the second feature map FM_2, so that a nonlinear characteristic can be imparted to the second feature map FM_2.

[0067] As described above, according to the present disclosure, by performing group convolution on the input feature map IMF via the first group convolution block 510a and the second group convolution block 520a, the number of parameters and the amount of computation can be reduced and channels having high correlation can be trained for each group compared to a conventional CNN.

[0068] The point-wise convolution block 530a performs point-wise convolution on the second feature map FM_2 using point-wise convolution filters having a size of 1x1 to generate a third feature map FM_3. Unlike the first group convolution block 510a and the second group convolution block 520a, the point-wise convolution block 530a does not handle spatial characteristics and performs an operation only between channels. Accordingly, the point-wise convolution block 530a uses point-wise convolution filters having a fixed size of 1x1, so that the size of the output feature map is not changed and only the number of channels is adjusted.

[0069] As described above, according to the present disclosure, by combining the first group convolution block 510a and the second group convolution block 520a and the point-wise convolution block 530a, the amount of computation can be reduced compared to a conventional convolution operation.

[0070] The operation unit 540a calculates a sum of the third feature map FM_3 and the input feature map IMF to generate a fourth feature map FM_4. In the present disclosure, the reason why the sum of the third feature map FM_3 and the input feature map IMF is calculated by the operation unit 540a is to prevent a vanishing problem in which features become blurred as the depth in the CNN increases, and at the same time to simplify matters to be trained by allowing a difference between the input feature map IMF and the third feature map FM_3 to be trained.

[0071] The output activation layer 550a applies an activation function to the fourth feature map FM_4 outputted from the operation unit 540a to nonlinearize the fourth feature map FM_4 and generate a first residual feature map RFM_1. In an embodiment, the output activation layer 550a applies a ReLu function in which positive pixel values among pixel values of the fourth feature map FM_4 are outputted without change and negative pixel values are outputted as zero to the fourth feature map FM_4, so that a nonlinear characteristic can be imparted to the fourth feature map FM_4.

[0072] Referring again to Figure 4 , the first concatenation layer 410 concatenates the input feature map IFM and the first residual feature map RFM_1, and inputs the concatenated feature map to the first dimension reduction layer 420. For example, when the input feature map IFM has 64 channels and the first residual feature map RFM_1 has 64 channels, the first concatenation layer 410 concatenates the input feature map IFM and the first residual feature map RFM_1 to generate a concatenated result of 128 channels, and inputs the generated concatenated result to the first dimension reduction layer 420.

[0073] The first dimension reduction layer 420 reduces the dimension of the concatenated result generated by the first concatenation layer 410 to the same dimension as the input feature map IFM to generate a first dimension-reduced feature map DRFM_1. In an embodiment, the first dimension reduction layer 420 performs a convolution operation on the concatenated result using a dimension reduction convolution filter having a size of 1x1 and the same number of channels as the input feature map IFM to generate the first dimension-reduced feature map DRFM_1. For example, when the concatenated result generated by the first concatenation layer 410 has 128 channels, the first dimension reduction layer 420 applies a dimension reduction convolution filter to the concatenated result to reduce the dimension of the concatenated result to 64 channels.

[0074] The second residual block 430 sequentially performs a group convolution operation using a third type of convolution filter and a point-wise convolution operation using a second type of convolution filter on the first dimension-reduced feature map DRFM_1 input from the first dimension reduction layer 420, and performs a residual operation based on the results of the operations to generate a second residual feature map RFM_2. In an embodiment, the third type of convolution filter can be a filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field.

[0075] Unlike the first residual block 400, the second residual block 430 according to the disclosure uses a third type of convolution filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field. This is because when a filter in which the size of the vertical receptive field is smaller than the size of the horizontal receptive field is repeatedly used, the performance of the CNN 130 can be degraded.

[0076] However, in the second residual block 430 according to the disclosure, even when a filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field is used, only horizontal padding is performed without performing vertical padding when performing a convolution operation to minimize an increase in the number of line memories, and thus the size of a feature map can be reduced in the vertical direction every time the feature map passes through the third type of convolution filter.

[0077] Hereinafter, the second residual block according to the disclosure will be described in more detail with reference to Figure 5B FIG. 2.Figure 5B is a block diagram illustrating a configuration of a second residual block according to an embodiment of the disclosure. As Figure 5B indicated, the second residual block 430 according to an embodiment of the disclosure includes a first group convolution block 510b, a second group convolution block 520b, a point-wise convolution block 530b, an operation unit 540b, and an output activation layer 550b.

[0078] In Figure 5B , for convenience of description, the first group convolution block 510b and the second group convolution block 520b are illustrated as using a third type convolution filter having a size of 3x3, but a filter having a size of a vertical receptive field equal to that of a horizontal receptive field is sufficient as the third type convolution filter, and thus the third type convolution filter can have any size other than the size of 3x3.

[0079] The first group convolution block 510b performs a group convolution operation on the first dimension-reduced feature map DRFM_1 and includes a first convolution layer 512b and a first activation layer 514b.

[0080] The first convolution layer 512b performs group convolution on the first dimension-reduced feature map DRFM_1 using a third type convolution filter having a size of 3x3 to generate a fifth feature map FM_5, and the first activation layer 514b applies an activation function to the fifth feature map FM_5 to nonlinearize the fifth feature map FM_5. In this case, when the group convolution is performed, the first convolution layer 512b performs only horizontal padding without performing vertical padding.

[0081] In an embodiment, as Figure 5B indicated, the first activation layer 514b applies a ReLu function in which positive pixel values among pixel values of the fifth feature map FM_5 are output without change and negative pixel values are output as zero to the fifth feature map FM_5, so that a nonlinear characteristic can be imparted to the fifth feature map FM_5.

[0082] The second group convolution block 520b performs a group convolution operation on the fifth feature map FM_5 generated by the first group convolution block 510b and includes a second convolution layer 522b and a second activation layer 524b.

[0083] The second convolution layer 522b performs group convolution on the fifth feature map FM_5 using a third type convolution filter having a size of 3x3 to generate a sixth feature map FM_6, and the second activation layer 524b applies an activation function to the sixth feature map FM_6 to nonlinearize the sixth feature map FM_6. In this case, when the group convolution is performed, the second convolution layer 522b performs only horizontal padding without performing vertical padding.

[0084] In an embodiment, the second activation layer 524b applies a ReLu function in which positive pixel values among pixel values of the sixth feature map FM_6 are outputted without change and negative pixel values are outputted as zero to the sixth feature map FM_6, so that a non-linear characteristic can be imparted to the sixth feature map FM_6.

[0085] As described above, according to the present disclosure, by performing group convolution on the first size-reduced feature map DRFM_1 through the first group convolution block 510b and the second group convolution block 520b, the number of parameters and the amount of calculation can be reduced and channels having high correlation can be trained for each group compared to a conventional CNN.

[0086] The point-wise convolution block 530b performs point-wise convolution on the sixth feature map FM_6 using a point-wise convolution filter having a size of 1x1 to generate a seventh feature map FM_7. Unlike the first group convolution block 510b and the second group convolution block 520b, the point-wise convolution block 530b does not handle spatial characteristics and performs an operation only between channels. Accordingly, the point-wise convolution block 530b uses a point-wise convolution filter having a fixed size of 1x1, so that the size of the output feature map does not change and only the number of channels is adjusted.

[0087] As described above, according to the present disclosure, by combining the first group convolution block 510b and the second group convolution block 520b and the point-wise convolution block 530b, the amount of calculation can be reduced compared to a conventional convolution operation.

[0088] The operation unit 540b calculates a sum of the seventh feature map FM_7 and the first size-reduced feature map DRFM_1 to generate an eighth feature map FM_8. In the present disclosure, the reason why the sum of the seventh feature map FM_7 and the first size-reduced feature map DRFM_1 is calculated by the operation unit 540b is to prevent a vanishing problem that causes features to become blurred as the depth in the CNN increases, and at the same time to simplify matters to be trained by allowing training on a difference between the seventh feature map FM_7 and the first size-reduced feature map DRFM_1.

[0089] The output activation layer 550b applies an activation function to the eighth feature map FM_8 outputted from the operation unit 540b to non-linearize the eighth feature map FM_8 and generate a second residual feature map RFM_2. In an embodiment, the output activation layer 550b applies a ReLu function in which positive pixel values among pixel values of the eighth feature map FM_8 are outputted without change and negative pixel values are outputted as zero to the eighth feature map FM_8, so that a non-linear characteristic can be imparted to the eighth feature map FM_8.

[0090] In the above-described embodiments, the first residual block 400 and the second residual block 430 are described as using a ReLu function as an activation function for imparting a non-linear characteristic to a feature map, but this is merely exemplary, and the first residual block 400 and the second residual block 430 can impart a non-linear characteristic to a feature map using another activation function other than the ReLu function.

[0091] Referring again to Figure 4 , the second concatenation layer 440 concatenates the input feature map IFM, the first residual feature map RFM_1, and the second residual feature map RFM_2, and inputs the concatenated feature map to the second dimension reduction layer 450. For example, when the input feature map IFM has 64 channels, the first residual feature map RMF_1 has 64 channels, and the second residual feature map RFM_2 has 64 channels, the second concatenation layer 440 concatenates the input feature map IFM, the first residual feature map RFM_1, and the second residual feature map RFM_2 to generate a concatenation result of 192 channels, and inputs the generated concatenation result to the second dimension reduction layer 450.

[0092] The second dimension reduction layer 450 reduces the dimension of the concatenation result generated by the second concatenation layer 440 to the same dimension as the input feature map IFM to generate an output feature map OFM. In an embodiment, the second dimension reduction layer 450 performs a convolution operation on the concatenation result using a dimension reduction convolution filter having a size of 1x1 and the same number of channels as the input feature MIFM to generate the output feature map OFM. For example, when the concatenation result generated by the second concatenation layer 440 has 192 channels, the second dimension reduction layer 450 applies a dimension reduction convolution filter to the concatenation result to reduce the dimension of the concatenation result to 64 channels.

[0093] Further, the upgrading device 100 according to the present disclosure can further include a plurality of line memories 350a to 350d each disposed at at least one of an input end and an output end of the cascade blocks 320a to 320c, as shown in Figure 3 Each of the line memories 350a to 350d can store at least one of a feature map input to each of the cascade blocks 320a to 320c and a feature map output from each of the cascade blocks 320a to 320c.

[0094] For example, a feature map input to the first cascade block 320a is stored in the first line memory 350a in a line unit, a feature map output from the first cascade block 320a is stored in the second line memory 350b, a feature map output from the second cascade block 320b is stored in the third line memory 350c, and a feature map output from the third cascade block 320c is stored in the fourth line memory 350d.

[0095] In this case, as described above, each of the cascade blocks 320a to 320c uses the first type convolution filter having a size of 1x5 when performing the convolution operation to reduce the size of the feature map in the vertical direction, and thus the number of line memories 350a to 350d can be reduced, and each of the cascade blocks 320a to 320c does not perform the vertical padding when performing the convolution operation using the third type convolution filter having a size of 3x3, and thus an increase in the number of line memories 350a to 350d can be minimized.

[0096] Referring again to Figure 3 , the upgrade block 330 upgrades the output feature map OFM output from the third cascade block 320c. Hereinafter, the upgrade block according to the present disclosure will be described in detail with reference to Figure 6 .

[0097] Figure 6 is a block diagram illustrating a configuration of the upgrade block according to an embodiment of the present disclosure. As Figure 6 indicated, the upgrade block 330 according to the present disclosure can include a first upgrade block 610 and a second upgrade block 620.

[0098] The first upgrade block 610 performs a convolution operation using the first type convolution filter and a first shuffle operation on the output feature map OFM to upgrade the output feature map OFM by a factor p. In an embodiment, as Figure 6 indicated, the first shuffle operation can be a pixel shuffle operation capable of upgrading the output feature map OFM by a factor of two, and the output feature map OFM is upgraded by the first shuffle operation by a factor of two.

[0099] Specifically, the first upgrade block 610 increases the number of output feature maps OFM by a square of p through a first upgrade convolution layer 612, and performs upgrading by arranging pixels included in the output feature maps OFM of the square of p in a feature map upgraded by a factor of p through a first shuffle layer 614.

[0100] For example, when p is 2, the first upsampling block 610 can increase the number of output feature maps OFM to four, which is a square of p, by the first upsampling convolution layer 612, and perform upsampling by arranging a pixel at a position (1, 1) of a first output feature map among the four output feature maps OFM as a pixel at a position (1, 1) of a feature map to be output, a pixel at a position (1, 1) of a second output feature map as a pixel at a position (1, 2) of the feature map to be output, a pixel at a position (1, 1) of a third output feature map as a pixel at a position (2, 1) of the feature map to be output, and a pixel at a position (1, 1) of a fourth output feature map as a pixel at a position (2, 2) of the feature map to be output, through the first permutation layer 614.

[0101] The second upsampling block 620 performs a convolution operation using a first type convolution filter and a second permutation operation on the output feature maps OFM to upsample the output feature maps OFM by a factor of q. In an embodiment, as shown in FIG. 6B, the second permutation operation can be a pixel permutation operation capable of upsampling the output feature maps OFM by a factor of three, and thus the output feature maps OFM are upsampl ed by a factor of three. Figure 6

[0102] Specifically, the second upsampling block 620 increases the number of output feature maps OFM to a square of q by the second upsampling convolution layer 622, and performs upsampling by arranging pixels included in the output feature maps OFM of the square of q in a feature map upsampl ed by a factor of q through the second permutation layer 624.

[0103] For example, when q is 3, the second upsampling block 620 can increase the number of output feature maps OFM to nine, which is a square of q, by the second upsampling convolution layer 622, and perform upsampling by arranging a pixel at a position (1, 1) of a first output feature map among the nine output feature maps OFM as a pixel at a position (1, 1) of a feature map to be output, a pixel at a position (1, 1) of a second output feature map as a pixel at a position (1, 2) of the feature map to be output, and a pixel at a position (1, 1) of a third output feature map as a pixel at a position (1, 3) of the feature map to be output, through the second permutation layer 624.

[0104] In addition, the second upsampling block 620 can perform upsampling by arranging a pixel at a position (1, 1) of a fourth output feature map as a pixel at a position (2, 1) of the feature map to be output, a pixel at a position (1, 1) of a fifth output feature map as a pixel at a position (2, 2) of the feature map to be output, and a pixel at a position (1, 1) of a sixth output feature map as a pixel at a position (2, 3) of the feature map to be output, through the second permutation layer 624. ​

[0105] Further, the second upsampling block 620 can perform upsampling by arranging a pixel at a position (1, 1) of the seventh output feature map as a pixel at a position (3, 1) of the feature map to be output, arranging a pixel at a position (1, 1) of the eighth output feature map as a pixel at a position (3, 2) of the feature map to be output, and arranging a pixel at a position (1, 1) of the ninth output feature map as a pixel at a position (3, 3) of the feature map to be output via the second permutation layer 624.

[0106] In Figure 6 , the upsampling block 330 is described as including the first upsampling block 610 and the second upsampling block 620, but the upsampling block 330 according to the disclosure can further include a third upsampling block 630 that upsamples the output feature map by connecting a plurality of upsampling blocks in series, each of the plurality of upsampling blocks including at least one of the first upsampling block 610 and the second upsampling block 620.

[0107] For example, when the third upsampling block 630 upsamples the output feature map OFM by a factor of four, the third upsampling block 630 sequentially performs both a convolution operation using a first type of convolution filter and a second permutation operation twice on the output feature map OFM to upsample the output feature map OFM by a factor of four, as shown in Figure 7 .

[0108] Referring again to Figure 3 , the output convolution layer 340 performs a convolution operation using a predetermined output convolution filter on the upscaled output feature map OFM_US to reduce the number of channels of the upscaled output feature map OFM_US. Accordingly, the output slice images 220a to 220n having the same number of channels as the input slice images are generated.

[0109] In an embodiment, the output convolution filter can be a convolution filter having a square size. For example, the output convolution filter can be a convolution filter having a size of 3x3, as shown in Figure 3 .

[0110] According to the above-described embodiment, when the upscaled output feature map OFM_US is composed of j channels and the output convolution filter is composed of i channels, the output slice images 220a to 220n are composed of i channels. For example, when the upscaled output feature map OFM_US is composed of 64 channels and the output convolution filter is composed of three channels, the output slice images 220a to 220n having three channels are generated by the output convolution layer 340.

[0111] Referring again to Figure 1When the output slice images 220a to 220n respectively corresponding to the input slice images 210a to 210n are output from the convolution operation unit 120, the image concatenation unit 140 concatenates the output slice images 220a to 220n in order to generate an HR output image 220. For example, as shown in FIG. 3, the image concatenation unit 140 concatenates the plurality of output slice images 220a to 220n output from the convolution operation unit 120 in order to generate the HR output image 220. Figure 2

[0112] The training unit 150 trains the CNN 130 using predetermined training images to optimize parameters of the convolution filters of each layer constituting the CNN 130. In this case, the training unit 150 can use an image patch of k x k size as the training image.

[0113] In an embodiment, when the training unit 150 trains the CNN 130, the training unit 150 can use two loss functions, i.e., L pixel as a first loss function and L relation as a second loss function, as described in the following Equation 1, and trains the CNN 130 so as to reduce a difference between an output image acquired based on the training image and an HR original image corresponding to the training image.

[0114] [Equation 1]

[0115] L total = L pixel + λ x L relation

[0116] In Equation 1, L pixel denotes a loss function in which the output image acquired based on the training image and the HR original image corresponding to the training image are compared in pixel units so that the CNN 130 is trained to reduce a difference therebetween, and a smoothL1 function described in the following Equation 2 can be used.

[0117] [Equation 2]

[0118]

[0119] In Equation 2, x denotes a pixel value difference between the HR original image and the output image. In the loss function described in Equation 2, a region in which a difference in pixel values between the HR original image and the output image is less than one (i.e., a region having a small amount of error) is a curve, and other regions are straight lines. Thus, when the number of errors is small, the loss value rapidly decreases.

[0120] As described above, in the present disclosure, the smoothL1 function is used as the first loss function, and thus a delay in the CNN 130 can be minimized.​

[0121] Further, in Equation 1, L relarion represents a loss function for which the CNN 130 is trained in order to reduce a difference between a first similarity between pixels of the HR original image and a second similarity between pixels of an HR output image output from the convolutional network based on the training image, and can be defined as in Equation 3 below.

[0122] [Equation 3]

[0123]

[0124] In Equation 3, represents the first similarity between pixels included in the HR original image and is defined as in Equation 4 below, and represents the second similarity between pixels of the HR output image output from the CNN 130 based on the training image and is defined as in Equation 5 below.

[0125] [Equation 4]

[0126]

[0127] [Equation 5]

[0128]

[0129] In Equations 4 and 5, C(x HR ) and C(x SR ) represent a normalization factor, i represents a specific pixel in each image, and j represents all possible pixels in each image.

[0130] Further, in Equation 1, the weight λ reflected in the second loss function is set to a value smaller than a value that stabilizes the image expression, and thus can allow the reflectance of the second loss function to be greater than the reflectance of the first loss function.

[0131] As described above, according to the present disclosure, the CNN 130 can be trained in order to reduce a difference between the first similarity obtained from the HR original image and the second similarity obtained from the HR output image by additionally using the second loss function, and thus can minimize performance degradation without a separate additional module.

[0132] Figure 8A and Figure 8B are example graphs illustrating a comparison of an output image converted to have a high resolution according to the present disclosure with an HR original image and an output image converted by another algorithm. As Figure 8A and Figure 8BAs shown, it can be seen that the method according to the disclosure has excellent image quality compared to a method using bicubic interpolation as another algorithm.

[0133] The upgrading apparatus 100 described above can be applied to a display apparatus. In this case, the upgrading apparatus 100 can be included in a timing controller of the display apparatus, or can be installed together with the timing controller on a board in which the timing controller is installed.

[0134] Hereinafter, a method of upgrading resolution based on a slice image according to the disclosure will be described with reference to Figure 9 A method of upgrading resolution based on a slice image according to the disclosure will be described.

[0135] Figure 9 is a flowchart illustrating a method of upgrading resolution based on a slice image according to an embodiment of the disclosure. Figure 9 The method of upgrading resolution based on a slice image as shown can be performed by an upgrading apparatus in Figure 1 as shown.

[0136] First, the upgrading apparatus divides an LR input image to obtain a plurality of input slice images (S900). In an embodiment, the upgrading apparatus can divide the LR input image into units of horizontal lines (for example, 15 horizontal lines) to obtain input slice images composed of a plurality of horizontal lines.

[0137] As described above, according to the disclosure, since a slice image is used as an input image, it is possible to replace frame memory with line memory, thereby making it easier to make a CNN lightweight and easier to implement an SoC.

[0138] Thereafter, the upgrading apparatus generates an output feature map using at least one of the cascade blocks included in the CNN (S910). Specifically, the upgrading apparatus obtains an input feature map from the input slice image, and performs a convolution operation and a residual operation on the obtained input feature map using a convolution filter having a predetermined size to generate an output feature map.

[0139] A method of generating an output feature map using a cascade block performed by an upgrading apparatus according to the disclosure will be described in more detail with reference to Figure 10 A method of generating an output feature map using a cascade block performed by an upgrading apparatus according to the disclosure will be described in more detail with reference to

[0140] Figure 10 is a flowchart illustrating a method of generating an output feature map using a cascade block performed by an upgrading apparatus according to the disclosure.

[0141] First, the upgrading apparatus performs a first set of convolution operations using a first type of convolution filter on the input feature map (S1010). In an embodiment, the first type of convolution filter can be a filter in which the size of a vertical receptive field is smaller than the size of a horizontal receptive field. For example, the first type of convolution filter can be a filter having a size of 1x5.

[0142] The reason why the upgrading device uses the first type of convolution filter in which the size of the vertical receptive field is smaller than the size of the horizontal receptive field when performing the first group of convolution operations is to reduce the size of the feature map in the vertical direction by reducing the size of the vertical receptive field whenever the feature map passes through the first type of convolution filter, so that the amount of line memory required to store the feature map is reduced.

[0143] In an embodiment, the upgrading device can repeatedly perform the first group of convolution operations multiple times, and can non-linearize the result of the group convolution operation using an activation function such as a ReLu function after performing each first group of convolution.

[0144] Thereafter, the upgrading device performs a point-wise convolution operation using a second type of convolution filter having a square size on the result of the first group of convolution operations (S1020). In an embodiment, the second type of convolution filter can be a filter having a square size in which the size of the vertical receptive field is equal to the size of the horizontal receptive field. For example, the second type of convolution filter can be a filter having a size of 1x1.

[0145] Thereafter, the upgrading device calculates a sum of the result of the point-wise convolution operation acquired in S1020 and the input feature map to generate a first residual feature map (S1030). As described above, by calculating a sum of the result of the point-wise convolution operation and the input feature map, it is possible to prevent the vanishing problem in which features become blurred as the depth of the CNN increases.

[0146] Thereafter, the upgrading device concatenates the input feature map and the first residual feature map (S1040), and then down-sizes the size of the concatenated result to the same size as the input feature map to generate a first down-sized feature map (S1050).

[0147] Thereafter, the upgrading device performs a second group of convolution operations using a third type of convolution filter on the first down-sized feature map (S1060). In an embodiment, the third type of convolution filter can be a filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field. For example, the third type of convolution filter can be a filter having a size of 3x3.

[0148] In the present disclosure, when performing the second group of convolution operations, a third type of convolution filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field is used. This is because when a filter in which the size of the vertical receptive field is smaller than the size of the horizontal receptive field is repeatedly used, the performance of the CNN can be reduced.

[0149] However, in the present disclosure, even when a filter in which the size of the vertical receptive field is equal to the size of the horizontal receptive field is used, only horizontal padding is performed without performing vertical padding when the second set of convolution operations is performed to minimize an increase in the number of line memories, and thus the size of the feature map can be reduced in the vertical direction each time the feature map passes through the third type convolution filter.

[0150] Thereafter, the upgrading device performs a point-wise convolution operation using the second type convolution filter (S1070), and calculates a sum of results of the point-wise convolution operation acquired in S1070 and the first size-reduced feature map to generate a second residual feature map (S1080).

[0151] Thereafter, the upgrading device concatenates the input feature map, the first residual feature map, and the second residual feature map (S1090), and reduces the size of the concatenated result to the same size as the input feature map to generate an output feature map (S1110).

[0152] Further, in the above-described embodiment, when the CNN includes the first to third cascade blocks, the first cascade block performs a convolution operation and a residual operation on an input feature map acquired from an input slice image to generate a first output feature map, the second cascade block performs a convolution operation and a residual operation on the first output feature map to generate a second output feature map, and the third cascade block performs a convolution operation and a residual operation on the second output feature map to generate a final output feature map.

[0153] Referring again to Figure 9 , the upgrading device upscales the output feature map generated in S910 by a factor of a predetermined multiple to generate an HR output slice image (S920).

[0154] In an embodiment, the upgrading device can sequentially perform a convolution operation using a first type convolution filter and a shuffling operation on the output feature map to upscale the output feature map by a factor of a predetermined multiple. The first type convolution filter has a size of a vertical receptive field that is smaller than a size of a horizontal receptive field.

[0155] For example, when upsampling by a factor of p is required, the upgrading device can increase the number of output feature maps to the square of p by using a convolution operation of the first type convolution filter, and perform upsampling by arranging pixels included in the output feature maps of the square of p in an upscaled feature map through a shuffling operation.

[0156] Thereafter, the upgrading device concatenates the output slice images corresponding to the input slice image in sequence to generate an HR output image (S930).

[0157] Further, although in Figure 9As not shown, the method of upgrading resolution based on a slice image according to the disclosure can further include a process of training a CNN using a predetermined training image.

[0158] In this case, the CNN can be trained using the loss function defined by the above-described Equations 1 to 4 in order to reduce a difference between an output image acquired based on a training image and an HR original image corresponding to the training image. Since the content of the loss function has been described in detail in Equations 1 to 4 described above, a detailed description thereof will be omitted.

[0159] According to the disclosure, since the LR input image is divided into a plurality of input slice images, and the divided input slice images are upgraded and then concatenated to acquire the HR output image, it is possible to implement the system using only line memory without frame memory, and thus, it is possible to make the convolution network lightweight, so that it is possible to easily implement SoC.

[0160] Further, according to the disclosure, since the residual block performs a convolution operation using a convolution filter in which the size of a vertical receptive field is smaller than the size of a horizontal receptive field, it is possible to reduce the size of the vertical receptive field. Thus, it is possible to reduce the number of line memories required to implement the system, so that it is possible to maximize the reduction of the weight of the convolution network.

[0161] Further, according to the disclosure, when the convolution layer in the residual block uses a convolution filter having a size of 3x3, vertical padding is not performed, so that it is possible to reduce the number of line memories that increase as more pass through the convolution layer.

[0162] Further, according to the disclosure, when the convolution network is trained, a SmoothL1 function is used as a loss function, so that it is possible to minimize network latency, and it is possible to improve the performance of the convolution network without an additional separate module by additionally using a similarity loss function defined as a difference between a first similarity between pixels included in an HR original image and a second similarity between pixels included in an HR output image output from the convolution network.

[0163] It will be understood by those skilled in the art that the present disclosure can be implemented in other specific forms without changing the technical idea and essential characteristics of the present disclosure.

[0164] All of the disclosed methods and processes described herein can be at least partially implemented using one or more computers or components thereof. The components can be provided as a series of computer instructions in any conventional computer readable medium or machine readable medium, including volatile and non-volatile memory, such as random access memory (RAM), read only memory (ROM), flash memory, magnetic or optical disks, optical memory, or other storage media. The instructions can be provided as software or firmware, and can be implemented with hardware configuration such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), or any other like device. The instructions can be configured to be executed by one or more processors or other hardware configurations, and when executing the series of computer instructions, the processor or other hardware configuration is allowed to perform all or part of the methods and processes disclosed herein.

[0165] Therefore, the above-described embodiments should be understood not to be limiting in every aspect. The scope of the disclosure will be defined by the appended claims rather than the above description, and all changes and modifications that come within the meaning and range of equivalency of the claims and their equivalents are to be embraced by the scope of the disclosure.

[0166] Cross Reference to Related Applications

[0167] This application claims the benefit of Korean Patent Application No. 10-2020-0038173, filed March 30, 2020, which is incorporated by reference as if fully set forth herein in its entirety.

Claims

1. An apparatus for upgrading resolution based on slice images, the apparatus comprising: a convolution operation unit configured to convert a low-resolution input slice image into a high-resolution output slice image using a convolutional neural network, wherein the convolutional neural network comprises: a cascaded block configured to perform a convolution operation using a convolution filter having a predetermined size and a residual operation on an input feature map generated from the low-resolution input slice image to generate an output feature map; and an upgrading block configured to upgrade the output feature map to generate the high-resolution output slice image, wherein the upgrading block is configured to perform the convolution operation using a first type of convolution filter and a shuffle operation on the output feature map to upgrade the output feature map by a factor of a predetermined multiple, and wherein the first type of convolution filter has a size of a vertical receptive field that is smaller than a size of a horizontal receptive field.

2. The apparatus of claim 1, wherein, the cascaded block comprises: a first residual block configured to sequentially perform a group convolution operation using the first type of convolution filter and a point-wise convolution operation using a second type of convolution filter having a square size on the input feature map and perform the residual operation based on a result of the operations to generate a first residual feature map; and a first size reduction layer configured to concatenate the input feature map with the first residual feature map and reduce a size of a result of the concatenation to a same size as the input feature map to generate a first size reduced feature map. the first type of convolution filter is a filter having a size of 1×5, and 3. The apparatus of claim 2, wherein, the second type of convolution filter is a filter having a size of 1×1. 4.The apparatus of claim 2, further comprising: a second residual block connected in series to the first size reduction layer, sequentially performing a group convolution operation using a third type of convolution filter having a square size and a point-wise convolution operation using the second type of convolution filter on the first size reduced feature map and performing the residual operation based on a result of the operations to generate a second residual feature map; and a second size reduction layer configured to concatenate the input feature map, the first residual feature map, and the second residual feature map and reduce a size of a result of the concatenation to a same size as the input feature map to generate the output feature map. the third type of convolution filter is a filter having a size of 3×3, and when the second residual block performs the group convolution operation using the third type of convolution filter, the second residual block performs horizontal padding and does not perform vertical padding on the first size reduced feature map.

5. The apparatus of claim 4, wherein, the first residual block and the second residual block repeatedly perform the group convolution operation a plurality of times and nonlinearize a result obtained by performing the group convolution operation each time using an activation function. the convolutional neural network comprises n cascaded blocks, and 6. The apparatus of claim 4, wherein, ​ 7. The apparatus of claim 1, wherein, ​ The n cascade blocks are connected in series such that an output feature map of an (n-1)th cascade block becomes an input feature map of an nth cascade block.

8. The apparatus of claim 7, further comprising: A line memory is provided at at least one of an input end and an output end of each cascade block, and at least one of an input feature map input to the cascade block and an output feature map output from the cascade block is stored in a line unit.

9. The apparatus of claim 1, wherein, The upgrade block includes: a first upgrade block configured to perform a first convolution operation using the first type of convolution filter and a first shuffling operation on the output feature map to upgrade the output feature map by a factor p; and a second upgrade block configured to perform a second convolution operation using the first type of convolution filter and a second shuffling operation on the output feature map to upgrade the output feature map by a factor q.

10. The apparatus of claim 9, wherein, The upgrade block further includes a third upgrade block that upgrades the output feature map by a factor r by connecting a plurality of upgrade blocks in series, and wherein the plurality of upgrade blocks includes at least one of the first upgrade block and the second upgrade block. 11.The apparatus of claim 1, further comprising: an image dividing unit configured to divide a low-resolution input image in a plurality of horizontal line units to obtain a plurality of low-resolution input slice images composed of the plurality of horizontal lines; and an image concatenating unit configured to concatenate high-resolution output slice images output from the convolution operation unit into high-resolution output images corresponding to the low-resolution input slice images.

12. The device of claim 1, further comprising: a training unit configured to train the convolutional neural network using predetermined training images, wherein the training unit trains the convolutional neural network such that a difference between a first similarity between pixels included in a high-resolution original image with respect to the training images and a second similarity between pixels included in a high-resolution output image output from the convolutional neural network based on the training images is reduced.

13. The apparatus of claim 12, wherein, The training unit trains the convolutional neural network using a first loss function defined as the formula and a second loss function defined as the formula ​ In the formula, x represents a difference in pixel values between the high-resolution original image and the high-resolution output image, represents the first similarity and is defined as the formula represents the second similarity and is defined as the formula C(x HR ) and C(x SR ) represent normalization factors, i represents a particular pixel in each image, and j represents all possible pixels in each image. 14.A method of upgrading resolution based on slice images, the method comprising the steps of: dividing a low-resolution input image and obtaining a plurality of low-resolution input slice images; performing a convolution operation using a convolution filter having a predetermined size and a residual operation on an input feature map generated from the low-resolution input slice images by a convolutional neural network and generating an output feature map; upgrading the output feature map and generating high-resolution output slice images; and concatenating the high-resolution output slice images corresponding to the low-resolution input slice images in sequence and generating a high-resolution output image, wherein the step of generating the high-resolution output slice images includes performing a convolution operation using a first type of convolution filter and a shuffling operation on the output feature map to upgrade the output feature map by a factor of a predetermined multiple, and wherein the first type of convolution filter has a size of a vertical receptive field that is smaller than a size of a horizontal receptive field.

15. The method of claim 14, wherein, The step of generating the output feature map comprises the steps of: performing a first convolution operation and a first residual operation on the input feature map using a first type of convolution filter to generate a first residual feature map; performing a second convolution operation and a second residual operation on a feature map generated based on a result of concatenation of the input feature map and the first residual feature map using a second type of convolution filter having a square size to generate a second residual feature map; and generating the output feature map based on a result of concatenation of the input feature map, the first residual feature map, and the second residual feature map.

16. The method of claim 15, wherein, The steps of performing the first convolution operation and the second convolution operation comprise performing at least one of a group convolution operation and a point-wise convolution operation.

17. The method of claim 15, wherein, The first type of convolution filter is a filter having a size of 1x5, and The second type of convolution filter is a filter having a size of 3x3.

18. The method of claim 15, wherein, The step of performing the second convolution operation comprises performing a horizontal padding operation without performing a vertical padding operation.

19. The method of claim 14, further comprising the step of: The convolutional neural network is trained using predetermined training images, wherein, when training the convolutional neural network, the convolutional neural network is trained using a first loss function defined as and a second loss function defined as ​ In the formula, x represents a difference between pixel values between a high-resolution original image of the training image and the high-resolution output image output from the convolutional neural network based on the training image, represents a first similarity between pixels included in the high-resolution original image and is defined as a formula represents a second similarity between pixels included in the high-resolution output image and is defined as a formula C(x HR ) and C(x SR ) represent normalization factors, i represents a specific pixel in each image, and j represents all possible pixels in each image.

Citation Information

Patent Citations

  • Plasma processing apparatus and plasma processing method

    KR1020200038173A

  • Flexible semiconductor device, method for manufacturing the same, and image display device

    CN104425516A

  • Segmentation method of pathological section unconventional cells based on multi-scale hybrid segmentation model

    CN108447062A