IMAGE PROCESSING APPARATUS, LEARNING APPARATUS, METHOD, AND PROGRAM

The learning device optimizes neural network parameters to align coded and decoded images with original images, addressing the issue of low-quality images in existing pixel count conversion technologies by effectively considering image coding effects.

JP7678719B2Active Publication Date: 2025-05-16NIPPON HOSO KYOKAI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021107111
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-28
Publication Date
2025-05-16
Estimated Expiration
2041-06-28

AI Technical Summary

Technical Problem

Existing pixel count conversion technologies do not effectively consider the effects of image coding, resulting in low-quality images when pixel count conversion is performed in conjunction with encoding techniques.

Method used

A learning device that includes a frame acquisition unit, an image magnification unit, a learning neural network unit, a coding unit, and a decoding unit, which performs machine learning to optimize neural network parameters such that the coded and decoded image coincides with the original image, thereby achieving efficient pixel count conversion.

Benefits of technology

This approach enables efficient pixel count conversion while considering the effects of image coding, resulting in high-quality images when performed in conjunction with encoding techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678719000001
    Figure 0007678719000001
  • Figure 0007678719000002
    Figure 0007678719000002
  • Figure 0007678719000003
    Figure 0007678719000003
Patent Text Reader

Abstract

To provide an image processing device, a learning device, a method, and a program that enable efficient pixel number conversion considering the influence of image coding, and obtain high-quality images when pixel number conversion is performed in combination with an encoding technology.SOLUTION: An image processing device includes a frame memory that holds an input image, and a learning unit configured by a neural network and outputs a reduced image obtained by reducing the number of pixels of the input image from information on the input image, and the process of reducing the number of pixels in the learning unit is machine-learned such that the image that has undergone an image enlargement process and an encoding / decoding process matches the original image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device, a learning device, a method, and a program, and more particularly to an image processing device, a learning device, a method, and a program that realizes conversion of the number of pixels using machine learning. [Background technology]

[0002] In recent years, video resolution has been increasing. 8K Super Hi-Vision, which has a resolution 16 times that of current Hi-Vision broadcasts, began broadcasting in December 2018. Uncompressed 8K requires a massive amount of data, at 72Gbps (in the case of 8K / 4:4:4 / 12bit / 60Hz format), making compression essential for broadcasting.

[0003] To achieve ultra-high compression transmission of ultra-high definition video such as 8K Super Hi-Vision, there is a technology called super-resolution restoration video encoding system that combines conventional video encoding technology with super-resolution technology. On the sending side, the video resolution is reduced (pixel count conversion) and the encoding process is carried out before transmission. Meanwhile, on the receiving side, after decoding, the original resolution is restored using super-resolution technology. By sharing the compression using different mechanisms of resolution reduction and encoding, it is possible to reduce unsightly degradation such as block distortion.

[0004] Among conventional pixel count conversion techniques, well-known image enlargement techniques include the nearest neighbor method, bilinear interpolation, and bicubic interpolation. However, these have the drawback that the image becomes blurred overall during enlargement processing. On the other hand, super-resolution is widely used as a technique for increasing the number of pixels while suppressing image blur. Super-resolution methods include methods that use linear filters such as wavelet super-resolution, and methods that use nonlinear filters such as total variation and non-local super-resolution. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Toshie Misu et al., "Super-resolution restoration type video coding system", NHK STRL R&D, January 2016, internet,<URL:https: / / www.nhk.or.jp / strl / publica / rd / 155 / 5.html> Summary of the Invention [Problem to be solved by the invention]

[0006] Pixel count conversion technology is increasingly being used in the pre- or post-stages of image processing involving encoding and decoding. However, pixel count conversion technology has been researched as a standalone technology to date, and the effects of encoding have not been fully considered. As a result, it has not been possible to obtain high-quality images for images that have undergone both encoding and pixel count conversion processing.

[0007] Therefore, in consideration of the above-mentioned problems, an object of the present invention is to provide an image processing device, a learning device, a method, and a program that enable efficient pixel count conversion that takes into account the effects of image encoding, and that can obtain high-quality images when pixel count conversion is performed in conjunction with encoding technology. [Means for solving the problem]

[0010] above In order to solve the above problem, the learning device of the present invention includes a frame acquisition unit that acquires an input original training image, an image enlargement unit that generates an enlarged image by enlarging the original training image, a training neural network unit that generates a reduced image by reducing the number of pixels of the enlarged image, and an encoding unit and a decoding unit that encode and decode the reduced image to generate an encoded / decoded image, wherein the training neural network unit performs machine learning using the original training image as teacher data so that the encoded / decoded image matches the original training image, and outputs optimized neural network parameters.

[0013] In addition, in order to solve the above problem, a method according to the present invention is a method for generating parameters of a neural network to be applied to an image processing device, comprising the steps of: generating an enlarged image by enlarging an original training image; generating a reduced image by reducing the number of pixels of the enlarged image using a training neural network unit; encoding and decoding the reduced image to generate an encoded / decoded image; calculating an error between the encoded / decoded image and the original training image; and performing machine learning by modifying parameters of the training neural network unit based on the error so that the encoded / decoded image and the original training image match.

[0017] In order to solve the above-mentioned problems, the present invention provides a program for causing a computer to function as the learning device. Effect of the Invention

[0018] The image processing device, learning device, method, and program of the present invention enable efficient pixel count conversion that takes into account the effects of encoding, and can provide higher quality images when pixel count conversion is performed in conjunction with encoding technology. [Brief description of the drawings]

[0019] [Figure 1] FIG. 1 is a diagram illustrating an example of a pixel number reduction device as an image processing device according to a first embodiment. [Diagram 2] FIG. 1 is a diagram showing an example of a conceptual diagram of a neural network. [Diagram 3] FIG. 1 is a conceptual diagram of machine learning when reducing the number of pixels while taking into account the effects of encoding. [Figure 4] FIG. 13 is a block diagram illustrating an example of a learning device according to a second embodiment. [Diagram 5] 10 is a flowchart showing a process flow of a learning device according to a second embodiment. [Figure 6] FIG. 13 is another conceptual diagram of machine learning when reducing the number of pixels while taking into account the effects of encoding. [Figure 7]FIG. 13 is a block diagram showing an example of a modification of the learning device according to the second embodiment. [Figure 8] 13 is a flowchart showing the flow of processing in a modified example of the learning device of the second embodiment. [Figure 9] FIG. 13 is a diagram illustrating an example of a pixel number expansion device as an image processing device according to a third embodiment. [Figure 10] FIG. 1 is a conceptual diagram of machine learning when increasing the number of pixels while taking into account the effects of encoding. [Figure 11] FIG. 13 is a block diagram illustrating an example of a learning device according to a fourth embodiment. [Figure 12] 13 is a flowchart showing a process flow of a learning device according to a fourth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0021] (Pixel reduction technology) 1 shows an example of a pixel count reduction device as an image processing device according to the first embodiment. The pixel count reduction device 10 is composed of a frame memory 11 and a learning unit 12, and receives an image as an input, and outputs an image in which the number of pixels of the input image has been reduced (decreased). In the present invention, the term "image" includes a moving image (video).

[0022] The frame memory 11 holds the input original image and adjusts the timing of inputting the image to the learning unit 12. The learning unit 12 has a learning function, and converts the input image information into a reduced image (e.g., a pixel value string) in which the number of pixels of the input image is reduced, and outputs the reduced image. As will be described later, the pixel number reduction process of the learning unit 12 is machine-learned so that an image that has further undergone image enlargement processing and encoding / decoding processing matches the original image. The pixel number reduction device 10 in FIG. 1 can be realized, for example, by a computer including an input / output unit, a storage unit (memory), and a processing unit (CPU, etc.).

[0023] The learning unit 12 can be configured by a neural network. FIG. 2 shows an example of a conceptual diagram of a convolutional neural network that realizes the learning unit 12, which is configured by an input layer, an intermediate layer, and an output layer. Here, an image is input to the input layer. For example, data of each pixel value of the input image from the frame memory 11 is input to each node of the input layer. The intermediate layer can have a general configuration including one or more convolutional layers, pooling layers, activation layers, etc., and performs image conversion processing. In addition, the output layer performs processing using a softmax function, etc., and outputs an output image with a converted number of pixels. The learning unit 12 in this embodiment is a neural network that reduces the number of pixels in consideration of the effects of encoding, and it is desirable that the parameters in the network between each layer in FIG. 2 are optimized by prior learning, which will be described later. Note that, in addition to pixel values, multiple data converted into frequency decomposition values, image features, motion vectors, etc. may be used as the video data used in the convolutional neural network.

[0024] Next, the machine learning of the learning unit 12 of the pixel number reduction device 10 will be described. Fig. 3 is a conceptual diagram of machine learning when reducing the number of pixels while taking into account the effects of encoding. In the conceptual diagram of Fig. 3, the image is shown to be transformed in sequence from the left to the right. The flow of image transformation and learning will be described.

[0025] First, the number of pixels of the original image is increased using a technique such as super-resolution to generate an enlarged image. The technique used for this image enlargement process is not limited to super-resolution, and any pixel number conversion technique can be used. A reduced image of a desired size is created for this enlarged image by a process G using machine learning. This reduced image by machine learning can have, for example, the same number of pixels as the original image. The machine learning process G in FIG. 3 reduces the number of pixels while taking into account the effects of encoding, and is equivalent to the process of the learning unit 12 of the pixel number reduction device 10. The reduced image generated by machine learning is encoded and decoded to generate an encoded and decoded image (an image that has undergone encoding and decoding processes). At this time, the encoded and decoded image has the same number of pixels as the original image. The encoded and decoded image is compared with the original image, and the machine learning process G is optimized (adjusted parameters of the neural network, etc.) so that the two images become the same. This enables optimal image reduction while taking into account the effects of encoding.

[0026] Next, a learning device of the second embodiment will be described. The learning device of this embodiment is a parameter learning device that generates parameters suitable for the learning unit 12 of the pixel number reduction device 10 of the first embodiment. Fig. 4 is an example of a block diagram of the learning device, and Fig. 5 is a flowchart showing the flow of processing of the learning device.

[0027] The learning device 20 includes a frame acquisition unit 21, an image enlargement unit 22, a learning neural network unit 23, an encoding unit 24, a decoding unit 25, and an error extraction unit (subtraction processing unit) 26, to which an original image for learning is input and which outputs optimal parameters for the learning unit 12 of the pixel number reduction device 10. The learning device 20 can be realized by, for example, a computer including an input / output unit, a storage unit (memory), and a processing unit (CPU, etc.). Each block will be described below.

[0028] The frame acquisition unit 21 acquires the input original learning image, adjusts the timing as necessary, and outputs the original learning image to the image enlargement unit 22 and the error extraction unit 26 .

[0029] The image enlargement unit 22 uses a method such as super-resolution to improve (increase) the number of pixels and generates an enlarged image by enlarging the original learning image. Note that the image enlargement unit 22 may use a method other than super-resolution to improve the number of pixels. The image enlargement unit 22 outputs the enlarged image of the original learning image to the learning neural network unit 23, for example, as information of pixel values.

[0030] The learning neural network unit 23 inputs the enlarged image (e.g., information such as pixel values) of the learning original image from the image enlargement unit 22 to an input layer, generates a reduced image by reducing (reducing) the number of pixels by neural network processing, and outputs the reduced image to the encoding unit 24. Here, it is preferable that the reduced image has the same image size (number of pixels) as the learning original image. The learning neural network unit 23 has the same configuration as the learning unit 12 in the pixel number reduction device 10. Here, the image is treated as information of pixel values, but for example, a frequency decomposition value based on a discrete cosine transform or a discrete wavelet transform may be treated as image information and processed. In addition, the learning neural network unit 23 uses the learning original image as teacher data, and performs machine learning (parameter optimization) based on an error (difference) between the learning original image and an encoded / decoded image described later so that the encoded / decoded image and the learning original image match, and outputs the optimized parameters. Therefore, the image processing learned by the learning neural network unit 23 is image enlargement processing by the image enlargement unit 22 and image reduction processing that reflects the influence of the encoding method by the encoding unit 24.

[0031] The encoding unit 24 performs encoding on the reduced image (image data such as pixel values) from the learning neural network unit 23. The encoding method used in the encoding can be any encoding method, such as MPEG-2 (Moving Picture Experts Group -2), AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), etc. The encoded data generated by the encoding process is output to the decoding unit 25.

[0032] The decoding unit 25 decodes the coded data from the coding unit 24 based on a method corresponding to the coding, and generates a coded and decoded image. The generated coded and decoded image is output to the error extraction unit 26. The coded and decoded image has the same image size (number of pixels) as the original image for learning.

[0033] The error extraction unit 26 obtains the difference (error) between the original learning image (original image representing the correct answer) input from the frame acquisition unit 21 and the encoded / decoded image (image after encoding and decoding) from the decoding unit 25. The obtained difference (error) data is output to the training neural network unit 23.

[0034] Each block functions as described above to configure the learning device 20. Next, the flow of machine learning processing by the learning device 20 will be described with reference to the flowchart in Fig. 5. Through this machine learning, parameters of a neural network to be applied to an image processing device (pixel reduction device) are generated.

[0035] Step S11: An arbitrary one of the prepared learning original images is selected and input to the learning device 20. The learning device 20 acquires the learning original image in the frame acquisition unit 21, and sends the learning original image to the image enlargement unit 22.

[0036] Step S12: The learning device 20 enlarges the learning original image by the image enlargement unit 22. The enlargement method may be super-resolution or any other method. The image enlargement unit 22 outputs the learning original image (enlarged image) with an increased number of pixels to the learning neural network unit 23.

[0037] Step S13: The learning device 20 causes the learning section (learning neural network section 23) to perform a reduction process (reduction in the number of pixels) of the enlarged image based on the processing of the neural network that has been learned up to that point.

[0038] Step S14: The encoding unit 24 of the learning device 20 performs encoding on the reduced image reduced by the learning neural network unit 23 using a predetermined encoding method (for example, AVC, HEVC, etc.).

[0039] Step S15: Next, the decoding unit 25 of the learning device 20 decodes the coded data from the coding unit 24 to generate a coded / decoded image. This coded / decoded image is a restored image of the original learning image that has been subjected to encoding / decoding.

[0040] Step S16: The learning device 20 compares the input learning original image with the encoded / decoded image, and calculates an error (difference).

[0041] Step S17: The learning device 20 corrects the parameters of the learning neural network unit 23 based on the calculated error so as to reduce the error.

[0042] Step S18: It is determined whether or not all of the previously prepared learning data has been used for learning. If the determination is YES, the process ends, and if the determination is NO, the process proceeds to step S19.

[0043] Step S19: Determine whether learning has been performed a pre-specified number of times. If the determination is YES, the process ends, and if the determination is NO, the process returns to step S11.

[0044] Note that here, the end of the learning process is determined when all learning data has been completed or when the number of learning times reaches a predetermined value (e.g., 1 million times). However, learning may also be terminated when the error value falls below a predetermined value, or upon completion of a mini-batch (e.g., parameter correction using 100 pieces of data is repeated 10,000 times).

[0045] The above is an overview of the processing of the learning device 20. After completing the learning processing based on all data of the learning original image, the learning device 20 outputs optimal parameters (trained parameters) from the learning neural network unit 23. These parameters are trained parameters that enable optimal image reduction for the image to be encoded, and by transferring these parameters to the learning unit 12 of the pixel number reduction device 10, it is possible to realize a pixel number reduction device 10 that takes into account the effects of encoding.

[0046] The pixel count reduction process in the learning unit (learning neural network unit 23) may be performed on the entire image at once, or may be performed on a block basis, such as 8×8 pixels or 16×16 pixels. In this embodiment, the information input to the learning neural network unit 23 is pixel values, but other information may be input as additional information. For example, the coding method (AVC, HEVC, etc.) and bit rate used in the encoding / decoding unit, resolution information, frequency decomposition value, image feature amount, motion vector, frame rate, size of the reduced image, etc. may be added as input information. By inputting such information, the learning accuracy is improved.

[0047] Next, another example of machine learning by the learning unit 12 of the pixel number reduction device 10 will be described. Fig. 6 is another conceptual diagram of machine learning when reducing the number of pixels while taking into account the effects of encoding. In the conceptual diagram of Fig. 6, the image is shown to be transformed in sequence from the left to the right. The flow of image transformation and learning will be described.

[0048] First, a reduced image of a desired size is created for the original image by a process G using machine learning. This machine learning process G reduces the number of pixels while taking into account the effects of encoding, and is equivalent to the process of the learning unit 12 of the pixel number reduction device 10. The reduced image generated by machine learning is encoded and decoded to generate an encoded and decoded image. Next, the number of pixels of the encoded and decoded image is increased using a predetermined method such as super-resolution to generate an enlarged image. At this time, the enlarged image has the same number of pixels as the original image. Then, the enlarged image and the original image are compared, and the machine learning process G is optimized (adjusted parameters of the neural network, etc.) so that the two images become the same. This enables optimal image reduction while taking into account the effects of encoding.

[0049] Next, a learning device based on the conceptual diagram of Fig. 6 will be described. Fig. 7 is a diagram showing an example of a block diagram of a modified example of the learning device of the second embodiment. The learning device 30 is a modified example of a parameter learning device that generates parameters suitable for the learning unit 12 of the pixel number reduction device 10 of the first embodiment. Fig. 8 is a flowchart showing the flow of processing of the modified example of the learning device.

[0050] The learning device 30 includes a frame acquisition unit 31, a learning neural network unit 32, an encoding unit 33, a decoding unit 34, an image enlargement unit 35, and an error extraction unit (subtraction processing unit) 36, to which an original image for learning is input and which outputs optimal parameters for the learning unit 12 of the pixel number reduction device 10. Each block will be explained below, but the explanation of parts common to the learning device 20 in FIG. 4 will be simplified.

[0051] The frame acquisition unit 31 acquires the input original learning image, and outputs the original learning image to the learning neural network unit 32 and the error extraction unit .

[0052] The learning neural network unit 32 generates a reduced image by reducing (reducing) the number of pixels from the learning original image by neural network processing, and outputs the reduced image to the encoding unit 33. The learning neural network unit 32 has the same configuration as the learning unit 12 in the pixel number reduction device 10. The learning neural network unit 32 also uses the learning original image as teacher data, and performs machine learning (parameter optimization) based on the output (error between the learning original image and the enlarged image) from an error extraction unit 36 ​​described later, so that the enlarged image and the learning original image match, and outputs the optimized parameters.

[0053] The encoding unit 33 performs encoding on the reduced image from the learning neural network unit 32. The encoding method used in the encoding process may be any encoding method, such as AVC, HEVC, etc. The encoded data generated by the encoding process is output to the decoding unit 34.

[0054] The decoding unit 34 decodes the coded data from the coding unit 33 to generate a coded / decoded image, and outputs it to the image enlargement unit 35.

[0055] The image enlargement unit 35 improves (increases) the number of pixels using a technique such as super-resolution, generates an enlarged image of the encoded / decoded image, and outputs it to the error extraction unit 36. Note that it is desirable for the enlarged image to have the same image size (number of pixels) as the original image for learning.

[0056] The error extraction unit 36 ​​obtains the difference (error) between the original learning image (original image representing the correct answer) input from the frame acquisition unit 31 and the enlarged image after encoding and decoding from the image enlargement unit 35. The obtained error data is output to the training neural network unit 32.

[0057] Each block functions as described above to configure the learning device 30. Next, the flow of machine learning processing by the learning device 30 will be described with reference to the flowchart in Fig. 8. Through this machine learning, parameters of a neural network to be applied to an image processing device (pixel reduction device) are generated.

[0058] Step S21: An arbitrary one of the prepared learning original images is selected and input to the learning device 30. The learning device 30 uses the frame acquisition unit 31 to acquire the learning original image.

[0059] Step S22: The learning device 30, in the learning section (the learning neural network section 32), performs a reduction process (reduction in the number of pixels) of the acquired learning original image based on the processing of the neural network that has been learned up to that point.

[0060] Step S23: The encoding unit 33 of the learning device 30 performs encoding on the reduced image reduced by the learning neural network unit 32 using a predetermined encoding method (for example, AVC, HEVC, etc.).

[0061] Step S24: Next, the decoding unit 34 of the learning device 30 decodes the coded data from the coding unit 33 to generate a coded / decoded image.

[0062] Step S25: The learning device 30 enlarges the encoded / decoded image in the image enlargement unit 35. The enlargement method may be super-resolution or any other method.

[0063] Step S26: The learning device 30 compares the input original learning image with the image that has been subjected to image reduction, encoding / decoding, and image enlargement, and calculates an error (difference).

[0064] Step S27: The learning device 30 corrects the parameters of the learning neural network unit 32 based on the calculated error so as to reduce the error.

[0065] Step S28: It is determined whether or not all of the prepared learning data has been used for learning. If the determination is YES, the process ends, and if the determination is NO, the process proceeds to step S29.

[0066] Step S29: Determine whether learning has been performed a predetermined number of times. If the determination is YES, the process ends, and if the determination is NO, the process returns to step S21. As described above, the determination of the end of the entire learning process can be made using various indicators.

[0067] The above is an overview of the processing of the learning device 30. After completing the learning processing based on all data of the learning original image, the learning device 30 outputs optimal parameters (trained parameters) from the learning neural network unit 32. These parameters are trained parameters that enable optimal image reduction for the image to be encoded, and by transferring these parameters to the learning unit 12 of the pixel number reduction device 10, it is possible to realize a pixel number reduction device 10 that takes into account the effects of encoding.

[0068] (Pixel count expansion technology) Next, an image processing device and a learning device that increase the number of pixels by machine learning that takes into account the effects of encoding will be described.

[0069] 9 shows an example of a pixel number expansion device as an image processing device according to the third embodiment. The pixel number expansion device 40 is composed of a frame memory 41 and a learning unit 42, and inputs an image (video) and outputs an image in which the number of pixels of the input image is expanded.

[0070] The frame memory 41 holds the input original image and adjusts the timing of inputting the image to the learning unit 42. The learning unit 42 has a learning function, and converts the input original image information into an enlarged image (e.g., a pixel value string) in which the number of pixels of the input image is enlarged, and outputs the enlarged image. As will be described later, the pixel number enlargement process of the learning unit 42 is machine-learned so that an image that has further undergone image reduction processing and encoding / decoding processing matches the original image. The pixel number enlargement device 40 in FIG. 9 can be realized by, for example, a computer including an input / output unit, a storage unit (memory), and a processing unit (CPU, etc.).

[0071] The learning unit 42 can be configured by a neural network (for example, a convolutional neural network). As described in FIG. 2, the convolutional neural network can have a general configuration including one or more convolutional layers, a pooling layer, an activation layer, a softmax function, and the like. The learning unit 42 in this embodiment is a neural network that expands the number of pixels while taking into account the effects of encoding, and is desirably optimized by prior learning. Note that the video data used in the convolutional neural network may be a plurality of data converted into frequency decomposition values, image features, motion vectors, and the like in addition to pixel values.

[0072] Next, the machine learning of the learning unit 42 of the pixel number expansion device 40 will be described. Fig. 10 is a conceptual diagram of machine learning when expanding the number of pixels while taking into account the influence of encoding. In the conceptual diagram of Fig. 10, the image is shown to be transformed in sequence from the left to the right. The flow of image transformation and learning will be described.

[0073] First, the number of pixels of the original image is reduced (decreased) to generate a reduced image. Any pixel number conversion technology can be used as a method used for this image reduction process. This reduced image is encoded and decoded to generate an encoded / decoded image (an image that has undergone encoding and decoding processes). This encoded / decoded image is enlarged to a desired size by a process G using machine learning to generate an enlarged image. This enlarged image by machine learning can have the same number of pixels as the original image, for example. The machine learning process G in FIG. 10 has a learning function and enlarges the number of pixels while taking into account the effects of encoding, and is equivalent to the process of the learning unit 42 of the pixel number enlargement device 40. The enlarged image obtained by enlarging the encoded / decoded image by machine learning is compared with the original image, and the machine learning process G is optimized (adjusted parameters of the neural network, etc.) so that the two images become the same. This enables optimal image enlargement while taking into account the effects of encoding.

[0074] Next, a learning device according to a fourth embodiment will be described. The learning device according to this embodiment is a parameter learning device that generates parameters suitable for the learning unit 42 of the pixel number expansion device 40 according to the third embodiment. Fig. 11 is an example of a block diagram of the learning device, and Fig. 12 is a flowchart showing the flow of processing performed by the learning device.

[0075] The learning device 50 includes a frame acquisition unit 51, an image reduction unit 52, an encoding unit 53, a decoding unit 54, a learning neural network unit 55, and an error extraction unit (subtraction processing unit) 56, to which an original image for learning is input and which outputs optimal parameters for the learning unit 42 of the pixel number expansion device 40. The learning device 50 can be realized by, for example, a computer including an input / output unit, a storage unit (memory), and a processing unit (CPU, etc.). Each block will be described below.

[0076] The frame acquisition unit 51 acquires the input original learning image, adjusts the timing as necessary, and outputs the original learning image to the image reduction unit 52 and the error extraction unit 56 .

[0077] The image reduction unit 52 reduces (reduces) the number of pixels to generate a reduced image of the original image for learning. Any pixel number conversion method can be used for the pixel number reduction process. The image reduction unit 52 outputs the reduced image of the original image for learning to the encoding unit 53, for example, as information of pixel values.

[0078] The encoding unit 53 performs encoding on the reduced image (image data such as pixel values) from the image reducing unit 52. The encoding method used in the encoding process may be any encoding method such as AVC, HEVC, VVC, etc. The encoded data generated by the encoding process is output to the decoding unit 54.

[0079] The decoding unit 54 decodes the coded data from the coding unit 53 based on a method corresponding to the coding, and generates a coded and decoded image. The generated coded and decoded image is output to the learning neural network unit 55.

[0080] The learning neural network unit 55 inputs the encoded / decoded image (information such as pixel values) that has been encoded / decoded to an input layer, generates an enlarged image with an increased number of pixels by neural network processing, and outputs the enlarged image to the error extraction unit 56. It is preferable that the enlarged image has the same image size (number of pixels) as the original image for learning. The learning neural network unit 55 has the same configuration as the learning unit 42 in the pixel number enlargement device 40. The learning neural network unit 55 also uses the original image for learning as teacher data, and performs machine learning (parameter optimization) based on the error (difference) between the enlarged image and the original image for learning so that the enlarged image and the original image for learning match, and outputs the optimized parameters. Therefore, the image processing learned by the learning neural network unit 55 is image enlargement processing that reflects the influence of the image reduction processing of the image reduction unit 52 and the encoding method of the encoding unit 53.

[0081] The error extraction unit 56 obtains the difference (error) between the original learning image (original image representing the correct answer) input from the frame acquisition unit 51 and the enlarged image generated by the learning neural network unit 55. The obtained error data is output to the learning neural network unit 55.

[0082] Each block functions as described above to configure the learning device 50. Next, the flow of machine learning processing by the learning device 50 will be described with reference to the flowchart in Fig. 12. Through this machine learning, parameters of a neural network to be applied to an image processing device (pixel number expansion device) are generated.

[0083] Step S31: An arbitrary one of the prepared learning original images is selected and input to the learning device 50. The learning device 50 acquires the learning original image at the frame acquisition unit 51, and sends the learning original image to the image reduction unit 52.

[0084] Step S32: The learning device 50 reduces the learning original image in the image reducing unit 52. Any pixel number conversion method can be used as the reduction method. The image reducing unit 52 outputs the learning original image (reduced image) with the reduced (reduced) number of pixels to the encoding unit 53.

[0085] Step S33: The encoding unit 53 of the learning device 50 performs encoding processing on the reduced image of the learning original image using a predetermined encoding method (for example, AVC, HEVC, etc.).

[0086] Step S34: Subsequently, the decoding unit 54 of the learning device 50 decodes the coded data from the coding unit 53 to generate a coded / decoded image.

[0087] Step S35: The learning device 50, in the learning section (learning neural network section 55), performs enlargement processing (increase in the number of pixels) of the encoded / decoded image based on the processing of the neural network that has been trained up to that point, and generates an enlarged image by machine learning.

[0088] Step S36: The learning device 50 compares the input original learning image with the enlarged image generated by the learning neural network unit 55, and calculates an error (difference).

[0089] Step S37: The learning device 50 corrects the parameters of the learning neural network unit 55 based on the calculated error so as to reduce the error.

[0090] Step S38: It is determined whether or not all of the previously prepared learning data has been used for learning. If the determination is YES, the process ends, and if the determination is NO, the process proceeds to step S39.

[0091] Step S39: Determine whether or not learning has been performed a pre-specified number of times. If the determination is YES, the process ends, and if the determination is NO, the process returns to step S31.

[0092] Note that here, the end of the learning process is determined when all learning data has been completed or when the number of learning times reaches a predetermined value (e.g., 1 million times). However, learning may also be terminated when the error value falls below a predetermined value, or upon completion of a mini-batch (e.g., parameter correction using 100 pieces of data is repeated 10,000 times).

[0093] The above is an overview of the processing of the learning device 50. After completing the learning processing based on all data of the learning original image, the learning device 50 outputs optimal parameters (trained parameters) from the learning neural network unit 55. These parameters are trained parameters that enable optimal image enlargement for the image to be encoded, and by transferring these parameters to the learning unit 42 of the pixel number enlargement device 40, it is possible to realize a pixel number enlargement device 40 that takes into account the effects of encoding.

[0094] The pixel number enlargement process in the learning unit (learning neural network unit 55) may be performed for the entire image at once, or may be performed for each block of 8×8 pixels or 16×16 pixels. In this embodiment, the information input to the learning neural network unit 55 is pixel values, but other information may be input as additional information. For example, the coding method (AVC, HEVC, etc.) and bit rate used in the encoding / decoding unit, resolution information, frequency decomposition value, image feature amount, motion vector, frame rate, size of the enlarged image, etc. may be added as input information. By inputting such information, the learning accuracy is improved.

[0095] The image processing device (pixel number reduction device, pixel number expansion device) of the present invention can be used as pre-processing or post-processing for encoding / decoding and transmission.

[0096] In the above embodiment, the configuration and operation of the image processing device (pixel number reduction device 10, pixel number increase device 40) have been described, but the present invention is not limited to this, and may be configured as a method for reducing or enlarging (increasing) the number of pixels of an image. In other words, the present invention may be configured as a method including a step of holding an input original image and a step of converting the input image into an image with the number of pixels reduced or enlarged based on the information of the input image, and outputting the image.

[0097] Furthermore, in the above embodiment, the configuration and operation of the learning devices 20, 30, and 50 have been described, but the present invention is not limited to this, and may be configured as a method for generating parameters of a neural network to be applied to an image processing device. That is, the present invention may be configured as a method including each step according to the flow of the flowcharts in Figures 5, 8, 12, etc.

[0098] A computer can be suitably used to function as the image processing device or learning device described above, and such a computer can be realized by storing a program describing the processing contents for realizing each function of the image processing device or learning device in the memory unit of the computer, and reading and executing this program by the CPU of the computer. Note that this program can be recorded on a computer-readable recording medium.

[0099] Although the above-mentioned embodiment has been described as a representative example, it is clear to those skilled in the art that many modifications and substitutions can be made within the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited by the above-mentioned embodiment, and various modifications or changes can be made without departing from the scope of the claims. For example, the functions included in each block, each step, etc. described in the embodiment can be rearranged so as not to be logically inconsistent, and multiple configuration blocks, steps, etc. can be combined into one or divided. [Explanation of symbols]

[0100] 10 Image processing device (pixel reduction device) 40 Image processing device (pixel expansion device) 11,41 Frame Memory 12,42 Learning Department 20,30,50 Learning Device 21,31,51 Frame acquisition section 22,35 Image enlargement section 52 Image reduction section 23,32,55 Learning neural network section 24,33,53 Encoding section 25,34,54 Decoding section 26,36,56 Error extraction part

Claims

1. A frame acquisition unit that acquires an input original image for learning; an image enlargement unit for generating an enlarged image by enlarging the learning original image; a learning neural network unit for generating a reduced image by reducing the number of pixels of the enlarged image; an encoding unit and a decoding unit that encode and decode the reduced image to generate an encoded and decoded image; Equipped with The learning device is characterized in that the learning neural network unit performs machine learning using the original learning image as teacher data so that the encoded / decoded image matches the original learning image, and outputs optimized neural network parameters.

2. A method for generating parameters of a neural network to be applied to an image processing device, comprising the steps of: A step of generating an enlarged image by enlarging the original learning image; generating a reduced image by reducing the number of pixels of the enlarged image using a learning neural network unit; A step of encoding and decoding the reduced image to generate an encoded and decoded image; calculating an error between the encoded / decoded image and the learning original image; a step of performing machine learning by correcting parameters of the learning neural network unit based on the error so that the encoded / decoded image and the learning original image match; 23. A method comprising:

3. A program causing a computer to function as the learning device according to claim 1.

Citation Information

Patent Citations

  • Electronic apparatus, system and controlling method thereof

    US20210166343A1