Information processing apparatus, information processing method, and program

By acquiring and converting teacher data to match target data characteristics and updating neural network parameters, the method addresses the issue of reduced accuracy in low-precision quantization, maintaining image quality in high-quality image processing.

JP2025099497APending Publication Date: 2025-07-03CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023216193
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Quantizing neural network weights and feature amounts to low-precision numerical values results in decreased accuracy of data, particularly in high-image-quality image processing, leading to rough gradation and deteriorated image quality when bit depth is reduced.

Method used

Acquire teacher data that has undergone depth conversion processing to match the characteristics of the target data, and use this data to learn and update neural network parameters to minimize the error between output and teacher data, followed by quantizing the weights and feature amounts to a lower bit depth.

Benefits of technology

This approach suppresses the deterioration of final data quality even when quantizing with a bit depth smaller than the original data, maintaining image gradation and overall quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099497000001_ABST
    Figure 2025099497000001_ABST
Patent Text Reader

Abstract

To prevent the ultimate deterioration of data even if weight or the like is quantified with a bit depth smaller than a bit depth of data to be processed, in a NN intended to perform data processing.SOLUTION: An information processing apparatus processes target data with a neural network, and the information processing apparatus comprises: input data acquisition means that acquires target data; teacher data acquisition means that acquires teacher data; and learning means that performs learning so as to reduce the difference between the teacher data and output data obtained by inputting the target data to the neural network and performing processing, and updates a parameter of the neural network. When the bit depth of the teacher data is a second bit depth smaller than a first bit depth of the target data, the teacher data acquisition means acquires the teacher data on which depth conversion processing is performed, which is to convert the value of the teacher data with a resolution according to the characteristics of the target data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] In recent years, in data processing such as image processing technology for improving the image quality, techniques using a neural network (Neural Network, hereinafter NN) have been actively developed. For example, there are techniques for realizing high-image-quality image processing such as noise removal, blur removal, and super-resolution using an NN (Non-Patent Document 1).

[0003] In recent NNs, since the number of layers is large and the amount of calculation is large, a high-speed computer is used during learning. However, in data processing during inference after learning, there are many cases where computing resources are limited, and a more efficient operation method is required.

[0004] As an efficient operation method during inference, a method of quantizing the weights and feature amounts of an NN into low-precision numerical values and performing operations is known (Non-Patent Document 2). By quantization, it is possible to operate even on a device with scarce computing resources such as an embedded device. Also, even in a general-purpose computer, by quantizing the weights of an NN, etc., there are cases where high-throughput operation instructions such as SIMD (Single Instruction Multiple Data) instructions can be used, and speedup can be expected.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, by quantizing the weights and feature amounts of the NN to low-precision numerical values, the accuracy of data such as the resolution of the image output by the NN generally decreases. In particular, when quantizing the weights and the like with a bit depth smaller than the bit depth of the original data, the accuracy of data such as the gradation of the output image becomes rough, and the deterioration becomes significantly apparent. For example, in the case of an NN for high-image-quality image processing, if the input image to the target NN is a 12-bit to 14-bit RAW image, when the weights and feature amounts of the NN are quantized to 8 bits, the high-image-quality image output by the NN is also output in 8 bits, so the gradation of the original image cannot be expressed. Although the RAW image is finally converted to an 8-bit JPEG image or the like by development processing, since the gradation of the RAW image, which is the original image, has become rough, the resulting 8-bit JPEG image also has a rough gradation and outputs an image with deteriorated image quality.

[0007] An object of the present invention is to suppress deterioration of final data even when quantizing with a bit depth smaller than the bit depth of data to be processed in an NN for data processing such as high-image-quality image processing.

Means for Solving the Problem

[0008] To solve this problem, for example, the information processing apparatus of the present invention has the following configuration. That is, An information processing apparatus that processes target data by a neural network, Input data acquisition means for acquiring the target data, Teacher data acquisition means for acquiring teacher data, Learning means for learning so that the error between the output data obtained by inputting and processing the target data to the neural network and the teacher data becomes small, and updating the parameters of the neural network, and comprising When the bit depth of the teacher data is a second bit depth smaller than the first bit depth of the target data, the teacher data acquisition means acquires the teacher data that has undergone depth conversion processing for converting the value of the teacher data with a resolution according to the characteristics of the target data. It is characterized by this.

Effect of the Invention

[0009] According to the present invention, in an NN for the purpose of data processing such as high-quality image processing, even if quantization is performed with a bit depth smaller than the bit depth of the data to be processed, deterioration of the final data can be suppressed.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are given the same reference numerals, and redundant explanations are omitted.

[0012] <First Embodiment> Hereinafter, the present invention will be described in detail based on its preferred embodiments with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.

[0013] FIG. 1 is a block diagram showing an example of the hardware configuration in the present embodiment. The information processing apparatus 1 may be, for example, a computer. As shown in the figure, the information processing apparatus 1 includes a CPU 11, a ROM 12, a RAM 13, a secondary storage device 14, an input device 15, a display device 16, and a connection bus 17.

[0014] The CPU 11 is an abbreviation for Central Processing Unit, and controls the information processing apparatus 1 by reading out the control program stored in the ROM 12, expanding it in the RAM 13, and executing it. Further, the CPU 11 includes SIMD instructions for collectively performing arithmetic operations on 8-bit integer types, and is used in the inference processing described later. The information processing apparatus 1 may have other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), and a QPU (Quantum Processing Unit) instead of or in addition to the CPU 11.

[0015] The ROM 12 is an abbreviation for Read Only Memory and is a non-volatile memory. The ROM 12 stores a control program and various parameter data necessary for program execution. The control program is executed by the CPU 11 to realize each process described later.

[0016] RAM13 is an abbreviation for Random Access Memory and is a volatile memory. RAM13 temporarily stores images, control programs, and their execution results.

[0017] The secondary storage device 14 stores data such as various programs and image data used in the processing of this embodiment in a rewritable manner. The secondary storage device 14 is a non-volatile storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), and a flash memory, for example. The secondary storage device 14 stores, for example, images used for the calculation of the NN, control programs such as the model of the NN, and the processing results of the control programs. The information stored in the secondary storage device 14 is output to the RAM13 in response to a request from the CPU11 or the like, and is used by the CPU11 for program execution.

[0018] The input device 15 serves as an interface with the outside such as a user. The input device 15 may be a mouse, a keyboard, or the like that acquires input from the user.

[0019] The display device 16 is a monitor such as a liquid crystal display and an organic EL (Electro Luminescence) display, for example. The display device 16 displays the processing results of the program and images and the like.

[0020] The connection bus 17 connects the CPU11, the ROM12, the RAM13, the secondary storage device 14, the input device 15, and the display device 16 that constitute the information processing apparatus 1 and performs data communication with each other.

[0021] In this embodiment, the CPU 11 executes software or a program to implement the processes and functions described below, but part or all of the processes and functions may be implemented in hardware. Examples of the hardware include dedicated circuits (ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array)), and processors (reconfigurable processors, DSP (Digital Signal Processor)).

[0022] Further, the information processing apparatus 1 may acquire software or a program that describes the functions and processes described below via a network or various storage media, and execute the software or program using a processing apparatus (processors such as a CPU and a GPU) of a personal computer or the like.

[0023] FIG. 2 is a functional block diagram showing an example of the functional configuration of the information processing apparatus 2 related to learning and the information processing apparatus 3 related to inference according to this embodiment. The hardware configurations of the information processing apparatus 2 and the information processing apparatus 3 are the same as that of the information processing apparatus 1. The information processing apparatus 2 and the information processing apparatus 3 may be implemented by one information processing apparatus.

[0024] The information processing apparatus 2 learns an NN model based on image data, updates parameters such as the weights of the NN model, and quantizes the model. The image data is an example of the target data. Note that the term "image" may be used as a term including videos, still images, images, and their data. The information processing apparatus 2 has functions of an input data acquisition unit 201, a model acquisition unit 202, a learning unit 203, a teacher data acquisition unit 204, and a quantization unit 205.

[0025] The input data acquisition unit 201 acquires input data to be input to the NN model acquired by the model acquisition unit 202. For example, the input data acquisition unit 201 acquires an image as the input data. For example, the input data acquisition unit 201 acquires, as the input data, an image obtained by converting a 14-bit RAW image into a 16-bit unsigned integer type image.

[0026] The model acquisition unit 202 acquires an NN model. FIG. 4 shows an example of an NN model. The model acquired by the model acquisition unit 202 has, for example, one or more layers 401, 402, 403 in which a CNN layer and a ReLU layer are combined. Here, CNN is an abbreviation for Convolutional Neural Network (convolutional NN) and is a type of NN. Also, ReLU is an abbreviation for Rectified Linear Unit and is a type of activation function. Here, the unit of a layer is defined as a combination from an NN (here, a CNN layer) to an activation function (here, a ReLU layer). For example, layer 401 includes a CNN layer 404 and a ReLU layer 405, and this combination of layers is defined as one unit of a layer. Layers 402 and 403 are the same as layer 401, and two layers of a CNN layer and a ReLU layer are regarded as one unit of a layer. The output of a layer to be handled below refers to the output of this one unit of layers 401, 402, 403. Also, when referring to layer i, i refers to the index of one unit of a layer. In the example of FIG. 4, layer 1 corresponds to layer 401, layer 2 corresponds to layer 402, and layer 3 corresponds to layer 403. Note that layer 401 is an input layer and performs a convolution operation on the input image. Layer 403 is an output layer and outputs an image with noise removed from the input image. Note that the configuration of the NN is not limited to the number and types of layers in FIG. 4. The number of layers may be, for example, four or more. For the types of layers, a pooling layer, a skip connection, etc. may be applied. Also, the NN model does not necessarily have to be a learned model. If it is not a learned model, the NN model may be initialized by a known NN weight initialization method such as the initialization method in Non-Patent Document 3.

[0027] The learning unit 203 has a function for learning the NN model acquired by the model acquisition unit 202. The learning unit 203 has a quantization parameter acquisition unit 206 and a weight determination unit 207.

[0028] The quantization parameter acquisition unit 206 acquires the quantization parameters q used for quantization of each layer of the NN model acquired by the model acquisition unit 202. i is obtained. When it is not necessary to distinguish which layer's quantization parameter is, it is written as quantization parameter q. In the example of Figure 4, each layer is output through a ReLU layer, so the output value is 0 or more. In this case, for example, if q = 4, the upper limit of the output range of each layer is 4, and the output range of each layer of a single-precision 32-bit NN is [0, 4]. If the precision is set to 8 bits, the output value of the layer is quantized at (bin width) = 4 / 256 = 0.015625. For example, if the output value of a layer of a single-precision 32-bit NN is 3.1, the output value when quantized to 8 bits is 3.09375. When this output value is converted to an 8-bit integer in the range of [0, 255], it becomes 3.09375 × 256 / 4 = 198. The quantization parameter q is determined from the statistics of the output value of each layer, after the NN model is inferred in advance using the input data. For example, the quantization parameter q may be the maximum value of the output value of the layer, or may be a value corresponding to 99.9% of the output values ​​of the layer arranged in ascending order. In this embodiment, each layer of the NN model is quantized to 8 bits.

[0029] The weight determination unit 207 updates and determines the weights of the NN model acquired by the model acquisition unit 202 through learning. For example, the weight determination unit 207 performs inference processing on the image acquired by the input data acquisition unit 201 for the NN model to output an output image (output data), and calculates the loss (objective function) between the teacher data acquired by the teacher data acquisition unit 204 and the output image. In this embodiment, the objective function is the mean squared error, which is an example of an error. Then, the weight determination unit 207 calculates the gradient using the error backpropagation method with the calculated loss, and learns to minimize the loss to calculate the update amount of the model weights. The weight determination unit 207 updates the weights of the NN model according to the update amount. As the learning method for weight update, a known NN learning method can be applied, and detailed description is omitted.

[0030] The teacher data acquisition unit 204 acquires and converts teacher data. The teacher data acquisition unit 204 includes a 16-bit teacher data acquisition unit 209 and a teacher data bit depth conversion unit 210.

[0031] The 16-bit teacher data acquisition unit 209 acquires the data to be input to the teacher data bit depth conversion unit 210. The 16-bit teacher data acquisition unit 209 acquires, for example, an image obtained by converting a 14-bit RAW image into an unsigned integer type 16-bit image. The image is a high-quality image with noise removed from the image acquired by the input data acquisition unit 201.

[0032] The teacher data bit depth conversion unit 210 executes depth conversion processing to convert the unsigned integer type 16-bit image acquired by the 16-bit teacher data acquisition unit 209 into an unsigned integer type 8-bit image. 16 bits is an example of the first bit depth, and 8 bits is an example of the second bit depth. The teacher data bit depth conversion unit 210 performs conversion from unsigned integer type 16-bit to unsigned integer type 8-bit using, for example, a preset look-up table.

[0033] FIG. 3(a) is a look-up table used for bit depth conversion in this embodiment. The values in FIG. 3(a) indicate the correspondence of pixel values (luminance) at each bit depth. FIG. 3(b) is a tone curve applied to tone adjustment in the tone adjustment unit 214 described later. The horizontal axis of FIG. 3(b) is the luminance of the input image (image before conversion), and the vertical axis is the luminance of the output image (image after conversion). The look-up table is created by converting and reflecting a tone curve of 14-bit gradation into 8-bit gradation. For example, the left 16-bit column of the look-up table in FIG. 3(a) is created based on the values in FIG. 3(b). The right 8-bit column of the look-up table in FIG. 3(a) is created based on the value obtained by converting the maximum value on the vertical axis of FIG. 3(b) to 255, which is the maximum value of 8 bits. As shown in the tone curve of FIG. 3(b) and the look-up table of FIG. 3(a), in the relatively low-luminance region from pixel value 0 to about pixel value 4000, the gradation is finely converted, and in the region after the low-luminance region, the gradation is coarsely converted up to 16383, which is the upper limit value of 14 bits. FIGS. 3(a) and 3(b) show that the resolution changes with respect to the pixel value. Therefore, by using the look-up table in FIG. 3(a), the teacher data bit depth conversion unit 210 converts unsigned integer type 16 bits to unsigned integer type 8 bits according to the resolution for each pixel value. The look-up table in FIG. 3(a) assigns the numbers of 8 bits (0 to 256) more finely to the relatively low-luminance region from pixel value 0 to about pixel value 4000 than to the other regions. That is, instead of uniformly quantizing into 256 levels, the resolution is increased for the relatively low-luminance region according to the characteristics of the data. Also, the tone adjustment unit 214 can convert the image with the same gradation resolution as the tone curve. Therefore, the teacher data bit depth conversion unit 210 converts the bit depth of the teacher data based on the resolution corresponding to the resolution used by the tone adjustment unit 214 for tone adjustment, here the same resolution. Here, although the image data is 16 bits, since it was originally a 14-bit RAW image, the pixel values only exist up to a maximum of 16383. Therefore, there is no table of values of 16383 or more in the look-up table.The teacher data acquisition unit 204 finally outputs, using this lookup table, an image obtained by converting a 16-bit unsigned integer type image into an 8-bit unsigned integer type image as teacher data to the weight determination unit 207.

[0034] The quantization unit 205 quantizes the output values such as the weights and feature amounts of each layer of the NN learned by the learning unit 203 using quantization parameters. For example, the quantization unit 205 quantizes the weights and output values to the same bit depth as the bit depth of the image data output by the teacher data bit depth conversion unit 210. Therefore, in the present embodiment, the quantization unit 205 quantizes the weights and output values to 8 bits. Details may be obtained by applying a known NN quantization method, and the description thereof is omitted.

[0035] The information processing device 3 uses the model of the NN learned by the information processing device 2 to execute an inference process on the RAW image, and then converts the image of the inference result with the changed bit depth into an image in a format such as a JPEG image, that is, develops it. The information processing device 3 includes an inference data acquisition unit 215, a quantization model acquisition unit 211, an inference data bit depth conversion unit 212, and a development unit 213. The inference data bit depth conversion unit 212 is an example of a depth conversion means.

[0036] The inference data acquisition unit 215 acquires the image data used for inference. For example, the inference data acquisition unit 215 acquires an 8-bit or 16-bit RAW image and delivers it to the quantization model acquisition unit 211.

[0037] The quantization model acquisition unit 211 acquires the model of the NN quantized by the quantization unit 205. For example, the quantization model acquisition unit 211 acquires the model of the NN including 8-bit weights quantized by the quantization unit 205. The quantization model acquisition unit 211 executes an inference process on the RAW image acquired by the inference data acquisition unit 215 using the acquired NN model, and outputs an 8-bit image as the inference result.

[0038] The inference data bit-depth conversion unit 212 performs a depth conversion process of converting the output value (e.g., pixel value of an image) of the quantized NN model acquired by the quantization model acquisition unit 211 from unsigned integer type 8-bit to unsigned integer type 16-bit. The inference data bit-depth conversion unit 212 performs the bit-depth conversion from unsigned integer type 8-bit to unsigned integer type 16-bit using the same look-up table as that used by the teacher data bit-depth conversion unit 210. Note that the look-up tables used for the conversions of both bit-depths may be substantially the same. Here, the teacher data bit-depth conversion unit 210 performs the conversion from the "unsigned integer type 16-bit" column to the "unsigned integer type 8-bit" column in FIG. 3(a), while the inference data bit-depth conversion unit 212 performs the conversion from the "unsigned integer type 8-bit" column to the "unsigned integer type 16-bit" column. In other words, the inference data bit-depth conversion unit 212 executes the depth conversion process in the reverse procedure to that of the teacher data bit-depth conversion unit 210.

[0039] The developing unit 213 performs image processing such as tone adjustment on the RAW image and finally converts it into an image in any format such as a JPEG image and a PNG image. For example, the developing unit 213 converts the RAW image of unsigned integer type 16-bit whose bit-depth has been converted by the inference data bit-depth conversion unit 212 and outputs a 16-bit image. The developing unit 213 includes, for example, a tone adjustment unit 214. The tone adjustment unit 214 is an example of data changing means.

[0040] The tone adjustment unit 214 performs tone adjustment processing within the developing unit 213. The tone adjustment unit 214 performs tone adjustment processing on the 16-bit image using the tone curve shown in FIG. 3(b).

[0041] The flowchart of FIG. 5 shows an example of the procedure of the learning process until the information processing apparatus 2 in the present embodiment determines the weights of the NN during one learning. Hereinafter, with reference to the same figure, the processing contents of the weight determination unit 207 of the learning unit 203 will be described.

[0042] In S501, the model acquisition unit 202 acquires the model of the NN and outputs it to the weight determination unit 207.

[0043] In S502, the quantization parameter acquisition unit 206 acquires the quantization parameter q from the model acquisition unit 202 and outputs it to the weight determination unit 207.

[0044] In S503, the input data acquisition unit 201 acquires the mini-batch image of the learning data set as a 16-bit unsigned integer type input image and outputs it to the weight determination unit 207.

[0045] In S504, the 16-bit teacher data acquisition unit 209 acquires the 16-bit unsigned integer type image data corresponding to the input image acquired in S503 as teacher data and outputs it to the teacher data bit depth conversion unit 210.

[0046] In S505, the teacher data bit depth conversion unit 210 executes a depth conversion process for converting the bit depth of the teacher data. For example, the teacher data bit depth conversion unit 210 uses the look-up table in Fig. 3(a) to convert the 16-bit unsigned integer type teacher data acquired from the 16-bit teacher data acquisition unit 209 into 8-bit unsigned integer type. The teacher data bit depth conversion unit 210 outputs the 8-bit unsigned integer type teacher data to the weight determination unit 207.

[0047] In S506, the weight determination unit 207 performs an inference process on the model of the NN acquired from the model acquisition unit 202 using the mini-batch image acquired from the input data acquisition unit 201. The weight determination unit 207 calculates the loss (objective function) with the 8-bit unsigned integer type teacher data acquired from the teacher data bit depth conversion unit 210 using the inference result.

[0048] In S507, the weight determination unit 207 calculates the gradient using the calculated loss and the error backpropagation method, and calculates the update amount of the weight of the model of the NN.

[0049] In S508, the weight determination unit 207 updates the weights of the NN based on the calculated weight update amount.

[0050] In S509, the weight determination unit 207 generates and outputs a model of the NN with updated weights.

[0051] As described above, the information processing apparatus 2 repeats the processing from S501 to S509 until the learning loss converges, and determines the weights of the model of the NN. When the learning loss converges and the learning is completed, the weight determination unit 207 outputs the model of the NN to the quantization unit 205. The quantization unit 205 quantizes the model of the NN obtained from the weight determination unit 207 using quantization parameters. For example, the quantization unit 205 quantizes the weights and feature amounts of the model of the NN obtained from the weight determination unit 207 to 8 bits using quantization parameters.

[0052] The flowchart of FIG. 6 shows an example of the procedure from the inference by the quantized NN model executed by the information processing apparatus 3 in the present embodiment to the development process.

[0053] In S601, the quantization model acquisition unit 211 acquires a model of the NN quantized to, for example, 8 bits by the quantization unit 205.

[0054] In S602, the inference data acquisition unit 215 acquires, as an input image, a RAW image of, for example, unsigned integer type 16 bits for inference and outputs it to the quantization model acquisition unit 211. As a result, the model of the NN acquired by the quantization model acquisition unit 211 executes an inference process on the inference image acquired by the inference data acquisition unit 215. The model of the NN outputs an image of unsigned integer type 8 bits as an inference result to the inference data bit depth conversion unit 212.

[0055] In S603, the inference data bit-depth conversion unit 212 performs a depth conversion process that converts the bit-depth of the inference result of the quantized NN model. For example, the inference data bit-depth conversion unit 212 converts an unsigned integer type 8-bit image output from the NN model into an unsigned integer type 16-bit image by means of a depth conversion process using the look-up table in Fig. 3(a).

[0056] In S604, the developing unit 213 performs a developing process on the unsigned integer type 16-bit image and converts it into an image in any image format such as a JPEG image and a PNG image. In this conversion process, the tone adjustment unit 214 performs tone adjustment using the tone curve in Fig. 3(b).

[0057] As described above, the information processing apparatus 3 performs the processes from inference to development using the learned quantized NN model in the procedures of S601 to S604.

[0058] [Effects of this Embodiment] The unsigned integer type 16-bit image output from the inference data bit-depth conversion unit 212 finally undergoes tone processing by the tone curve in Fig. 3(b) in the developing process of the developing unit 213 and is converted into an image in a format such as a JPEG image and a PNG image. Since the unsigned integer type 16-bit image output from the inference data bit-depth conversion unit 212 is generated by converting an unsigned integer type 8-bit image using a look-up table, there are actually only 256 tones. Since the tones of a 14-bit RAW image are reduced to 256 tones, image quality degradation would originally occur. However, in this embodiment, since the look-up table corresponds to the tone curve in Fig. 3(b), degradation is reduced in the developed images such as JPEG images and PNG images. The reason for this is described below.

[0059] In the process of converting 16-bit unsigned integer data to 8-bit unsigned integer data, the teacher data bit-depth conversion unit 210 converts the pixel values according to the resolution for the pixel values. Specifically, in the bit-depth conversion process, the teacher data bit-depth conversion unit 210 performs conversion with fine gradation in the low-luminance region where fine gradation is required during development, and performs conversion with coarse gradation after the low-luminance region to generate teacher data. Then, the NN model is trained to output an 8-bit unsigned integer image with the same gradation using the 8-bit teacher data whose bit-depth has been converted. When the inference data bit-depth conversion unit 212 converts the 8-bit unsigned integer image output from the NN model to 16-bit unsigned integer in the inference data bit-depth conversion unit 212, it performs the conversion to 16-bit unsigned integer in the reverse procedure using a similar look-up table. Therefore, the gradation in the low-luminance region of the 16-bit unsigned integer data converted by the inference data bit-depth conversion unit 212 becomes finer and corresponds to the gradation of the tone curve in the development process after the inference process. As a result, the fine gradation change part in the tone curve gradation process maintains the fine gradation, and the final image quality degradation can be suppressed.

[0060] Note that in this embodiment, in order to explain a general development processing procedure in which the developing unit 213 performs development processing on a 16-bit image, an example is given in which the inference data bit-depth conversion unit 212 converts 8-bit unsigned integer data to 16-bit unsigned integer data and then the developing unit 213 performs development processing. However, if the developing unit 213 is designed to process 8-bit unsigned integer data, the effects of the present proposal can be obtained even without the inference data bit-depth conversion unit 212 and the gradation change unit 214.

[0061] In addition, in this embodiment, after the model of the NN with quantization parameters was learned, the quantization unit 205 performed 8-bit quantization on the model of the NN. However, even if the quantization unit 205 performs 8-bit quantization on the model of the ordinary NN model without quantization parameters after learning, it is possible to obtain the effect of this proposal. Also, even if the model of the NN that has already been quantized to 8 bits is learned while remaining 8 bits, it is possible to obtain the effect of this proposal.

[0062] <Second Embodiment> In the first embodiment described above, a look-up table was used for the bit-depth conversion process of the teacher data bit-depth conversion unit 210, but a mathematical formula may also be used.

[0063] FIG. 7(b) is a tone curve used for the tone processing of the tone change unit 214 in this embodiment, and is a graph of Equation (1).

[0064]

Equation

[0065] Since the target of the tone processing is an unsigned integer type 16-bit image, the value calculated by Equation (1) is converted into an integer value.

[0066] FIG. 7(a) is a graph of Equation (2) when the teacher data bit-depth conversion unit 210 in this embodiment converts an unsigned integer type 16-bit image into an unsigned integer type 8-bit image.

[0067]

Equation

[0068] Equation (2) is an equation for converting pixel values according to the resolution with respect to pixel values (luminance) when converting a 16-bit image into an 8-bit image by bit-depth conversion processing. Since Equation (2) is an equation for converting an unsigned integer type 16-bit integer value into an unsigned integer type 8-bit integer value, an unsigned integer type 16-bit image is converted into an unsigned integer type 8-bit image by Equation (2). Since the unsigned integer type 16-bit image is originally an image obtained by type-converting a 14-bit RAW image, the maximum value of x in Equation (2) is 16383. As can be seen from comparing Equation (1) and Equation (2), by using Equation (2), the teacher data bit-depth conversion unit 210 can convert an unsigned integer type 16-bit image into an unsigned integer type 8-bit image, while the tone adjustment unit 214 can convert the image with the same tone resolution as the tone curve of Equation (1) in the present embodiment.

[0069] FIG. 7(c) is a graph of Equation (3) when the inference data bit-depth conversion unit 212 in the present embodiment converts an unsigned integer type 8-bit image into an unsigned integer type 16-bit image.

[0070]

Number

[0071] Equation (3) is an equation for converting pixel values according to the resolution with respect to pixel values (luminance) when converting an 8-bit image into a 16-bit image by bit-depth conversion processing. Since Equation (3) is an equation for converting an unsigned integer type 8-bit integer value into an unsigned integer type 16-bit integer value, the integer value of an unsigned integer type 8-bit image is converted into an unsigned integer type 16-bit integer value. Since the unsigned integer type 16-bit image is assumed to be an image obtained by type-converting a 14-bit RAW image, the calculated value of Equation (3) takes only a value of 16383 at most. Since Equation (3) is the inverse function of Equation (1), when using Equation (3), the inference data bit-depth conversion unit 212 can convert an unsigned integer type 8-bit image into an unsigned integer type 16-bit image, while the tone adjustment unit 214 can convert the image with the same tone resolution as the tone curve of Equation (1) in this embodiment.

[0072] As described above, according to this embodiment, the effect of the present proposal can also be obtained by implementing the bit-depth conversion of the teacher data bit-depth conversion unit 210 and the inference data bit-depth conversion unit 212 by mathematical formulas.

[0073] Note that the mathematical formulas and look-up tables of the teacher data bit-depth conversion unit 210 and the inference data bit-depth conversion unit 212 may be acquired by learning.

[0074] Also, in the embodiment of the present proposal, although the luminance of the image has been described as the processing target, it is also applicable to frequencies and can also be applied to processing such as audio and regression tasks other than images.

[0075] Also, the above-described embodiments may be combined. For example, the user may be able to select the look-up table of the first embodiment and the mathematical formula of the second embodiment.

[0076] In the above-described embodiment, 8-bit and 16-bit images have been described, but the number of bits of the image may be appropriately changed.

[0077] (Other Embodiments) The present invention can also be implemented by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be implemented by a circuit (for example, ASIC) that realizes one or more functions.

[0078] The disclosure of this specification includes the following information processing apparatus, information processing method, and program. (Item 1) An information processing apparatus that processes target data by a neural network, input data acquisition means for acquiring the target data, teacher data acquisition means for acquiring teacher data, learning means for learning so that the error between the output data obtained by inputting and processing the target data to the neural network and the teacher data is reduced, and updating the parameters of the neural network, comprising: when the bit depth of the teacher data is a second bit depth smaller than the first bit depth of the target data, the teacher data acquisition means acquires the teacher data that has been subjected to depth conversion processing for converting the value of the teacher data with a resolution according to the characteristics of the target data An information processing apparatus characterized by the above. (Item 2) The information processing apparatus according to Item 1, further comprising quantization means for quantizing the neural network whose parameters have been updated by the learning means. (Item 3) comprising quantization parameter acquisition means for acquiring quantization parameters of the neural network, the quantization means quantizes the neural network based on the quantization parameters An information processing apparatus according to Item 2, characterized by the above. (Item 4) the learning means updates the parameters of the quantized neural network The information processing apparatus according to item 2, characterized in that... (Item 5) When converting the bit depth of the target data inferred by the neural network from the second bit depth to the first bit depth, depth conversion means for converting the value of the target data in a procedure reverse to the depth conversion process; Data change means for changing the value of the target data of the first bit depth according to the resolution; The information processing apparatus according to any one of items 1 to 4, characterized by comprising the above. (Item 6) In the depth conversion process, the value of the target data is changed based on the resolution corresponding to the resolution used by the data change means. The information processing apparatus according to item 5, characterized in that... (Item 7) The teacher data acquisition means changes the value of the target data by a look-up table reflecting the resolution adapted to the characteristics of the value of the target data. The information processing apparatus according to any one of items 1 to 6, characterized by comprising the above. (Item 8) The teacher data acquisition means changes the value of the teacher data by a mathematical formula according to the resolution of the value of the target data. The information processing apparatus according to any one of items 1 to 7, characterized by comprising the above. (Item 9) The quantization means quantizes the neural network to the second bit depth. The information processing apparatus according to any one of items 2 to 8, characterized by comprising the above. (Item 10) The target data and the teacher data are image data. The information processing apparatus according to any one of items 1 to 9, characterized by comprising the above. (Item 11) The teacher data acquisition means acquires the conversion of the value of the teacher data according to the resolution by learning. The information processing apparatus according to any one of items 1 to 10, characterized by comprising the above. (Item 12) An information processing method for processing target data by a neural network, an input data acquisition step of acquiring the target data, a teacher data acquisition step of acquiring teacher data, a learning step of learning so that the error between the output data obtained by inputting and processing the target data into the neural network and the teacher data is reduced, and updating the parameters of the neural network, comprising: in the teacher data acquisition step, when the bit depth of the teacher data is a second bit depth smaller than the first bit depth of the target data, the teacher data obtained by performing depth conversion processing for converting the value of the teacher data with a resolution according to the characteristics of the target data is acquired An information processing method characterized by the above. (Item 13) A program for causing a computer to function as each means of the information processing apparatus according to any one of Items 1 to 11.

[0079] The invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Signs

[0080] 201 ··· Input data acquisition unit, 202 ··· Model acquisition unit, 203 ··· Learning unit, 204 ··· Teacher data acquisition unit, 205 ··· Quantization unit, 206 ··· Quantization parameter acquisition unit, 210 ··· Teacher data bit depth conversion unit, 214 ··· Tone change unit, 212 ··· Inference data bit depth conversion unit.

Claims

1. An information processing apparatus that processes target data by a neural network, comprising: input data acquisition means for acquiring the target data; teacher data acquisition means for acquiring teacher data; learning means for learning so that the error between the output data obtained by inputting and processing the target data to the neural network and the teacher data is reduced, and updating the parameters of the neural network; wherein when the bit depth of the teacher data is a second bit depth smaller than the first bit depth of the target data, the teacher data acquisition means acquires the teacher data that has been subjected to depth conversion processing for converting the value of the teacher data with a resolution according to the characteristics of the target data An information processing apparatus characterized by the above.

2. The information processing apparatus according to claim 1, further comprising quantization means for quantizing the neural network whose parameters have been updated by the learning means.

3. The information processing apparatus according to claim 2, further comprising quantization parameter acquisition means for acquiring quantization parameters of the neural network, wherein the quantization means quantizes the neural network based on the quantization parameters. An information processing apparatus characterized by the above.

4. The information processing apparatus according to claim 2, wherein the learning means updates the parameters of the quantized neural network. An information processing apparatus characterized by the above.

5. When converting the bit depth of the target data inferred by the neural network from the second bit depth to the first bit depth, depth conversion means for converting the value of the target data in a procedure reverse to the depth conversion processing, and data change means for changing the value of the target data with the first bit depth according to the resolution. An information processing apparatus according to claim 1, characterized by comprising the above.

6. In the depth conversion processing, the value of the target data is changed based on a resolution corresponding to the resolution used by the data change means. An information processing apparatus according to claim 5, characterized by the above.

7. The teacher data acquisition means changes the value of the target data by a look-up table reflecting the resolution according to the characteristics of the value of the target data. An information processing apparatus according to claim 1, characterized by the above.

8. The teacher data acquisition means changes the value of the teacher data by a mathematical formula according to the resolution of the value of the target data. An information processing apparatus according to claim 1, characterized by the above.

9. The quantization means quantizes the neural network to the second bit depth. The information processing apparatus according to claim 2, characterized in that.

10. The target data and the teacher data are image data. The information processing apparatus according to claim 1, characterized in that.

11. The teacher data acquisition means acquires, by learning, a conversion of the value of the teacher data according to the resolution. The information processing apparatus according to claim 1, characterized in that.

12. An information processing method for processing target data by a neural network, comprising: an input data acquisition step of acquiring the target data; a teacher data acquisition step of acquiring teacher data; a learning step of learning so that an error between output data obtained by inputting and processing the target data to the neural network and the teacher data is reduced, and updating parameters of the neural network; comprising: In the teacher data acquisition step, when the bit depth of the teacher data is a second bit depth smaller than the first bit depth of the target data, the teacher data obtained by performing depth conversion processing for converting the value of the teacher data with a resolution according to the characteristics of the target data is acquired. An information processing method characterized by that.

13. A program for causing a computer to function as each means of the information processing apparatus according to any one of claims 1 to 11.