Image processing apparatus, image processing method
By generating artificial teacher images with corrected hue distributions and combining with natural images, the method addresses biased teacher groups in deep learning demosaicing, reducing artifacts and improving inference accuracy.
Patent Information
- Application Number
- JP2021032033
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-01
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2041-03-01
AI Technical Summary
Conventional deep learning-based demosaicing methods using biased teacher image groups result in artifacts, particularly for hues that are less frequent, leading to inaccurate interpolation and false color patterns.
Generate artificial teacher images with a corrected probability distribution of hues to reduce bias, using a CNN for demosaicing inference, and combine with natural teacher images for robustness.
Suppresses artifacts in demosaicing inference results, especially for difficult hues, while maintaining robustness against natural images.
Smart Images

Figure 0007709289000011 
Figure 0007709289000012 
Figure 0007709289000013
Abstract
Description
Technical Field
[0001] The present invention relates to learning technology.
Background Art
[0002] An image sensor used in a digital imaging device such as a digital camera is equipped with a color filter composed of, for example, an RGB array, and is configured to incident light of a specific wavelength on each pixel. Specifically, for example, a color filter having a Bayer array is widely used. The captured image of the Bayer array is a so-called mosaic image in which only pixel values corresponding to any one of the RGB colors are set for each pixel. The development processing unit of the digital imaging device performs various signal processes such as demosaicing processing for interpolating the pixel values of the remaining two colors on this mosaic image, and generates and outputs a color image. As a conventional method of demosaicing processing, there is a method in which a linear filter is applied to the sparse pixel values of each RGB color, and linear interpolation of the pixel values of the same color in the surroundings is executed to calculate and set each RGB color corresponding to each pixel. Since this method has low interpolation accuracy, many non-linear interpolation methods have been proposed so far. However, in any method, there has been a problem that false colors and artifacts occur in the image area that each method is not good at.
[0003] Therefore, in recent years, a data-driven interpolation method applying deep learning technology has been proposed. Non-Patent Document 1 discloses a method of training a CNN (Convolutional Neural Network)-based demosaic network. In this method, learning is performed using a group of teacher images collected from nature. Then, using the learning result, an inference (a regression task for input data) is performed in which a mosaic image is input to the CNN and converted into an RGB image.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
[0005] However, in the conventional technology, although a sufficient amount of data can be secured, there is a problem that the distribution of hues in the teacher image group is biased. When deep learning is performed using such a teacher image group, it may not be possible to generate a learning model with high robustness. For example, when learning the CNN-based demosaic network described in Non-Patent Document 1, if there is a bias in the hue distribution of the teacher image group used for learning, artifacts such as false patterns that do not actually exist will occur when the mosaic image is demosaic using the learned model. This phenomenon is noticeable in hues that occur less frequently in the teacher image group.
[0006] The present invention provides a technique for suppressing the occurrence of artifacts in the demosaic inference results, even when performing demosaic inference on an image having a hue that is difficult to infer. [Means for solving the problem]
[0007] One aspect of the present invention is a method for detecting a probability distribution of colors in a group of teacher images, comprising: A generating means for generating an image having a color sampled based on the probability distribution as an artificial teaching image; A learning means for learning a learning model that performs demosaic inference using the artificial teacher image; Equipped with 、 The acquisition means corrects the variance in the probability distribution of hue to a larger variance, The artificial teacher image has one or more connected regions having the same pixel value as the color It is characterized by: Effect of the Invention
[0008] According to the configuration of the present invention, even when inferring the demosaicing of an image having a hue that is difficult to infer, the occurrence of artifacts in the demosaicing inference result can be suppressed.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0011] [First Embodiment] In this embodiment, an artificial teacher image, which is an artificial teacher image with reduced hue bias, is generated, and learning is performed using the artificial teacher image as learning data, thereby reducing artifacts in the inference result of demosaicing for a mosaic image.
[0012] (Principle of Artifact Generation and Countermeasure Policy) First, the principle of artifact generation in the prior art will be described with reference to FIGS. 2 and 3. A mosaic image (here, a Bayer image) is an image in which pixels of three colors, R (red), G (green), and B (blue), are arranged according to a color filter array 201. Here, consider the case where a subject 204 with one side being magenta is imaged. At this time, the obtained mosaic image has large R and B pixel values and small G pixel values, that is, a checkerboard-shaped mosaic image 202 in which pixels with large pixel values and small pixels are arranged alternately. This mosaic image 202 can also be seen as having pixels with large pixel values arranged diagonally from the upper right to the lower left. That is, when imaging a subject 203 with a diagonal stripe pattern, the mosaic image 202 is also obtained.
[0013] When inputting the mosaic image 202 into the CNN for mosaic inference, since the mosaic image 202 can correspond to both the subject 203 (pattern image) and the subject 204 (magenta image), it is difficult to uniquely determine the inference result. Therefore, it is considered to output the result of alpha-blending the two images (pattern image, magenta image), which are candidates for the inference result, according to the likelihood. Therefore, even when the correct image is the subject 204 (magenta image), the subject 203 (pattern image) is superimposed and output, and this is perceived as an artifact such as a false pattern.
[0014] As described above, due to the characteristics of its color filter array, it is difficult to make inferences about the magenta image. In other words, depending on the characteristics of the color filter array, there are hues for which inference is difficult. For an image of such a hue, it is difficult to make a correct inference when referring to a local area such as the subject 203 (pattern image), but it may be possible to make a correct inference by referring to a wider area and considering the consistency with the surroundings. In order to effectively utilize the wide-area image features, a large number of teacher images are required during learning, and teacher images of that hue are required in abundance. However, simply increasing the total number of teacher images does not necessarily result in sufficient collection of data for that hue.
[0015] Figure 3(a) is an example of a histogram (hue distribution) created by converting the color space of a group of teacher images obtained from nature from the RGB color space to the HSV color space and extracting only the hue (H) values. The horizontal axis represents the position of the hue on the spectrum, represented by an angle of 0 to 180 degrees. The vertical axis represents the appearance frequency of each hue. From the hue distribution in Figure 3(a), it can be seen that the hue of 150 degrees corresponding to the magenta image is less than other hues. Thus, there is a bias in the hues existing in nature, and the number of images of difficult hues can be insufficient. A countermeasure is to generate artificial teacher images in which all hues are equalized, as shown in Figure 3(c), and use these as learning data for learning.
[0016] When generating the artificial teacher image, it is necessary to determine not only the hue but also the luminance of the artificial teacher image. Unlike the hue, there is no difficulty in inference derived from the characteristics of the color filter array for luminance. Therefore, it is preferable to determine the luminance of the artificial teacher image according to the luminance distribution (pixel value distribution) of the mosaic image that can be input to the CNN during inference. FIG. 3(d) is a histogram (luminance distribution) generated by extracting only the luminance (V) values for the same group of teacher images as in FIG. 3(a), and the luminance is determined based on this.
[0017] Thus, in this embodiment, in order to address artifacts, an artificial teacher image is generated for learning. When determining the pixel values of the artificial teacher image, the luminance value is determined to be equivalent to the luminance distribution of the group of teacher images, and the hue value is determined according to a distribution with less bias than the hue distribution of the group of teacher images.
[0018] (Regarding the configuration of the image processing apparatus) First, a hardware configuration example of the image processing apparatus 100 according to this embodiment will be described with reference to the block diagram of FIG. 1. As the image processing apparatus 100 according to this embodiment, a computer device such as a PC (personal computer), a tablet terminal device, or a smartphone can be applied.
[0019] The CPU 101 executes various processes using computer programs and data stored in the RAM 102 and the ROM 103. Thereby, the CPU 101 controls the operation of the entire image processing apparatus 100 and executes or controls each process described as being performed by the image processing apparatus 100.
[0020] The RAM 102 has an area for storing computer programs and data loaded from the ROM 103, the secondary storage device 104, the external storage device 108, etc., and an area for storing information such as an input image (RAW image) output from the imaging device 111. Further, the RAM 102 has a work area used when the CPU 101 and the GPU 110 execute various processes. Thus, the RAM 102 can appropriately provide various areas.
[0021] The ROM 103 stores the setting data of the image processing apparatus 100, computer programs and data related to the startup of the image processing apparatus 100, computer programs and data related to the basic operations of the image processing apparatus 100, and the like.
[0022] The secondary storage device 104 is a non-volatile memory such as a hard disk drive. The secondary storage device 104 stores an OS (Operating System), computer programs and data for causing the CPU 101 and the GPU 110 to execute or control various processes described as being performed by the image processing apparatus 100, and the like. The computer programs and data stored in the secondary storage device 104 are appropriately loaded into the RAM 102 according to the control by the CPU 101 and become processing targets for the CPU 101 and the GPU 110. In addition to the hard disk drive, various storage devices such as an optical disk drive and a flash memory can be used for the secondary storage device 104.
[0023] The GPU 110 operates based on the computer programs and data loaded into the RAM 102, performs various arithmetic processes on the data received from the CPU 101, and notifies the CPU 101 of the result of the arithmetic operation.
[0024] The imaging device 111 has an image sensor equipped with a color filter having an array such as a Bayer array, and outputs a RAW image output from the image sensor to the system bus 107.
[0025] The input interface 105 is a serial bus interface such as USB or IEEE 1394. The image processing apparatus 100 acquires data, commands, and the like from the outside via the input interface 105.
[0026] The output interface 106 is a serial bus interface such as USB or IEEE 1394, similar to the input interface 105. Note that the output interface 106 may also be a video output terminal such as DVI or HDMI (registered trademark). The image processing apparatus 100 outputs data and the like to the outside via the output interface 106.
[0027] The CPU 101, RAM 102, ROM 103, secondary storage device 104, GPU 110, imaging device 111, input interface 105, and output interface 106 are all connected to the system bus 107.
[0028] The operation unit 112 is a user interface such as a keyboard, mouse, or touch panel. By operating the user, various instructions can be input to the CPU 101 via the input interface 105.
[0029] The external storage device 108 is a memory device connected to / attached to the image processing apparatus 100, such as a hard disk drive, memory card, CF card, SD card, or USB memory. The computer programs and data read from the external storage device 108 are input to the image processing apparatus 100 via the input interface 105 and stored in the RAM 102 or secondary storage device 104. Also, the computer programs and data to be saved in the external storage device 108 are written to the external storage device 108 via the output interface 106.
[0030] The display device 109 has a liquid crystal screen or a touch panel screen and displays the processing results by the CPU 101 or GPU 110 as images, characters, etc. Note that the display device 109 may also be a projection device such as a projector that projects images and characters.
[0031] Note that the configuration shown in FIG. 1 is an example of the configuration of an apparatus capable of realizing each process described below, and the configuration capable of realizing each process described below is not limited to the configuration shown in FIG. 1. For example, in FIG. 1, the imaging device 111 is incorporated into the image processing device 100 as a device built into the image processing device 100. However, it is not limited to this. For example, such an imaging device 111 may be connected to the input interface 105 as an external device of the image processing device 100.
[0032] In the present embodiment, the image processing device 100 operates as follows by executing an image processing application. That is, the image processing device 100 divides an input image (RAW image) output from the imaging device 111 to generate a plurality of pixel blocks, and performs demosaicing inference on each of the plurality of pixel blocks to generate a plurality of inference result blocks. Then, the image processing device 100 combines each inference result block to generate a "combined image having the same size as the input image".
[0033] (Regarding CNN) In the present embodiment, demosaicing inference for a pixel block is performed using a convolutional neural network (CNN). Here, CNN, which is used in all image processing technologies applying deep learning techniques including Non-Patent Document 1, will be described.
[0034] CNN is a learning-based image processing technology that repeatedly performs non-linear operations after convolving an image with a filter generated by training (either training or learning). The filter is also called a Local Receptive Field (LRF). The image obtained by performing non-linear operations after convolving a filter on an image is called a feature map. Also, the training is performed using training data (training images or data sets) consisting of pairs of input images and output images. Briefly, learning is to generate the values of the filter that can convert the input image to the corresponding output image with high precision from the training data. Details of this will be described later.
[0035] When the image has RGB color channels or when the feature map is composed of multiple images, the filter used for convolution also has multiple channels accordingly. That is, the convolutional filter is represented as a 4D array with the number of channels added in addition to the vertical and horizontal sizes and the number of sheets. The process of performing non-linear operations after convolving a filter on an image (or feature map) is represented in units of layers. For example, it is called the n-th layer feature map or the n-th layer filter. Also, for example, a CNN that repeats the convolution and non-linear operation of the filter three times has a three-layer network structure. This process can be formulated as shown in the following equation (1).
[0036]
Equation
[0037] In Equation (1), Wn is the filter of the n-th layer, bn is the bias of the n-th layer, G is a non-linear operator, Xn is the feature map of the n-th layer, and * is the convolution operator. Note that the (l) on the upper right indicates that it is the l-th filter or feature map. The filters and biases are generated by the learning described later and are collectively called network parameters. For non-linear operations, for example, the sigmoid function or ReLU (Rectified Linear Unit) is used. ReLU is given by the following Equation (2).
[0038] [Number]
[0039] That is, it is a non-linear process that sets the negative elements of the input vector X to zero and leaves the positive elements as they are. Next, the learning of CNN will be explained. The learning of CNN is generally performed by minimizing the objective function represented by the following Equation (3) for the learning data consisting of a pair of an input learning image (student image) and a corresponding output learning image (teacher image).
[0040] [Number]
[0041] Here, L is a loss function that measures the error between the correct answer and its estimation. Also, Yi is the i-th output learning image, and Xi is the i-th input learning image. Also, F is a function that collectively represents Equation (1) performed in each layer of CNN. Also, θ is the network parameter (filter and bias). Also,
[0042] [Number]
[0043] is the L2 norm, which is simply the square root of the sum of the squares of the elements of vector Z. Also, n is the total number of training data used for learning. However, since the total number of training data is generally large, in Stochastic Gradient Descent (SGD), a part of the training images is randomly selected and used for learning. This can reduce the computational load in learning using a large amount of training data. Also, as methods for minimizing (optimizing) the objective function, various methods such as the momentum method, AdaGrad method, AdaDelta method, and Adam method are known. The Adam method is given by the following equation (4).
[0044]
Equation
[0045] In equation (4), θi t is the i-th network parameter at the t-th iteration, and g is the gradient of the loss function L with respect to θi t . Also, m and v are momentum vectors, α is the base learning rate, β1 and β2 are hyperparameters, and ε is a small constant. Note that since there is no selection guideline for the optimization method in learning, basically anything can be used, but it is known that there are differences in convergence for each method, resulting in differences in learning time.
[0046] As networks using CNN, ResNet in the field of image recognition and its application RED-Net in the field of super-resolution are well-known. In both cases, by stacking multiple layers of CNN and performing convolution of filters many times, high-precision processing is achieved. For example, ResNet features a network structure with a path that shortcuts the convolutional layer, thereby realizing a 152-layer multi-layer network and achieving high-precision recognition approaching the human recognition rate. The reason why processing is made more precise by multi-layer CNN is simply that by repeating non-linear operations many times, a non-linear relationship between input and output can be expressed.
[0047] The CNN according to this embodiment is a trained CNN that is trained to output an inference result (inference result block) of demosaicing for a pixel block when the pixel block is input.
[0048] (Functional configuration example of the image processing apparatus) An example of the functional configuration of the image processing apparatus 100 is shown in the block diagram of FIG. 4(a). Also, the learning of demosaicing inference (learning of the CNN that performs demosaicing inference) by the image processing apparatus 100 will be described according to the flowchart of FIG. 5. In FIG. 5, steps S501 to S507 are processes related to the learning of demosaicing inference by the image processing apparatus 100, and steps S508 to S509 are demosaicing processes by a processing apparatus separate from the image processing apparatus 100. Note that steps S508 to S509 are not limited to being performed by a processing apparatus separate from the image processing apparatus 100, and the image processing apparatus 100 may perform the demosaicing processes of steps S508 to S509. In this embodiment, the hardware configuration of the processing apparatus is the same as that of the image processing apparatus 100 (that is, has the configuration shown in FIG. 1), but it may be different from the image processing apparatus 100. An example of the functional configuration of the processing apparatus is shown in the block diagram of FIG. 4(c).
[0049] Hereinafter, each functional unit shown in FIGS. 4(a), (b), and (c) will be described with the processing entity. However, in reality, the functions of the functional units are realized by the CPU 101 or the GPU 110 executing a computer program for causing the CPU 101 or the GPU 110 to execute the functions of the functional units. Note that one or more of the functional units shown in FIGS. 4(a), (b), and (c) may be implemented in hardware.
[0050] In step S501, the acquisition unit 401 acquires a plurality of teacher images, which are images (RGB format images) in which each pixel has respective pixel values of RGB. For example, the acquisition unit 401 acquires teacher images according to the method described in Non-Patent Document 1. Specifically, as shown in FIG. 8, the acquisition unit 401 applies simple demosaicing to the mosaic image 801 obtained by imaging with the imaging device 111 to generate an RGB image 802, and generates a reduced image obtained by reducing the RGB image 802 as the teacher image 803. Bilinear interpolation is used for the simple demosaicing, but other demosaicing methods may also be used. Also, here, a Bayer array is shown as the color filter array of the imaging element in the imaging device 111, but other color filter arrays such as X-Trans may also be used.
[0051] Note that the above method for acquiring the teacher image is an example, and the method for acquiring the teacher image by the acquisition unit 401 is not limited to a specific acquisition method. For example, the acquisition unit 401 may generate a teacher image by a method other than the method described in Non-Patent Document 1. Also, for example, teacher images generated in advance by some method may be stored in advance in the secondary storage device 104 or the external storage device 108. In this case, the acquisition unit 401 may read out the teacher image stored in the secondary storage device 104, or may read out the teacher image stored in the external storage device 108 via the input interface 105. Also, teacher images generated in advance by some method may be registered in an external device connected via a network (a wired / wireless network such as the Internet or a LAN) to the image processing device 100. In this case, the acquisition unit 401 may acquire the teacher image from the external device via the network. Also, an RGB format image may be acquired as the teacher image by imaging while shifting the position of the imaging element of the imaging device 111. Thus, any acquisition method may be adopted, and in step S501, the acquisition unit 401 acquires a plurality of teacher images (a group of teacher images) by any acquisition method.
[0052] In step S502, the statistical processing unit 402 performs statistical analysis on the plurality of teacher images acquired in step S501 to obtain statistical quantities for each channel (hue (H), saturation (S), brightness (V)). Here, the details of the processing in step S502 will be described. An example of the functional configuration of the statistical processing unit 402 is shown in the block diagram of FIG. 4(b).
[0053] For each of the plurality of teacher images acquired in step S501, the statistic calculation unit 410 converts the pixel value of each pixel in the teacher image into a pixel value in the HSV color space (pixel value of hue (H), pixel value of saturation (S), pixel value of brightness (V)).
[0054] Then, based on the pixel values of hue (H) collected from all the teacher images acquired in step S501, the statistic calculation unit 410 generates "a histogram of the pixel values of hue (H) (for example, a histogram representing the number of pixels corresponding to each class of hue (H) at 5-degree intervals).
[0055] Also, based on the pixel values of saturation (S) collected from all the teacher images acquired in step S501, the statistic calculation unit 410 generates "a histogram of the pixel values of saturation (S) (for example, a histogram representing the number of pixels corresponding to each class of saturation (S) at 5% intervals).
[0056] Also, based on the pixel values of brightness (V) collected from all the teacher images acquired in step S501, the statistic calculation unit 410 generates "a histogram of the pixel values of brightness (V) (for example, a histogram representing the number of pixels corresponding to each class of brightness (V) at 5% intervals).
[0057] Next, the calculation unit 411 converts each of the histogram ν of the pixel values of hue (H), H the histogram ν of the pixel values of saturation (S), S the histogram ν of the pixel values of brightness (V), V into a probability distribution function according to the following formula (5).
[0058]
Equation
[0059] Here, p c (x) is a probability distribution function representing the occurrence probability of class x in channel c, and ν c (x) is a histogram representing the frequency of class x in channel c.
[0060] Here, the probability distribution function is calculated from the actual measurement of the teacher image group, but it may be obtained by other methods. For example, the statistic calculation unit 410 obtains the average μ and variance σ of the pixel values in each channel from the teacher image group acquired in step S501. Then, the calculation unit 411 obtains a Gaussian distribution G(μ, σ) according to the average μ and variance σ obtained for each channel as the probability distribution function of the channel. Also, the average of the Gaussian distribution may be used as another statistic such as the mode or median of the teacher image group, and it may be fitted to a probability distribution function other than the Gaussian distribution. Further, processing such as smoothing, shifting, or linear transformation may be performed on the obtained probability distribution function.
[0061] Next, the distribution correction unit 412 corrects the probability distribution function p H of the hue (H) so that the variance in the probability distribution function p H increases. The distribution correction unit 412 corrects the probability distribution function p H of the hue (H) according to the following formula (6) using, for example, preset coefficients t and u.
[0062]
Equation
[0063] By such correction, a new probability distribution function p H can be obtained. Note that when t = 0, a uniform probability distribution function p HIt may also be obtained. Taking the processing by the above statistical processing unit 402 as an example in FIG. 3, first, histograms for each channel such as a histogram 301 of pixel values of hue (H) and a histogram 304 of pixel values of luminance (V) are obtained. Then, a probability distribution function 302 obtained by correcting the probability distribution function obtained by converting the histogram and a uniform probability distribution function 303 with t = 0 are obtained.
[0064] Note that the correction method for correcting the above probability distribution function is an example and is not limited to a specific correction method. For example, smoothing may be applied to p H (x), or the variance may be increased after fitting to a Gaussian distribution. Also, for the probability distribution function p S of saturation (S), the same correction as the above correction may or may not be performed.
[0065] In step S503, the generation unit 403 generates a plurality of teacher images based on the statistics (probability distribution functions p H of hue (H), probability distribution functions p S of saturation (S), and probability distribution functions p V of luminance (V)) obtained for each channel in step S502. Hereinafter, the processing for generating one teacher image will be described.
[0066] First, the generation unit 403 selects an object from the object database stored in the secondary storage device 104 or the external storage device 108 according to a specified rule or randomly. The object database is a database storing objects such as figures, symbols, characters, and repeating patterns (objects with simple shapes). Then, the generation unit 403 generates an image having the selected object as the foreground (that is, the area other than the object area in the image is the background) as a teacher image. At this time, the generation unit 403 determines the colors of the foreground and background in the teacher image based on the probability distribution function p c (c = H, S, V). First, the generation unit 403 samples a random variable xc according to the probability distribution function p c (x) as follows.
[0067] [Number]
[0068] That is, the generation unit 403 samples a random variable x H that follows the probability distribution function p H (x), samples a random variable x S that follows the probability distribution function p S (x), and samples a random variable x V that follows the probability distribution function p V (x). Then, the generation unit 403 sets the value of the random variable x H as the pixel value of the hue (H) of the foreground, sets the value of the random variable x S as the pixel value of the saturation (S) of the foreground, and sets the value of the random variable x V as the pixel value of the brightness (V) of the foreground.
[0069] The same process is also performed for the background. That is, the generation unit 403 samples a random variable x H that follows the probability distribution function p H (x), samples a random variable x S that follows the probability distribution function p S (x), and samples a random variable x V that follows the probability distribution function p V (x). Then, the generation unit 403 sets the value of the random variable x H as the pixel value of the hue (H) of the background, sets the value of the random variable x S as the pixel value of the saturation (S) of the background, and sets the value of the random variable x V as the pixel value of the brightness (V) of the background.
[0070] In this way, the generation unit 403 samples a random variable xc that follows the probability distribution function p c (x) for each of the foreground and the background, and sets the color corresponding to the sampled random variable c. Then, the generation unit 403 generates a plurality of artificial teacher images by repeating the above process a plurality of times.
[0071] Next, for each of the plurality of generated artificial teacher images, the generation unit 403 performs subsampling of pixel values according to the color filter array (the color filter array of the imaging device 111 that captured the imaging image that is the source of the artificial teacher image) from the artificial teacher image to generate a mosaic image (student image). Here, the process for generating a student image from an artificial teacher image will be described with reference to FIG. 10.
[0072] The R component image 1001 is an image of the R plane of the artificial teacher image (an image composed of the pixel values of the R components of each pixel in the artificial teacher image). The G component image 1002 is an image of the G plane of the artificial teacher image (an image composed of the pixel values of the G components of each pixel in the artificial teacher image). The B component image 1003 is an image of the B plane of the artificial teacher image (an image composed of the pixel values of the B components of each pixel in the artificial teacher image).
[0073] The generation unit 403 generates a student image 1004 in which pixel values are subsampled and arranged according to the color filter array 1005 of the imaging device 111 from the R component image 1001, the G component image 1002, and the B component image 1003.
[0074] More specifically, when the student image 1004 is divided into four "2-pixel x 2-pixel regions" (divided regions of the same size as the color filter array 1005) and the color filter array 1005 is overlaid on each divided region, the pixel in the upper left corner of each divided region corresponds to the channel "R", the pixels in the upper right corner and the lower left corner correspond to the channel "G", and the pixel in the lower right corner corresponds to the channel "B". Therefore, the generation unit 403 sets, as the pixel value of each pixel of the student image 1004, the pixel value corresponding to the pixel position of the pixel in the component image corresponding to the channel of the pixel among the three component images (the R component image 1001, the G component image 1002, and the B component image 1003).
[0075] In this way, the generation unit 403 generates a plurality of sets of an image set of an artificial teacher image and a student image generated based on the artificial teacher image by generating a student image for each of the plurality of artificial teacher images.
[0076] Note that the object includes at least one or more connected regions having pixel values of the same degree, the size of each connected region is larger than the filter size of the CNN, and it is desirable that the shape of the hue histogram of all the connected regions is bimodal. Also, it is okay even if noise is added. What is important is the variety of variations in the boundary (edge) shape of the two types of hues assigned to each connected region.
[0077] An example of the artificial teacher image is shown in FIG. 9. FIG. 9(a) shows an example of the artificial teacher image generated when a symbol is selected as the object, and FIG. 9(b) shows an example of the artificial teacher image generated when a figure is selected as the object. Also, FIG. 9(c) shows an example of the artificial teacher image generated when a repeating pattern is selected as the object.
[0078] Next, in step S504, the acquisition unit 404 acquires the parameters (network parameters) that define the CNN. The network parameters include the coefficients of each filter in the CNN. The network parameters are set as random numbers following the normal distribution of He. The normal distribution of He is a normal distribution with a mean of 0 and a variance of the following σ h and is such a normal distribution.
[0079]
Equation
[0080] Here, m N indicates the number of neurons of the filter in the CNN. Note that the method for determining the network parameters is not limited to the above method, and other determination methods may be adopted. Also, instead of or in addition to the coefficients of each filter, other types of parameters may be acquired as the network parameters.
[0081] The method for obtaining network parameters is not limited to a specific method. For example, the acquisition unit 404 may read the network parameters stored in the secondary storage device 104, or may read the network parameters stored in the external storage device 108 via the input interface 105. Further, network parameters may be registered in an external device connected to the image processing apparatus 100 via a network (a wired / wireless network such as the Internet or a LAN). In this case, the acquisition unit 404 may acquire the network parameters from the external device via the network. Thus, any acquisition method may be adopted, and in step S504, the acquisition unit 404 acquires network parameters by any acquisition method.
[0082] In step S505, the learning unit 405 configures a CNN (initializes the weights of the CNN) according to the network parameters acquired in step S504. Then, the learning unit 405 uses the plurality of image sets generated in step S503 as learning data to perform learning of "a CNN for performing demosaicing inference". For the learning, a CNN disclosed in Non-Patent Document 1 is used. The structure and learning process of this CNN will be described with reference to FIG. 11.
[0083] The CNN has a plurality of filters 1102 that perform the operation of Equation (1). When inputting the student image 1004 to this CNN, it is converted into a 3-channel missing image 1101. In the missing image 1101 of the R channel, only the pixels of the R component of the student image 1004 are included, and the pixel values of the other pixels are set to missing values (0). In the missing image 1101 of the G channel, only the pixels of the G component of the student image 1004 are included, and the pixel values of the other pixels are set to missing values (0). In the missing image 1101 of the B channel, only the pixels of the B component of the student image 1004 are included, and the pixel values of the other pixels are set to missing values (0). Note that the missing values may be interpolated by a method such as bilinear interpolation. Next, the filter 1102 is sequentially applied to this missing image 1101 to calculate a feature map. Subsequently, the concatenation layer 1103 of the CNN concatenates the calculated feature map and the missing image 1101 in the channel direction. If the number of channels of the feature map and the missing image 1101 are n1 and n2, respectively, the number of channels of the concatenation result is (n1 + n2). Subsequently, the filter 1102 is applied to this concatenation result, and the final filter 1102 outputs 3 channels to obtain an inference result 1104. Then, the residual between the obtained inference result 1104 and the artificial teacher image corresponding to the student image 1004 is calculated, and the average of the residuals for all image sets is obtained as the loss function value. Then, the CNN is trained by updating the network parameters from the loss function value by the error backpropagation method or the like. Such a series of CNN training (updating of network parameters) is performed by the learning unit 405.
[0084] In step S506, the inspection unit 406 determines whether the learning end condition is satisfied. For example, the inspection unit 406 obtains a mosaic image chart including objects such as figures and symbols having a hue with low statistical frequency in an image group such as landscape photos and portrait photos that are not used for learning. Examples of hues with low frequency include those in a complementary color relationship such as green / magenta in particular. In this embodiment, such a mosaic image chart is created in advance and stored in the secondary storage device 104, the external storage device 108, etc., and the inspection unit 406 obtains the mosaic image chart from the secondary storage device 104 or the external storage device 108. However, the method for obtaining the mosaic image chart is not limited to a specific obtaining method. Then, the inspection unit 406 uses a CNN according to the current network parameters to demosaic the mosaic image chart to generate a demosaiced image. And when the degree of occurrence of artifacts in the demosaiced image is less than the threshold value, the inspection unit 406 determines that "the learning end condition is satisfied", and when the degree of occurrence of artifacts is greater than or equal to the threshold value, it determines that "the learning end condition is not satisfied".
[0085] Note that the learning end condition described here is an example and is not limited thereto. For example, the learning end condition may be "the amount of change from the network parameters before the update of the updated network parameters is less than the threshold value". Also, the learning end condition may be "the residual between the inference result by the CNN and the artificial teacher image is below the threshold value". Also, the learning end condition may be "the number of iterations of learning (update of network parameters) has reached the threshold value". Also, the learning end condition may be a combination of two or more conditions, and when all of the two or more conditions are satisfied, it may be determined that the learning end condition is satisfied.
[0086] As a result of such determination, when the learning end condition is satisfied, the process proceeds to step S507. On the other hand, when the learning end condition is not satisfied, the process proceeds to step S503 to generate a new artificial teacher image group and perform learning again.
[0087] In step S507, the inspection unit 406 stores the latest network parameters updated by the learning unit 405 in the storage unit 407. Note that the learning process of "inference of the demosaicing result for the mosaic image" is completed in step S507.
[0088] In step S508, the acquisition unit 408 acquires a mosaic image (RAW image) to be demosaiced as an input image. The method for acquiring the input image by the acquisition unit 408 is not limited to a specific acquisition method. For example, the acquisition unit 408 may control the imaging device 111 and acquire the RAW image captured by the imaging device 111 as the input image by this control. Also, for example, the acquisition unit 408 may acquire the RAW image stored in the secondary storage device 104 as the input image, or may acquire the RAW image stored in the external storage device 108 as the input image via the input interface 105. Further, when the processing device 499 is connected to a network (a wired / wireless network such as the Internet or a LAN), the acquisition unit 408 may acquire a RAW image from an external device as the input image via the network.
[0089] In step S509, the inference unit 409 acquires the network parameters stored in the storage unit 407 from the image processing device 100, and constructs a learned CNN based on the network parameters. Then, the inference unit 409 inputs the input image acquired in step S508 into the learned CNN and obtains the output of the learned CNN as an inference result image that is the inference result of the demosaicing for the input image. And the inference unit 409 outputs the inference result image, but the output destination of the inference result image is not limited to a specific output destination.
[0090] For example, the inference unit 409 may cause the display device 109 to display the inference result image by outputting the inference result image to the display device 109 via the output interface 106. Also, for example, the inference unit 409 may save the inference result image in the secondary storage device 104, or may output the inference result image to the external storage device 108 via the output interface 106 and cause the external storage device 108 to save the inference result image. Further, when the processing device 499 is connected to a network (a wired / wireless network such as the Internet or a LAN), the inference unit 409 may transmit the inference result image to an external device via the network.
[0091] As described above, according to the present embodiment, even when performing inference of the mosaic of an input image having a hue that is difficult to infer, it is possible to suppress the occurrence of artifacts in the inference result of the mosaic.
[0092] <Modification Example> In the first embodiment, the statistical processing unit 402 obtains probability distribution functions for each color component of H, S, and V from the teacher image group, and sets the colors of the foreground and background in the artificial teacher image based on the probability distribution functions. However, the method for determining the colors of the foreground and background in the artificial teacher image is not limited to a specific determination method. In this modification example, several examples of this determination method will be described. Hereinafter, the differences from the first embodiment will be described, and it is assumed that the same as the first embodiment unless otherwise specifically mentioned below.
[0093] The statistics calculation unit 410 generates a "histogram of R pixel values" (e.g., a histogram showing the number of pixels corresponding to a class in increments of 5 (pixel values of R)) based on the pixel values of R collected from all teacher images acquired in step S501. The statistics calculation unit 410 also generates a "histogram of G pixel values" (e.g., a histogram showing the number of pixels corresponding to a class in increments of 5 (pixel values of G)) based on the pixel values of G collected from all teacher images acquired in step S501. The statistics calculation unit 410 also generates a "histogram of B pixel values" (e.g., a histogram showing the number of pixels corresponding to a class in increments of 5 (pixel values of B)) based on the pixel values of B collected from all teacher images acquired in step S501.
[0094] Next, the calculation unit 411 converts the histogram v R Let p be the probability distribution function R (x) is transformed into the probability distribution function p R (x) is the probability distribution function that represents the probability of occurrence of class x in R.
[0095] The calculation unit 411 also calculates the histogram v G Let p be the probability distribution function G Probability distribution function p G (x) is the probability distribution function that represents the occurrence probability of class x in G.
[0096] In addition, the calculation unit 411 converts the histogram v B Let p be the probability distribution function B Probability distribution function p B (x) is the probability distribution function that represents the occurrence probability of class x in B.
[0097] The generation unit 403 generates, as a teacher image, an image with the selected object as the foreground (in this image, the area other than the object area is the background), similar to the first embodiment. At this time, the generation unit 403 determines the colors of the foreground and the background in the teacher image based on the probability distribution functions p c (c = R, G, B).
[0098] That is, the generation unit 403 samples a random variable x R that follows the probability distribution function p R (x), samples a random variable x G that follows the probability distribution function p G (x), and samples a random variable x B that follows the probability distribution function p B (x). Then, the generation unit 403 sets the value of the random variable x R as the R pixel value of the foreground, sets the value of the random variable x G as the G pixel value of the foreground, and sets the value of the random variable x B as the B pixel value of the foreground.
[0099] The same process is performed for the background. That is, the generation unit 403 samples a random variable x R that follows the probability distribution function p R (x), samples a random variable x G that follows the probability distribution function p G (x), and samples a random variable x B that follows the probability distribution function p B (x). Then, the generation unit 403 sets the value of the random variable x R as the R pixel value of the background, sets the value of the random variable x G as the G pixel value of the background, and sets the value of the random variable x B as the B pixel value of the background.
[0100] Note that, as shown in the following formula (9), the probability distribution functions p R (x), p G (x), and p B (x) are integrated into one probability distribution function p(x), and based on this probability distribution function p(x), x R , xB , x G may be sampled.
[0101] [Number]
[0102] In addition, the method for calculating the probability distribution function is not limited to the above calculation method. For example, a probability distribution function in which each channel has a correlation may be obtained. This probability distribution function can be expressed as, for example, p(x R , x G , x B ). Also, the color space is not limited to the HSV color space or the RGB color space, and color spaces such as the YUV color space, the L*a*b* color space, and the YCbCr color space may be used. In that case, the hue component and the color difference component are corrected so that the variance becomes small in the distribution correction unit 412.
[0103] [Second Embodiment] Hereinafter, differences from the first embodiment will be described, and it is assumed that they are the same as the first embodiment unless otherwise specifically mentioned. In the first embodiment, an example of learning using an artificial teacher image as learning data was described. However, overfitting may occur only with learning using an artificial teacher image, and the robustness against natural images (non-artificial images acquired by imaging in the real world) may be lost. Therefore, in this embodiment, an example of learning the inference of demosaicing using both a difficult teacher image generated from a natural image and an artificial teacher image as learning data will be described. More specifically, in this embodiment, as shown in FIG. 12, first, pre-learning using a difficult teacher image generated from a natural image as learning data is performed to obtain network parameters of the CNN by the pre-learning. Then, using the CNN according to the "network parameters of the CNN by pre-learning", main learning using the artificial teacher image as learning data is performed to obtain network parameters of the CNN of the main learning.
[0104] An example of the functional configuration of the image processing apparatus 600 according to the present embodiment is shown in the block diagram of FIG. 6. Further, the learning of the inference of demosaicing by the image processing apparatus 600 will be described according to the flowchart of FIG. 7. In FIG. 7, steps S701, S702, S504, S703, S704, S502, S503, S705, S706, S506, and S507 are processes related to the learning of the inference of demosaicing by the image processing apparatus 100. And steps S508 to S509 are demosaicing processes by a processing apparatus separate from the image processing apparatus 600. Note that steps S508 to S509 are not limited to being performed by a processing apparatus separate from the image processing apparatus 600, and the image processing apparatus 600 may perform the demosaicing processes of steps S508 to S509.
[0105] Hereinafter, each functional unit shown in FIG. 6 will be described with the processing entity. However, in actuality, the functions of the functional units are realized by the CPU 101 or the GPU 110 executing a computer program for causing the CPU 101 or the GPU 110 to execute the functions of the functional units. Note that one or more of the functional units shown in FIG. 6 may be implemented in hardware.
[0106] In step S701, the acquisition unit 601 acquires a general teacher image, which is a teacher image based on a natural image. The teacher image based on a natural image is a teacher image generated from a natural image, and the acquisition method thereof may be the same as the acquisition method of the teacher image in step S501 above, and is not limited to a specific acquisition method. In step S701, the acquisition unit 601 acquires a plurality of general teacher images (general teacher image group) by any acquisition method.
[0107] In step S702, the extraction unit 602 extracts, from the general teacher image group obtained in step S701, an image for which demosaicing inference is difficult as a difficult teacher image. The method for extracting the difficult teacher image from the general teacher image group is not limited to a specific method, but here, as an example, the method described in Non-Patent Document 1 is used. In this method, a mosaic image obtained by mosaicing the general teacher image is generated, a simple demosaicing method is applied to the mosaic image to generate a demosaiced image, and the error between the general teacher image and the demosaiced image is obtained. The error is, for example, generating a difference image between the general teacher image and the demosaiced image, and the sum of the pixel values of each pixel in the difference image (the total of the differences in pixel values between corresponding pixel positions in the general teacher image and the demosaiced image). Then, when the error between the general teacher image and the demosaiced image is equal to or greater than the threshold value θ, the general teacher image is extracted as the difficult teacher image. As the above simple demosaicing, bilinear interpolation or the result of training the CNN of Non-Patent Document 1 for a small number of epochs using the general teacher image is used. Note that the threshold value θ may be set to -∞, and in this case, the general teacher image group and the difficult teacher image group coincide.
[0108] Furthermore, for each of the extracted difficult teacher images, the extraction unit 602 generates a corresponding student image in the same manner as the generation unit 403. In this way, the extraction unit 602 generates a plurality of sets of image sets (difficult image sets) of the difficult teacher image and the student image generated based on the difficult teacher image.
[0109] In step S703, the learning unit 405 configures a CNN (initializes the weights of the CNN) according to the network parameters obtained in step S504. Then, the learning unit 405 uses the plurality of difficult image sets generated in step S702 as learning data, and performs learning (pre-learning) of demosaicing inference in the same manner as in step S505 above. Since images for which demosaicing inference is difficult are images with high learning efficiency, performing learning of demosaicing inference using these difficult teacher images leads to an improvement in the performance of demosaicing inference by the CNN.
[0110] In step S704, the inspection unit 406 stores the latest network parameters (network parameters of the CNN by pre-training) updated by the learning unit 405 in the storage unit 407.
[0111] In step S502 according to the present embodiment, the statistical processing unit 402 performs statistical analysis on the general teacher image instead of the teacher image to obtain the statistic for each channel. And in step S503 according to the present embodiment, the generation unit 403 generates a plurality of artificial teacher images based on the statistics obtained for each channel in step S502, and generates student images for each of the plurality of artificial teacher images, in the same manner as in the first embodiment. That is, also in the present embodiment, in step S503, the generation unit 403 generates a plurality of sets of image sets of the artificial teacher image and the corresponding student image.
[0112] In step S705, the acquisition unit 404 acquires the "network parameters of the CNN obtained by pre-training" stored in the storage unit 407. In step S706, the learning unit 405 configures the CNN (initializes the weights of the CNN) according to the network parameters acquired in step S705. Then, the learning unit 405 performs learning of demosaicing inference (this learning) using the plurality of image sets generated in step S503 as learning data.
[0113] In the learning of demosaicing inference according to this embodiment, difficult teacher images were used as learning data for pre-training, and artificial teacher images were used as learning data for this learning. However, this is not the only way. Artificial teacher images may be used as learning data for pre-training, and difficult teacher images may be used as learning data for this learning. Also, in pre-training, an image group in which a part of the difficult teacher image and a part of the artificial teacher image are mixed may be used as learning data, and in this learning, an image group in which the remaining part of the difficult teacher image and the remaining part of the artificial teacher image are mixed may be used as learning data. Further, the mixing ratio of the difficult teacher image and the artificial teacher image may be controlled in each of the pre-training and this learning. Also, pre-training and this learning may be alternately repeated. Whichever learning method is adopted, both the difficult teacher image and the artificial teacher image are used for the learning of demosaicing inference.
[0114] Also, in this embodiment, a distribution function was calculated by actual measurement based on a general teacher image group, and an artificial teacher image was generated according to it. This is because it is considered that the statistical luminance distribution of the input image input to the CNN at the time of inference becomes equivalent to the luminance distribution of the general teacher image. Note that the population for measuring the distribution function does not have to be the general teacher image group, and measurements may be performed from another population, such as a difficult teacher image group or another previously prepared image database.
[0115] Thus, according to this embodiment, even when performing demosaicing inference on an input image having a hue difficult to infer, it is possible to suppress the occurrence of artifacts in the demosaicing inference result while ensuring robustness against natural images.
[0116] [Third Embodiment] The configurations shown in FIGS. 4 and 6 can be appropriately modified / changed. For example, one functional unit may be divided into a plurality of functional units according to functions, or two or more functional units may be integrated into one functional unit. Also, the configurations in FIGS. 4 and 6 may be constituted by two or more devices. In that case, each device is connected via a circuit or a wired or wireless network, and performs data communication with each other to perform a cooperative operation, thereby realizing each process described above as being performed by the image processing apparatus 100 or the image processing apparatus 600.
[0117] In addition, in each of the above-described embodiments and modification examples, CNN is used as an example of the learning model, but other types of learning models may be used instead of CNN. In that case, the network parameters will use the parameters that define the learning model to be used.
[0118] Also, the numerical values, processing timings, processing order, processing entities, configurations / sending destinations / sending sources / storage locations of data (information), etc. used in each of the above-described embodiments and modification examples are given as examples for the purpose of specific explanation, and are not intended to be limited to such examples.
[0119] In addition, some or all of the above-described embodiments and modification examples may be used in appropriate combinations. Also, some or all of the above-described embodiments and modification examples may be selectively used.
[0120] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Also, it can be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0121] The invention is not limited to the above-described embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.
Explanation of Symbols
[0122] 401: Acquisition Unit 402: Statistical Processing Unit 403: Generation Unit 404: Acquisition Unit 405: Learning Unit 406: Inspection Unit 407: Memory Unit
Claims
1. An acquisition means for acquiring a probability distribution of colors in a group of teacher images; A generation means for generating an artificial teacher image having an image with a color sampled based on the probability distribution; A learning means for performing learning of a learning model that performs demosaicing inference using the artificial teacher image and comprising: The acquisition means corrects the variance in the probability distribution of hue to a larger variance, The artificial teacher image has one or more connected regions having the same pixel value as the color An image processing apparatus characterized by that.
2. The acquisition means generates a histogram of colors in the group of teacher images, and acquires the probability distribution of the colors from the histogram. The image processing apparatus according to claim 1, characterized by that.
3. The acquisition means obtains the average and variance of colors in the group of teacher images, and acquires a probability distribution based on the obtained average and variance. The image processing apparatus according to claim 1, characterized by that.
4. The probability distribution is a probability distribution for each color in the group of teacher images. The image processing apparatus according to any one of claims 1 to 3, characterized by that.
5. The probability distribution is a probability distribution obtained by integrating the probability distributions of each color in the group of teacher images. The image processing apparatus according to any one of claims 1 to 3, characterized by that.
6. The generation means generates an artificial teacher image including a foreground having a color sampled based on the probability distribution and a background having a color sampled based on the probability distribution. The image processing apparatus according to any one of claims 1 to 5, characterized by that.
7. The learning means performs the learning using the artificial teacher image and a pupil image corresponding to the artificial teacher image. The image processing apparatus according to any one of claims 1 to 6, characterized by that.
8. Furthermore, The image processing apparatus according to claim 7, characterized by comprising means for generating the pupil image by performing color subsampling from the artificial teacher image according to a color filter array of an imaging device that captured the image that is the source of the group of teacher images.
9. Furthermore, An extraction means for extracting, from the group of teacher images, a teacher image for which demosaicing inference is difficult as a difficult teacher image is provided, The learning means further performs the learning using the difficult teacher image. The image processing apparatus according to any one of claims 1 to 8, characterized by that.
10. The learning means performs pre-training on the learning model using the difficult teacher image, and then performs main learning using the artificial teacher image. The image processing apparatus according to claim 9, characterized in that.
11. The learning means performs pre-training on the learning model using the artificial teacher image, and then performs main learning using the difficult teacher image. The image processing apparatus according to claim 9, characterized in that.
12. The learning means performs pre-training on the learning model using an image group in which a part of the difficult teacher image and a part of the artificial teacher image are mixed, and then performs main learning using an image group in which the remaining part of the difficult teacher image and the remaining part of the artificial teacher image are mixed. The image processing apparatus according to claim 9, characterized in that.
13. The learning means controls the ratio of mixing the artificial teacher image and the difficult teacher image in the pre-training and the main learning. The image processing apparatus according to claim 12, characterized in that.
14. The learning means alternately repeats the pre-training and the main learning. The image processing apparatus according to any one of claims 10 to 13, characterized in that.
15. Furthermore, Obtain a mosaic image chart including an object having a hue with statistically low frequency, generate a demosaicked image obtained by demosaicking the mosaic image chart using the learning model, and when the degree of occurrence of artifacts in the demosaicked image is less than a threshold value, it is determined that the end condition of the learning is satisfied, and when the degree of occurrence is equal to or greater than the threshold value, it is determined that the end condition of the learning is not satisfied. The image processing apparatus according to any one of claims 1 to 14, characterized in that it comprises means for.
16. Furthermore, The image processing apparatus according to any one of claims 1 to 15, characterized in that it comprises means for obtaining an inference result image which is an inference result of demosaicking for an input image using the learning model.
17. An image processing method performed by an image processing apparatus, An acquisition step in which the acquisition means of the image processing apparatus acquires the probability distribution of colors in the teacher image group; A generation step in which the generation means of the image processing apparatus generates an image having a color sampled based on the probability distribution as an artificial teacher image; A learning step in which the learning means of the image processing apparatus performs learning of a learning model that performs demosaicking inference using the artificial teacher image Comprising, In the acquisition step, the variance in the probability distribution of the hue is corrected to a larger variance, The artificial teacher image has one or more connected regions having the same pixel value as the color An image processing method characterized by this.
18. A computer program for causing a computer to function as each means of the image processing apparatus according to any one of claims 1 to 16.
Citation Information
Patent Citations
CFA image demosaicing method based on generative adversarial neural network
CN111383200A