Method for generating learning model, program, and information processing apparatus

By using a neural network with zero-sum convolutional filters in the first layer, the method addresses the issue of brightness affecting image recognition, ensuring stable accuracy across varying lighting conditions.

JP7700790B2Active Publication Date: 2025-07-01SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022536232
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-14
Filing Date
2021-06-30
Publication Date
2025-07-01
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

The brightness of an image significantly affects the accuracy of image recognition using existing learning models, particularly in varying lighting conditions.

Method used

A neural network with a zero-sum convolutional filter in the first layer is employed, where the sum of coefficients in each channel of the convolution filters is adjusted to zero, minimizing the influence of brightness variations in the image.

Benefits of technology

This approach enables consistent and accurate image recognition regardless of changes in brightness, improving recognition accuracy across different lighting conditions without the need for additional signal level corrections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700790000001
    Figure 0007700790000001
  • Figure 0007700790000002
    Figure 0007700790000002
  • Figure 0007700790000003
    Figure 0007700790000003
Patent Text Reader

Abstract

This technology relates to: a method for generating a learning model that makes it possible to perform image recognition in which the impact of the brightness of a target image is reduced; a program; and an information processing device. Given a neural network which is to be applied to a learning model that performs recognition processing on input data and which has a plurality of convolution filters, the learning model is learned such that the total sum of coefficients for one or more channels of one or more convolution filters, from among the convolution filters of a first layer of the neural network, approaches zero.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a method for generating a learning model, a program, and an information processing apparatus, and more particularly, to a method for generating a learning model, a program, and an information processing apparatus that can perform image recognition with reduced influence of the brightness of a target image.

Background Art

[0002] Patent Document 1 discloses performing learning by imposing a constraint such that the sum of the weights at the same element position in a plurality of channels of a convolution filter becomes zero in a learning model having a structure of a convolutional neural network.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When performing image recognition of a target image using a learning model, the brightness of the image affects the recognition accuracy.

[0005] The present technology has been made in view of such a situation, and enables image recognition with reduced influence of the brightness of a target image.

Means for Solving the Problems

[0006] A method for generating a learning model according to a first aspect of the present technology is a neural network applied to a learning model that performs recognition processing on input data, and is a neural network having a plurality of convolution filters. The learning is performed on the learning model such that the sum of the coefficients in one or more channels of at least one or more of the convolution filters in the first layer of the convolution filters approaches zero.

[0007] The program of the first aspect of the present technology causes a computer to function as a processing unit for performing learning of the learning model so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, which is a neural network applied to the learning model for performing recognition processing on input data, approaches zero. It is a program for this purpose.

[0008] In the first aspect of the present technology, learning of the learning model is performed so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, which is a neural network applied to the learning model for performing recognition processing on input data, approaches zero.

[0009] The information processing apparatus of the second aspect of the present technology is an information processing apparatus having a processing unit that executes an operation of the learning model learned so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, which is a neural network applied to the learning model for performing recognition processing on input data, approaches zero.

[0010] In the second aspect of the present technology, an operation of the learning model learned so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, which is a neural network applied to the learning model for performing recognition processing on input data, approaches zero is executed.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present technology will be described with reference to the drawings.

[0013] <Embodiments of an Image Recognition Apparatus to which the Present Technology is Applied> (Overall Configuration of Image Recognition Apparatus 11) FIG. 1 is a block diagram showing a configuration example of an image recognition apparatus to which the present technology is applied.

[0014] The image recognition apparatus 11 in FIG. 1 captures sensor data (image data) output from an imaging sensor such as a CMOS (Complementary MOS) sensor or a CCD (Charge Coupled Device) sensor as an image to be recognized (hereinafter referred to as a target image), and performs image recognition on the target image. The image recognition apparatus 11 performs classification of the subject of the target image, classification and detection of the subject included in the target image, etc. by image recognition, and outputs the recognition result to an external processing unit, device, etc. However, the input data to the image recognition apparatus 11 is not limited to sensor data from a sensor.

[0015] Although details will be described later, the image recognition apparatus 11 suppresses changes in the result of image recognition even when the brightness of the target image changes when the global illumination of the shooting environment changes. The image recognition apparatus 11 improves the image recognition accuracy when the target image is dark.

[0016] The image recognition device 11 uses, for example, an image captured by a camera used outdoors such as an in-vehicle camera, a camera mounted on a drone, or a surveillance camera as a target image. The captured images of cameras used outdoors have different brightness levels between day and night. Generally, the image recognition accuracy for captured images at night is lower than that during the day. The image recognition device 11 can obtain the same image recognition accuracy as during the day even for captured images at night.

[0017] The image recognition device 11 is not limited to using an image captured by a camera used indoors as a target image. The image recognition device 11 may also use an image captured indoors by a camera mounted on a mobile terminal such as a smartphone, a webcam, or a camera used in a game UI (User Interface) such as a digital camera, or any camera. The image recognition device 11 can obtain stable image recognition accuracy regardless of changes in indoor lighting. The image recognition device 11 can be applied as an image recognition device in a game UI or the like.

[0018] The image recognition device 11 may also use an image captured by a hand-held camera or a camera mounted on a moving body as a target image. The image recognition device 11 can obtain stable image recognition accuracy regardless of changes in the amount of light due to adjustment of the shutter speed of the camera. For example, when the shutter speed is shortened for anti-blur purposes, even when the captured image becomes dark, a decrease in image recognition accuracy can be suppressed.

[0019] Note that although the result of image recognition may change when the location or number of lighting sources changes, the image recognition accuracy improves.

[0020] The image recognition device 11 includes a preprocessing unit 31 and a recognizer 32 (recognition unit).

[0021] The preprocessing unit 31 performs log conversion (logarithmic conversion) on the target image from the imaging sensor and supplies it to the image recognition device 11. Log conversion is a process of converting the pixel value x of each pixel of the target image into a pixel value y according to the following formula.

[0022] y = a×log x + b

[0023] However, a and b are constants.

[0024] The recognizer 32 performs arithmetic processing of a learning model learned by a deep learning method. A convolutional neural network (CNN) is applied to the learning model.

[0025] The recognizer 32 performs image recognition on the target image from the preprocessing unit 31 using the learning model (CNN), and supplies the result to an external processing unit or device. Note that the image recognition in the recognizer 32 may be a process of extracting arbitrary information from the target image, and is not limited to a specific process.

[0026] Figures 2 and 3 are diagrams for explaining the operation of the preprocessing unit 31.

[0027] In the graph of Figure 2, the image signals C1 and C2 represent the image signals of the target image input to the preprocessing unit 31. The image signal C2 exemplifies the image signal of the target image taken in an environment N times brighter than in the case of the image signal C1. According to this, when the shooting environment of the target image becomes brighter, not only the offset but also the amplitude changes greatly corresponding to N times the brightness.

[0028] In the graph of Figure 3, the image signals D1 and D2 exemplify the image signals after the image signals C1 and C2 in Figure 2 are respectively log-transformed by the preprocessing unit 31. In the log transformation, a change in brightness N times that of the shooting environment is converted from an action of multiplying the image signal by N to an action of adding a value corresponding to log N. The log transformation also has the effect of compressing the image signal. Therefore, although the offset of the image signal D2 changes with respect to the image signal D1, the change in amplitude is suppressed.

[0029] The image recognition device 11 may not have the preprocessing unit 31, or the preprocessing unit 31 may perform a conversion other than the log conversion on the target image.

[0030] When the preprocessing unit 31 performs a conversion similar to the log conversion, the conversion may be performed using a piecewise linear function approximated to the log function. In this case, the amount of calculation is reduced.

[0031] The preprocessing unit 31 may perform a conversion using a gamma curve. The preprocessing unit 31 can perform a conversion approximated to the log function by using a gamma curve with I 2.2 as the output for the input I. Since the conversion by the gamma curve is common in image signal processing, existing technologies and resources can be diverted.

[0032] (Details of the recognizer 32) The recognizer 32 performs image recognition by CNN as described above on the target image after the log conversion process from the preprocessing unit 31. In the present embodiment, the CNN has a plurality of convolutional layers.

[0033] The recognizer 32 has a first CNN processing unit 41 and a second CNN processing unit 42.

[0034] The first CNN processing unit 41 represents a processing unit that performs CNN processing up to the convolutional layer of the first layer among the CNN-based processes (CNN processes). The second CNN processing unit 42 represents a processing unit that performs CNN processing subsequent to the convolutional layer of the first layer.

[0035] Note that, as described later, the CNN has multiple stages of combinations of convolutional layers and pooling layers, and the convolutional layers and pooling layers exist alternately. Hereinafter, the n-th layer represents the n-th layer among the layers of the same type. Therefore, the pooling layer next to the first convolutional layer counted from the input side is the first pooling layer, and the convolutional layer next to the first pooling layer is the second convolutional layer. The first CNN processing unit 41 is configured to perform only the processing (convolution operation) of the first convolutional layer, but may perform the processing up to the previous stage of the second convolutional layer (processing up to the first pooling layer). Hereinafter, the convolutional filter used in the n-th convolutional layer is also referred to as the n-th convolutional filter.

[0036] The first CNN processing unit 41 performs the convolution operation in the first convolutional layer using a zero-sum convolutional filter.

[0037] In this specification, the zero-sum convolutional filter refers to a filter that is learned so that the sum of the weights (coefficients) arranged in a grid pattern of the convolutional filter (kernel) approaches 0 (becomes 0). The sum of the weights of the zero-sum convolutional filter is ideally 0, but does not necessarily have to be 0. When the convolutional filter has a plurality of channels (described later), the sum of the weights for each channel is learned to approach 0 respectively.

[0038] To explain the operation of the zero-sum convolutional filter, assume that the first convolutional filter is a one-dimensional filter with three weights as elements. Assume one is a non-zero-sum convolutional filter (-1, 4, -1), and the other is a zero-sum convolutional filter (-1, 2, -1).

[0039] FIGS. 4 to 6 are diagrams for explaining the operation of the first zero-sum convolutional filter.

[0040] In the graph of FIG. 4, the image signals E1 and E2 represent the image signals of the target image input to the recognizer 32. The image signals E1 and E2 respectively correspond to the image signals D1 and D2 in FIG. 3. The image signal E2 is the image signal of the target image captured in an environment brighter than that of the image signal E1.

[0041] In the graph of FIG. 5, the image signals F1 and F2 respectively represent the response signals when convolution operations are performed on the image signals E1 and E2 using a non-zero-sum convolution filter (-1, 4, -1). According to this, an offset remains in the image signal F2 with respect to the image signal F1. Therefore, the results of the respective convolution operations on the image signal E1 and the offset image signal E2 are different.

[0042] In the graph of FIG. 6, the image signals G1 and G2 respectively represent the response signals when convolution operations are performed on the image signals E1 and E2 using a zero-sum convolution filter (-1, 2, -1). According to this, the offset disappears in the image signal G2 with respect to the image signal G1. Therefore, the results of the convolution operations on each of the image signal E1 and the offset image signal E2 substantially coincide. That is, by using a zero-sum convolution filter as the first-layer convolution filter, an image with reduced brightness influence is calculated as the image (feature map) of the first-layer convolution layer for target images (target images with different brightnesses) captured in shooting environments with different brightnesses. As a result, the influence of the brightness of the target image on the image recognition by the recognizer 32 is reduced.

[0043] Patent Document 1 (Japanese Unexamined Patent Application Publication No. 2019-87021) discloses that for the convolution filters for the three channels of R (red), G (green), and B (blue) of a color image, the sum of the weights at the same position is set to 0. However, Patent Document 1 is not a zero-sum convolution filter that sets the sum of the weights in each channel of one convolution filter to 0. Patent Document 1 aims to improve the learning efficiency. In these respects, Patent Document 1 is different from the recognizer 32 that uses a zero-sum convolution filter.

[0044] In the recognizer 32 of FIG. 1, the second CNN processing unit 42 performs CNN processing after the first layer convolutional layer. In the second CNN processing unit 42, convolution using convolutional filters from the second layer onwards is performed. The convolutional filters from the second layer onwards are not limited to filters learned as zero-sum convolutional filters, and convolutional filters learned by any well-known method are used. Note that the convolutional filters from the second layer onwards are described as non-zero-sum convolutional filters.

[0045] (Explanation of CNN) The CNN used by the recognizer 32 as a learning model may have a well-known CNN structure and is not limited to a CNN with a specific structure.

[0046] FIG. 7 is a diagram illustrating the structure of the CNN in the recognizer 32. Note that the size of the target image, the size of the convolutional filter, etc. in FIG. 7 do not necessarily match the actual values. Also, since the CNN structure in FIG. 7 is a general CNN structure, it will be briefly described.

[0047] In FIG. 7, a target image as an input to the recognizer 32 is input to the input layer 51 of the CNN. The target image is, for example, a color image in which each pixel has pixel values of R (red), G (green), and B (blue) (hereinafter referred to as RGB). However, the target image may be a grayscale image in which each pixel has only a luminance pixel value.

[0048] The image size (width (W) × height (H)) of the target image is, for example, 28 × 28 and consists of three RGB channels. When one target image is represented as one volume by combining the three RGB channels, the volume size (W × H × number of channels) of the target image is 28 × 28 × 3.

[0049] In the convolutional layer 52, a convolutional operation using six types of convolutional filters is performed to generate an image (feature map) for six channels. The convolutional filters have a structure in which 5×5 filter sizes (W×H) for each RGB channel of the target image in the input layer 51 are arranged in the depth direction for three channels. Therefore, the filter size (W×H×number of channels) of the convolutional filter used for the target image is 5×5×3. Note that in this embodiment, since the first-layer convolutional filter is a zero-sum convolutional filter, learning is performed so that the sum of the weights of each 5×5 (W×H) convolutional filter for each RGB channel approaches 0.

[0050] The convolutional operation extracts pixel values within a window of the same size (5×5×3) as the convolutional filter from the volume of the target image in the input layer 51, multiplies each pixel value in the window by each weight of the convolutional filter for elements at the same position, and then sums them all. The summed value becomes the pixel value for one pixel of the image in the convolutional layer 52. Such a convolutional operation is performed while sliding the position of the window for the target image by one pixel at a time, and an image of 24×24 (W×H) is generated in the convolutional layer 52. Note that the slide width is not limited to one pixel.

[0051] Since there are six types of convolutional filters, six-channel images are generated in the convolutional layer 52 by the convolutional operation. Therefore, the volume size (W×H×number of channels) of the image in the convolutional layer 52 is 24×24×6. Note that in the first CNN processing unit 41 of FIG. 1, processing is performed until an image in the convolutional layer 52 is generated.

[0052] In the pooling layer 53, maximum value pooling and activation processing are performed on the 24×24×6 (W×H×number of channels) image generated in the convolutional layer 52 to generate an image (feature map) of 12×12×6 (W×H×number of channels).

[0053] The process of max pooling extracts the pixel values within a 2×2 (W×H) window for each channel of the image in the convolutional layer 52, and sets the maximum value among the pixel values in the window as the pixel value of one pixel. Such a process is performed while sliding by two pixels for each channel of the image, generating images for six channels with a size of 12×12 (W×H). Other pooling processes such as average pooling may be performed instead of max pooling.

[0054] The activation process converts the pixel value of each pixel in the image generated by the max pooling process using an activation function such as the ReLu function.

[0055] Through these processes of max pooling and activation, in the pooling layer 53, an image with a volume size of 12×12×6 (W×H×number of channels) is generated.

[0056] The convolutional layer 54 serves as the second convolutional layer with respect to the first convolutional layer 52. In the convolutional layer 54, a convolutional operation using 16 types of convolutional filters is performed on the image of 12×12×6 (W×H×number of channels) generated by the pooling layer 53, generating images (feature maps) for 16 channels. The filter size (W×H×number of channels) of the convolutional filter is 5×5×6. The convolutional operation in the convolutional layer 54 is performed in the same manner as in the convolutional layer 52, generating an image of 8×8×16 (W×H×number of channels).

[0057] The pooling layer 55 serves as the second pooling layer with the pooling layer 53 as the first layer. In the pooling layer 55, the same max pooling and activation processes as in the pooling layer 53 are performed on the image of 8×8×16 generated by the convolutional layer 54, generating an image of 4×4×16 (W×H×number of channels).

[0058] Note that in the CNN of FIG. 7, the case where two sets of a convolutional layer and a pooling device are provided is illustrated, but three or more sets of a convolutional layer and a pooling device may be provided.

[0059] In the fully connected layer 56, the pixel values of all the pixels of the 4×4×16 (W×H×number of channels) image generated by the pooling layer 55 are input to each of the 10 nodes. For each node in the fully connected layer 56, a weight to multiply with the input value and a bias to add to the product of the input value and the weight are assigned in the same way as in a normal neural network. Each pixel value of the image of the pooling layer 55 input to each node in the fully connected layer 56 is converted by those weights and biases and then added as the output value of each node.

[0060] In the output layer 57, the output values of the 10 nodes of the fully connected layer 56 are converted into the output values of the 10 nodes of the output layer 57 by the softmax function. The softmax function converts the output value of each node in the fully connected layer 56 into a value representing the probability corresponding to the class corresponding to each node in the output layer 57.

[0061] The learning of the learning model (CNN) will be described.

[0062] For the learning of the parameters included in the CNN, such as the weights of the convolutional filters in the CNN, the weights and biases of the fully connected layer 56, a dataset (learning data) consisting of a large number of learning images (input data) with correct labels (correct outputs) is used. When the output of the CNN when the learning image (input data) of the learning data is input to the CNN is Y and the correct output of that learning image is Y gt , the parameter W included in the CNN is updated so as to satisfy the following formula (1).

[0063] W = argmin W {E(Y, Y gt ) + λ·R(W)} ···(1)

[0064] In formula (1), the function argmin W represents the value of the parameter W when minimizing the value of the expression inside the parentheses. The first term E(Y, Y W ) of the mathematical formula inside the argument of argmin gt is the error term (the output Y and the correct output Y gtis the difference), and the second term λ·R(W) represents the regularization term.

[0065] The function E(Y, Y gt ) of the error term represents a loss function that indicates the degree of difference between the output Y of the CNN and the correct output Y gt . As the loss function, well-known functions for deriving the sum of squared errors, cross-entropy errors, etc. can be used.

[0066] The regularization term is generally added to the error term to prevent overfitting by preventing the parameter W from becoming too large, and to achieve learning stabilization and accuracy improvement. As regularization, L1 regularization that uses a function R(W) for calculating the sum of the absolute values of the parameter W as the regularization term, and L2 regularization that uses a function R(W) for calculating the sum of squares of the parameter W as the regularization term are well-known. Note that λ of the regularization term is an adjustment value determined in advance by user specification or the like.

[0067] In the learning of the CNN, the parameter W is calculated so that the sum of the error term and the regularization term becomes the minimum value.

[0068] In the learning of the CNN in the present embodiment, R(W) of the regularization term is represented by the following formulas (2) to (4) based on L1 regularization.

[0069] R(W) = R1(W1) + Σ m R2(W m )(m is an integer of 2 or more) ···(2)

[0070] However, W1 represents the parameters (weights) of all the convolutional filters in the first layer of the CNN. W m represents the parameters (weights) of the convolutional filters in the m-th layer (m is an integer of 2 or more) included in the CNN. The second term on the right side of formula (2) is R2(W m ) for the parameter W m of the convolutional filter in the m-th layer, and represents the total value for all layers other than the first layer of the convolutional layers included in the CNN.

[0071] R1(W1), and R2(W m) are respectively represented by the following formulas (3) and (4).

[0072] R1(W1)=Σ n,x,y,c |w n (x,y,c)|+α·Σ n,c |Σ x,y w n (x,y,c)| ···(3)

[0073] R2(W m )=Σ n,x,y,c |w n (x,y,c)| ···(4)

[0074] However, in formula (3), n represents the number uniquely assigned to all the convolutional filters in the first layer. x and y represent the horizontal and vertical positions (x - coordinate and y - coordinate) within the convolutional filter. c represents the depth - direction position (channel number) within the convolutional filter. w n (x,y,c) represents the parameter (coefficient, i.e., weight) at the position (x,y,c) in the n - th convolutional filter. α represents an adjustment value determined in advance by user specification or the like.

[0075] The first term on the right - hand side of formula (3) represents the sum of the absolute values of all the weights of all the convolutional filters in the first layer. The second term on the right - hand side of formula (3) represents the value obtained by summing, for all channels of all convolutional filters in the first layer, the absolute value of the sum of the weights at all positions (x - coordinate and y - coordinate) in each channel of each convolutional filter in the first layer.

[0076] In formula (4), n represents the number assigned to the convolutional filter in the m - th layer (m is an integer greater than or equal to 2). The other variables x, y, c are the same as in formula (3). The first term on the right - hand side of formula (4) represents the sum of the absolute values of all the weights of the convolutional filter in the m - th layer.

[0077] Note that the regularization term R(W) may be interpreted as the sum of the absolute values of all the weights of all the convolutional filters in the first layer and all subsequent layers, plus the second term on the right - hand side of formula (3).

[0078] According to the regularization term \(R(W)\) represented by these formulas (2) to (4), the first term on the right side of formula (3) and the first term on the right side of formula (4) are well-known L1 regularization terms. The weights of each convolutional filter are learned so as not to become too large by the L1 regularization term.

[0079] The second term on the right side of formula (3) is a regularization term for making the first-layer convolutional filter zero-sum (hereinafter referred to as the zero-sum regularization term). The weights are learned so that the sum of the weights for each channel of each convolutional filter in the first layer approaches 0 (becomes 0), that is, so that the convolutional filter becomes a zero-sum convolutional filter.

[0080] The regularization term \(R(W)\) in formula (1) is taken as the value obtained by adding the zero-sum regularization term to the L1 regularization term based on L1 regularization, but it is not limited to this. The regularization term \(R(W)\) may be the value obtained by adding the zero-sum regularization term to the L2 regularization term based on L2 regularization.

[0081] When based on L2 regularization, \(R1(W1)\) and \(R2(W\) m ) in formula (2) representing the regularization term \(R(W)\) are respectively represented by the following formulas (5) and (6).

[0082] \(R1(W1)=\sum\) n,x,y,c \(\{w\) n (x,y,c)\} 2 +\(\alpha\cdot\sum\) n,c |\(\sum\) x,y w n (x,y,c)|\ ···(5)

[0083] \(R2(W\) m )=\sum\) n,x,y,c \(\{w\) n (x,y,c)\} 2 ···(6)

[0084] However, \(n\), \(x\), \(y\), \(c\), and \(\alpha\) in formula (5) are the same as those in the case of formula (3). \(n\), \(x\), \(y\), and \(c\) in formula (6) are the same as those in the case of formula (4).

[0085] The first term on the right side of Equation (5) represents the sum of the squared values of all the weights of all the convolutional filters in the first layer. The second term on the right side of Equation (5) is the zero-sum regularization term, which is the same as the second term on the right side of Equation (3).

[0086] The first term on the right side of Equation (6) represents the sum of the squared values of all the weights of the convolutional filters in the m-th layer.

[0087] The zero-sum regularization term in the second term on the right side of Equation (3) or Equation (5) is the sum of the absolute values of the sums of all the weights in each channel of each convolutional filter in the first layer, summed over all the channels of all the convolutional filters in the first layer (L1 norm), but it may also be the L2 norm. When the zero-sum regularization term is the L2 norm, the zero-sum regularization term in the second term on the right side of Equation (3) or Equation (5) is changed to the following Equation (7).

[0088] Zero-sum regularization term = α·Σ n,c {Σ x,y w n (x,y,c)} 2 ···(7)

[0089] However, n, x, y, c, and α in Equation (7) are the same as in the case of Equation (3).

[0090] The zero-sum regularization term in Equation (7) is the sum of the squared values of the sums of all the weights in each channel of each convolutional filter in the first layer, summed over all the channels of all the convolutional filters in the first layer (L2 norm).

[0091] When using the zero-sum regularization term of the L2 norm as compared to the zero-sum regularization term of the L1 norm, as the sum of all weights in each channel of each convolutional filter in the first layer approaches 0 (when the zero-sum regularization term approaches 0), the slope of the zero-sum regularization term becomes gentle. Conversely, as the sum of all weights in each channel of each convolutional filter in the first layer moves away from 0 (when the zero-sum regularization term moves away from 0), the slope of the zero-sum regularization term becomes steep. Therefore, it is possible to weaken the regularization effect near 0 of the zero-sum regularization term and strengthen the regularization effect away from 0 of the zero-sum regularization term. For the purpose of stabilizing learning and adjusting the regularization near 0 of the zero-sum regularization term, etc., the zero-sum regularization term may be used separately between the L1 norm and the L2 norm.

[0092] Figures 8 and 9 are diagrams for explaining the action of the regularization term.

[0093] In Figures 8 and 9, focus on the weights of a predetermined channel of the convolutional filter numbered 0 in the first layer. Assume that the filter size (W×H) of the focused channel is 2×2. The focused weight w n (x,y,c), omitting the channel number, is represented by w0(0,0), w0(0,1), w0(1,0), and w0(1,1).

[0094] In the graph 71 of Figure 8, the values of each weight w0(0,0), w0(0,1), w0(1,0), and w0(1,1) when learning is performed without adding the regularization term λ·R(W) to the error term in Equation (1) are represented by a bar graph.

[0095] In the graph 72 of Figure 8, when the regularization term λ·R(W) is added to the error term as in Equation (1), the values of each weight w0(0,0), w0(0,1), w0(1,0), and w0(1,1) when learning is performed without including the second term on the right side of Equation (3) (the zero-sum regularization term) in the R(W) of the regularization term represented by Equations (2) to (4) are represented by a bar graph.

[0096] When comparing Graph 71 and Graph 72, by including the regularization term λ·R(W) in the error term as in Equation (1), the absolute values of the weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) are suppressed in the direction of decreasing. However, since both positive and negative values of the weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) approach 0, there is little direct effect of making the sum of these weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) approach 0. When the λ of the regularization term formula λ·R(W) is set to a large value to strengthen the regularization effect, the sum can be made to approach 0, but all the weights will approach 0. In that case, since only the feature quantity information is lost from the target image by the convolution operation, λ cannot be set to such a large value to make the sum approach 0.

[0097] The graph 71 in FIG. 9 shows the values of the respective weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) when learning is performed under the same conditions as the graph 71 in FIG. 8, represented as a bar graph. However, FIG. 9 shows the average values of the weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1).

[0098] The graph 73 in FIG. 9 shows the values of the respective weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) when learning is performed using the R(W) of the regularization term (including the zero-sum regularization term) as represented by Equations (2) to (4) in the case where the regularization term λ·R(W) is added to the error term as in Equation (1), represented as a bar graph.

[0099] When comparing Graph 71 and Graph 73, by including the regularization term λ·R(W) in the error term as in Equation (1) and including the second term on the right side of Equation (3) (zero-sum regularization term) in R(W), the weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) are adjusted so that their sum (average value) approaches 0. That is, when the average of the weights is a positive value as in Graph 71, the weights w0(0,0), w0(0,1), w0(1,0), and w0(1,1) as in Graph 73 are shifted to negative values overall.

[0100] In this way, by using Equations (2) to (4) as the R(W) of the regularization term, each convolutional filter in the first layer is learned to be a zero-sum convolutional filter for each channel.

[0101] (Procedure of CNN learning process) FIG. 10 is a flowchart illustrating the procedure of the learning process of the learning model (CNN) executed by the recognizer 32 in FIG. 1. Note that the CNN learning process is performed by the recognizer 32. However, an arbitrary device other than the recognizer 32 (such as a computer shown in FIG. 22 described later) may perform the CNN learning process. When an apparatus other than the recognizer 32 performs the learning process, data such as the parameters of the CNN set by the learning process is supplied to the recognizer 32.

[0102] In step S11 of FIG. 10, the recognizer 32 reads learning data. The process proceeds from step S11 to step S12.

[0103] In step S12, the recognizer 32 inputs the target image (input data) of the learning data read in step S11 into the CNN and calculates the output of the CNN. Note that the initial values of various parameters of the CNN set by learning are set to random values, for example. The recognizer 32 calculates the error term of the above Equation (1) based on the calculated output Y of the CNN and the correct output Y for the input target image. The process proceeds from step S12 to step S13. gt Based on this, the error term of the above Equation (1) is calculated. The process proceeds from step S12 to step S13.

[0104] In step S13, the recognizer 32 calculates the regularization term λ·R(W) of the above formula (1) based on the currently set parameter W. The R(W) of the regularization term is calculated based on the above formulas (2) to (4). The adjustment value λ is set to a predetermined value or a value specified by the user. The process proceeds from step S13 to step S14.

[0105] In step S14, the recognizer 32 updates the parameter W of the CNN so as to minimize the sum of the error term calculated in step S12 and the regularization term calculated in step S13. The process proceeds from step S14 to step S15.

[0106] In step S15, the recognizer 32 determines whether the reading of all the learning data has been completed. Note that there are, for example, tens of thousands of pieces of learning data, and the processing from step S11 to step S15 is repeated the number of times of the learning data.

[0107] In step S15, if it is determined that the reading of all the learning data has been completed, the process returns to step S11 and repeats from step S11.

[0108] In step S15, if it is determined that the reading of all the learning data has not been completed, the process ends.

[0109] Note that the recognizer 32 may repeat the processing of this flowchart a predetermined number of times using the same learning data (data set).

[0110] FIG. 11 is a flowchart illustrating the calculation process of the regularization term in step S13 of FIG. 10.

[0111] In step S31, the recognizer 32 calculates the sum of the absolute values of the coefficients (weights) of all the convolutional filters of the CNN. The calculated sum corresponds to the sum of the value of the first term on the right side of formula (3) and the first term on the right side of formula (4). The process proceeds from step S31 to step S32.

[0112] In step S32, the recognizer 32 calculates the absolute value of the sum of the coefficients (weights) of the convolutional filter (sum for each channel) for all channels of all the convolutional filters in the first layer. The value obtained by summing up these calculated absolute values corresponds to the second term on the right side of Equation (4). The process proceeds from step S32 to step S33.

[0113] In step S33, the recognizer 32 adds together the sum calculated in step S31 and the absolute value calculated in step S32. Thereby, the regularization term R(W) is calculated, and by multiplying it by a predetermined adjustment value λ, the regularization term λ·R(W) is calculated.

[0114] In the CNN learning process of FIGS. 10 and 11 above, the coefficients (weights) of the convolutional filter are updated for each piece of learning data, but it is not limited to this. For example, the error term E(Y, Y gt ) in Equation (1) is taken as the sum (or average value) of the error terms (Y, Y gt ) for a plurality of learning data, and the coefficients of the convolutional filter may be updated for each plurality of learning data.

[0115] In the above CNN learning process, learning of the coefficients (weights) in the zero-sum convolutional filter is performed by the second term on the right side of Equation (3) or Equation (5), or the zero-sum regularization term of Equation (7), but it is not limited to this.

[0116] As another example 1, in the CNN learning process, every time the weights of the convolutional filter (parameters of the CNN) are updated, in conjunction with this, each weight may be updated so that the sum of the weights in each channel of the zero-sum convolutional filter becomes 0.

[0117] As another example 2, in the CNN learning process, every time the weights of the convolutional filter (parameters of the CNN) are updated, in conjunction with this, the average of the weights in each channel of the zero-sum convolutional filter may be subtracted from each weight so that the sum of the weights in each channel becomes 0.

[0118] The update of the convolutional filter is performed using the regularization term λ·R(W) including the zero-sum regularization term, and other Examples 1 and 2 may also be performed together.

[0119] (Procedure of the image recognition process of the image recognition apparatus 11) FIG. 12 is a flowchart illustrating the processing procedure of the image recognition process performed by the image recognition apparatus 11.

[0120] In step S51, the preprocessing unit 31 performs log conversion on the input target image (input data). The process proceeds from step S51 to step S52.

[0121] In step S52, the recognizer 32 performs CNN processing on the target image log-converted in step S51. The process proceeds from step S52 to step S53.

[0122] In step S53, the recognizer 32 outputs the recognition result by the CNN processing in step S52.

[0123] The processes of steps S51 to S53 above are executed each time a target image is input to the image recognition apparatus 11.

[0124] FIG. 13 is a flowchart illustrating the processing procedure of the CNN processing in step S52 of FIG. 12.

[0125] In step S71, the recognizer 32 (the first CNN processing unit 41) performs a convolution operation on the target image from the preprocessing unit 31 using the zero-sum convolutional filter of the first layer of the CNN (the convolutional filter whose weights are set by zero-sum regularization), and calculates the output of the first layer of the CNN. The process proceeds from step S71 to step S72.

[0126] In step S72, the recognizer 32 (the second CNN processing unit 42) calculates the output of the CNN at a stage later than the first layer with respect to the CNN output calculated in step S71.

[0127] According to the above image recognition device 11, in the CNN of the learning model used for image recognition, by setting the first-layer convolutional filter as a zero-sum convolutional filter, even when the brightness is different, image recognition can be performed on target images of the same subject with equivalent recognition accuracy. That is, image recognition with reduced influence of the brightness of the target image can be performed. Regarding what kind of images are actually generated by the zero-sum convolutional filter, examples are shown in FIGS. 19 to 21.

[0128] According to the image recognition device 11, correction of the signal level becomes unnecessary for target images from an imaging sensor including pixels with different spectral characteristics. Generally, the outputs of pixels with different filters have different signal levels. As a specific example, the pixels of G have higher sensitivity and higher signal levels than the pixels of R and B. Therefore, for the outputs of pixels with different filters, gain correction is performed to correct the signal level.

[0129] In contrast, in the image recognition device 11, since the influence of the gain is removed, correction of the signal level becomes unnecessary.

[0130] Image recognition can be performed with high accuracy without performing signal level correction even for images from an imaging sensor to which clear pixels with increased signal levels such as RGBW are added.

[0131] According to the image recognition device 11, even when the spectral characteristics approach flatness due to aging changes, the influence is reduced by the first-layer zero-sum convolutional filter.

[0132] (Modification Example 1 of Recognizer 32) In the recognizer 32 of FIG. 1, all of the first-layer convolutional filters were zero-sum convolutional filters, but at least one or more of the first-layer convolutional filters may be zero-sum convolutional filters. When one convolutional filter includes a plurality of channels (3 channels in the example of FIG. 7), some (at least one or more channels) of the plurality of channels may be zero-sum convolutional filters. In this case, in the learning of the CNN used by the recognizer 32, in the zero-sum regularization term in the second term on the right side of Equation (3) or Equation (5), or in the zero-sum regularization term of Equation (7), values related to weights that are non-zero-sum rather than zero-sum convolutional filters are excluded. When the learning model is a neural network having only one convolutional layer and a plurality of convolutional filters, at least one or more of the plurality of convolutional filters may be zero-sum convolutional filters.

[0133] When a part of the first-layer convolutional filter is a non-zero-sum convolutional filter, the DC component included in the target image is transmitted to the subsequent stage without being completely lost. If it is better to use the signal level itself of the DC component for recognition, the image recognition accuracy may be improved by making a part of the first-layer convolutional filter a non-zero-sum convolutional filter.

[0134] For example, by inputting a temperature image from an infrared camera as the target image of the image recognition device 11 into the image recognition device 11, a person can be detected from the spatial temperature change. In this case, by using not only zero-sum convolutional filters but also non-zero-sum convolutional filters in the first layer of the CNN used by the recognizer 32, the temperature can be directly detected. By detecting the temperature, even if the silhouette of an object that is not a person has a human shape, if the temperature is not in the human temperature range (about 20 to 40 degrees), misrecognition as a person can be avoided. By using a non-zero-sum convolutional filter for a part of the first-layer convolutional filter to transmit the DC component of the target image to the subsequent stage, it is possible to train the CNN so that such image recognition is performed.

[0135] (Modification Example 2 of Recognizer 32) In the recognizer 32 of FIG. 1, it was assumed that only the first-layer convolutional filter was a zero-sum convolutional filter, but a part of the second-layer convolutional filter may also be a zero-sum convolutional filter. That is, a zero-sum convolutional filter may be used in one or more channels of one or more convolutional filters among the second-layer convolutional filters. By adding a zero-sum convolutional filter to the second layer in this way, the influence of noise is reduced.

[0136] FIGS. 14 to 17 are diagrams for explaining the operation of the CNN when there is a zero-sum convolutional filter in the second layer.

[0137] In the graphs 91 and 92 of FIG. 14, the image signals H1 and H2 represent the image signals of the target image input to the recognizer 32. The image signal H1 in graph 91 represents the case where there is no noise, and the image signal H2 in graph 92 represents the case where there is noise with respect to the image signal H1.

[0138] In the graphs 93 and 94 of FIG. 15, the image signals I1 and I2 represent the response signals (first-layer output) when convolution operations are performed on the image signals H1 and H2, respectively, using a zero-sum convolutional filter (-1, 2, -1). According to this, an offset occurs in the image signal H2 with respect to the image signal H1 due to the influence of noise.

[0139] In the graphs 95 and 96 of FIG. 16, the image signals J1 and J2 represent the response signals (second-layer output) when convolution operations are performed on the image signals I1 and I2, respectively, using a non-zero-sum convolutional filter (-1, 3, -1). According to this, the offset generated in the image signal I2 due to the influence of noise remains in the image signal J2, and an offset occurs with respect to the image signal J1.

[0140] In the graphs 97 and 98 of FIG. 17, the image signals K1 and K2 represent the response signals (output of the second layer) when convolution operations are performed on the image signals I1 and I2 respectively using a zero-sum convolution filter (-1, 2, -1). According to this, in the image signal K2, the influence of the offset generated in the image signal I2 due to the influence of noise is reduced, and the image signal K2 becomes a signal substantially similar to the image signal K1. Therefore, by using a zero-sum convolution filter in the second layer, the influence of noise included in the target image is reduced. When there are three or more convolution filters in the CNN, a zero-sum convolution filter may be used in one or more channels of one or more of the convolution filters after the second layer. In the learning of the CNN used by the recognizer 32, for the channels of the zero-sum convolution filter for the layers after the second layer, a zero-sum regularization term in the second term on the right side of Equation (3) or Equation (5), or a zero-sum term similar to the zero-sum regularization term of Equation (7) is included in the regularization term R(W). When the learning model has a plurality of convolution filters, it is also conceivable to adopt a zero-sum convolution filter for one or more channels of any one or more of the convolution filters.

[0141] (Another configuration example of the image recognition device 11) FIG. 18 is a block diagram showing another configuration example of the image recognition device 11. In the figure, the parts corresponding to the image recognition device 11 in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted.

[0142] In the image recognition device 11 of FIG. 18, the preprocessing unit 31 in the image recognition device 11 of FIG. 1 and the first CNN processing unit 41 of the recognizer 32 are incorporated in the stacked sensor 102. The stacked sensor 102 is a sensor in which a signal processing circuit is stacked on an imaging sensor such as a CMOS sensor. The second CNN processing unit 42 of the recognizer 32 in FIG. 1 is incorporated in an external sensor device 103 with respect to the stacked sensor 102. Therefore, the first CNN processing unit 41 and the second CNN processing unit 42 constituting the recognizer 32 are separated and arranged in the stacked sensor 102 and the external sensor device 10.

[0143] The image captured by the imaging sensor of the stacked sensor 102 is supplied as a target image to the preprocessing unit 31 within the stacked sensor 102. The target image supplied to the preprocessing unit 31 is log-transformed in the preprocessing unit 31 and then supplied to the first CNN processing unit 41.

[0144] The target image supplied to the first CNN processing unit 41 undergoes a convolution operation in the first CNN processing unit 41 using a zero-sum convolution filter in the first layer. As a result, the generated image (feature map) is transmitted to the second CNN processing unit 42 of the external sensor device 103. In the second CNN processing unit 42, CNN processing subsequent to the first CNN processing unit 41 is executed based on the image from the first CNN processing unit 41.

[0145] However, the sensor in which the preprocessing unit 31 and the first CNN processing unit 41 are incorporated does not necessarily have to be the stacked sensor 102, and it may also be the case where it is an imaging sensor or is incorporated on the same chip as the imaging sensor.

[0146] According to the configuration of the image recognition device 11 in FIG. 18, since the DC component of the target image is cut by the zero-sum convolution filter, unnecessary bit allocation within the imaging sensor may be eliminated and the bit length may be reduced. In particular, when the imaging sensor for capturing the target image is an HDR (High Dynamic Range) sensor, since the high-bit-length sensor output is compressed, bandwidth reduction and power saving can be achieved.

[0147] (Measured results of the convolution operation using the zero-sum convolution filter) FIGS. 19 to 21 are diagrams for explaining the results of the first CNN processing unit 41 in the recognizer 32 in FIGS. 1 and 19 performing a convolution operation using a zero-sum convolution filter.

[0148] The target images 111 and 112 in FIG. 19 are a bright image when the same subject (scene) is photographed in a bright shooting environment and a dark image when photographed in a dark shooting environment, respectively.

[0149] Images 113 and 114 in FIG. 20 represent images obtained when convolution operations are performed on target images 111 and 112 in FIG. 19 using non-zero-sum convolution filters, respectively. In image 114 obtained for dark target image 112, subject information is almost lost.

[0150] Images 115 and 116 in FIG. 21 represent images obtained when convolution operations are performed on target images 111 and 112 in FIG. 19 using zero-sum convolution filters, respectively. Even in image 116 obtained for dark target image 112, information such as the outline of the subject is extracted in the same way as in image 115 obtained for bright target image 111. As a result, in recognizer 32, appropriate image recognition can be performed without the image recognition accuracy being affected by the brightness of the target image.

[0151] <Program> Some or all of the series of processes in the above-described preprocessing unit 31 and recognizer 32, and some or all of the series of processes of the learning process of the learning model executed by recognizer 32 can be executed by hardware or by software. When a series of processes are executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, and, for example, a general-purpose personal computer that can execute various functions by installing various programs.

[0152] FIG. 22 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.

[0153] In a computer, a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, and a RAM (Random Access Memory) 203 are interconnected by a bus 204.

[0154] The bus 204 is further connected to an input / output interface 205. The input / output interface 205 is connected to an input unit 206, an output unit 207, a storage unit 208, a communication unit 209, and a drive 210.

[0155] The input unit 206 includes a keyboard, a mouse, a microphone, etc. The output unit 207 includes a display, a speaker, etc. The storage unit 208 includes a hard disk, a non-volatile memory, etc. The communication unit 209 includes a network interface, etc. The drive 210 drives a removable medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0156] In the computer configured as described above, the CPU 201 loads and executes, for example, a program stored in the storage unit 208 via the input / output interface 205 and the bus 204 into the RAM 203, thereby performing the above-described series of processes.

[0157] The program executed by the computer (CPU 201) can be recorded and provided, for example, on a removable medium 211 as a package medium or the like. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0158] In the computer, the program can be installed in the storage unit 208 via the input / output interface 205 by mounting the removable medium 211 on the drive 210. Further, the program can be received by the communication unit 209 via a wired or wireless transmission medium and installed in the storage unit 208. Additionally, the program can be installed in advance in the ROM 202 or the storage unit 208.

[0159] Note that the program executed by the computer may be a program in which processing is performed in time series in the order described in this specification, or a program in which processing is performed in parallel or at a necessary timing such as when a call is made.

[0160] The present technology can also adopt the following configurations. (1) A neural network applied to a learning model that performs recognition processing on input data, wherein learning of the learning model is performed such that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters approaches zero. A method for generating a learning model. (2) Learning of the learning model is performed such that the sum of coefficients approaches zero for each of all the channels of at least one of the convolution filters in the first layer of the convolution filters. The method for generating a learning model according to (1) above. (3) Learning of the learning model is performed such that the sum of an error term based on the difference between the output of the learning model and the correct output when input data with a correct output is input to the learning model and a regularization term based on the coefficients of the convolution filters included in the neural network is minimized. The method for generating a learning model according to (1) or (2) above. (4) The regularization term includes a value corresponding to the absolute value of the sum of coefficients in the channel of the convolution filter that makes the sum of coefficients approach zero. The method for generating a learning model according to any one of (1) to (3) above. (5) The regularization term includes a value proportional to the value obtained by squaring the sum of coefficients in the channel of the convolution filter that makes the sum of coefficients approach zero. The method for generating a learning model according to any one of (1) to (3) above. (6) The regularization term includes a value corresponding to the sum of the absolute values of all the coefficients in all the convolutional filters included in the neural network. The method for generating a learning model according to any one of (3) to (5) above. (7) The regularization term includes a value corresponding to the sum of the squared values of all the coefficients in all the convolutional filters included in the neural network. The method for generating a learning model according to any one of (3) to (5) above. (8) Training the learning model so that the sum of the coefficients in at least one or more channels of the convolutional filters in the second layer of the neural network approaches zero. The method for generating a learning model according to any one of (1) to (7) above. (9) When updating the coefficients in the channels of the convolutional filters that bring the sum of the coefficients closer to zero by the minimization, setting the sum of the coefficients in the channels of the convolutional filters that bring the sum of the coefficients closer to zero to zero. The method for generating a learning model according to any one of (3) to (8) above. (10) When updating the coefficients in the channels of the convolutional filters that bring the sum of the coefficients closer to zero by the minimization, subtracting the average of the coefficients in the channels of the convolutional filters that bring the sum of the coefficients closer to zero from the coefficients. The method for generating a learning model according to any one of (3) to (9) above. (11) The input data is image data. The method for generating a learning model according to any one of (1) to (10) above. (12) A computer, A neural network applied to a learning model that performs recognition processing on input data, the processing unit for performing learning of the learning model so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters approaches zero A program for causing it to function as. (13) A neural network applied to a learning model that performs recognition processing on input data, the processing unit for executing the operation of the learning model learned so that the sum of coefficients in one or more channels of at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters approaches zero An information processing apparatus having. (14) Having, in front of the processing unit, a preprocessing unit that converts the input data by a predetermined function The information processing apparatus according to (13) above. (15) The preprocessing unit converts the input data by a log function The information processing apparatus according to (14) above. (16) The processing unit converts the input data by a piecewise linear function The information processing apparatus according to (14) above. (17) The processing unit converts the input data by a gamma curve The information processing apparatus according to (14) above. (18) The processing unit executes the operation of the learning model learned so that the sum of coefficients in one or more channels of at least one of the convolution filters in the second layer of the neural network approaches zero The information processing apparatus according to any one of (13) to (17) above.

Explanation of Signs

[0161] 11 Image recognition device, 31 Preprocessing unit, 32 Recognizer, 41 First CNN processing unit, 42 Second CNN processing unit

Claims

**Claim 1** A neural network applied to a learning model that performs recognition processing on input data, wherein, for at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, the sum of the coefficients in one or more channels of the convolution filter approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and the regularization term based on the coefficients of the convolution filters included in the neural network, the learning of the learning model is performed so that the sum is minimized. The regularization term includes a value corresponding to the absolute value of the sum of the coefficients in the channels of the convolution filter that makes the sum of the coefficients approach zero. A method for generating a learning model. **Claim 2** A neural network applied to a learning model that performs recognition processing on input data, wherein, for at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, the sum of the coefficients in one or more channels of the convolution filter approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and the regularization term based on the coefficients of the convolution filters included in the neural network, the learning of the learning model is performed so that the sum is minimized. The regularization term includes a value proportional to the value obtained by squaring the sum of the coefficients in the channels of the convolution filter that makes the sum of the coefficients approach zero. A method for generating a learning model. **Claim 3** A neural network applied to a learning model that performs recognition processing on input data, wherein, for at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters, the sum of the coefficients in one or more channels of the convolution filter approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and the regularization term based on the coefficients of the convolution filters included in the neural network, the learning of the learning model is performed so that the sum is minimized. When updating the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero by the minimization, set the sum of the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero to zero. Method for generating a learning model. **Claim 4**: A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolutional filters in the first layer of the neural network having a plurality of convolutional filters is such that the sum of the coefficients in one or more channels approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting input data with correct output into the learning model, and the regularization term based on the coefficients of the convolutional filters included in the neural network, the learning of the learning model is performed so that the sum is minimized. When updating the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero by the minimization, subtract the average of the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero from the coefficients. Method for generating a learning model. **Claim 5** Perform learning of the learning model so that the sum of the coefficients approaches zero for each of all the channels of at least one of the convolutional filters in the first layer of the convolutional filter. The method for generating a learning model according to any one of claims 1 to 4. **Claim 6** The regularization term includes a value corresponding to the sum of the absolute values of all the coefficients in all the convolutional filters included in the neural network. The method for generating a learning model according to any one of claims 1 to 4. **Claim 7** The regularization term includes a value corresponding to the sum of the squared values of all the coefficients in all the convolutional filters included in the neural network. The method for generating a learning model according to any one of claims 1 to 4. **Claim 8** Perform learning of the learning model so that the sum of the coefficients in one or more channels of at least one of the convolutional filters in the second layer of the neural network approaches zero. The method for generating a learning model according to any one of claims 1 to 4. **Claim 9** The input data is image data. The method for generating a learning model according to any one of claims 1 to 4.

10. A computer, A neural network applied to a learning model that performs recognition processing on input data, among the first-layer convolutional filters of the neural network having a plurality of convolutional filters, at least one or more of the coefficients in one or more channels of the convolutional filters approach zero, and based on the error term based on the difference between the output of the learning model and the correct output when input data with correct output is input to the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network, the learning of the learning model is performed so that the sum is minimized, and a processing unit including a value corresponding to the absolute value of the sum of the coefficients in the channel of the convolutional filter that makes the sum of the coefficients approach zero in the regularization term A program for functioning as.

11. A computer, A neural network applied to a learning model that performs recognition processing on input data, among the first-layer convolutional filters of the neural network having a plurality of convolutional filters, at least one or more of the coefficients in one or more channels of the convolutional filters approach zero, and based on the error term based on the difference between the output of the learning model and the correct output when input data with correct output is input to the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network, the learning of the learning model is performed so that the sum is minimized, and a processing unit including a value proportional to the value obtained by squaring the sum of the coefficients in the channel of the convolutional filter that makes the sum of the coefficients approach zero in the regularization term A program for functioning as.

12. A computer, A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters is such that the sum of the coefficients in one or more channels approaches zero, and an error term based on the difference between the output of the learning model and the correct output when input data with a correct output is input to the learning model, and a regularization term based on the coefficients of the convolution filters included in the neural network are minimized, and when updating the coefficients in the channels of the convolution filter that bring the sum of the coefficients closer to zero by the minimization, a processing unit that sets the sum of the coefficients in the channels of the convolution filter that bring the sum of the coefficients closer to zero to zero A program for causing the same to function **Claim 13** A computer A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolution filters in the first layer of the neural network having a plurality of convolution filters is such that the sum of the coefficients in one or more channels approaches zero, and an error term based on the difference between the output of the learning model and the correct output when input data with a correct output is input to the learning model, and a regularization term based on the coefficients of the convolution filters included in the neural network are minimized, and when updating the coefficients in the channels of the convolution filter that bring the sum of the coefficients closer to zero by the minimization, a processing unit that subtracts the average of the coefficients in the channels of the convolution filter that bring the sum of the coefficients closer to zero from the coefficients A program for causing the same to function **Claim 14** A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolutional filters in the first layer of the neural network having a plurality of convolutional filters is such that the sum of the coefficients in one or more channels approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network, a processing unit that executes the operation of the learning model learned so that the sum is minimized having including in the regularization term a value corresponding to the absolute value of the sum of the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero An information processing apparatus.

15. A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolutional filters in the first layer of the neural network having a plurality of convolutional filters is such that the sum of the coefficients in one or more channels approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network, a processing unit that executes the operation of the learning model learned so that the sum is minimized having including in the regularization term a value proportional to the value obtained by squaring the sum of the coefficients in the channels of the convolutional filter that brings the sum of the coefficients closer to zero An information processing apparatus.

16. A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolutional filters in the first layer of the neural network having a plurality of convolutional filters is such that the sum of the coefficients in one or more channels approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting the input data with the correct output into the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network, a processing unit that executes the operation of the learning model learned so that the sum is minimized having In the learning of the learning model, when the coefficient in the channel of the convolutional filter that brings the sum of the coefficients closer to zero by the minimization is updated, the sum of the coefficients in the channel of the convolutional filter that brings the sum of the coefficients closer to zero is set to zero Information processing apparatus.

17. A neural network applied to a learning model that performs recognition processing on input data, wherein at least one of the convolutional filters in the first layer of the neural network having a plurality of convolutional filters is such that the sum of the coefficients in one or more channels approaches zero, and based on the error term based on the difference between the output of the learning model and the correct output when inputting input data with correct output to the learning model, and a regularization term based on the coefficients of the convolutional filters included in the neural network A processing unit that executes the operation of the learning model learned so that the sum of the two is minimized having In the learning of the learning model, when the coefficient in the channel of the convolutional filter that brings the sum of the coefficients closer to zero by the minimization is updated, the average of the coefficients in the channel of the convolutional filter that brings the sum of the coefficients closer to zero is subtracted from the coefficients Information processing apparatus.

Citation Information

Patent Citations

  • Gesture recognition method based on deep learning

    CN108537147A

  • Capsule endoscopy image recognition model based on neural network feature fusion

    CN110705440A

  • Video high dynamic range inverse tone mapping model construction and mapping method and device

    CN110717868A

  • Convolutional neural network device and its manufacturing method

    JP2019087021A

  • Method and System for Approximating Deep Neural Networks for Anatomical Object Detection

    US20160328643A1