Image processing apparatus, image processing method, and program

By associating feature information from non-distorted and distorted images to match high-contribution neuron firing states, the method addresses error propagation issues in CNNs, enhancing accuracy on distorted images.

JP7712638B2Active Publication Date: 2025-07-24NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024015227
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-07-24
Estimated Expiration
2040-05-20

AI Technical Summary

Technical Problem

Existing neural networks, particularly CNNs, struggle with maintaining accuracy in the presence of image distortions due to insufficient error propagation during fine-tuning, leading to neurons losing responsiveness to original image features.

Method used

An image processing method that associates feature information from both non-distorted and distorted images to match the firing states of high-contribution neurons, using intermediate similarity loss to enhance learning and improve robustness.

Benefits of technology

Enhances the accuracy of image processing on distorted images by ensuring error information propagates effectively, allowing neurons to maintain responsiveness to original features, thereby improving overall image processing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712638000001
    Figure 0007712638000001
  • Figure 0007712638000002
    Figure 0007712638000002
  • Figure 0007712638000003
    Figure 0007712638000003
Patent Text Reader

Abstract

To improve the accuracy of image processing of CNN on a distorted image.SOLUTION: An image processing apparatus includes: a storage unit which stores image processing information obtained by associating characteristic information that characterizes a non-distorted image with characteristic information that characterizes a distorted image; and an image processing unit which acquires a target image, which is an image to be processed, and performs predetermined image processing on the acquired target image using the image processing information, to obtain an image processing result. The non-distorted image is an image obtained by performing first processing that receives a non-distorted image to identify a subject in the image. The distorted image is an image obtained by performing second processing that receives a distorted image to identify the subject in the image. The image processing unit performs processing, as the predetermined image processing, to identify the subject in the image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] In this description, a subscript in an actual mathematical formula is described using an underscore "_". For example, when indicating a string in which "n" is added as a subscript to the character "X", it is described as "X_n" using full-width or half-width characters.

[0003] In recent years, the accuracy improvement of image processing technology using machine learning technology has been developing. Specific examples of image processing include a process of identifying a subject in an image (hereinafter referred to as "image identification process"), a process of detecting a subject in an image, a process of dividing an area of an image based on a specific criterion, and the like. Among such machine learning technologies for performing such processes, the accuracy improvement of a convolutional neural network (CNN) is particularly remarkable. Technologies for automating visual processes in various operations by using these machine learning technologies have attracted particular attention.

[0004] When promoting the automation of the visual process as described above, it is assumed that the quality of the captured image is not necessarily high. For example, an image with reduced contrast or an image with camera noise may be obtained by shooting in a dark place. Also, when the shooting target moves at high speed, blur may occur. Furthermore, when the captured image is non-reversibly compressed, compression distortion may also occur. In the image identification accuracy using CNN, it is known that noise and blur in particular among image distortions cause a significant decrease in accuracy (see Non-Patent Document 1). Furthermore, in prior art documents that investigated the robustness of humans and CNN against distortion, it is known that the robustness of CNN is significantly inferior to that of humans (see Non-Patent Document 2). Therefore, in the automation of the human visual process by CNN, there is a risk that behaviors that humans do not expect may be shown.

[0005] Regarding such problems of CNN, methods for realizing distortion-robust image processing have been proposed. As one of such methods, a method based on fine-tuning has been proposed (Non-Patent Documents 3 and 4). Fine-tuning is a technique of re-training CNN with a dataset composed of distorted images, using the weights of a CNN that has been trained with a standard dataset containing almost no distortion as the initial values. By performing such processing, it is possible to enhance the robustness against distortion. It is expected that the fine-tuned CNN will have higher robustness against distortion compared to the CNN before fine-tuning.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

[0007] However, with regard to fine-tuning, sufficient robustness might not be obtained for the following reasons. In the fine-tuning-based method, learning is performed by the error backpropagation method used in the learning of a normal CNN. FIG. 11 is a diagram showing the process of recognition by the forward propagation of a general CNN and the corresponding learning by the backpropagation. In the non-linear transformation here, the processing by ReLU generally used in a CNN is assumed. 1(X_(n + 1)>0) in the backpropagation is a matrix that has 1 for the components of the matrix X_(n + 1) having values larger than 0 and 0 for the other components. Therefore, the error is propagated only for the components that did not become 0 in the ReLU processing during the forward propagation. FIG. 12 is a schematic diagram simply showing the propagation of the error shown in FIG. 11. Note that not only in the ReLU processing but also in the max-pooling processing, error information is not propagated from the unselected components (Non-Patent Document 5). That is, there is a problem that error information passing through neurons that selectively respond to specific image features cannot be propagated to the lower layers unless the neurons show positive outputs.

[0008] Thus, when considering fine-tuning for an image containing distortion, due to the distortion, the features of the original image (such as texture and curvature) that the CNN has been trained to respond to may disappear. In that case, the neurons that selectively responded to such image features will stop responding. As described above, since the error information is no longer propagated, backpropagation learning may become insufficient. Such a problem is not limited to CNNs but is a common problem for all neural networks, such as neural networks where convolution is performed.

[0009] In view of the above circumstances, an object of the present invention is to provide a technique for improving the accuracy of image processing of a neural network for an image in which distortion has occurred.

Means for Solving the Problems

[0010] One aspect of the present invention includes an acquisition step of acquiring a target image that is an image to be processed, and an image processing step of obtaining an image processing result by performing predetermined image processing on the acquired target image. In the image processing, image processing information obtained by associating feature information characterizing an image without distortion and feature information characterizing an image with distortion is used.

[0011] One aspect of the present invention is an image processing apparatus including a storage unit that stores image processing information obtained by associating feature information characterizing an image without distortion and feature information characterizing an image with distortion, and an image processing unit that acquires a target image that is an image to be processed and obtains an image processing result by performing predetermined image processing on the acquired target image using the image processing information. One aspect of the present invention includes a storage unit that stores image processing information obtained by associating feature information characterizing an image of a non-distorted image with feature information characterizing an image of a distorted image, and an image processing unit that acquires a target image that is an image to be processed and performs predetermined image processing on the acquired target image using the image processing information to obtain an image processing result. The non-distorted image is an image obtained by performing a first process of identifying a subject in the image with the non-distorted image as an input, and the distorted image is an image obtained by performing a second process of identifying a subject in the image with the distorted image as an input. The image processing unit performs a process of identifying a subject in the image as the predetermined image processing, and is an image processing apparatus. One aspect of the present invention includes an acquisition step of acquiring a target image that is an image to be processed, and an image processing step of obtaining an image processing result by performing predetermined image processing on the acquired target image. In the image processing, image processing information obtained by associating feature information characterizing an image of a non-distorted image obtained by performing a first process of identifying a subject in the image with the non-distorted image as an input, and feature information characterizing an image of a distorted image obtained by performing a second process of identifying a subject in the image with the distorted image as an input is used, and is an image processing method.

[0012] One aspect of the present invention is a program for causing a computer to execute the above image processing method.

Effects of the Invention

[0013] According to the present invention, it is possible to improve the accuracy of image processing of a neural network for an image in which distortion has occurred.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0015] First, the outline of this embodiment will be described. In this embodiment, among the features of an image without distortion, the firing of neurons of features with a high contribution degree (hereinafter referred to as "high contribution features") for characterizing the image is matched. The high contribution features are, for example, large textures or features indicating boundaries. If the firing states of neurons corresponding to such high contribution features can be made to match between a distorted image and an undistorted image, it is considered that the error in the result of image processing performed on each image can be reduced. Therefore, in the present invention, learning is performed so that the firing states of neurons corresponding to high contribution features match between a distorted image and an undistorted image. In the CNN described below, neurons corresponding to highly contributing features are often higher-order neurons. Therefore, in one embodiment, when using a CNN, the firing states of higher-order neurons are matched. When using a neural network other than a CNN, instead of matching the firing states of higher-order neurons, the firing states of neurons corresponding to highly contributing features should be matched.

[0016] [First Embodiment] FIG. 1 is a diagram showing a functional configuration example of the first embodiment of the present invention. The learning device 10 is configured using an information processing device such as a personal computer or a server device. The learning device 10 includes a learning image storage unit 11, a parameter storage unit 12, and a control unit 13. The learning device 10 executes learning processing using a neural network. As a specific example of the neural network, for example, a neural network in which convolution processing is performed may be applied. A more specific example is a CNN. In the following description, an embodiment in the case where a CNN is applied as a specific example of the neural network will be described.

[0017] The learning image storage unit 11 is configured using a storage device such as a magnetic hard disk device or a semiconductor storage device. The learning image storage unit 11 stores learning image data. The learning image data is so-called teacher data used for supervised learning. The learning image data includes, for example, a combination of image data used for learning and a correct label indicating the correct answer in the image data. It is desirable that the image data included in the learning image is an image without distortion.

[0018] What kind of information the correct label is depends on what kind of learning processing is performed. For example, in the case of learning processing for determining the type of subject shown in the image data, a correct label indicating the correct answer of the type of subject shown in the image data is given. In such a case, the correct label may be given, for example, as a vector sequence indicating the type of subject.

[0019] For example, in the case of a learning process for detecting a predetermined object from among the subjects depicted in the image data, a correct label indicating the correct position of a specific object depicted in the image data is given. In such a case, the correct label may be given, for example, as an array indicating the position of the specific object.

[0020] For example, in the case of a learning process for performing region division of image data, a correct label indicating which region each pixel of the image data belongs to is given. In such a case, the correct label may be given, for example, as an array indicating the region to which each pixel belongs.

[0021] The parameter storage unit 12 is configured using a storage device such as a magnetic hard disk device or a semiconductor storage device. The parameter storage unit 12 stores image processing parameters indicating a learned model obtained by a learning process based on a CNN executed in advance. The learned model related to the image processing parameters stored in the parameter storage unit 12 is obtained by a learning process of the same type as the learning process performed using the teacher data stored in the learning image storage unit 11. For example, when the learned model stored in the parameter storage unit 12 is a learned model for determining the type of the subject depicted in the image data, the teacher data stored in the learning image storage unit 11 is the teacher data used in the learning process for determining the type of the subject depicted in the image data.

[0022] The control unit 13 is composed of a processor such as a CPU (Central Processing Unit) and a memory. By the processor executing a learning program, the control unit 13 functions as a distortion image generation unit 131, an original image processing unit 132, a distortion image processing unit 133, an intermediate similarity loss calculation unit 134, and an optimization unit 135. Note that all or part of each function of the control unit 13 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The above learning program may be recorded on a computer-readable recording medium. A computer-readable recording medium is, for example, a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, a semiconductor memory device (e.g., SSD: Solid State Drive), or a storage device such as a hard disk or a semiconductor memory device built into a computer system. The above learning program may be transmitted via a telecommunication line.

[0023] First, an outline of the processing of each functional unit will be described. The distortion image generation unit 131 acquires learning image data from the learning image storage unit 11 and generates a distortion image by performing distortion image processing. Distortion image processing is a process of applying some kind of distortion to an image. The distortion image processing may be, for example, a process of applying blur, a process of applying noise, or a process of applying compression distortion performed in the process of generating a JPEG image.

[0024] The original image processing unit 132 acquires learning image data from the learning image storage unit 11. Also, the original image processing unit 132 acquires image processing parameters from the parameter storage unit 12. The original image processing unit 132 performs image processing on the acquired learning image data according to the image processing parameters. The image processing performed by the original image processing unit 132 is an inference process performed according to the image processing parameters (trained model). The original image processing unit 132 acquires the intermediate processing result of the inference process according to the image processing parameters. The intermediate processing result is, for example, the output of any intermediate layer in the case of an NN performing image processing. In the case of a CNN, in many cases, the response to some of the features constituting the image as features can be obtained as the output of the intermediate layer. For example, the intermediate processing result may indicate the output in any hidden layer (the firing state of higher-order neurons) instead of the output in the output layer.

[0025] The distorted image processing unit 133 acquires learning image data from the learning image storage unit 11. At this time, the distorted image processing unit 133 may acquire only the correct label without acquiring the image data. Also, the distorted image processing unit 133 acquires image processing parameters from the parameter storage unit 12. The distorted image processing unit 133 performs image processing on the distorted image acquired by the distorted image generation unit 131. The image processing performed by the distorted image processing unit 133 is the same as the processing of the original image processing unit 132 except that the image to be processed is different. That is, the image to be processed by the original image processing unit 132 is the image stored in the learning image storage unit 11, while the image to be processed by the distorted image processing unit 133 is the distorted image generated by performing distorted image processing on the same image by the distorted image generation unit 131. The distorted image processing unit 133 acquires the intermediate processing result by performing image processing on the distorted image. Further, the distorted image processing unit 133 calculates an image processing loss for minimizing the difference between the image processing result obtained by the image processing and the correct label.

[0026] The intermediate similarity loss calculation unit 134 acquires the intermediate processing results obtained respectively from the original image processing unit 132 and the distorted image processing unit 133. The intermediate similarity loss calculation unit 134 calculates an intermediate similarity loss using the obtained intermediate processing results. The intermediate similarity loss is information used to solve the problem that error information does not propagate in the learning process of the CNN. The intermediate similarity loss is information that makes the firing patterns of the neurons in the intermediate processing results match. Specific examples of the intermediate similarity loss include the sigmoid cross-entropy loss of the intermediate processing result O_dist of the distorted image, with the firing pattern 1(O_clean) of the intermediate processing result O_clean obtained by image processing on the original image as the teacher. For example, the cross-entropy L represented by the following Equation 1 may be calculated.

[0027] L = -t log y - (1 - t) log (1 - y) ···(Equation 1) t is an element of 1(O_clean). y is an element of sigmoid(O_dist). Note that the intermediate similarity loss does not necessarily have to be limited to such specific examples. Any information may be used as long as the intermediate similarity loss is a loss that makes the firing patterns of the intermediate processing results match.

[0028] The optimization unit 135 updates the image processing parameters output from the distorted image processing unit 133 within the framework of optimization using the image processing loss generated by the distorted image processing unit 133 and the intermediate similarity loss generated by the intermediate similarity loss calculation unit 134. The updated image processing parameters (post-update image processing parameters) are input to the distorted image processing unit 133 when the learning process in the learning device 10 continues. On the other hand, when the learning process in the learning device 10 ends, the post-update image processing parameters are output as post-learning parameters.

[0029] The processing flow of the control unit 13 will be described below. FIG. 2 is a diagram showing a specific example of the overall processing flow of the control unit 13. First, the original image processing unit 132 and the distorted image processing unit 133 acquire image processing parameters for a predetermined image processing from the parameter storage unit 12 (step S001). The distorted image generation unit 131 acquires learning image data to be processed from the learning image storage unit 11 and executes distorted image processing on the acquired image (step S002). The original image processing unit 132 obtains an intermediate processing result by performing learning processing on an image (original image) read from the learning image storage unit 11 on which distorted image processing has not been performed (step S003). The distorted image processing unit 133 obtains an intermediate processing result by performing learning processing on an image (distorted image) on which distorted image processing has been performed, and further obtains an image processing loss (step S004). The intermediate similarity loss calculation unit 134 calculates an intermediate similarity loss using the intermediate processing result of the original image and the intermediate processing result of the distorted image in the same learning image data (step S005). The optimization unit 135 updates the image processing parameters using the intermediate similarity loss and the image processing loss (step S006).

[0030] If the learning process has not ended (step S007 - NO), the optimization unit 135 outputs the updated image processing parameters to the distorted image processing unit 133 (step 008). In the next process performed by the distorted image processing unit 133, the updated image processing parameters are used. If the learning process has ended (step S007 - YES), the optimization unit 135 outputs the updated image processing parameters as post-learning parameters (step 009).

[0031] FIG. 3 is a flowchart showing a specific example of the processing of the distortion image generation unit 131. First, the distortion image generation unit 131 selects a distortion d from a set of distortions D including no distortion, noise, and blur (step S101). This selection may be random or in a predetermined order. The set of distortions D may include other distortions. For example, a distortion such as JPEG compression distortion may be included. If the selected distortion d is "no distortion" (step S102 - no distortion), the distortion image generation unit 131 outputs the learning image as the distortion image as it is (step S103).

[0032] If the selected distortion d is "noise" (step S102 - noise), the distortion image generation unit 131 selects a parameter p for determining the noise intensity (step S104). The parameter p may be a value of about 10 to 50 for an image with 8-bit depth, for example. The parameter p may be selected according to the noise intensity assumed for the imaging target. Next, the distortion image generation unit 131 generates a distortion image by superimposing noise obtained from a Gaussian distribution with the parameter p as the standard deviation on the learning image (step S105). Note that the type of noise to be superimposed may be determined according to the type of noise assumed for the imaging target. For example, noise such as salt-and-pepper noise may be selected. The distortion image generation unit 131 outputs the generated distortion image (step S106).

[0033] If the selected distortion d is "blur" (step S102 - Blur), the distortion image generation unit 131 selects the content of the blur process (step S107). For example, the distortion image generation unit 131 selects a parameter p that determines the blur intensity and a size k of the blur kernel. The parameter p may be a value of about 1 to 5, for example. The size k may be calculated by an equation such as k = 4*p - 1. These values may be selected according to the blur intensity assumed for the imaging target. The distortion image generation unit 131 creates a Gaussian filter of size k*k with the parameter p as the standard deviation and convolves it with the learning image (step S108). Note that different blurs such as motion blur may be superimposed according to the type of blur assumed for the imaging target. The distortion image generation unit 131 outputs the obtained blurred image as a distorted image (step S109).

[0034] Figure 4 is a flowchart showing a specific example of the processing of the original image processing unit 132. First, the original image processing unit 132 acquires image processing parameters from the parameter storage unit 12 (step S201). The original image processing unit 132 acquires a learning image from the learning image storage unit 11 (step S202). The original image processing unit 132 executes image processing based on the image processing parameters on the acquired learning image (step S203). The original image processing unit 132 outputs the intermediate processing result of the image processing to the intermediate similarity loss calculation unit 134 (step S204).

[0035] Figure 5 is a flowchart showing a specific example of the processing of the distortion image processing unit 133. First, the distortion image processing unit 133 acquires image processing parameters from the parameter storage unit 12 (step S301). The distortion image processing unit 133 acquires the correct label from the learning image storage unit 11 (step S302). The distortion image processing unit 133 acquires the distorted image generated by the distortion image generation unit 131 (step S303).

[0036] The distortion image processing unit 133 performs image processing based on image processing parameters on the acquired distorted image (step 304). The distortion image processing unit 133 outputs the intermediate processing result obtained by executing the image processing in step S304 to the intermediate similarity loss calculation unit 134 (step S305). The distortion image processing unit 133 calculates an image processing loss that reduces (e.g., minimizes) the difference between the image processing result x and the correct label y (step S306). As the image processing loss, for example, cross-entropy L_dist(x',y)=-Σy_q log(x'_q) may be used, or the mean squared error may be used, or other objective functions may be used. For example, any function may be used as long as it is appropriate for the executed image processing. The distortion image processing unit 133 outputs the calculated image processing loss to the optimization unit 135 (step S307).

[0037] Figure 6 is a flowchart showing a specific example of the processing of the intermediate similarity loss calculation unit 134. First, the intermediate similarity loss calculation unit 134 acquires the respective intermediate processing results from the original image processing unit 132 and the distortion image processing unit 133 (step S401). The intermediate similarity loss calculation unit 134 calculates an intermediate similarity loss such that the intermediate processing result of the distorted image and the intermediate processing result of the original image are similar (step S402). The intermediate similarity loss is used to solve the problem that error information does not propagate. Constraints are imposed on the intermediate similarity loss to match the firing patterns of the neurons of the intermediate processing results. The intermediate similarity loss calculation unit 134 outputs the calculated intermediate similarity loss to the optimization unit 135 (step S403).

[0038] FIG. 7 is a flowchart showing a specific example of the processing of the optimization unit 135. First, the optimization unit 135 acquires the calculated image processing loss (step S501). The optimization unit 135 acquires the calculated intermediate similarity loss (step S502). The optimization unit 135 acquires the image processing parameters from the distortion image processing unit 133 (step S503). The optimization unit 135 updates the image processing parameters based on the image processing loss and the intermediate similarity loss (step S504). For example, the optimization unit 135 may update the image processing parameters by linearly combining the image processing loss and the intermediate similarity loss with a combined weight λ. This combined weight λ may be, for example, a ratio of about 1:0.1. The combined weight λ may be manually tuned while observing the transition of the entire loss function. For the update of the image processing parameters, a stochastic gradient descent method such as SGD or Adam may be used. For the update of the image processing parameters, other optimization algorithms such as the Newton method may be used.

[0039] In this repetition, the optimization unit 135 determines whether learning has ended (step S505). The determination of the end of learning may be based on any criterion. For example, the number of learning times (number of repetitions) may be determined in advance as a fixed value. For example, the end condition based on the loss function may be determined in advance. If learning has not ended (step S505-NO), the optimization unit 135 outputs the updated image processing parameters to the distortion image processing unit 133 (step S506). If learning has ended (step S505-YES), the optimization unit 135 outputs the updated image processing parameters as the post-learning parameters (step S507).

[0040] Through such processing, in the intermediate processing results O_clean of the CNN for an image without distortion image processing (original image) and the intermediate processing results O_dist of the CNN for an image with distortion image processing (distorted image), learning processing is performed such that the respective firing patterns 1 (O_clean) and 1 (O_dist) match. Specifically, learning may be performed using the sigmoid cross-entropy between 1 (O_clean) and O_dist. That is, when the response of O_dist is converted by sigmoid, the learning process is advanced so that it becomes the same as 1 (O_clean). Therefore, the response to the original image and the response to the distorted image become closer responses. Therefore, in the learning process using the learning results (parameters after learning) obtained by such learning processing, even if a distorted image is input, a value closer to the estimation result obtained when no distortion occurs is obtained.

[0041] Also, through the above-described processing, it is expected that the neurons in the hidden layer will respond to features such as a mixture of textures and a plurality of boundary regions within the image. Learning is performed so that such higher-order neurons show the same response for the distorted image and the original image. As a result, even when the complex features to which the higher-order neurons respond are distorted, robustness is shown and processing can be performed accurately.

[0042] [Second Embodiment] FIG. 8 is a diagram showing a functional configuration example of the second embodiment of the present invention. The learning device 20 is configured using an information processing device such as a personal computer or a server device. The learning device 20 includes a learning image storage unit 11, a parameter storage unit 12, and a control unit 23. The control unit 23 functions as a distortion image generation unit 231, an original image processing unit 232, a distortion image processing unit 233, an intermediate similarity loss calculation unit 234, an optimization unit 235, and a weight adjustment unit 236. The distortion image generation unit 231, the original image processing unit 232, the distortion image processing unit 233, the intermediate similarity loss calculation unit 234, and the optimization unit 235 in the second embodiment perform the same processing as the distortion image generation unit 131, the original image processing unit 132, the distortion image processing unit 133, the intermediate similarity loss calculation unit 134, and the optimization unit 135 in the first embodiment, respectively.

[0043] The intermediate similarity loss used for optimization learning in the optimization unit 235 is expected to contribute to reducing the image processing loss. However, there are cases where it cannot necessarily be said that the image processing loss is reduced. For example, when the influence of distortions such as noise and blur is very large, it may be more reasonable to set a new learning pattern without matching the reaction pattern rather than matching it with the reaction pattern when processing the distortion-free image. Therefore, when it is assumed that the distortion becomes severe, it is necessary to appropriately adjust the weights between the image processing loss and the intermediate similarity loss during optimization learning. This weight adjustment is performed based on, for example, the method described in the following reference.

[0044] Reference: Y. Du, W. M. Czarnecki, S. M. Jayakumar, R. Pascanu, B. Lakshminarayanan, “Adapting Auxiliary Losses Using Gradient Similarity”, 2018.

[0045] The weight adjustment unit 236 calculates the weights between the image processing loss and the intermediate similarity loss. The weight adjustment unit 236 outputs the adjusted weights to the optimization unit 235. In the optimization unit 235, during the optimization learning in step S504, the adjusted weights are learned as the combined weight λ.

[0046] FIG. 9 is a flowchart showing a specific example of the processing of the weight adjustment unit 236. The weight adjustment unit 236 acquires the calculated image processing loss (step S601). The weight adjustment unit 236 acquires the calculated intermediate similarity loss (step S602). The weight adjustment unit 236 acquires the image processing parameters (step S603). The weight adjustment unit 236 obtains gradients by backpropagating the image processing loss and the intermediate similarity loss under the image processing parameters respectively (step S604). At this time, generally, the error backpropagation method is used to obtain the gradient information of the CNN, but other methods may be used as long as they are appropriate gradient acquisition methods. The weight adjustment unit 236 calculates the similarity of the two obtained gradients (step S605). Cosine similarity may be used for calculating the similarity, or other similarity metrics may be used. The weight adjustment unit 236 outputs the calculated gradient similarity to the optimization unit 235 as the adjusted weight.

[0047] By performing such processing, even when the influence of distortions such as noise and blur is very large, more appropriate processing can be performed, and it becomes possible to improve the accuracy of image processing.

[0048] [Third Embodiment] FIG. 10 is a diagram showing a functional configuration example of the third embodiment of the present invention. The image processing apparatus 30 is configured using an information processing apparatus such as a personal computer or a server apparatus. The image processing apparatus 30 is an apparatus that performs image processing using the learned parameters generated by the learning apparatus of the first embodiment or the second embodiment. The image processing apparatus 30 includes a target image storage unit 31, a parameter storage unit 32, and a control unit 33.

[0049] The target image storage unit 31 is configured using a storage device such as a magnetic hard disk device or a semiconductor storage device. The target image storage unit 31 stores data of an image (target image) that is the target of image processing in the image processing apparatus 30.

[0050] The parameter storage unit 32 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The parameter storage unit 32 stores image processing parameters (post-learning parameters) indicating a learned model obtained by a learning process previously executed by the learning device 10 or the learning device 20.

[0051] The control unit 33 is configured using a processor such as a CPU and a memory. The control unit 33 functions as an image processing unit 331 when the processor executes an image processing program. Note that all or part of each function of the control unit 33 may be realized using hardware such as an ASIC, a PLD, or an FPGA. The above image processing program may be recorded on a computer-readable recording medium. A computer-readable recording medium is, for example, a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, a semiconductor storage device (e.g., SSD), or a storage device such as a hard disk or a semiconductor storage device built into a computer system. The above image processing program may be transmitted via a telecommunication line.

[0052] The image processing unit 331 performs image processing on the target image stored in the target image storage unit 31 based on the image processing parameters (post-learning parameters) stored in the parameter storage unit 32. The target image to be processed by the image processing unit 331 does not necessarily have to be limited to the one stored in the target image storage unit 31. For example, the image processing unit 331 may execute image processing on a target image transmitted together with a request for image processing from another information processing device.

[0053] When the image processing unit 331 acquires the result of the image processing, it outputs the result of the image processing to a predetermined device.

[0054] According to the image processing apparatus of the third embodiment configured as described above, by performing image processing using the post-learning parameters obtained by the learning device of the first embodiment or the second embodiment, it is possible to perform robust image processing on a distorted image.

[0055] What the correct label is depends on what learning process is performed. For example, in the case of a learning process for determining the type of subject depicted in image data, a correct label indicating the correct answer of the type of subject depicted in the image data is given. In such a case, the correct label may be given, for example, as a vector sequence indicating the type of subject.

[0056] For example, in the case of a learning process for detecting a predetermined object from among the subjects depicted in image data, a correct label indicating the correct position of the specific object depicted in the image data is given. In such a case, the correct label may be given, for example, as an array indicating the position of the specific object.

[0057] For example, in the case of a learning process for performing region division of image data, a correct label indicating which region each pixel of the image data belongs to is given. In such a case, the correct label may be given, for example, as an array indicating the region to which each pixel belongs.

[0058] The learning device 10 may be implemented using a plurality of information processing devices communicably connected via a network. In this case, each functional unit included in the learning device 10 may be implemented in a distributed manner across a plurality of information processing devices. For example, the learning image storage unit 11, the parameter storage unit 12, and the control unit 13 may be implemented in different information processing devices. Also, the functions implemented in the control unit 13 may be implemented in a distributed manner across a plurality of information processing devices.

[0059] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.

Explanation of Reference Numerals

[0060] 10…Learning device, 11…Learned image memory unit, 12…Parameter memory unit, 13, 23…Control unit, 131, 231…Distorted image generation unit, 132, 232…Original image processing unit, 133, 233…Distorted image processing unit, 134, 234…Intermediate similarity loss calculation unit, 135, 235…Optimization unit, 236…Weight adjustment unit

Claims

1. A storage unit that stores image processing information obtained by associating feature information characterizing an image of an undistorted image with feature information characterizing an image of a distorted image; An image processing unit that acquires a target image that is an image to be processed, and obtains an image processing result by performing predetermined image processing on the acquired target image using the image processing information; and has, The feature information characterizing the image of the undistorted image is feature information obtained by performing an image identification process for identifying a subject in the image with the undistorted image as an input, The feature information characterizing the image of the distorted image is feature information obtained by performing an image identification process with the distorted image as an input, The image processing unit performs a process of identifying a subject in the image to be processed as the predetermined image processing, The image processing information is information that matches the firing state of neurons of the feature information characterizing the image of the non-distorted image and the firing state of neurons corresponding to the neurons in the distorted image, an image processing apparatus.

2. The association is a process of matching the firing states of higher-order neurons, The image processing apparatus according to claim 1.

3. The feature information is an intermediate processing result obtained by executing the image identification process on each of the undistorted image and the distorted image using image processing parameters obtained by performing predetermined machine learning processing on a learning image associated with a correct label, The image processing information is an intermediate similarity loss calculated using the intermediate processing result, the image processing apparatus according to claim 1 or 2.

4. An acquisition step of acquiring a target image that is an image to be processed; An image processing step of obtaining an identification result of a subject in the image to be processed by performing predetermined image processing on the acquired target image; and has, In the image processing, image processing information obtained by associating feature information characterizing an image of an undistorted image obtained by performing an image identification process for identifying a subject in the image with the undistorted image as an input, and feature information characterizing an image of a distorted image obtained by performing an image identification process with the distorted image as an input is used. The image processing information is information that matches the firing state of neurons of feature information characterizing the image of the image without the distortion and the firing state of neurons corresponding to the neurons in the image with the distortion. Image processing method.

5. A program for causing a computer to function as the image processing apparatus according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Image processing apparatus, image processing system, and image processing method

    JP2018060296A

  • Data augmentation techniques using style transformation with neural network

    JP2019032821A

  • Information processing method, information processing device, and program

    JP2020042760A