Image processing device, image processing method and image processing program

The image processing device optimizes CNN parameters to address the lack of shift invariance in CNNs by calculating and minimizing image identification and shift-invariant losses, enhancing classification accuracy and invariance.

JP2025120049APending Publication Date: 2025-08-15NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024015266
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Conventional Convolutional Neural Networks (CNNs) lack sufficient shift invariance, leading to inconsistent classification results when the position of the object in an image changes.

Method used

An image processing device that calculates an image identification loss and a shift-invariant loss for classification results, optimizing CNN parameters to minimize both losses, thereby enhancing shift invariance and overall performance.

Benefits of technology

Improves the classification accuracy and shift invariance of CNNs, resulting in enhanced performance by reducing both image identification and shift-invariant losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025120049000001_ABST
    Figure 2025120049000001_ABST
Patent Text Reader

Abstract

To improve performance of a CNN.SOLUTION: An image processing device 10 calculates an image identification loss being a loss of an identification result by a CNN to an input image. Also, the image processing device 10 calculates a shift invariable loss based on shift invariance of the identification result by the CNN to each of a plurality of images which are obtained by shifting the input image and are different from each other. Also, the image processing device 10 optimizes a parameter of the CNN so as to make an image identification loss and the shift invariable loss small.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an image processing method, and an image processing program. [Background technology]

[0002] Convolutional Neural Networks (CNNs) have been known as a machine learning method for identifying, detecting, and segmenting objects in images. For example, CNNs contribute to the automation of visual inspection processes in business.

[0003] When promoting the automation of visual inspection processes in business by processing captured images, it is desirable for the image processing of CNN to be in line with the intuition of the humans who are originally performing the visual inspection. For example, it is considered desirable for CNN to have the property that the classification results do not change as much as possible even if the position of the object to be classified during image classification shifts (hereinafter referred to as shift invariance).

[0004] On the other hand, it has been pointed out that CNNs lack shift invariance (see, for example, Non-Patent Document 1). Also, an index called CADD (Correctness-Aware Distribution Distance) has been proposed as an index for evaluating the shift invariance of CNNs (see, for example, Non-Patent Document 2). CADD is obtained by performing image classification using a CNN model on images given two different shifts, and measuring the degree to which the respective outputs deviate from each other. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] R. Zhang, “Making Convolutional Networks Shift-Invariant Again”, 2019. [Non-patent document 2] H. Higuchi, S. Suzuki, H. Shouno, “Measuring Shift-invariance of Convolutional Neural Network with a Probability-incorporated Metric”, 2022. Summary of the Invention [Problem to be solved by the invention]

[0006] However, conventional techniques may not be able to sufficiently improve the performance of CNNs. For example, conventional techniques cannot solve the problem of the lack of shift invariance in CNNs. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, the image processing device is characterized by having a first loss calculation unit that calculates a first loss, which is the loss of a classification result by an image classification model, based on a first image; a second loss calculation unit that calculates a second loss based on the shift invariance of the classification result by the image classification model for each of a plurality of mutually different images obtained by shifting the first image; and an optimization unit that optimizes parameters of the image classification model so that the first loss and the second loss are small. [Effects of the Invention]

[0008] According to the present invention, the performance of CNN can be improved. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an image processing apparatus according to the first embodiment. [Figure 2] FIG. 2 is a flowchart showing the flow of processing by the image processing apparatus according to the first embodiment. [Figure 3] FIG. 3 is a flowchart showing the flow of processing by the shifted image identifying unit. [Figure 4] FIG. 4 is a flowchart showing the flow of processing by the non-shifted image classification unit. [Figure 5] FIG. 5 is a flowchart showing the flow of processing by the image discrimination loss calculation unit. [Figure 6] FIG. 6 is a flowchart showing the flow of processing by the shift-invariant loss calculation unit. [Figure 7] FIG. 7 is a flowchart showing the flow of processing by the optimization unit. [Figure 8] FIG. 8 is a diagram illustrating an example of the configuration of an image processing apparatus according to the second embodiment. [Figure 9] FIG. 9 is a flowchart showing the flow of processing by the image processing apparatus according to the second embodiment. [Figure 10] FIG. 10 shows the results of the experiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a computer that executes an image processing program. DETAILED DESCRIPTION OF THE INVENTION

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An image processing apparatus, an image processing method, and an image processing program according to the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention is not limited to the following embodiments.

[0011] [First embodiment] First, the configuration of an image processing device according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of an image processing device according to the first embodiment. Note that the CNN in this embodiment is an example of an image recognition model. The CNN may be replaced as appropriate with another model (e.g., Transformer) capable of executing an image recognition task.

[0012] As shown in FIG. 1, the image processing device 10 includes a shifted image classification unit 101, a non-shifted image classification unit 102, a shift-invariant loss calculation unit 103, an image classification loss calculation unit 104, an optimization unit 105, a training image storage unit 111, and an initial parameter storage unit 112.

[0013] The image processing device 10 performs a process of optimizing the parameters of the CNN, i.e., a learning process. The image processing device 10 performs the learning process using input learning images, and outputs the optimized CNN parameters as learned parameters 20.

[0014] Images for learning are stored in a learning image storage unit 111. Furthermore, parameters before optimization are stored in an initial parameter storage unit 112. For example, the parameters of a CNN are the weights and biases of a neural network.

[0015] The processing flow of the image processing device 10 will be described with reference to Fig. 2. Fig. 2 is a flowchart showing the processing flow of the image processing device according to the first embodiment.

[0016] 2, first, the image processing device 10 sends an input image from the training image storage unit 111 to the shifted image classification unit 101 and the non-shifted image classification unit 102, and sends a correct label to the image classification loss calculation unit 104 (step S11). The input image is an example of a first image.

[0017] Next, the image processing device 10 generates two images by randomly shifting the input image using the shift image classification unit 101, performs classification using the image classification parameters, and then sends each classification result to the shift invariant loss calculation unit 103 (step S12).

[0018] Furthermore, the image processing device 10 performs classification on the input image using the image classification parameters in the non-shifted image classification unit 102, and sends the result to the image classification loss calculation unit 104 (step S13).

[0019] Next, the image processing device 10 calculates the image discrimination loss for the non-shifted image in the image discrimination loss calculation unit 104, and sends it to the optimization unit 105 (step S14). The image discrimination loss is an example of a first loss.

[0020] Furthermore, the image processing device 10 calculates a shift-invariant loss using the shift-invariant loss calculation unit 103 and sends it to the optimization unit 105 (step S15). Here, the image processing device 10 counts the distance for each correct answer and incorrect answer using the optimization unit 105 (step S16). The shift-invariant loss is an example of a second loss.

[0021] Then, if the condition for ending the learning is satisfied (step S17, Yes), the image processing device 10 acquires the optimized parameters obtained by the optimization unit 105 as the learned parameters 20 (step S18). Thereafter, the image processing device 10 may output the acquired learned parameters 20.

[0022] If the condition for ending the learning is not met (No at step S17), the image processing device 10 returns to step S11 and repeats the process.

[0023] The conditions for ending learning are that the parameter update amount has converged, that the processing from steps S11 to S16 has been repeated a predetermined number of times, that a predetermined time has elapsed, and so on.

[0024] Here, the classification using image classification parameters in steps S12 and S13 is a process in which a CNN is constructed based on the parameters acquired by the image processing device 10, an image is input to the constructed CNN, and classification is performed.

[0025] The image processing device 10 may acquire initial parameters from the initial parameter storage unit 112, or may acquire optimized parameters from the optimization unit 105. The optimized parameters are parameters that have been optimized (updated) at least once with respect to the initial parameters.

[0026] The CNN also outputs a vector as the classification result. Each element of the vector corresponds to a class to be classified. The sum of the values of the vector's elements is assumed to be 1. The classification result is the class to be classified that corresponds to the element with the largest value among the elements of the vector output by the CNN.

[0027] The identification target includes any object that appears in the image, such as a person, an animal, a plant, a car, a structure, or a building.

[0028] Furthermore, a class is previously associated with each learning image as a correct label. The image processing device 10 can determine whether the classification result is correct or incorrect depending on whether the classification result by CNN and the class of the correct label are the same.

[0029] In step S15, the image processing device 10 uses D JS According to the method described in Non-Patent Document 2, D is calculated for each of the images for which the classification result is correct and for which the classification result is incorrect. JS Then, the image processing device 10 calculates the distance D JS Based on this, CADD is calculated in Non-Patent Document 2.

[0030] The flow of processing by the shifted image identifying unit 101 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of processing by the shifted image identifying unit.

[0031] 3, if it is the start of learning (when the initial parameters have never been optimized) (Step S101, Yes), the shifted image identification unit 101 acquires the initial parameters (Step S102). If it is not the start of learning (Step S101, No), the shifted image identification unit 101 acquires the optimized parameters (Step S103).

[0032] Next, the shifted image identification unit 101 receives an input image (step S104). The shifted image identification unit 101 shifts the input image based on two different random number seeds (step S105). As a result, the shifted image identification unit 101 obtains shifted images corresponding to the two random number seeds, respectively.

[0033] For example, the random number seed determines the number of pixels by which the input image is shifted up, down, left, and right using a uniform distribution or the like. In other words, the shift image classification unit 101 performs image processing to shift the classification target in a predetermined direction by a random number of pixels. At this time, if the random number seed is different, the shift direction and the shift amount (number of pixels) will differ.

[0034] The shifted image identifying unit 101 may perform the shift by cutting and pasting the image, or may perform the shift by a method using a machine learning model.

[0035] The shifted image classification unit 101 applies image classification based on model parameters (image classification parameters) to each of the two shifted images (step S106). That is, the shifted image classification unit 101 obtains classification results for each of the two shifted images using CNN.

[0036] The shifted image classifying unit 101 outputs the output distribution obtained as a result of image classification for each of the two images to the shift-invariant loss calculating unit 103 (step S107). For example, the shifted image classifying unit 101 outputs the vector obtained as the output of the CNN to the shift-invariant loss calculating unit 103.

[0037] The flow of processing by the non-shifting image classification unit 102 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the flow of processing by the non-shifting image classification unit.

[0038] 4, if it is the start of learning (step S201, Yes), the non-shifted image classification unit 102 acquires initial parameters (step S202). If it is not the start of learning (step S201, No), the non-shifted image classification unit 102 acquires optimized parameters (step S203).

[0039] Next, the non-shifted image classification unit 102 acquires an input image from the training image storage unit 111 (step S204). Then, the non-shifted image classification unit 102 applies image classification based on the model parameters to the input image (step S205). The non-shifted image classification unit 102 outputs the obtained output result to the image classification loss calculation unit 104 (step S206).

[0040] The flow of processing by the image identification loss calculation unit 104 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of processing by the image identification loss calculation unit.

[0041] 5, the image discrimination loss calculation unit 104 acquires the classification result of the unshifted image from the unshifted image classification unit 102 (step S301). In addition, the image discrimination loss calculation unit 104 acquires the correct label from the training image storage unit 111 (step S302).

[0042] The image discrimination loss calculation unit 104 outputs a loss that reduces the difference between the model output and the correct label as the image discrimination loss to the optimization unit 105 (step S303).

[0043] For example, if the output of the CNN is x and the correct label is y, the image recognition loss calculation unit 104 calculates the cross entropy L(x, y)=-Σy q log(x q ) is calculated as the image identification loss. The image identification loss calculation unit 104 may calculate the image identification loss using an appropriate function (for example, mean square error) according to the classification task, instead of using the cross entropy.

[0044] The flow of processing in the shift-invariant loss calculation unit 103 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the flow of processing in the shift-invariant loss calculation unit.

[0045] 6, the shift-invariant loss calculation unit 103 acquires a classification result for the shifted image from the shifted image classification unit 101 (step S401). The shift-invariant loss calculation unit 103 calculates a loss that matches the classification result for the shifted image, and outputs the loss as a shift-invariant loss to the optimization unit 105 (step S402).

[0046] The shift-invariant loss calculation unit 103 can calculate the CADD as the shift-invariant loss as described in Non-Patent Document 2. Furthermore, the shift-invariant loss calculation unit 103 may calculate the shift-invariant loss not only by CADD but also by using a function that increases as the classification result of one shifted image and the classification result of the other shifted image become more different from each other.

[0047] The processing flow of the optimization unit 105 will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the processing flow of the optimization unit.

[0048] 7, if it is the start of learning (step S501, Yes), the optimization unit 105 acquires initial parameters (step S502). If it is not the start of learning (step S501, No), the optimization unit 105 proceeds to step S503 without acquiring post-optimization parameters.

[0049] The optimization unit 105 acquires the image discrimination loss from the image discrimination loss calculation unit 104 (step S503). Also, the optimization unit 105 acquires the shift invariant loss from the shift invariant loss calculation unit 103 (step S504).

[0050] The optimization unit 105 optimizes the parameters so as to reduce the image discrimination loss and the shift invariant loss (step S505). That is, the optimization unit 105 updates the parameters of the CNN so as to simultaneously optimize both the image discrimination loss and the shift invariant loss. For example, the optimization unit 105 updates the parameters using backpropagation or gradient descent.

[0051] Then, if the condition for ending the learning is satisfied (Yes at step S506), the optimization unit 105 outputs the optimized parameters as the learned parameters 20 (step S508).

[0052] If the condition for terminating the learning is not satisfied (step S506, No), the optimization unit 105 outputs the optimized parameters to the shifted image classification unit 101 and the non-shifted image classification unit 102 (step S507). After that, the shifted image classification unit 101 and the non-shifted image classification unit 102 further repeat the process using the optimized parameters.

[0053] [Advantages of the first embodiment] As described above, the image processing device 10 calculates an image identification loss, which is the loss of a classification result by a CNN (an example of an image classification model) for an input image. The image processing device 10 also calculates a shift invariance loss based on the shift invariance of the classification result by the CNN for each of a plurality of different images obtained by shifting the input image. The image processing device 10 also optimizes the parameters of the CNN so as to reduce the image identification loss and the shift invariance loss.

[0054] The image processing device 10 can improve the classification accuracy of the CNN by reducing the image classification loss, and can also improve the shift invariance of the CNN by reducing the shift invariance loss. As a result, according to this embodiment, the performance of the CNN is improved.

[0055] [Second embodiment] The image processing device 10 of the first embodiment calculates the image identification loss based on the classification result of the CNN for the input image. On the other hand, the image identification loss may be calculated for an image obtained by shifting the input image.

[0056] Fig. 8 is a diagram showing an example of the configuration of an image processing apparatus according to the second embodiment. In Fig. 8, parts common to those in the first embodiment are given the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0057] As shown in Fig. 8, the image processing device 10a includes an image classification unit 102a. The image classification unit 102a calculates an image classification loss, which is the loss of classification results by CNN for an image obtained by shifting an input image. This is expected to improve shift invariance through optimization using the image classification loss.

[0058] The flow of processing by the image identification unit 102a will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of processing by the image processing device according to the second embodiment.

[0059] 9, if it is the start of learning (step S601, Yes), the image classification unit 102a acquires initial parameters (step S602). If it is not the start of learning (step S601, No), the image classification unit 102a acquires optimized parameters (step S603).

[0060] Next, the image classification unit 102a acquires an input image from the training image storage unit 111 (step S604). The image classification unit 102a shifts the input image based on one random number seed to obtain a shifted image (step S605).

[0061] Then, the image classification unit 102a applies image classification based on the model parameters to the shifted image (step S606).The image classification unit 102a outputs the obtained output result to the image classification loss calculation unit 104 (step S607).

[0062] [experiment] An experiment conducted to confirm the effects of the embodiment will be described. In the experiment, a CNN model called VGG-11 was trained using CIFAR-10, an image classification dataset. The training was performed using momentum stochastic gradient descent with a learning rate of 0.001.

[0063] FIG. 10 shows the results of the experiment. Each line represents CADLoss (shift-invariant loss) for each epoch number (number of learning iterations). Line 51 is the result of the first embodiment when shift-invariant loss is not used for parameter optimization. Line 52 is the result of the second embodiment when shift-invariant loss is not used for parameter optimization. Line 54 is the result of the first embodiment when shift-invariant loss is used for parameter optimization. Line 53 is the result of the second embodiment when shift-invariant loss is used for parameter optimization.

[0064] That is, lines 51 and 52 correspond to the prior art. Lines 54 and 53 correspond to the first and second embodiments, respectively. The smaller the shift invariance loss on the vertical axis, the greater the shift invariance of the CNN. From FIG. 10, it can be said that the shift invariance of the CNN improves during the training process according to the first and second embodiments.

[0065] [System configuration, etc.] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic. Note that the program may be executed not only by the CPU but also by other processors such as a GPU.

[0066] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.

[0067] [program] In one embodiment, the image processing device 10 can be implemented by installing an image processing program that executes the above-described processes as package software or online software on a desired computer. For example, by executing the image processing program on an information processing device, the information processing device can function as the image processing device 10. The information processing device referred to here includes desktop and notebook personal computers. Furthermore, other information processing devices also include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants).

[0068] The image processing device 10 can also be implemented as a server device that provides services related to the above processing to a client terminal device used by a user. For example, the server device is implemented as a server device that provides a service that outputs trained CNN parameters. In this case, the server device may be implemented as a web server or as a cloud that provides services related to the above processing through outsourcing.

[0069] 11 is a diagram showing an example of a computer that executes an image processing program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0070] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0071] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the image processing device 10 is implemented as a program module 1093 in which computer-executable code is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configuration of the image processing device 10 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0072] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0073] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070. [Explanation of symbols]

[0074] 10, 10a Image processing device 20 Trained parameters 101 Shift image identification unit 102 Non-shift image classification unit 102a Image identification unit 103 Shift-invariant loss calculation part 104 Image recognition loss calculation unit 105 Optimization Department 111 Learning image memory unit 112 Initial parameter storage section

Claims

1. a first loss calculation unit that calculates a first loss, which is a loss of a classification result by an image classification model, based on the first image; a second loss calculation unit that calculates a second loss based on shift invariance of a classification result by the image classification model for each of a plurality of mutually different images obtained by shifting the first image; an optimization unit that optimizes parameters of the image recognition model so that the first loss and the second loss are small; 1. An image processing device comprising:

2. The first loss calculation unit calculates, as the first loss, a loss in a classification result by the image classification model for the first image or an image obtained by shifting the first image.

2. The image processing device according to claim 1, wherein:

3. An image processing method executed by an image processing device, comprising: a first loss calculation step of calculating a first loss, which is a loss of a classification result by the image classification model, based on the first image; a second loss calculation step of calculating a second loss based on shift invariance of a classification result by the image classification model for each of a plurality of mutually different images obtained by shifting the first image; an optimization step of optimizing parameters of the image recognition model so that the first loss and the second loss are reduced; An image processing method comprising:

4. a first loss calculation step of calculating a first loss, which is a loss of a classification result by the image classification model, based on the first image; a second loss calculation step of calculating a second loss based on shift invariance of a classification result by the image classification model for each of a plurality of mutually different images obtained by shifting the first image; an optimization step of optimizing parameters of the image identification model so that the first loss and the second loss are small; An image processing program characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Learning device and learning method

    JP2024006730A

  • Image processing method, learning device, and image processing device

    WO2021106174A1