Information processing apparatus, information processing method, and program

The information processing apparatus addresses the labor-intensive nature of subjective evaluation by dividing images into local regions, calculating similarities, and enabling user evaluation, thus reducing the user's burden and enhancing the learning model's performance.

JP2025089055APending Publication Date: 2025-06-12CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023204013
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing subjective evaluation methods for assessing the performance of learning models in image processing are labor-intensive and time-consuming, particularly when evaluating local regions within images.

Method used

An information processing apparatus that divides reference and comparison images into local regions, calculates the similarity between corresponding regions, and displays this information to the user for evaluation, thereby facilitating the selection of images for re-learning the learning model.

Benefits of technology

This approach reduces the user's evaluation burden by providing a structured method for assessing image similarities and selecting appropriate images for re-learning, thereby improving the performance of the learning model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089055000001_ABST
    Figure 2025089055000001_ABST
Patent Text Reader

Abstract

To reduce a burden on a user in evaluating an image obtained through processing with a learning model.SOLUTION: An information processing apparatus divides a reference image, and a comparison image obtained by processing, with a learning model, another image of the same scene as the reference image, into local areas, acquires the similarity between the local areas at the same coordinate positions of the reference image and the comparison image, and on the basis of the similarity, determines the details of screen output in displaying the local areas of the reference image and the comparison image. The information processing apparatus acquires a result of a user's evaluation of the local areas displayed according to the details of screen output, and on the basis of the evaluation result, selects an image used for re-training of the learning model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing technology for handling images used in machine learning.

Background Art

[0002] There is known an image generation technology that generates an image with noise removed or reduced from an image having noise or a super-resolution image with high resolution from an image with low resolution by inference processing using a learning model. Further, in such an image generation technology, there is known a method of improving the performance of image generation processing by learning a learning model using an image and updating the parameters of the learning model. In order to efficiently improve the performance of the learning model, it is important to correctly evaluate the performance of the learning model and appropriately select an image to be used for re-learning based on the evaluation result. As a method for evaluating the performance of the learning model, there is known a subjective evaluation method in which a human visually compares the difference between a reference image and an image obtained by inference processing using the learning model. As a result of this subjective evaluation, when it is evaluated (judged) that it is necessary to improve the performance of the learning model for a specific local region in the image, an image similar to the image in that local region is acquired, and the learning model is re-learned using that image as re-learning data. By doing so, the performance for a specific region that the learning model is not good at is improved, and the non-uniformity of performance occurring for each local region in the image is improved. Further, as a method for acquiring an image to be used for re-learning, for example, there is a method of searching for and displaying a plurality of images similar to a specific region respectively, and selecting an image to be used for re-learning data by the user from among the images of the plurality of search results.

[0003] Note that as a technology for displaying a plurality of searched images, there is a technology as disclosed in Patent Document 1. Patent Document 1 discloses a technology in which a user selects a partial region to be noted from a target image, assigns a weight (importance) to the partial region, calculates the similarity between the target image and the search image, and rearranges and displays the search results in the order of the similarity.

Prior Art Documents

Patent Document

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] As described above, in subjective evaluation, a human visually compares the difference between a reference image and an image obtained by inference processing using a learning model. However, in order to correctly evaluate the performance of the learning model, it is necessary to perform evaluation using a huge number of images, and the time and labor required for the evaluation impose a great burden on the user. In particular, when performing subjective evaluation for each local region within an image, even more time, labor, and effort are required.

[0006] Therefore, an object of the present invention is to reduce the burden on the user when evaluating an image processed by a learning model.

Means for Solving the Problems

[0007] The information processing apparatus of the present invention includes: a dividing means for dividing a reference image and a comparison image obtained by processing, using a learning model, another image of the same scene as the reference image into local regions; a similarity acquisition means for acquiring the similarity between local regions at the same coordinate positions of the reference image and the comparison image; an output content determination means for determining, based on the similarity, the screen output content when displaying the local regions of the reference image and the comparison image; a result acquisition means for acquiring the evaluation result of the user with respect to the local regions included in the displayed screen output content; and a selection means for selecting, based on the evaluation result of the user, an image to be used for re-learning the learning model.

Effects of the Invention

[0008] According to the present invention, it is possible to reduce the burden on the user when evaluating an image processed by a learning model.

Brief Description of Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each of the embodiments described hereinafter does not limit the present invention, and not all of the plurality of features described in this embodiment are essential for the solution means of the present invention, and those plurality of features may be arbitrarily combined. The configuration of the embodiment can be appropriately modified or changed according to the specifications of the device to which the present invention is applied and various conditions (usage conditions, usage environment, etc.). Also, a configuration may be formed by appropriately combining a part of each of the embodiments described later. In the following embodiments, the same or similar configurations and processing steps are denoted by the same reference numerals, and duplicate explanations are omitted.

[0011] <First Embodiment> FIG. 1 is a diagram showing an example of the functional configuration of an information processing apparatus 100 according to this embodiment. In this embodiment, a case of re-learning a learning model used for image generation processing for generating an image with reduced noise from an image with noise or generating a super-resolution image with high resolution from an image with low resolution will be described as an example. In the embodiments described hereinafter, a case of re-learning a learning model used for inference processing for generating an image with reduced noise from an image with noise will be described as an example. Also, in this embodiment, a neural network model (hereinafter referred to as an NN model) is cited as the learning model.

[0012] The display unit 140 is a display device such as a liquid crystal display or an organic EL display externally attached to the information processing apparatus 100. Note that the display unit 140 may be a display device built in the information processing apparatus 100. The user input unit 130 is an input device such as a mouse or a keyboard externally attached to the information processing apparatus 100. Note that the user input unit 130 may be an input device built in the information processing apparatus 100, for example, a touch panel provided on the screen of the display unit 140. In the case of this embodiment, the user of the information processing apparatus 100 can input various instructions such as instructions for acquiring and selecting an image described later and instructions for display items on the display screen through the user input unit 130.

[0013] The model storage unit 123 of the learning processing unit 120 stores the NN model. The inference unit 124 generates an image with reduced noise from an image with noise by executing inference processing using the NN model stored in the model storage unit 123. The model learning unit 122 learns the NN model used for inference processing that generates an image with reduced noise from an image with noise. Further, the model learning unit 122 uses the image data selected by the image selection unit 110 described later as re-learning data to re-learn the NN model. Then, the model learning unit 122 updates the NN model stored in the model storage unit 123 with the re-learned NN model.

[0014] The data storage unit 121 stores a plurality of learning data sets used for pre-learning and re-learning of the NN model. The learning data set is a data set consisting of a reference image (GT) that is an image with very little noise or a noise-free image with no noise, and another image with noise that is an image of the same scene as the reference image (hereinafter referred to as a noise image). In the case of this embodiment, the reference image is, for example, an image taken with an ISO sensitivity of ISO 100, and the noise image is, for example, an image taken under shooting setting conditions where the ISO sensitivity is sufficiently higher than ISO 100. Note that the noise image does not necessarily have to be an image acquired by shooting, and may be an image generated by performing predetermined image processing (in this embodiment, image processing for adding noise) on the reference image.

[0015] Also, in this embodiment, as will be described later, the data storage unit 121 also stores a comparison image generated by performing inference processing using the NN model on the noise image. The comparison image is associated with the learning data set consisting of the reference image and the noise image and is stored in the data storage unit 121. Note that the comparison image may be generated in advance or may be generated each time.

[0016] The image acquisition unit 101 of the image selection unit 110 reads and acquires the reference image stored in the data storage unit 121 of the learning processing unit 120 and the comparison image associated with the reference image. For example, when a list of a plurality of reference images stored in the data storage unit 121 is displayed on the display unit 140 and the user selects a reference image from the list, the image acquisition unit 101 acquires the reference image and the comparison image associated therewith from the data storage unit 121.

[0017] The region division unit 102 divides the reference image and the comparison image acquired by the image acquisition unit 101 into predetermined local regions respectively. In the reference image and the comparison image, the range for dividing into local regions may be a predetermined range determined in advance or a range arbitrarily specified by the user. Details of the division process of the local regions for the reference image and the comparison image will be described later.

[0018] The similarity acquisition unit 103 performs, for all local regions, the process of acquiring the similarity between the corresponding local regions at the same coordinate positions of the reference image and the comparison image. In this embodiment, the similarity acquisition unit 103 is assumed to perform, as the similarity acquisition process, a process of calculating the similarity based on the image feature amounts of the corresponding local regions at the same coordinate positions. Details of the similarity calculation process based on the image feature amounts will be described later.

[0019] The output content determination unit 104 determines the content to be output (displayed) on the screen of the display unit 140 based on the similarity calculated by the similarity acquisition unit 103, that is, the similarity between all corresponding local regions at the same coordinate positions of the reference image and the comparison image. Although details will be described later, in the case of this embodiment, the screen output content is determined to include at least the images of the local regions at the same coordinate positions in the reference image and the comparison image and the values of the similarities between those local regions.

[0020] The screen output unit 105 outputs screen display data corresponding to the screen output content determined by the output content determination unit 104 to the display unit 140. Thereby, on the display unit 140, a screen display corresponding to the screen output content is performed. Details of the display screen of the display unit 140 corresponding to the screen output content will be described later.

[0021] In the case of this embodiment, the user can evaluate whether it is necessary to perform relearning for any specific local region by looking at the similarity values between the reference image and the images of the local regions in the comparison image, which are displayed as the screen output content on the display unit 140. The evaluation of whether relearning is necessary is performed by checking whether there is, for example, image degradation or noise in the corresponding local regions of the reference image and the comparison image. Then, when there is a local region evaluated as requiring relearning, the user designates that local region via the user input unit 130.

[0022] The evaluation result acquisition unit 106 acquires information indicating the local region designated by the user via the user input unit 130 after evaluating that the user needs relearning. The similar data selection unit 107 selects, based on the information indicating the local region acquired by the evaluation result acquisition unit 106, image data similar to the image of that local region from the data storage unit 121 of the learning processing unit 120. Then, the image data selected by the similar data selection unit 107 is used as relearning data for the relearning of the NN model by the model learning unit 122 of the learning processing unit 120. The details of the relearning of the NN model using the image data (relearning data) selected by the similar data selection unit 107 will be described later.

[0023] <Hardware Configuration> FIG. 2 is a diagram showing an example of the hardware configuration applicable to the information processing apparatus 100 according to this embodiment, and here, the basic configuration of a computer is shown. In FIG. 2, the processor 201 is, for example, a CPU and controls the overall operation of the computer. The memory 202 is, for example, a RAM and temporarily stores the information processing program, image data, learning dataset, etc. according to this embodiment. The storage medium 203 is a computer-readable storage medium, such as an HDD, SSD, CD-ROM, etc., and stores various programs including the information processing program according to this embodiment, image data, learning dataset, NN model, etc. in the long term. The information processing program of this embodiment stored in the storage medium 203 is read into the memory 202, and when the processor 201 executes the program, each functional unit shown in FIG. 1 is realized.

[0024] The input IF 204 is an interface for acquiring information from an external device. A mouse, keyboard, etc. of the user input unit 130 shown in FIG. 1 are connected to the input IF 204, and information such as an instruction by the user is input. Also, the output IF 205 is an interface for outputting information to an external device. A liquid crystal display, organic EL display, etc. of the display unit 140 shown in FIG. 1 are connected to the output IF 205, and screen display data, etc. output from the screen output unit 105 are output. The bus 206 connects the above-described units and enables data exchange.

[0025] In FIG. 1, the information processing apparatus 100 including the image selection unit 110 and the learning processing unit 120 is taken as an example, but the image selection unit 110 and the learning processing unit 120 may be separate devices. When the image selection unit 110 and the learning processing unit 120 are separate devices, those devices may be realized by the hardware configuration of FIG. 2 respectively.

[0026] FIG. 3 is a flowchart showing the overall flow of information processing in the information processing apparatus 100 of this embodiment shown in FIG. 1. First, as the process of step S301, the image acquisition unit 101 of the image selection unit 110 acquires the reference image selected by the user from the data storage unit 121. In the present embodiment, the reference image selected by the user is assumed to be, for example, an image for which the user himself / herself expects image quality improvement and desires to perform image quality evaluation by subjective evaluation. In the case of an example assuming improvement of the accuracy of the NN model as in the present embodiment, it is preferable to select a plurality of patterns of images as the reference image, but one image may also be used. The selection of the reference image is performed, for example, by displaying a list of a plurality of reference images stored in the data storage unit 121 on the display unit 140 and having the user select from the list through the user input unit 130.

[0027] Next, as the process of step S302, the inference unit 124 of the learning processing unit 120 acquires a noise image, which is another image taken of the same scene as the reference image selected in step S301, from the data storage unit 121. The inference unit 124 performs inference processing on the noise image acquired from the data storage unit 121 using the NN model stored in the model storage unit 123, and generates an image of the inference result as a comparison image. The NN model used at this time is a pre-trained NN model. The NN model may be a CNN model (Convolutional Neural Network) having a convolutional layer, or a model having a Transformer layer. Further, the NN model is not limited to the NN model pre-trained in the information processing apparatus 100, and may be, for example, a pre-trained NN model publicly released by a third party. Then, the comparison image generated in step S302 is stored in the data storage unit 121 in association with the reference image.

[0028] Next, as the process of step S303, the image selection unit 110 performs processing for selecting an image to be used for re-learning of the NN model based on the reference image selected in step S301 and the comparison image generated in step S302, as will be described later. The details of the image selection process for re-learning performed by the image selection unit 110 in step S303 will be described later. Subsequently, as the process of step S304, the learning processing unit 120 retrains the NN model using the image selected by the image selection unit 110 in step S303 as retraining data.

[0029] FIG. 4 is a flowchart showing the detailed flow of the selection process of the image used for retraining the NN model, which is performed by the image selection unit 110 in step S303. First, as the process of step S401, the image acquisition unit 101 of the image selection unit 110 acquires a reference image and a comparison image from the data storage unit 121. Here, the reference image acquired by the image acquisition unit 101 from the data storage unit 121 is the reference image selected by the user during the process of step S301 in FIG. 3. The comparison image is an image generated by performing an inference process using the NN model on a noise image obtained by photographing the same scene as the reference image in S302 or a noise image obtained by performing image processing to add noise to the reference image, and is stored in the data storage unit 121.

[0030] Next, as the process of step S402, the region division unit 102 acquires information on division settings for dividing the reference image and the comparison image into local regions respectively. In the present embodiment, the division settings include both a range setting indicating which range of the image is to be divided and a region setting indicating how to divide the inside of the range. In the case of the present embodiment, the information on the division settings is specified by the user arbitrarily or by the user selecting from a plurality of pre-prepared division settings. Note that if the image selection unit 110 is pre-provided with the information on the division settings, the process of this step S402 may be omitted. In the present embodiment, as an example of the division settings, an example will be described in which the entire image is used as the range setting, and the image (the entire image) within the range setting is divided into a total of nine local regions by a region setting in which the image is evenly divided into three parts in the vertical direction and evenly divided into three parts in the horizontal direction.

[0031] Next, as the process of step S403, the region division unit 102 divides the reference image and the comparison image into a plurality of local regions respectively based on the division setting information acquired in step S402. FIG. 5(a) is a diagram showing a plurality of local regions after the reference image and the comparison image are divided in step S403 based on the division setting information acquired in step S402 in the region division unit 102. In the case of this embodiment, since the range setting is the entire image and the region setting is an equal three-way division in both the vertical and horizontal directions, as shown in FIG. 5(a), the reference image 501 is divided into local regions E1 to E9, and similarly, the comparison image 502 is divided into local regions E1 to E9. The local regions E1 to E9 of the reference image 501 and the comparison image 502 are local regions at the same coordinate positions corresponding to each other.

[0032] Next, as the process of step S404, the similarity acquisition unit 103 determines whether there is a local region for which the similarity calculation process has not been performed among all the local regions divided by the region division unit 102. If the similarity acquisition unit 103 determines that the similarity calculation process has been performed for all the local regions, the process proceeds to step S406. On the other hand, if there is a local region for which the similarity calculation process has not been performed, the process proceeds to step S405.

[0033] When the process proceeds to step S405, the similarity acquisition unit 103 calculates the similarity between the local regions for which the similarity calculation process has not been performed. The similarity acquisition unit 103 in this embodiment calculates the similarity between the corresponding local regions located at the same coordinates on the reference image and the comparison image. In the case of this embodiment, the similarity between the local regions at the same coordinate positions of the reference image and the comparison image is calculated using the feature amounts of the images (referred to as local region images) of those corresponding local regions. Examples of the feature amounts of the image include one, or a combination of two or more, such as pixel values, lightness, chroma, and luminance values. Then, the similarity acquisition unit 103 calculates, as the similarity, one difference such as the difference in pixel values, the difference in lightness, the difference in chroma, or the difference in luminance values, or a combination of two or more differences between the local images at the same coordinate positions of the reference image and the comparison image. In addition, the similarity acquisition unit 103 may convert the image feature amounts into vector values and calculate the similarity as a vector distance such as cosine similarity or Euclidean distance.

[0034] In this embodiment, as an example of the similarity, a case will be described where the cosine similarity is used to evaluate the similarity between corresponding local regions at the same coordinate positions of a reference image and a comparison image. The cosine similarity is defined by the following formula (1).

[0035]

Equation

[0036] Here, among two corresponding local regions at the same coordinate positions of the reference image and the comparison image, the feature amount of the local region of the reference image is an n-dimensional feature vector x = (x 1 , x 2 , ···, x n ). Similarly, the feature amount of the local region of the comparison image is an n-dimensional feature vector y = (y 1 , y 2 , ···, y n ). The cosine similarity is calculated by substituting the values of x and y into the above formula (1). At least one value among pixel values, brightness, and again, luminance values is stored in the feature vector. The cosine similarity becomes a larger value as the similarity between the two feature vectors in the two corresponding local regions at the same coordinate positions of the reference image and the comparison image is higher (the maximum value is 1). That is, the larger the value of the cosine similarity, the higher the similarity between the two local regions at the same coordinate positions of the reference image and the comparison image.

[0037] The similarity acquisition unit 103 repeats the processes of steps S404 to S405 described above until there are no local regions where the similarity calculation process has not been performed. FIG. 5(b) is a diagram used to explain the similarity calculation process for all local images of the reference image 501 and the comparison image 502 shown in FIG. 5(a). That is, the similarity acquisition unit 103 calculates the similarity 503 between the local regions E1 at the same coordinate positions of the reference image 501 and the comparison image 502. Similarly, hereinafter, for the local regions E2, E3,... at the same coordinate positions, the similarities 504, 505,... are calculated in order. Then, when the similarity calculation process is completed for all the local regions E1 to E9, the similarity acquisition unit 103 holds the values of all the similarities calculated for each corresponding local region in the internal similarity storage unit 506.

[0038] When the similarity calculation process is completed for all local regions in step S404 and the process proceeds to step S406, the output content determination unit 104 acquires the similarity information stored in the similarity storage unit 506. Then, based on the similarity information acquired from the similarity storage unit 506, the output content determination unit 104 determines what kind of screen output to perform on the display unit 140. In the case of this embodiment, the output content determination unit 104 determines the screen output content in which the corresponding local region images at the same coordinate positions of the reference image and the comparison image are arranged, for example, in the order of the local region images with smaller similarity values and where the images of the reference image and the comparison image are not similar. Note that the arrangement order of the local region images may be in the order of the local region images with larger similarity values and where the images of the reference image and the comparison image are similar, or may be in an order arbitrarily specified by the user.

[0039] Next, as the process of step S407, the screen output unit 105 outputs screen display data corresponding to the screen output content determined by the output content determination unit 104 to the display unit 140. FIG. 6 is a diagram showing an example of a display screen corresponding to the screen output content determined by the output content determination unit 104. As shown in FIG. 6, the display screen corresponding to the screen output content includes an image name area 601, an image / similarity area 602, an original image area 603, a parameter input area 604, and a process reception area 605.

[0040] In the image name area 601, the names of the images displayed in the original image area 603 and the image / similarity area 602 are shown. In the image name area 601, after the user checks an image, if it is evaluated that all the local area images displayed in relation to the name of that image need to be relearned, it also includes checkboxes for fully selecting those local area images.

[0041] In the image / similarity area 602, information is displayed such that the user can check the values of the similarity calculation results for the corresponding local area images at the same coordinate positions of the reference image and the comparison image. FIG. 7(a) is a diagram showing a specific example of the information displayed in the image / similarity area 602. In the image / similarity area 602, a list of local area comparison information 710 is arranged, which includes a set of local area images obtained by grouping the local areas to be compared at the same coordinate positions of the reference image and the comparison image, and the values of the similarities obtained by the similarity calculation process between those sets of local area images.

[0042] FIG. 7(b) is a diagram showing an enlarged view of one piece of local area comparison information 710 arranged and displayed in a plurality within the image / similarity area 602. As shown in FIG. 7(b), the local area comparison information 710 includes a similarity value area 721 where the calculated similarity value is shown, a tag name area 722 indicating which image, either the reference image or the comparison image, the local area belongs to, and an image display area 723 where a set of local area images is displayed. Also, in the similarity value area 721, checkboxes are provided so that the user can individually specify the images that are evaluated to need relearning after checking the images. When the user evaluates that relearning is necessary, the user can check the checkboxes provided in the similarity value area 721. Then, the evaluation result acquisition unit 106 saves information indicating that the image is an image evaluated to need relearning when it is checked. Here, checkboxes are taken as an example, but the user input indicating that an image is evaluated to need relearning may be obtained by a method other than checking the checkboxes.

[0043] Return to the explanation of FIG. 6. In the original image area 603, the image that is the source of the local area displayed in the image / similarity area 602 is displayed. The original image displayed in the original image area 603 may be the reference image, the comparison image, or both.

[0044] In the parameter input area 604 of FIG. 6, a plurality of parameters related to the local area, the similarity calculation process, and the change of the screen output content in the present embodiment are displayed, and this area is prepared as an area where the user can arbitrarily input or change these parameters. FIG. 8 is an enlarged view of the parameter input area 604 shown in FIG. 6.

[0045] The area size input area 841 is an area where the user can input an arbitrary numerical value as the size of the local area. The area division unit 102 acquires the numerical value input in the area size input area 841 as the local area size parameter. Then, the area division unit 102 divides the reference image and the comparison image into a plurality of local areas according to the numerical value acquired as the local area size parameter. For example, when the numerical value of (X, Y) = (1000, 1200) is input from the user to the area size input area 841, the area division unit 102 divides the reference image and the comparison image into local areas with a vertical × horizontal of 1000 × 1200 pixels each.

[0046] In this embodiment, an example is given in which the reference image and the comparison image are divided into local regions of the same shape and size according to the number of vertical and horizontal pixels input to the region size input region 841. However, the present invention is not limited to this example. The local region may be, for example, a region for each so-called semantic region, a region for each object region detected by object detection processing, a region arbitrarily selected by the user, or the like. Also, whether to use a local region of the same shape and size, a local region of a semantic region, a local region of an object region by object detection, or a local region arbitrarily selected by the user can be selected by the user. In this case, the output content determination unit 104 may include, in the screen output content, a pull-down menu item for the user to select which of these local regions to use. Thereby, the user can select any one of a local region of the same shape and size, a local region of a semantic region, a local region of an object region by object detection, and a local region arbitrarily selected by the user from the menu items prepared in the pull-down format.

[0047] The display number input region 842 is a region in which the user can input an arbitrary numerical value as the number of local regions to be listed in the image / similarity region 602 of FIG. 6. The output content determination unit 104 acquires the numerical value input to the display number input region 842 as a display number parameter. The output content determination unit 104 determines the number of local regions to be listed in the image / similarity region 602 according to the numerical value acquired as the display number parameter. That is, in other words, the output content determination unit 104 restricts the number of local region comparison information 710 included in the screen output content based on the similarity calculated by the similarity acquisition unit 103. Then, the output content determination unit 104 determines screen output content in which a list table in which the local region comparison information 710 shown in FIG. 7 is arranged by the number indicated by the display number parameter is arranged in the image / similarity region 602. According to the present embodiment, the user can appropriately limit the number of local region comparison information 710 to be listed in the image / similarity region 602 by inputting a numerical value to the display number input region 842. Thereby, according to the present embodiment, it is possible to shorten the time required for screen output, and it is also possible to shorten the user's evaluation work time.

[0048] The rearrangement order input area 843 is an area where the user can arbitrarily specify the order in which a plurality of local area comparison information 710 arranged in the image / similarity area 602 is to be rearranged and displayed. In the case of this embodiment, it is assumed that, for example, a plurality of types of rearrangement orders are prepared as menu items in a pull-down format that can be selected by the user in the rearrangement order input area 843. The output content determination unit 104 acquires, as a rearrangement method parameter, the rearrangement order corresponding to the menu item selected by the user from among the plurality of rearrangement orders prepared in the rearrangement order input area 843. For example, when a menu item in ascending order of low similarity is selected by the user, the output content determination unit 104 determines the screen output content in which the local area comparison information 710 is arranged in ascending order of low similarity within the image / similarity area 602. When the local area comparison information 710 is arranged and displayed in ascending order of low similarity within the image / similarity area 602, the user can easily recognize what kind of images are the images that require re-learning, that is, the images with low similarity. Also, for example, when a menu item in ascending order of high similarity is selected by the user, the output content determination unit 104 determines the screen output content in which the local area comparison information 710 is arranged in ascending order of high similarity within the image / similarity area 602. Thus, when the local area comparison information 710 is rearranged in ascending order of high similarity, the user can analyze the characteristics of the areas in which the NN model excels. Further, the arrangement order of the local area comparison information 710 is not limited to the order according to the similarity value. For example, by arranging in the order according to hue or in chronological order, the user can analyze the characteristics of the NN model with respect to color, the resistance to temporal changes on the image, and the like.

[0049] The similarity calculation method input area 844 is an area where the user can arbitrarily specify what calculation method to use when calculating the similarity in step S405. In the case of this embodiment, it is assumed that in the similarity calculation method input area 844, for example, a plurality of types of similarity calculation methods are prepared as menu items in a pull-down format that can be selected by the user. The similarity acquisition unit 103 acquires, as a similarity calculation method parameter, a similarity calculation method corresponding to the menu item selected by the user from among the plurality of similarity calculation methods prepared in the similarity calculation method input area 844. In this embodiment, an example in which cosine similarity is calculated as the similarity has been described. However, for example, the similarity may be calculated based on the difference in pixel values or luminance values of the corresponding local region images at the same coordinate positions of the reference image and the comparison image. Also, the similarity calculation method input area 844 may be provided with menu items including a plurality of types of similarity calculation methods. When a menu item including a plurality of types of similarity calculation methods is selected by the user, the similarity acquisition unit 103 acquires a composite similarity using the plurality of similarity calculation results obtained by those plurality of types of similarity calculation methods. In this case, a composite analysis using the plurality of similarity calculation results obtained by the plurality of types of similarity calculation methods becomes possible.

[0050] The process reception area 605 in FIG. 6 is an area capable of acquiring user instructions that serve as triggers for executing the similarity calculation process, executing the storage of the similarity calculation result, and executing re-learning, respectively. FIG. 9 is a diagram showing an enlarged view of the process reception area 605 in FIG. 6. It is assumed that the user operations on the re-learning button 953, the save button 952, and the re-learning button 953 shown in FIG. 9 are operations such as mouse clicks and touches.

[0051] The calculation execution button 951 is a button that is operated by the user when instructing the execution of the similarity calculation process. When the user performs an execution instruction operation on the calculation execution button 951, the similarity acquisition unit 103 performs a similarity calculation process based on the similarity calculation method specified by the above-described similarity calculation method parameters. Then, the output content determination unit 104 determines the screen output content based on the similarity calculation result. If the screen output content has already been determined and the user instructs the execution of the similarity calculation process through the calculation execution button 951, the similarity acquisition unit 103 re-executes the similarity calculation process. Then, the output content determination unit 104 updates the screen output content based on the similarity calculation result of the re-execution.

[0052] The save button 952 is a button that is operated by the user when instructing the saving of the local area comparison information 710 displayed in the image / similarity area 602 of FIG. 6. When the user performs an execution instruction operation on the save button 952, the output content determination unit 104 saves the local area comparison information 710 displayed in the image / similarity area 602 internally. Note that the output content determination unit 104 may save the information including the image name area 601 and the parameter input area 604 described above. Further, the output content determination unit 104 may save only the local area comparison information 710 arbitrarily selected by the user from the local area comparison information 710 in the image / similarity area 602.

[0053] The re-learning button 953 is a button that is operated by the user when instructing the execution of re-learning. When the user performs an execution instruction operation from the re-learning button 953, the model learning unit 122 acquires re-learning data from the data storage unit 121 based on the user's evaluation result described later, and performs re-learning of the NN model in the model storage unit 123.

[0054] Return to the explanation of the flowchart in FIG. 4. After the above-described step S407, when proceeding to the process of step S408, the evaluation result acquisition unit 106 acquires the result of the subjective evaluation performed by the user by looking at the screen display corresponding to the screen output content determined in step S406 and displayed in step S407. In the case of this embodiment, the user checks the local area comparison information 710 of the local area image that requires relearning by subjective evaluation. The evaluation result acquisition unit 106 selects the local area image corresponding to the local area comparison information 710 selected by the user as the local area image evaluated as requiring relearning. In this way, by having the user perform subjective evaluation and select the area that needs to be relearned, it becomes possible to construct an NN model that conforms to the user's preference.

[0055] Next, as the process of step S409, the similar data selection unit 107 selects relearning data from the data storage unit 121 based on the local area image of the local area comparison information 710 acquired in step S408, that is, the local area image evaluated as requiring relearning by the user. For example, when the reference image is an image that can be used for relearning, the similar data selection unit 107 may select the reference image and the comparison image as the relearning data. Also, for example, if the dataset group is previously grouped using a classifier, the similar data selection unit 107 determines which group of images the image evaluated as requiring relearning by the user is similar to, and may select the images within the similar group.

[0056] FIG. 10 is a diagram used to explain a method of previously grouping the dataset group 1001 using the classifier 1002. The classifier 1002 groups the input dataset group 1001 and obtains the grouping result 1103. In the example of FIG. 10, as the grouping result 1103 by the classifier 1002, an example is shown in which the dataset group 1001 is classified into three groups GP1 to GP3. The classifier 1002 may use, for example, an NN model generated by supervised machine learning. For example, by using data with labels corresponding to a plurality of evaluation items as teacher data, it becomes easy to identify which evaluation item the image selected by user evaluation is effective for. Also, for example, by using data labeled by image characteristics (hue, lightness, luminance, etc.) or time series as teacher data, it becomes possible to identify data without bias regarding image characteristics and time series. As the classifier 1002, a classifier generated using unsupervised learning, hierarchical clustering, or non-hierarchical clustering may be used.

[0057] FIG. 11 is a diagram used to explain a method of determining which group GP1 to GP3 of the pre-grouped dataset group 1001 the image 1101 evaluated by the user as requiring re-learning belongs to and selecting re-learning data in the present embodiment. The image 1101 that requires re-learning is input to the same classifier 1002 as when the dataset group 1001 is grouped, and it is determined which group of groups GP1 to GP3 it belongs to. Then, the similar data selection unit 107 compares with the pre-grouped result 1103, and selects, as re-learning data, the result 1102 obtained by applying the image 1101 that requires re-learning to the classifier 1002 and the result 1103 belonging to the same group.

[0058] As described above, according to the first embodiment, it is possible to reduce the burden and time related to the subjective evaluation of the user, facilitate image selection for reflecting the subjective evaluation result in the NN model, and improve the performance of the NN model. Thus, according to the present embodiment, it becomes possible to re-learn the NN model into a model that conforms to the user's preference. In this embodiment, an example of image generation processing for generating an image with reduced noise from a noise image is given. However, the present embodiment is also applicable to other image generation processes such as super-resolution processing for generating a high-resolution image from a low-resolution image. In addition, the present embodiment is also applicable to image generation processing by style conversion, for example, image generation processing for converting a color image into a monochrome image.

[0059] <Second Embodiment> In the second embodiment, similar to the first embodiment, an application example to image generation processing for generating an image with reduced noise from a noise image will be described. In the first embodiment, an example of displaying local region images in ascending or descending order of the similarity between local regions at the same coordinate position of a reference image and a comparison image was given. In the second embodiment, an example of displaying a local region image where the similarity between local regions at the same coordinate position of a reference image and a comparison image is less than a predetermined threshold will be described. Note that in the second embodiment, the configurations of the image selection unit 110 and the learning processing unit 120 of the information processing apparatus 100 are substantially the same as those in FIG. 1 described above, and the hardware configuration is also the same as that in FIG. 2, so their illustrations are omitted. Also, since the flow of information processing in the second embodiment is also substantially the same as the flowcharts shown in FIGS. 4 and 5, their illustrations are also omitted. In the case of the second embodiment, mainly, the determination process of the screen output content based on the similarity performed by the output content determination unit 104 in FIG. 1 in step S407 of FIG. 4 is different from the example of the first embodiment.

[0060] FIG. 12 is a diagram used to explain the screen output content determination process performed by the output content determination unit 104 in the second embodiment. The histogram 1203 shown in Fig. 12(a) is a diagram showing the similarities calculated by the similarity acquisition unit 103 for each local region, arranged in the order of similarity as a histogram. Note that Fig. 12(a) shows an example of a histogram in which the similarity is lower on the left side of the similarity axis. The output content determination unit 104 of the second embodiment determines the number of local regions to be displayed in the screen output content based on the number of local regions less than a predetermined similarity threshold (less than the similarity threshold 1201 illustrated in Fig. 12(a)) based on this histogram 1203. In the case of the example in Fig. 12(a), the number of local regions corresponding to similarities less than the similarity threshold 1201, surrounded by the rectangle 1202, is determined as the number of local regions to be displayed in the screen output content. More specifically, the output content determination unit 104 counts the number of local regions having a similarity lower than the similarity threshold 1201, and determines the counted number of local regions as the number of display items. As a result, the local region comparison information 710 with a similarity less than the similarity threshold is listed and displayed in the image·similarity region 602 shown in Fig. 6.

[0061] Fig. 12(b) is a diagram showing an enlarged view of the parameter input region 604 of Fig. 6 in the case of the second embodiment. Since the region size input region 841, the sorting order input region 843, and the similarity calculation method input region 844 are the same as described above, the description thereof is omitted. In the case of the second embodiment, a threshold input region 2111 in which the user can input an arbitrary similarity threshold is provided in the parameter input region 604. The output content determination unit 104 acquires the numerical value input to this threshold input region 2111 as a similarity threshold parameter. Then, the output content determination unit 104 determines the number of local region images to be displayed on the screen based on the comparison between the histogram in which the similarities calculated for each local region by the similarity acquisition unit 103 are arranged in the order of similarity and the similarity threshold.

[0062] According to the second embodiment, by determining a predetermined similarity threshold, the number of local region images to be included in the screen output content is automatically determined, and by limiting the number of subjective evaluations before user evaluation, it is possible to reduce the evaluation load on the user. In the second embodiment, the number of local regions with a similarity less than a predetermined similarity threshold is used to determine the number of displayed items, but it is not limited to the example using the similarity threshold. For example, a display ratio may be specified in advance for the number of all local regions, and the number of displayed local region images may be determined so as to achieve the specified display ratio. Also, for example, the number of displayed items may be determined based on statistical information such as the average or variance of the similarity calculated for each local region.

[0063] In the second embodiment, when the similarity is less than a predetermined similarity threshold, the number of displayed local region images included in the screen output content is automatically determined. However, the relearning in the learning processing unit 120 may also be automatically performed. For example, when the number of local regions whose similarity calculated by the similarity acquisition unit 103 is less than a predetermined similarity threshold is further less than a predetermined number, the similar data selection unit 107 automatically selects image data similar to the images of those local regions with a number less than the predetermined number from the data storage unit 121. That is, when the number of local regions whose similarity is less than a predetermined similarity threshold is further less than a predetermined number, the similar data selection unit 107 selects, based on the images of the local regions with a number less than the predetermined number, the images to be used by the learning processing unit 120 for the relearning of the NN model. Then, the image data automatically selected by the similar data selection unit 107 is used as relearning data for the relearning of the NN model by the model learning unit 122 of the learning processing unit 120. In other words, in this example, when the number of local regions whose similarity is less than a predetermined similarity threshold is further less than a predetermined number, the model learning unit 122 of the learning processing unit 120 automatically performs the relearning of the NN model.

[0064] <The Third Embodiment> In the third embodiment, similar to the first embodiment, an application example to an image generation process for generating an image with reduced noise from a noise image will be described. In the first embodiment, an example in which, when dividing a reference image and a comparison image for each local region, division is performed based on a predetermined size determined in advance or a size arbitrarily specified by the user was described. In the third embodiment, an example in which the local regions for dividing the reference image and the comparison image are object regions detected by an object detection process or semantic regions will be described. Note that the configurations of the image selection unit 110 and the learning processing unit 120 in the information processing apparatus 100 of the third embodiment are substantially the same as those in FIG. 1 described above, and the hardware configuration is also the same as that in FIG. 2, so their illustrations are omitted. Also, since the flow of information processing in the third embodiment is also substantially the same as the flowcharts shown in FIGS. 4 and 5, their illustrations are also omitted.

[0065] In the case of the third embodiment, mainly, the division process performed by the region division unit 102 in FIG. 1, the division setting information acquired in step S402 of FIG. 4, and the division process performed in step S403 are different from the processes of the first embodiment. In the case of the third embodiment, the region division unit 102 in FIG. 1 includes an object detector that detects a specific object or a semantic region divider that divides an image into semantic regions.

[0066] Also, in the case of the third embodiment, in the parameter input area 604 of FIG. 8, instead of the region size input area 841, for example, an area for inputting the name of a specific object detected by the object detector or an area for specifying semantic region division by the semantic region divider is provided. That is, in the case of the third embodiment, in step S402 of FIG. 4, the region division unit 102 acquires the name of a specific object detected by the object detector or division setting information indicating that semantic region division is to be performed.

[0067] Then, in step S403, the region division unit 102 performs a division process on the local regions of the reference image and the comparison image based on the name of the specific object acquired in step S402 or the division setting information indicating semantic region division.

[0068] FIG. 13 is a diagram used to explain the division settings when dividing an image into local regions in the third embodiment. FIGS. 13(a) and 13(b) are explanatory diagrams of an example of performing a division process into local regions based on the region of a specific object detected by an object detector. FIG. 13(a) is a diagram showing an example of an image 1300 to be divided into local regions, and the image 1300 is a reference image and a comparison image. It is assumed that the image 1300 includes, for example, a person 1301a, a person 1302a, and a car 1303a.

[0069] FIG. 13(b) is a diagram showing an example of detecting a person from the image 1300 using an object detector that detects a person as a specific object. In the case of this example, the region division unit 102 in the third embodiment uses an object detector that detects a person to detect a person from the image, and assigns a local region to each detected person. In the case of the example in FIG. 13(b), the region division unit 102 detects the persons 1301a and 1302a from the image 1300, assigns a rectangular local region 1301b to the person 1301a, and similarly assigns a rectangular local region 1302b to the person 1302a. Then, in step S403, the region division unit 102 divides the reference image and the comparison image by the local regions 1301b and 1302b.

[0070] Next, in step S405, the similarity acquisition unit 103 calculates the similarity between the local regions 1301b at the same coordinate position in the reference image and the comparison image, and similarly calculates the similarity between the local regions 1302b at the same coordinate position. And when the calculation of the similarity for all the local regions for each person detected from the reference image and the comparison image is completed, in step S406, the output content determination unit 104 determines the screen output content based on the calculated similarities. Note that in the examples of FIGS. 13(a) and 13(b), an object detector that detects a person as a specific object is taken as an example, but for example, an object detector that detects a vehicle such as a car or an object detector that detects a building on a road may be used.

[0071] FIG. 13(c) and FIG. 13(d) are explanatory diagrams of an example of performing a division process into local regions for each semantic region by a semantic region divider. FIG. 13(c) is a diagram showing an example of an image 1310 to be divided into local regions, and the image 1310 is a reference image and a comparison image. It is assumed that the image 1310 includes, for example, a person 1311c, a road 1312c, and a background 1313c.

[0072] FIG. 13(d) is a diagram showing an example in which a person 1311c, a road 1312c, and a background 1313c are detected from the image 1310 using a semantic region divider. In the case of this example, the region division unit 102 of the third embodiment uses a semantic region divider to detect a semantic region from the image and assigns a local region to each of the detected semantic regions. In the case of FIG. 13(d), the region division unit 102 assigns a local region 1311d to the person 1311c detected from the image 1310, a local region 1312d to the road 1312c, and a local region 1313d to the background 1313c, respectively. Then, in step S403, the region division unit 102 divides the reference image and the comparison image by the local region 1311d, the local region 1312d, and the local region 1313d.

[0073] Next, in step S405, the similarity acquisition unit 103 calculates the similarity between the local regions 1311d at the same coordinate positions in the reference image and the comparison image, and similarly calculates the similarity between the local regions 1312d and the similarity between the local regions 1313d. When the calculation of the similarity for all the local regions for each semantic region detected from the reference image and the comparison image is completed, in step S406, the output content determination unit 104 determines the screen output content based on the calculated similarities.

[0074] In the third embodiment, the local regions can be determined by dividing an image for each object or semantic region in the image, so that the user does not need to arbitrarily specify the size or number of local regions, and the user's workload is reduced. In addition, according to the third embodiment, the searchability for local regions is improved by dividing the image for each object or semantic region. For example, if the user wants to improve the accuracy of the NN model especially for people, the image selection unit 110 is made to acquire local region images and similarities focusing on people, and the images and similarities are displayed on the screen of the display unit 140, thereby reducing the burden of subjective evaluation by the user. In this embodiment, an example is given in which the user himself sets the objects and semantic regions to be detected, but it is also possible to set these settings in advance in the image selection unit 110, in which case the user does not need to perform the setting work.

[0075] <Other embodiments> The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions. The above-mentioned embodiments are merely examples of the implementation of the present invention, and the technical scope of the present invention should not be interpreted as being limited by these. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.

[0076] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) A division means for dividing a reference image and a comparison image obtained by processing another image of the same scene as the reference image using a learning model into local regions; a similarity obtaining means for obtaining a similarity between local regions at the same coordinate position of the reference image and the comparison image; Output content determination means for determining the screen output content when displaying the local regions of the reference image and the comparison image based on the similarity; Result acquisition means for acquiring the user's evaluation result for the displayed local region according to the screen output content; Selection means for selecting an image to be used for re-training the learning model based on the user's evaluation result; An information processing apparatus characterized by comprising the same. (Configuration 2) The splitting means splits the reference image and the comparison image into a plurality of the local regions, The similarity acquisition means acquires the similarity for each local region at the same coordinate position of the reference image and the comparison image, The information processing apparatus according to Configuration 1, wherein the output content determination means determines the screen output content including regions in which the plurality of local regions are arranged in an order based on the similarity. (Configuration 3) The information processing apparatus according to Configuration 2, wherein the output content determination means determines the screen output content including regions in which the plurality of local regions are arranged in an order of decreasing similarity, an order of increasing similarity, or an order specified by the user. (Configuration 4) The information processing apparatus according to any one of Configurations 1 to 3, wherein the selection means selects, as an image to be used for re-training the learning model, an image similar to the image of the local region evaluated by the user as requiring re-training. (Configuration 5) The information processing apparatus according to any one of Configurations 1 to 4, wherein the other image is an image obtained by photographing the same scene under conditions different from those of the reference image, or an image obtained by subjecting the reference image to predetermined image processing. (Configuration 6) The division means uses, as the local area, an area based on a user's designation, an object area at the same coordinate position as a result of performing object detection on the reference image and the comparison image, or a semantic area at the same coordinate position as a result of performing semantic area division on the reference image and the comparison image. The information processing apparatus according to any one of Configurations 1 to 5, characterized in that. (Configuration 7) The similarity acquisition means acquires the similarity using image feature amounts. The information processing apparatus according to any one of Configurations 1 to 6, characterized in that. (Configuration 8) The image feature amounts used by the similarity acquisition means are at least any of pixel values of luminance, lightness, and chroma. The information processing apparatus according to Configuration 7, characterized in that. (Configuration 9) The similarity acquisition means converts the image feature amounts into feature vectors and acquires the similarity. The information processing apparatus according to Configuration 8, characterized in that. (Configuration 10) The output content determination means determines the screen output content by arranging local area comparison information having at least the images of the local areas of the reference image and the comparison image, the value of the similarity, and an area for receiving a selection instruction by the user. The information processing apparatus according to any one of Configurations 1 to 9, characterized in that. (Configuration 11) The output content determination means restricts the number of the local area comparison information arranged in the screen output content based on the similarity. The information processing apparatus according to Configuration 10, characterized in that. (Configuration 12) The output content determination means determines the number of the local area comparison information arranged in the screen output content according to the number of the local areas where the similarity is less than a predetermined similarity threshold value. The information processing apparatus according to Configuration 11, characterized in that. (Configuration 13) The output content determination means determines the number of the local area comparison information arranged in the screen output content based on a predetermined display ratio with respect to the number of all the local areas. The information processing apparatus according to Configuration 10, characterized in that. (Configuration 14) The output content determination means determines the number of local area comparison information to be arranged in the screen output content based on the statistical information of the similarity obtained for each local area, in the information processing apparatus according to Configuration 10. (Configuration 15) The output content determination means determines the number of local area comparison information to be arranged in the screen output content based on the number of local areas less than a predetermined similarity threshold in the histogram of the order based on the similarity, in the information processing apparatus according to Configuration 10. (Configuration 16) The information processing apparatus according to any one of Configurations 1 to 15, further comprising learning means for re-learning the learning model using the image selected by the selection means. (Configuration 17) When the number of local areas where the similarity is less than a predetermined similarity threshold is further less than a predetermined number, the selection means selects an image for the learning means to use for re-learning the learning model based on the images of the local areas less than the predetermined number, in the information processing apparatus according to Configuration 16. (Method 1) A dividing step of dividing a reference image and a comparison image, which is an image obtained by processing another image of the same scene as the reference image with a learning model, into local areas; A similarity acquisition step of acquiring the similarity between local areas at the same coordinate positions of the reference image and the comparison image; An output content determination step of determining the screen output content when displaying the local areas of the reference image and the comparison image based on the similarity; A result acquisition step of acquiring a user's evaluation result for the displayed local areas according to the screen output content; A selection step of selecting an image to be used for re-learning the learning model based on the user's evaluation result; An information processing method, characterized by comprising: (Program 1) A program for causing a computer to function as the information processing apparatus according to any one of Configurations 1 to 17.

Explanation of Signs

[0077] 100: Information processing device, 110: Image selection unit, 120: Learning processing unit, 101: Image acquisition unit, 102: Region division unit, 103: Similarity acquisition unit, 104: Output content determination unit, 105: Screen output unit, 106: Evaluation result acquisition unit, 107: Similar data acquisition unit, 121: Data storage unit, 122: Model learning unit, 124: Inference unit, 125: Model storage unit

Claims

1. A dividing means for dividing a reference image and a comparison image obtained by processing another image of the same scene as the reference image with a learning model into local regions; A similarity acquisition means for acquiring the similarity between local regions at the same coordinate positions of the reference image and the comparison image; An output content determination means for determining the screen output content when displaying the local regions of the reference image and the comparison image based on the similarity; A result acquisition means for acquiring the evaluation result of the user with respect to the displayed local regions according to the screen output content; A selection means for selecting an image to be used for re-learning the learning model based on the evaluation result of the user; An information processing apparatus characterized by comprising the above.

2. The dividing means divides the reference image and the comparison image into a plurality of the local regions, The similarity acquisition means acquires the similarity for each local region at the same coordinate position of the reference image and the comparison image, The information processing apparatus according to claim 1, wherein the output content determination means determines the screen output content including a region in which the plurality of local regions are arranged in an order based on the similarity.

3. The information processing apparatus according to claim 2, wherein the output content determination means determines the screen output content including a region in which the plurality of local regions are arranged in an order of decreasing similarity, an order of increasing similarity, or an order designated by the user.

4. The information processing apparatus according to claim 1, wherein the selection means selects, as an image to be used for re-learning the learning model, an image similar to the image of the local region evaluated by the user as requiring re-learning.

5. The information processing apparatus according to claim 1, wherein the another image is an image obtained by photographing the same scene under conditions different from those of the reference image, or an image obtained by performing predetermined image processing on the reference image.

6. The information processing apparatus according to claim 1, wherein the dividing means uses, as the local regions, a region based on a user's designation, an object region at the same coordinate position as a result of performing object detection on the reference image and the comparison image, or a semantic region at the same coordinate position as a result of performing semantic region division on the reference image and the comparison image.

7. The information processing apparatus according to claim 1, wherein the similarity acquisition means acquires the similarity using image feature amounts.

8. The information processing apparatus according to claim 7, wherein the image feature amount used by the similarity acquisition means is at least one of luminance, lightness, chroma pixel values, and pixel values.

9. The information processing apparatus according to claim 8, wherein the similarity acquisition means converts the image feature amount into a feature vector to acquire the similarity.

10. The information processing apparatus according to claim 1, wherein the output content determination means determines the screen output content by arranging local region comparison information having at least the images of the local regions of the reference image and the comparison image, the value of the similarity, and a region for receiving a selection instruction by the user.

11. The information processing apparatus according to claim 10, wherein the output content determination means limits the number of the local region comparison information arranged in the screen output content based on the similarity.

12. The information processing apparatus according to claim 11, wherein the output content determination means determines the number of the local region comparison information arranged in the screen output content according to the number of the local regions where the similarity is less than a predetermined similarity threshold value.

13. The information processing apparatus according to claim 10, wherein the output content determination means determines the number of the local region comparison information arranged in the screen output content based on a predetermined display ratio with respect to the number of all the local regions.

14. The information processing apparatus according to claim 10, wherein the output content determination means determines the number of the local region comparison information arranged in the screen output content based on statistical information of the similarity acquired for each local region.

15. The information processing apparatus according to claim 10, wherein the output content determination means determines the number of the local region comparison information arranged in the screen output content based on the number of the local regions less than a predetermined similarity threshold value in a histogram of the order based on the similarity.

16. The information processing apparatus according to claim 1, further comprising a learning means for re-learning the learning model using the image selected by the selection means.

17. The information processing apparatus according to claim 16, wherein when the number of the local regions where the similarity is less than a predetermined similarity threshold value is further less than a predetermined number, the selection means selects an image for the learning means to use for re-learning the learning model based on the images of the local regions less than the predetermined number.

18. A dividing step of dividing a reference image and a comparison image obtained by processing another image in the same scene as the reference image with a learning model into local regions; A similarity acquisition step of acquiring the similarity between local regions at the same coordinate positions of the reference image and the comparison image; An output content determination step of determining the screen output content when displaying the local regions of the reference image and the comparison image based on the similarity; A result acquisition step of acquiring the evaluation result of the user for the displayed local region according to the screen output content; A selection step of selecting an image to be used for re-training the learning model based on the evaluation result of the user; An information processing method, characterized by comprising the above.

19. A computer, A dividing means for dividing a reference image and a comparison image obtained by processing another image in the same scene as the reference image with a learning model into local regions; A similarity acquisition means for acquiring the similarity between local regions at the same coordinate positions of the reference image and the comparison image; An output content determination means for determining the screen output content when displaying the local regions of the reference image and the comparison image based on the similarity; A result acquisition means for acquiring the evaluation result of the user for the displayed local region according to the screen output content; A selection means for selecting an image to be used for re-training the learning model based on the evaluation result of the user; A program for causing the above to function as an information processing apparatus.

Citation Information

Patent Citations

  • Similar area retrieval method, similar area retrieval device, and similar area retrieval program

    JP2010122931A