Image processing apparatus and its control method, imaging apparatus, program

The image processing apparatus addresses the challenge of selecting suitable machine learning dictionaries by generating and displaying characteristic training images, enhancing user understanding and performance in camera applications.

JP7894308B2Active Publication Date: 2026-07-23CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2022-11-25
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Users of cameras equipped with multiple machine learning dictionaries face difficulty in selecting a preferred dictionary due to the lack of methods to grasp the characteristics of each dictionary, such as differences in detecting specific subjects like horses or similar species like zebras.

Method used

An image processing apparatus that generates information about the characteristics of machine learning dictionaries by selecting and displaying representative and special training images associated with each dictionary, allowing users to understand their suitability for specific scenarios.

Benefits of technology

Enables users to grasp the characteristics of machine learning dictionaries in advance, facilitating informed selection for optimal performance in different photography contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894308000004
    Figure 0007894308000004
  • Figure 0007894308000005
    Figure 0007894308000005
  • Figure 0007894308000006
    Figure 0007894308000006
Patent Text Reader

Abstract

To enable generation of information for allowing users to grasp characteristics of a dictionary created through machine learning.SOLUTION: An image processing device disclosed herein is configured to acquire a dictionary created through machine learning and multiple learning images that were used for the machine learning of the dictionary, select one or more learning images to be used as information indicative of characteristics of the dictionary from among the acquired multiple learning images, and generate association information associating the selected one or more learning images and the dictionary.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing apparatus, a control method thereof, an imaging apparatus, and a program. [Background technology]

[0002] Cameras equipped with pre-programmed machine learning dictionaries (hereinafter referred to as ML dictionaries) are being sold by manufacturers. These types of cameras can use ML dictionaries to detect various subjects, such as people, dogs, and horses, from captured images. Furthermore, with the widespread adoption of machine learning technology, it has been proposed to equip cameras with multiple ML dictionaries. Regarding configurations using multiple ML dictionaries, Patent Document 1 discloses switching between multiple ML dictionaries based on past detection results using the ML dictionaries for a series of images. Patent Document 2 discloses analyzing the detection results for each of the multiple ML dictionaries and changing the analysis time for detection using each ML dictionary according to its respective accuracy. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-132369 [Patent Document 2] Japanese Patent Publication No. 2022-039667 [Non-patent literature]

[0004] [Non-Patent Document 1] S. Haykin, "Neural Networks: A Comprehensive Foundation, 2nd Edition", Prentice Hall, pp. 156–255, July 1998 (referenced in the embodiment). [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] However, it is not assumed that the user selects and uses a preferred ML dictionary from among a plurality of ML dictionaries mounted on the camera. When the user selects a preferred dictionary from a plurality of ML dictionaries, it is desirable that the user can grasp the characteristics of the selected ML dictionary in advance. For example, even an ML dictionary capable of detecting a horse may have differences in characteristics among a plurality of ML dictionaries, such as an ML dictionary suitable for detecting a special posture and an ML dictionary suitable for detecting similar species such as zebras. Conventionally, no method has been proposed for expressing the performance of an ML dictionary, and when a user tries to select and use one from a plurality of ML dictionaries, it may be difficult to grasp the characteristics of the ML dictionary in advance.

[0006] One aspect of the present invention aims to provide a configuration capable of generating information for allowing a user to grasp the characteristics of a dictionary obtained by machine learning.

Means for Solving the Problems

[0007] An image processing apparatus according to one aspect of the present invention has the following configuration. That is, This is a dictionary that outputs information about the subjects depicted in the input image. an acquisition unit that acquires a dictionary obtained by machine learning and a plurality of learning images used for the machine learning of the dictionary; a selection unit that selects one or more learning images used as information representing the characteristics of the dictionary from the plurality of learning images; and a generation unit that generates association information for associating the one or more learning images selected by the selection unit with the dictionary. The system comprises a detection means for detecting a subject from an image using one of several dictionaries obtained by machine learning, the selection means for selecting one or more training images based on a plurality of feature vectors obtained from the plurality of training images, and the selection means for obtaining intermediate data obtained from the detection means as the plurality of feature vectors by inputting the plurality of training images into the detection means. .

Effects of the Invention

[0008] According to one aspect of the present invention, information for allowing a user to grasp the characteristics of a dictionary obtained by machine learning can be generated.

Brief Description of the Drawings

[0009] [Figure 1A] Block diagram of an imaging apparatus in the first embodiment. [Figure 1B]Block diagram showing a functional configuration example of a dictionary switching unit in an imaging device. [Figure 2] Schematic diagram showing a configuration example of a CNN in the subject detection unit of the first embodiment. [Figure 3] Schematic diagram showing an example of a partial configuration of a CNN in the first embodiment. [Figure 4] Schematic diagram showing a configuration example of a CNN in the subject classification unit of the first embodiment. [Figure 5] Flowchart showing the switching process of the ML dictionary according to the first embodiment. [Figure 6] Flowchart showing the linking information generation process according to the first embodiment. [Figure 7A] Flowchart showing the target image selection process according to the first embodiment. [Figure 7B] Flowchart showing the target image selection process according to the first embodiment. [Figure 8] Schematic diagram showing an example of the display of dictionary characteristic expressions in the first embodiment. [Figure 9] Flowchart showing the target image selection process according to the second embodiment. [Figure 10] Flowchart showing the linking information generation process according to the third embodiment. [Figure 11] Diagram showing an example of the subject classification result according to the third embodiment. [Figure 12] Flowchart showing the target image selection process according to the third embodiment. [Figure 13] Schematic diagram showing an example of the display of dictionary characteristic expressions according to the third embodiment. [Figure 14] Block diagram showing a configuration example of an image processing system according to the fourth embodiment. [Figure 15] Block diagram showing a configuration example of a cloud system according to the fourth embodiment. [Figure 16A] Flowchart showing the process according to the fourth embodiment. [Figure 16B] Flowchart showing the process according to the fourth embodiment. [Figure 16C] Flowchart showing the process according to the fourth embodiment. [Modes for carrying out the invention]

[0010] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0011] <First Embodiment> (Configuration of the imaging device) Figure 1A is a block diagram showing an example configuration of an imaging device 100 according to the first embodiment. The imaging device 100 photographs a subject and records video and still image data onto various media such as tape, solid memory, optical disk, and magnetic disk. Examples of imaging devices 100 include digital still cameras and video cameras, but are not limited to these. Each part of the imaging device 100 is connected via a bus 160. Each part is controlled by a CPU 151 (Central Processing Unit). The imaging device 100 incorporates an image processing unit consisting of an image processing unit 152, an image compression / decompression unit 153, a subject detection unit 162, a subject classification unit 163, a dictionary switching unit 164, and the like.

[0012] The lens unit 101 comprises a fixed 1-group lens 102, a zoom lens 111, an aperture 103, a fixed 3-group lens 121, and a focus lens 131. The aperture control unit 105 adjusts the aperture diameter of the aperture 103 by driving the aperture 103 via the aperture motor 104 (AM) according to the command of the CPU 151, thereby adjusting the amount of light during shooting. The aperture control unit 105 performs exposure control using the brightness value of a specific subject area. The zoom control unit 113 changes the focal length by driving the zoom lens 111 via the zoom motor 112 (ZM). The focus control unit 133 determines the amount to drive the focus motor 132 (FM) based on the amount of deviation in the focus direction of the lens unit 101. The focus control unit 133 controls the focus adjustment state by driving the focus lens 131 via the focus motor 132 (FM). By controlling the movement of the focus lens 131 by the focus control unit 133 and the focus motor 132, AF control for a specific subject area is realized, for example. The focus lens 131 is a lens used for adjusting the focus, and although it is simply shown as a single lens in Figure 1A, it is usually composed of multiple lenses.

[0013] The image sensor 141 is a photoelectric conversion element that performs photoelectric conversion to convert the optical image of a subject into an electrical signal. The image sensor 141 has m pixels in the horizontal direction and n pixels in the vertical direction. The subject image formed on the image sensor 141 via the lens unit 101 is converted into an electrical signal by the image sensor 141. The image formed on the image sensor 141 and converted photoelectrically is processed into an image signal (image data) by the imaging signal processing unit 142 and acquired as an image on the imaging surface. The image data output from the imaging signal processing unit 142 is sent to the imaging control unit 143 and temporarily stored in the RAM 154 (random access memory). The image data stored in the RAM 154 is compressed by the image compression / decompression unit 153 and then recorded on the recording medium 157. In parallel with this, the image data stored in the RAM 154 is sent to the image processing unit 152.

[0014] The imaging control unit 143 receives instructions from the CPU 151 regarding the storage time of the image sensor 141 and the gain setting value when outputting from the image sensor 141 to the imaging signal processing unit 142, and controls the image sensor 141. The CPU 151 sets the storage time and gain setting value based on instructions from the operator input via the operation switch 156, or based on the magnitude of the pixel signal of the image data temporarily stored in the RAM 154.

[0015] The image processing unit 152 processes the image signal and performs operations such as scaling the image data to the optimal size and calculating the similarity between image data. The image data processed to the optimal size is sent to the display 150 as needed and displayed to enable preview image display and through image display. In addition, the subject detection results of the subject detection unit 162 can be superimposed on the image display on the display 150. For example, the subject detection results may be displayed as a rectangle on the display 150. Furthermore, by using the RAM 154 as a ring buffer, multiple image data captured within a predetermined period and the detection results of the subject detection unit 162 corresponding to each image data can be buffered. Similarly, image data used for training the subject detection unit 162 and the subject detection results corresponding to the image data can be buffered. The image processing unit 152 also performs gamma correction and white balance processing based on the subject area.

[0016] The operation switch 156 is an input interface that includes a touch panel and buttons. Various operations can be performed by selecting various function icons displayed on the display 150 using the operation switch 156. For example, the user can select images to be used for machine learning and perform operations necessary for machine learning (for example, specifying a two-dimensional ground truth region corresponding to an image) while viewing the captured images displayed on the display 150. In addition, the user can select ML dictionaries to download and give instructions for downloading them while viewing the GUI of the cloud system, which is acquired via the communication unit 161 and displayed on the display 150.

[0017] The recording medium 157 is a recording medium such as an SD card, and records image data output from the imaging signal processing unit 142, multiple ML dictionaries applicable to the subject detection unit 162 and the subject classification unit 163, etc. The ML dictionary applied to the subject detection unit 162 also records the training images used for each machine learning process and the association information described later. The ML dictionary applied to the subject classification unit 163 also records icon images corresponding to each classification and class of the subject classification result. The communication unit 161 connects to a cloud system, etc. via Ethernet or wireless and communicates information related to the ML dictionary, such as the ML dictionary and training images.

[0018] The subject detection unit 162 applies an ML dictionary selected from multiple ML dictionaries recorded on the recording medium 157 to determine the region where a subject, such as a horse, exists using the image signal. The subject detection process in the subject detection unit 162 is realized by feature extraction processing using a DNN (Deep Neural Network). The configuration of the subject detection unit 162 will be described in detail later.

[0019] The subject classification unit 163 applies the ML dictionary recorded on the recording medium 157 and uses the image signal to classify which of the predetermined classes the subject belongs to. The classification process in the subject classification unit 163 is realized by feature extraction processing using DNN (Deep Neural Networks). By switching the ML dictionary applied as appropriate, the subject classification unit 163 can perform various multi-class classifications such as classification of subject species, classification of posture, classification of presence or absence of ornaments, and classification of fur color. The configuration of the subject classification unit 163 will be described in detail later.

[0020] The dictionary switching unit 164 switches the ML dictionary used by the subject detection unit 162 to an ML dictionary selected from multiple ML dictionaries in response to a predetermined user operation on the operation switch 156. Figure 1B is a block diagram showing an example of the functional configuration of the dictionary switching unit 164. The switching control unit 201 receives user operations from the operation switch 156 and controls each functional unit of the dictionary switching unit 164, as well as controlling the display on the display 150 during dictionary switching operations. The characteristic acquisition unit 202 provides the characteristic representation of the ML dictionary to the switching control unit 201 based on the linking information recorded on the recording medium 157. The linking information generation unit 203 selects a learning image to be used for characteristic representation from a plurality of learning images used for machine learning of the ML dictionary, generates linking information that links the ML dictionary and the selected learning image, and records it on the recording medium 157. Details of the dictionary switching unit 164 will be described later.

[0021] Returning to Figure 1A, the battery 159 is managed by the power management unit 158, providing a stable power supply to the entire imaging device 100. The flash memory 155 stores the control program necessary for the operation of the imaging device 100, as well as parameters used for the operation of each part. When the imaging device 100 is started by user operation (when it transitions from the power-off state to the power-on state), the control program and parameters stored in the flash memory 155 are loaded into a part of the RAM 154. The CPU 151 controls the operation of the imaging device 100 according to the control program and parameters loaded into the RAM 154.

[0022] (Configuration of the subject detection unit 162) In this embodiment, the subject detection unit 162 is configured as a CNN (Convolutional Neural Network), but it is not limited to this; any DNN using machine learning techniques can be considered an embodiment of this disclosure. The basic configuration of the CNN will be explained using Figures 2 and 3. Figure 2 shows the basic configuration of a CNN that detects a subject from input 2D image data. In Figure 2, the left end is the input, and processing proceeds to the right. A CNN consists of two layers called a feature detection layer (S layer) and a feature integration layer (C layer), which are arranged hierarchically as a set.

[0023] In a CNN, the S layer first detects the next feature based on the features detected in the previous layer. The features detected in the S layer are then integrated in the C layer, and the detection results for that layer are sent to the next layer. The S layer consists of feature detection cell surfaces, and each feature detection cell surface detects a different feature. The C layer consists of feature integration cell surfaces, and the detection results from the previous feature detection cell surfaces are pooled. Hereafter, unless there is a need to distinguish between them, the feature detection cell surfaces and feature integration cell surfaces will be collectively referred to as feature surfaces. In this embodiment, the output layer, which is the final stage layer, does not use the C layer and consists only of the S layer.

[0024] The details of the feature detection process on the feature detection cell surface and the feature integration process on the feature integration cell surface will be explained using Figure 3. The feature detection cell surface is composed of multiple feature detection neurons, which are connected to the preceding layer C in a predetermined structure. The feature integration cell surface is composed of multiple feature integration neurons, which are connected to the same layer S in a predetermined structure. In Figure 3, within the M-th cell surface of the L-th layer S, the output value of the feature detection neuron at position (ξ,ζ) is y LS M (ξ,ζ), within the M-th cell plane of the L-th layer C, the output value of the feature-integrating neuron at position (ξ,ζ) is y LC M Let (ξ,ζ) be denoted as such. At that time, the connection coefficient of each neuron is w LS M (n,u,v), w LCM If we set it as (u, v), each output value can be expressed as follows in [Equation 1] and [Equation 2].

[0025]

Equation

[0026]

Equation

[0027] In [Equation 1], f is an activation function, and it can be any sigmoid function such as the logistic function or the hyperbolic tangent function. For example, it can be realized by the tanh function. u LS M (ξ, ζ) is the internal state of the feature detection neuron at the position (ξ, ζ) on the M-th cell surface of the S layer in the L-th layer. Equation 2 takes a simple linear sum without using an activation function. When not using an activation function as in [Equation 2], the internal state u of the neuron LC M (ξ, ζ) and the output value y LC M (ξ, ζ) are equal. Also, y in [Equation 1] L-1C n (ξ + u, ζ + v), y in [Equation 2] LS M (ξ + u, ζ + v) are respectively called the output values of the connection destinations of the feature detection neuron and the feature integration neuron.

[0028] We will explain ξ, ζ, u, v, and n in [Equation 1] and [Equation 2]. The position (ξ, ζ) corresponds to the position coordinates in the input image. For example, y LS MA high output value for (ξ,ζ) indicates a high probability that the feature to be detected exists at the pixel position (ξ,ζ) of the input image in the M-th cell plane of the L-th layer S. In [Equation 2], n represents the n-th cell plane of the L-1-th layer C, and is called the integration target feature number. Basically, a sum-of-products operation is performed on all cell planes present in the L-1-th layer C. (u,v) is the relative position coordinate of the connection coefficient, and a sum-of-products operation is performed within a finite range (u,v) depending on the size of the feature to be detected. This finite range of (u,v) is called the receptive field. The size of the receptive field is referred to as the receptive field size below and is expressed as the number of horizontal pixels × the number of vertical pixels in the connected range.

[0029] Also, in [Equation 1], L=1, that is, in the very first S layer, y L-1C n (ξ+u,ζ+v) represents the input image y in_image (ξ+u,ζ+v) or input position map y in_posi_map The equation becomes (ξ+u, ζ+v). Incidentally, the distribution of neurons and pixels is discrete, and the feature numbers of the connected nodes are also discrete, so ξ, ζ, u, v, and n are not continuous variables but take discrete values. Here, ξ and ζ are non-negative integers, n is a natural number, and u and v are integers, all of which are within a finite range.

[0030] [Number 1] w LS M (n,u,v) is a coupling coefficient distribution for detecting a given feature, and by adjusting this to an appropriate value, it becomes possible to detect the given feature. This adjustment of the coupling coefficient distribution is learning, and in the construction of a CNN, various test patterns are presented, and y LS M The coupling coefficients are adjusted by repeatedly and gradually correcting them so that (ξ,ζ) becomes an appropriate output value.

[0031] Next, w in [Mathematics 2] LC M (u,v) is a two-dimensional Gaussian function and can be expressed as shown in [Equation 3] below.

[0032]

number

[0033] Here, (u,v) is given as a finite range, so, similar to the explanation of feature detection neurons, the finite range is called the receptive field, and the size of the range is called the receptive field size. This receptive field size can be set to an appropriate value according to the size of the Mth feature in the Lth layer S. In [Equation 3], σ is the feature size factor, and it can be set to an appropriate constant according to the receptive field size. Specifically, it is best to set it to a value such that the outermost value of the receptive field can be considered to be approximately 0.

[0034] In this embodiment, the subject detection unit 162 is configured to perform subject detection in the final S layer of the subject detection unit 162 by performing the above-described calculations at each layer.

[0035] (Learning method for subject detection unit 162) The specific learning method for the subject detection unit 162 will now be explained. In this embodiment, the coupling coefficients are adjusted by supervised learning. In supervised learning, a test pattern is provided to actually obtain the output value of the neuron, and the coupling coefficients w are determined from the relationship between that output value and the teacher signal (the desired output value that the neuron should output). LS M The only thing that needs to be corrected is (n,u,v). In the learning process of this embodiment, the least squares method is used for the feature detection layer of the final layer, and the backpropagation method is used for the feature detection layer of the intermediate layer to correct the coupling coefficients. For details on coupling coefficient correction methods such as the least squares method and the backpropagation method, please refer to Non-Patent Literature 1.

[0036] The subject detection unit 162 prepares numerous test patterns for learning, including specific patterns to be detected and patterns that should not be detected. Each test pattern consists of one set of image and training signal. When the tanh function is used as the activation function, when a specific pattern to be detected is presented, a training signal is given to the neurons in the region of the feature detection cell plane of the final layer where the specific pattern exists, so that the output is 1. Conversely, when a pattern that should not be detected is presented, a training signal is given to the neurons in the region of that pattern, so that the output is -1. In actual detection, the connection coefficient w constructed through learning is used. LS M The calculation is performed using (n,u,v), and if the neuron output on the feature detection cell surface of the final layer is greater than or equal to a predetermined value, it is determined that a subject is present there. In this way, the subject detection unit 162 is constructed to be able to detect a subject from a two-dimensional image.

[0037] (Configuration of subject classification unit 163) In this embodiment, the subject classification unit 163 is configured as a CNN, but it is not limited to this configuration; any DNN using machine learning techniques can be used as an embodiment of this disclosure. The basic configuration of the CNN will be explained using Figure 4. In the CNN of the subject classification unit 163, the number of outputs k in the n layers is configured to be the number of classes to be classified. The softmax function is used as the activation function. Other parts are the same as the configuration in the subject detection unit 162 and will be omitted. By performing the above-described calculations at each layer, subject classification is performed in the final S layer of the subject classification unit 163; this is the CNN configuration of the subject classification unit 163 in this embodiment.

[0038] (Learning method for subject classification unit 163) The specific learning method for the subject classification unit 163 will now be explained. The subject classification unit 163 prepares a large number of specific patterns to be classified into each of the many classes. Each test pattern consists of one set of image and training signal. When using the softmax function as the activation function, the training signal is provided so that the output for the correct class is 1 and the output for classes other than the correct class is 0. In addition, in actual classification, the coupling coefficient w constructed through learning is used. LS M The calculation is performed using (n,u,v), and it is determined that the subject belongs to a class where the neuron output on the feature detection cell surface of the final layer is greater than or equal to a predetermined value.

[0039] Multiple ML dictionaries, each created by performing the aforementioned learning process for multiple multi-class classifications, are prepared in advance and recorded on the recording medium 157. The rest of the process is the same as in the subject detection unit 162 and is therefore omitted. As a result, the subject classification unit 163 is constructed to perform multi-class classification of subjects from a two-dimensional image.

[0040] (Overall processing flow of dictionary switching unit 164) Figure 5 is a flowchart showing the overall processing of the dictionary switching unit 164 in this embodiment. Note that each functional unit of the dictionary switching unit 164 shown in Figure 1B may be implemented by the CPU 151 executing predetermined software, by dedicated hardware, or by the cooperation of software and hardware.

[0041] In S500, the switching control unit 201 displays an ML dictionary switching screen on the display 150 in response to a predetermined user operation (ML dictionary switching operation) using the operation switch 156. A specific example of the switching screen will be described later with reference to Figure 8. In S501, the switching control unit 201 determines whether the user has performed an operation to switch the ML dictionary displayed on the switching screen via the operation switch 156. If the switching control unit 201 determines that an operation to switch the ML dictionary has been performed (YES in S501), the process proceeds to S502. If the switching control unit 201 determines that an operation to switch the ML dictionary has been performed (NO in S501), the process proceeds to S505.

[0042] In S502, the characteristic acquisition unit 202 determines whether the association information for the new ML dictionary to be displayed, which was switched in S501, is recorded in the recording medium 157. If the characteristic acquisition unit 202 determines that the association information for the ML dictionary to be displayed is not recorded in the recording medium 157 (NO in S502), the process proceeds to S503. On the other hand, if the characteristic acquisition unit 202 determines that the association information for the ML dictionary to be displayed is recorded in the recording medium 157 (YES in S502), the process proceeds to S504.

[0043] In S503, the association information generation unit (hereinafter referred to as the generation unit) 203 reads the ML dictionary to be displayed after the display switching operation in S501 and the multiple training images used to train this ML dictionary from the recording medium 157, and generates association information for the ML dictionary. The generation unit 203 records the obtained association information on the recording medium 157. Here, the association information is information that represents a pair of an ML dictionary and a target image selected to represent the characteristics of that ML dictionary, and includes the target image itself. Details of the association information generation process in S503 will be described later with reference to Figures 6, 7A, and 7B.

[0044] Next, in S504, the characteristic acquisition unit 202 reads the association information related to the ML dictionary to be displayed after the switching operation in S501 from the recording medium 157, and acquires the learning images associated with the read association information. The switching control unit 201 displays the ML dictionary to be displayed and the multiple learning images acquired by the characteristic acquisition unit 202 on the display 150, thereby representing the characteristics of the ML dictionary to be displayed after the switch. Details of the display on the display 150 and the method of representing the characteristics of the ML dictionary will be described later with reference to Figure 8.

[0045] In S505, the switching control unit 201 determines whether the user has performed an operation to switch the ML dictionary to be applied to the subject detection unit 162 via the operation switch 156 for the ML dictionary currently displayed on the switching screen. If the switching control unit 201 determines that an operation to switch the ML dictionary has been performed (YES in S505), the process proceeds to S506. On the other hand, if the switching control unit 201 determines that an operation to switch the ML dictionary has not been performed (NO in S505), the process proceeds to S507. In S506, the switching control unit 201 switches the ML dictionary applied to the subject detection unit 162 to the ML dictionary currently displayed on the switching screen, according to the switching operation from the user. From this point onward, subject detection processing using the ML dictionary switched in S506 becomes possible for the through image and captured image from the imaging device 100. In S507, the switching control unit 201 determines whether the user has performed a termination operation via the operation switch 156. If the switching control unit 201 determines that a termination operation has been performed (YES in S507), this process ends. If the switching control unit 201 determines that the termination operation has not been performed (NO in S507), the process returns to S501 and the above process is repeated.

[0046] (Flow of the process for generating linked information) Figure 6 is a flowchart of the process for generating linking information, which is executed in S503 of the first embodiment.

[0047] In S600, the generation unit 203 calculates feature vectors for each of the multiple training images (hereinafter referred to as the "relevant image group") used for machine learning of the ML dictionary to be displayed after switching in S500. These relevant image groups are recorded on the recording medium 157 corresponding to the ML dictionary. It is desirable that the relevant image group includes all the training images used for machine learning of the ML dictionary, but it may also consist of training images randomly selected from all the training images used for machine learning, as long as the effects of this disclosure are not impaired. In S600, known methods for calculating feature vectors from images can be used. Alternatively, each training image may be input to the subject detection unit 162, and the output of the feature integration layer n-1, which is intermediate data obtained by the subject detection unit 162, may be used as the feature vector. Alternatively, each training image may be input to the subject classification unit 163 to which any of the ML dictionaries recorded on the recording medium 157 is applied, and the output of the feature integration layer n-1, which is intermediate data obtained by the subject classification unit 163, may be used as the feature vector.

[0048] In S601, the generation unit 203 calculates an average vector using the feature vector calculated in S600. In S602, the generation unit 203 determines whether the first target image has been selected. If the generation unit 203 determines that the first target image has not been selected (NO in S602), the process proceeds to S603. If the generation unit 203 determines that the first target image has been selected (YES in S602), the process proceeds to S604.

[0049] In S603, the generation unit 203 selects the first target image from the group of images. Details of the process for selecting the first target image will be described later with reference to the flowchart in Figure 7A. In S604, the generation unit 203 selects the second and subsequent target images from the group of images. Details of the process for selecting the second target image will be described later with reference to the flowchart in Figure 7B. Next, in S605, the generation unit 203 determines whether the termination condition is met. If the generation unit 203 determines that the termination condition is met (YES in S605), the process proceeds to S606. On the other hand, if the generation unit 203 determines that the termination condition is not met (NO in S605), the process returns to S602, and the above process is repeated. In this embodiment, the number of selected target images is used as the termination condition in S605. The number of selected target images is set considering the size that the user can see when the learning images are displayed on the display 150 in the display that expresses the characteristics of the ML dictionary (S504). In this embodiment, for example, there are four images (one target image for the first time and three target images for the second and subsequent times). In S606, the generation unit 203 records information that identifies the learning images selected in S603 and S604 on the recording medium 157 as association information relating to the ML dictionary to be displayed.

[0050] (Target image selection process) Figure 7A is a flowchart showing the process of selecting the first target image (S603). The generation unit 203 selects the training image with the feature vector that has the smallest distance from the average vector of the multiple feature vectors obtained from the group of images (multiple training images) as the first target image.

[0051] In S700, the generation unit 203 selects a learning image of interest from the group of images. In S701, the generation unit 203 determines whether the learning image of interest selected in S700 has already been selected as a target image. If the generation unit 203 determines that the learning image of interest has already been selected as a target image (YES in S701), the process returns to S700 and the next learning image of interest is selected. On the other hand, if the generation unit 203 determines that the learning image of interest has not already been selected as a target image (NO in S701), the process proceeds to S701. Since the selection of target images is performed repeatedly in the flowchart shown in Figure 6, the determination in this step is intended to prevent the selection of target images multiple times. Note that, as shown in Figure 6, if only one target image is selected using the process in S603, S701 can be omitted. However, if the process selects multiple target images in order of proximity to the mean vector (for example, if multiple target images are selected using the process in S603), S701 is required.

[0052] In S702, the generation unit 203 determines whether the focus learning image selected in S700 satisfies the conditions for being subject to linking processing. If the generation unit 203 determines that the focus learning image satisfies the conditions for being subject to linking processing (YES in S702), the process proceeds to S703. On the other hand, if the generation unit 203 determines that the focus learning image does not satisfy the conditions for being subject to linking processing (NO in S702), the process returns to S700. In machine learning, even if an image is included in the learning images, it may not have a sufficient impact on the performance of the ML dictionary. Therefore, as one of the conditions for being subject to linking processing (conditions for being selected as a target image), a condition is set to prevent learning images that have little impact on the characteristics of the ML dictionary from being selected as target images. More specifically, one of the conditions for being subject to linking processing is that when the focus learning image is input to the subject detection unit 162, subject detection is possible. At this time, the ML dictionary to be displayed is temporarily set in the subject detection unit 162. Alternatively, the proportion of feature vectors calculated in S600 that have a sufficiently small distance from the feature vector calculated from the training image of interest (feature vectors whose distance is below a predetermined threshold) may be greater than or equal to a predetermined proportion. Furthermore, the distance here can be any numerical value that represents the similarity of feature vectors, such as the Euclidean distance or cosine similarity between vectors.

[0053] In S703, the generation unit 203 calculates the distance between the feature vector calculated from the focus learning image among the feature vectors calculated in S600 and the mean vector calculated in S601. The distance between vectors is the same as the distance explained in S702. In S704, the generation unit 203 determines whether a candidate image has been selected in S705, which will be described later. If the generation unit 203 determines that a candidate image has not been selected (NO in S704), the process proceeds to S705. If the generation unit 203 determines that a candidate image has been selected (YES in S704), the process proceeds to S706.

[0054] In S705, the generation unit 203 selects the target learning image as a candidate image. Meanwhile, in S706, the generation unit 203 determines whether the distance between the feature vector and the mean vector of the target learning image is smaller than the distance between the feature vector and the mean vector of the candidate image selected in S705 or S707. If the generation unit 203 determines that the distance between the feature vector and the mean vector of the target learning image is smaller than the distance between the feature vector and the mean vector of the candidate image (YES in S706), the process proceeds to S707. On the other hand, if the generation unit 203 determines that the distance between the feature vector and the mean vector of the target learning image is not smaller than the distance between the feature vector and the mean vector of the candidate image (NO in S706), the process proceeds to S708. The determination in S706 is intended to select a learning image as a candidate image in which a feature vector with a smaller distance from the mean vector can be obtained. In S707, the generation unit 203 changes the candidate image to the current target learning image.

[0055] In S708, the generation unit 203 determines whether there are any training images in the group of images that have not been selected as the training images of interest. If the generation unit 203 determines that there are training images that have not been selected as the training images of interest (YES in S708), the process returns to S700 and the above process is repeated. On the other hand, if the generation unit 203 determines that there are no training images that have not been selected as the training images of interest (NO in S708), the process proceeds to S709. In S709, the generation unit 203 selects the training images of interest, which are candidate images, as the target images.

[0056] According to the processing shown in Figure 7A, the training image from which the feature vector with the smallest distance from the mean vector is obtained is selected as the target image. In other words, a representative training image that represents the performance of the ML dictionary for which a switching instruction was given in S501 is selected.

[0057] Figure 7B is a flowchart showing the process of selecting the second and subsequent target images, which is executed in S604 of the first embodiment. In this process, the generation unit 203 selects a predetermined number of feature vectors from a plurality of feature vectors obtained from the group of images, in descending order of their distance from the mean vector. Therefore, by repeating S604 a predetermined number of times, a predetermined number of training images are selected in descending order of the distance between the feature vectors and the mean vector. In Figure 7B, the processes from S710 to S719, excluding S716, are the same as the processes from S700 to S709 in Figure 7A, excluding S706. In S716, the generation unit 203 determines whether the distance between the feature vector and the mean vector of the training image of interest is greater than the distance between the feature vector and the mean vector of the candidate image selected in S715 or S717. If the generation unit 203 determines that the distance between the feature vector and the mean vector of the training image of interest is greater than the distance between the feature vector and the mean vector of the candidate image (YES in S716), the process proceeds to S717. On the other hand, if the generation unit 203 determines that the distance between the feature vector and the mean vector of the training image of interest is not greater than the distance between the feature vector and the mean vector of the candidate image (NO in S716), the process proceeds to S718. The determination in S716 is intended to select training images as candidate images from which feature vectors with a larger distance from the mean vector are calculated.

[0058] According to the processing shown in Figure 7B, the training image from which the feature vector with the largest distance from the mean vector is obtained is selected as the target image. In other words, a special training image is selected that represents the performance of the ML dictionary that was switched in S501.

[0059] (Example of dictionary characteristics display) Figure 8 shows an example of a switching screen displayed on the display 150 by the switching control unit 201 as a result of the processing in S500 and S504. When a dictionary switching mode is specified, the switching control unit 201 displays, for example, the switching screen 8a shown in Figure 8(a). Items 800 and 801 are GUIs for switching the ML dictionary to be displayed by operating the operation switch 156. When item 800 or 801 is operated, the switching control unit 201 determines in S501 that an operation to switch the ML dictionary to be displayed has occurred. Item 802a is the ML dictionary to be displayed, and is written together with "001", which is the file name of the ML dictionary recorded on the recording medium 157.

[0060] Area 803a is an area for representing the characteristics of the ML dictionary and is displayed by the processing in S504. In the example in Figure 8(a), training images 804a to 807a, which are target images selected to represent the characteristics of the ML dictionary "001", are displayed. Training image 804a is the training image selected as the first target image in S603. In S603, representative training images are likely to be selected, and in the example of training image 804a, it is a training image of a single horse standing on all fours. Training images 805a, 806a, and 807a are training images selected as the second and subsequent target images in S604. In S604, special training images are likely to be selected, and in the example of training images 805a to 807a, these are training images with a rider, ornaments, and multiple subjects. The switch button 808 is a GUI that accepts user operations to switch the ML dictionary applied to the subject detection unit 162. When the toggle button 808 is operated, it is determined that a switch in the ML dictionary to be applied in S505 has been instructed.

[0061] As described above, the switching control unit 201 functions as a display information generation unit that generates display information for associating a dictionary with one or more learning images based on the association information and displays this information on the display 150. This display information then presents the dictionary characteristics of each dictionary to the user. For example, from the display in Figure 8(a), the user can understand in advance that ML dictionary "001" is suitable for photography in places like horse racing tracks.

[0062] Next, Figure 8(b) will be explained. Items 800 and 801 and the toggle button 808 are the same as in Figure 8(a). Item 802b indicates the ML dictionary that has been selected for display by the operation of item 800 or 801, and is written together with "002", which is the file name of the ML dictionary recorded on the recording medium 157. Area 803b displays training images 804b to 807b, which are target images selected to represent the characteristics of the ML dictionary "002". Training image 804a is an example of a training image selected as the first target image in S603. Training images 805b, 806b, and 807b are examples of training images selected as the second and subsequent target images in S604. In S604, special training images are more likely to be selected, and examples of training images 805b to 807b are characteristic training images such as postures such as lowering the head or lying down, or similar species such as zebras.

[0063] As shown in Figure 8(b) above, the dictionary characteristics display allows the user to understand in advance that the ML dictionary is suitable for photography in places like zoos.

[0064] As described above, according to the first embodiment, training images possessing representative and special features of the ML dictionary are displayed. Therefore, by checking these, the user can grasp the overall dictionary characteristics in advance and select the ML dictionary to apply. In the above embodiment, the total number of target images to be selected was 4, the number of target images (representative training images) selected by the process in S603 (Figure 7A) was 1, and the number of target images (special training images) selected by the process in S604 (Figure 7B) was 3. However, this disclosure is not limited to this, and the total number of target images, the number of representative training images, and the number of special training images can be set arbitrarily. However, since using special training images can represent the characteristics of the ML dictionary more broadly, it is desirable from the viewpoint of grasping the overall characteristics that the number of special training images is greater than the number of representative training images.

[0065] <Second Embodiment> In the first embodiment, the distance between the feature vector and the mean vector of the focus learning image was used to determine whether or not to select the focus learning image as the second or subsequent target image. In the second embodiment, the distance to the feature vector of the first target image (representative learning image) is used to determine whether or not to select the focus learning image as the second or subsequent target image. The configuration, function, and processing of the imaging device 100 in the second embodiment are the same as in the first embodiment, except for the process of selecting the second or subsequent target images. The process of selecting the second or subsequent target images according to the second embodiment (process S604 in Figure 6) will be described below with reference to Figure 9.

[0066] (Selection process for images to be linked) Figure 9 is a flowchart showing the process of selecting the second and subsequent target images, which is performed in S604 of the second embodiment. The processes from S900 to S909, excluding S903 and S906, are the same as the processes from S710 to S719 of the first embodiment (Figure 7B), excluding S713 and S716.

[0067] In S903, the generation unit 203 calculates the distance between the feature vector calculated from the training image of interest and the feature vector calculated from the first target image selected in S603. These feature vectors are calculated in S600. In S906, the generation unit 203 determines whether the distance between the feature vector of the training image of interest and the feature vector of the first target image is greater than the distance between the feature vector of the candidate image selected in S905 or S907 and the feature vector of the first target image. If the generation unit 203 determines that the distance between the feature vector of the training image of interest and the feature vector of the first target image is greater than the distance between the feature vector of the selected candidate image and the feature vector of the first target image (YES in S906), the process proceeds to S907. If the generation unit 203 determines that the distance between the feature vector of the training image of interest and the feature vector of the first target image is not greater than the distance between the feature vector of the selected candidate image and the feature vector of the first target image (NO in S906), the process proceeds to S908. The S906 decision aims to select candidate images for training that produce feature vectors with a large distance from the feature vectors of the first selected image.

[0068] As described above, the second embodiment displays learning images with representative and special features, allowing users to understand the general dictionary characteristics in advance and select the appropriate ML dictionary by reviewing them.

[0069] <Third Embodiment> In the first and second embodiments, a training image (i.e., a training image linked to the ML dictionary) is selected to represent the characteristics of the ML dictionary applied to the subject detection unit 162, based on the feature vectors of the training images. In the third embodiment, a training image is selected to represent the characteristics of the ML dictionary applied to the subject detection unit 162, based on the results of classifying multiple training images used for machine learning of the ML dictionary into multiple classes by the subject classification unit 163. The overall processing of the dictionary switching unit 164 in the third embodiment is the same as in the first and second embodiments (Figure 5), except for the linking information generation process in S503 and the dictionary characteristic display method in S504.

[0070] (Generation of linked information (S503)) Figure 10 is a flowchart showing the linking information generation process according to the third embodiment. First, in S1000, the generation unit 203 selects a target ML dictionary from among a plurality of ML dictionaries (hereinafter referred to as the "relevant ML dictionary group") recorded on the recording medium 157 that are applied to the subject classification unit 163, and applies it to the subject classification unit 163. Alternatively, the relevant ML dictionary group may be selected according to the subject to which the ML dictionary applied to the subject detection unit 162 (the ML dictionary for which a switching instruction was given in S501) corresponds. For example, if the ML dictionary applied to the subject detection unit 162 is a dictionary for detecting "horses", the relevant ML dictionary group may be composed of ML dictionaries related to the classification of "horses". Next, in S1001, the generation unit 203 selects a target learning image from among a plurality of learning images (hereinafter referred to as the "relevant image group") recorded on the recording medium 157, corresponding to the ML dictionary for which a switching instruction was given in S501 (the ML dictionary applied to the subject detection unit 162). The group of images in question here preferably includes all the training images used in the machine learning of the ML dictionary, but randomly selected training images may be used as long as the effects of this disclosure are not impaired. In S1002, the generation unit 203 causes the subject classification unit 163 to process the training images of interest and perform subject classification by applying the training ML dictionary of interest. In S1003, the generation unit 203 determines whether there are any training images in the group of images that were not selected as training images of interest in S1001. If the generation unit 203 determines that there are training images in the group of images that were not selected as training images of interest (YES in S1003), the process returns to S1001 and the above process is repeated. On the other hand, if the generation unit 203 determines that there are no training images in the group of images that were not selected as training images of interest (NO in S1003), the process proceeds to S1004.

[0071] In S1004, the generation unit 203 determines whether there are any ML dictionaries in the group of ML dictionaries that were not selected as the ML dictionary of interest in S1000. If the generation unit 203 determines that there are ML dictionaries in the group of ML dictionaries that were not selected as the ML dictionary of interest (YES in S1004), the process returns to S1000 and the above process is repeated. On the other hand, if the generation unit 203 determines that there are no ML dictionaries in the group of ML dictionaries that were not selected as the ML dictionary of interest (NO in S1004), the process proceeds to S1005. In S1005, the generation unit 203 selects a target image from the training images and icon images recorded on the recording medium 157 to represent the characteristics of the ML dictionary, based on the subject classification results obtained from S1000 to S1004. The selection of the target image will be described later with reference to the flowchart in Figure 12. In S1006, the generation unit 203 records information identifying the target image selected in S1005 on the recording medium 157 as association information for the ML dictionary to be displayed. Here, the association information is information about the combination of target images selected in S1005 corresponding to the ML dictionary to be displayed for which the switching operation was performed in S501, and the target image itself.

[0072] (Example of subject classification results) Figure 11 shows an example of the subject classification results obtained from processing S1000 to S1004 in this embodiment. In the example in Figure 11, five subject classifications are performed on "species," "coat color," "posture," "mask," and "saddle," and the meaning of each class and the number of training images are shown. Examples of the "species" class include "horse," "zebra," "donkey," and "other." Examples of the "coat color" class include "bay," "black," "chestnut," and "other." Examples of the "posture" class include "standing on all fours," "standing on two legs," "lying down," and "sleeping." Examples of the "mask" and "saddle" classes include "present" and "absent."

[0073] (Target image selection process) Figure 12 is a flowchart showing the process of selecting a target image in the third embodiment.

[0074] In S1200, the generation unit 203 selects a focus classification from the subject classification results obtained in S1000 to S1004. The classification here is the classification obtained from the focus ML dictionary selected in S1000, and in the example in Figure 11, it is one of "species", "coat color", "posture", "mask", and "saddle". In S1201, the generation unit 203 selects a focus class from the classes included in the focus classification. For example, in the example in Figure 11, if the focus classification is "species", the focus class is determined from one of "horse", "zebra", "donkey", or "other".

[0075] In S1202, the generation unit 203 determines whether the class of interest satisfies the conditions for being linked. The condition here is that the proportion of training images belonging to the class of interest in the total number of training images is equal to or greater than a predetermined value. For example, if the predetermined value is 5%, in the example in Figure 11, the classes "zebra," "donkey," and "other" in the "species" classification are determined not to satisfy the condition. In CNN training, even if an image is included in the training images, it may not have a sufficient impact on the performance of the CNN. Therefore, the determination in S1202 is intended to prevent the selection of training images related to classes with little impact as target images. If the generation unit 203 determines that the class of interest satisfies the conditions for being linked (YES in S1202), the process proceeds to S1203. On the other hand, if the generation unit 203 determines that the class of interest does not satisfy the conditions for being linked (NO in S1202), the process proceeds to S1205.

[0076] In S1203, the generation unit 203 determines the training images classified as the focus class from the group of images as target images. Any training image classified as the focus class can be the target image. For example, in Figure 11, if the focus classification is "breed" and the focus class is "horse," any of the 49,000 training images classified as that class can be used. In S1204, the generation unit 203 further selects an icon image corresponding to the focus class selected in S1201 from the icon images recorded on the recording medium 157 as an additional target image. If icon images corresponding to multiple classes have been selected as target images for the focus classification, an icon image encompassing all of those classes may be selected. For example, in the example in Figure 11, if the focus classification is "coat color," all classes "bay," "black," "chestnut," and "other colored coats" satisfy the conditions for being linked. Therefore, instead of the icon images corresponding to these individual classes, an icon image representing the class of all coat colors may be selected.

[0077] In S1205, the generation unit 203 determines whether there are any classes that have not been selected as focus classes with respect to the classification determined as focus classification in S1200. If the generation unit 203 determines that there are classes that have not been selected as focus classes (YES in S1205), the process returns to S1201 and the above process is repeated. On the other hand, if the generation unit 203 determines that there are no classes that have not been selected as focus classes (NO in S1205), the process proceeds to S1206. In S1206, the generation unit 203 determines whether there are any classifications that have not been selected as focus classifications from the subject classification results obtained from S1000 to S1004. If the generation unit 203 determines that there are classes that have not been selected as focus classifications (YES in S1206), the process returns to S1200 and the above process is repeated. On the other hand, if the generation unit 203 determines that there are no classes that have not been selected as focus classifications (NO in S1206), this process ends.

[0078] As described above, according to the third embodiment, learning images and icon images corresponding to all classifications and all classes that satisfy the conditions for linking are selected as target images. Since icon images can represent many dictionary characteristics within the narrow display range of the display 150, it is desirable that all icon images that satisfy the conditions for linking are selected as target images. However, if the display range is limited, the determination in S1206 may be changed to a method that determines whether or not the termination condition is met, as in the determination in S605 of the first and second embodiments. In this case, it is desirable to select target images prioritizing classifications and classes that satisfy predetermined criteria. For example, in the process of selecting the first target image, it is desirable to prioritize classifications with many classes that satisfy the conditions in S1202, and further to prioritize classes that have many classified learning images among the classes belonging to that classification. Also, in the process of selecting the second and subsequent target images, it is desirable to prioritize classifications with few classes that satisfy the conditions for processing in S1202, and further to prioritize classes that have few classified learning images among the classes belonging to that classification. According to these criteria, when selecting the first target image, representative images (training images and icon images) are more likely to be selected, while when selecting subsequent target images, unusual images (training images and icon images) are more likely to be selected.

[0079] (Example of dictionary characteristics display) The dictionary characteristics can be displayed using the training image determined as the target image in S1203 described above. The display example in this case is the same as in the first embodiment (Figure 8). On the other hand, in the third embodiment, the dictionary characteristics can also be displayed using the icon image determined as the target image in S1204. Figure 13 is a diagram showing an example of the switching screen displayed on the display 150 by the switching control unit 201 in S504 of the third embodiment. Figure 13 shows an example of a display in which dictionary characteristics are represented using an icon image according to the third embodiment. Note that the display in which dictionary characteristics are represented using a training image as in Figure 8 and the display in which dictionary characteristics are represented using an icon image as in Figure 13 may be switched by the user at will.

[0080] First, Figure 13(a) will be explained. In the switching screen 13a, items 800, 801, and the switching button 808 are the same as in the first embodiment (Figure 8). Item 1002a is the ML dictionary to be displayed, and is written together with "001", which is the file name of the ML dictionary recorded on the recording medium 157. Area 1003a is an area that expresses the dictionary characteristics, and displays multiple icon images that express the dictionary characteristics, represented by icon image 1004a. Icon image 1004a is an example of an icon image selected as the target image in S1204. In the target image selection process in the third embodiment, since the determination of icon images corresponding to all classifications and all classes is performed without imposing constraints on the number of images, the characteristics of the dictionary can be expressed more comprehensively. Figure 13(b) is an example of the switching screen 13b in which the dictionary characteristics are displayed according to the third embodiment for a different ML dictionary than the switching screen 13a. The icon image 1004b for "All Postures" in area 1003a is an icon image that collectively represents all classes in Figure 11: "Quadruped Standing," "Two-Legged Standing," "Prone," and "Sleeping." Note that the icon image can be any image that allows the user to understand the class, such as a string representation of the class as shown in Figure 13, or a graphic or photograph representing the class.

[0081] As described above, in the third embodiment, multiple icon images are displayed, and by reviewing them, the user can grasp the overall dictionary characteristics in advance and select the ML dictionary to apply.

[0082] <Fourth Embodiment> (Overall structure) In the first to third embodiments, a configuration was described in which the characteristic representation of the ML dictionary is acquired by an image processing device (subject detection unit 162, subject classification unit 163, dictionary switching unit 164, etc.) within the imaging device 100. In the fourth embodiment, a configuration is described in which the characteristic representation of the ML dictionary is acquired by an external image processing device (e.g., a cloud system) and provided to the imaging device. Figure 14 is a diagram showing an example of the overall configuration of the image processing system 1400 according to the fourth embodiment. The image processing system 1400 comprises an imaging device 1401, an imaging device 1402, a cloud system 1403, and a network 1404. The imaging devices 1401, 1402, and 1403 are each connected to each other so as to be communicative via the network 1404. The detailed configurations of the imaging devices 1401 and 1402 are the same as those of the imaging device 100 in the first embodiment. However, the dictionary switching unit 164 does not need to have the function of acquiring a target image for representing the characteristics of the ML dictionary. The cloud system 1403 is an example of an image processing device that can communicate with imaging devices 1401 and 1402. Furthermore, the image processing device as an external device to the imaging device is not limited to the cloud system 1403, and may be implemented, for example, by a server device on a LAN.

[0083] The imaging device 1401 is configured similarly to the imaging device 100, and the user takes training images and performs machine learning on an ML dictionary applicable to the subject detection unit 162. The ML dictionary obtained by machine learning and the multiple training images used in machine learning (hereinafter referred to as the "relevant image group") are uploaded to the cloud system 1403 via the network 1404. It is desirable that the "relevant image group" here includes all the training images used in machine learning, but randomly selected training images may be used as long as the effects obtained by the embodiment of this disclosure are not impaired. The imaging device 1402 has the same configuration as the imaging device 100 and can utilize the environment provided by the cloud system 1403. For example, the imaging device 1402 downloads an ML dictionary from the cloud system 1403 and applies it to the subject detection unit 162 to utilize the subject detection function on the captured images. The cloud system 1403 records multiple ML dictionaries uploaded by the user and multiple training images for each ML dictionary. It also expresses the performance of the ML dictionaries and provides an environment in which the user can download multiple ML dictionaries. Note that while Figure 14 shows, for convenience, an imaging device 1401 that uploads one file at a time and an imaging device 1402 that downloads files, it is not limited to this configuration. For example, multiple ML dictionaries may be uploaded to the client system 1403 from multiple imaging devices.

[0084] (Cloud system configuration) Figure 15 is a block diagram showing an example configuration of the cloud system 1403 according to the fourth embodiment. Each functional unit of the cloud system 1403 is connected via the bus 1500. The control unit 1501 controls each functional unit. The recording unit 1502 is a large-capacity recording medium such as an HDD, and records multiple ML dictionaries uploaded from the imaging device 1401 and multiple ML dictionaries applicable to the subject classification unit 1506. In addition, the training images used for each machine learning and the associated information are also recorded in the multiple ML dictionaries uploaded from the imaging device 1401. Furthermore, the ML dictionary applicable to the subject classification unit 1506 also records icon images corresponding to each classification and class of the subject classification result. The communication unit 1503 connects to the imaging device 1401 and imaging device 1402 via Ethernet or wireless and communicates information related to the ML dictionaries, such as ML dictionaries and training images.

[0085] The display image generation unit 1504 generates a display image and provides the imaging device 1401 and its user with a GUI for uploading and downloading ML dictionaries via the communication unit 1503. The subject detection unit 1505 applies the ML dictionary recorded in the recording unit 1502 to determine the region in the image where a subject such as a horse exists. The subject detection unit 1505 is assumed to be composed of the same CNN as the subject detection unit 162 in the imaging device 1401 and imaging device 1402. The subject classification unit 1506 applies the ML dictionary recorded in the recording unit 1502 to classify which of the predetermined classes the subjects in the image belong to. The subject classification unit 1506 is assumed to be composed of the same CNN as the subject classification unit 163 in the imaging device 1401 and imaging device 1402.

[0086] (Overall processing flow) Figures 16A to 16C are flowcharts showing the overall processing of the cloud system 1403 according to the fourth embodiment. The processing of each of the imaging device 1401, the cloud system 1403, and the imaging device 1402 will be explained using Figures 16A, 16B, and 16C.

[0087] Figure 16A is a flowchart showing the overall processing of the imaging device 1401 according to the fourth embodiment. In S1600, the switching control unit 201 determines whether the user has performed an upload operation of the ML dictionary via the operation switch 156. If the switching control unit 201 determines that an upload operation has been performed (YES in S1600), the process proceeds to S1601. If the switching control unit 201 determines that an upload operation has not been performed (NO in S1600), the process proceeds to S1603.

[0088] In S1601, the switching control unit 201 sends an upload request to the cloud system 1403, an external device, via the communication unit 161. Next, in S1602, the switching control unit 201 transmits the ML dictionary recorded on the recording medium 157 and the multiple training images used to train the ML dictionary to the cloud system 1403 via the communication unit 161. In S1603, the switching control unit 201 determines whether there is a termination instruction from the user via the operation switch 156. If the switching control unit 201 determines that there is no termination instruction (NO in S1603), the process returns to S1600 and the above process is repeated. On the other hand, if the switching control unit 201 determines that there is a termination instruction (YES in S1603), this process ends.

[0089] Figure 16B is a flowchart showing the overall processing of the cloud system 1403 according to the fourth embodiment. First, in S1610, the control unit 1501 determines, via the communication unit 1503, whether there is an upload request for the ML dictionary from an external device (in this example, the imaging device 1401). If the control unit 1501 determines that there is an upload request (YES in S1610), the process proceeds to S1611. If the control unit 1501 determines that there is no upload request (NO in S1610), the process proceeds to S1613.

[0090] In S1611, the control unit 1501 receives the ML dictionary and multiple training images used to train the ML dictionary from the upload requesting imaging device 1401 via the communication unit 1503, and records them in the recording unit 1502. In S1612, the control unit 1501 reads the ML dictionary and multiple training images used to train the ML dictionary recorded in the recording unit 1502 in S1611 and generates association information. The details of the association information generation process in S1612 are the same as in the first or second embodiment (S503), except that it is performed by various parts of the cloud system 1403.

[0091] In S1613, the control unit 1501 determines, via the communication unit 1503, whether there is a request to display ML dictionary characteristics from an external device (in this example, the imaging device 1402). If the control unit 1501 determines that there is a request to display (YES in S1613), the process proceeds to S1614. If the control unit 1501 determines that there is no request to display (NO in S1613), the process returns to S1610. In S1614, the display image generation unit 1504 generates a characteristic representation image that represents the characteristics of the ML dictionary based on the association information generated in S1612, the ML dictionary, and a plurality of training images corresponding to the ML dictionary. The display image generation unit 1504 then transmits the generated characteristic representation image to the imaging device 1402, the source of the display request, via the communication unit 1503. The details of the process in S1614 are the same as in the first or second embodiment (S504), except that it is executed by various parts of the cloud system 1403.

[0092] In S1615, the control unit 1501 determines, via the communication unit 161, whether a download request for an ML dictionary has been received from an external device (in this example, the imaging device 1402). If the control unit 1501 determines that a download request has been received (YES in S1615), the process proceeds to S1616. If the control unit 1501 determines that no download request has been received (NO in S1615), the process returns to S1610. In S1616, the control unit 1501 transmits, via the communication unit 1503, the ML dictionary whose dictionary characteristics were expressed in S1614, and the association information generated for that ML dictionary in S1612, from among the multiple ML dictionaries recorded in the recording unit 1502, to the imaging device 1402 that made the download request.

[0093] Figure 16C is a flowchart showing the overall processing of the imaging device 1402 in the fourth embodiment. In S1620, it is determined whether the user has performed a display operation for the dictionary characteristics via the operation switch 156. If the switching control unit 201 determines that a display operation has been performed (YES in S1620), the process proceeds to S1621. If the switching control unit 201 determines that there has been no display operation (NO in S1620), the process proceeds to S1623.

[0094] In S1621, the switching control unit 201 transmits a request to display ML dictionary characteristics to the cloud system 1403, an external device, via the communication unit 161. In S1622, the switching control unit 201 receives a characteristic representation image of the ML dictionary from the cloud system 1403 via the communication unit 161 and displays it on the display 150. The characteristic representation image is similar to the display example of the switching screen shown in the first embodiment (Figure 8).

[0095] In S1623, the switching control unit 201 determines, via the operation switch 156, whether the user has requested a download operation for the ML dictionary whose characteristics are represented in the characteristic representation image. If the switching control unit 201 determines that a download operation has occurred (YES in S1623), the process proceeds to S1624. If the switching control unit 201 determines that there has been no download operation (NO in S1623), the process proceeds to S1626. In S1624, the switching control unit 201 sends a download request for the ML dictionary to the cloud system 1403, an external device, via the communication unit 161. In S1625, the switching control unit 201 receives the ML dictionary and the associated information (e.g., the target image) generated for the ML dictionary from the cloud system 1403 via the communication unit 161, and records them on the recording medium 157.

[0096] In S1626, the switching control unit 201 determines whether the user has performed an operation via the operation switch 156 to switch the ML dictionary recorded on the recording medium 157 in S1625 to be applied to the subject detection unit 162. If the switching control unit 201 determines that the switching operation has been performed (YES in S1626), the process proceeds to S1627. If the switching control unit 201 determines that the switching operation has not been performed (NO in S1626), the process proceeds to S1628. In S1627, the switching control unit 201 switches the ML dictionary applied to the subject detection unit 162 based on the switching operation from the user. From this point onward, subject detection processing using the ML dictionary switched in S1627 becomes possible for the through image and captured image from the imaging device 1402. In S1628, the switching control unit 201 determines whether the user has given a termination instruction via the operation switch 156. If the switching control unit 201 determines that a termination instruction has been given (YES in S1628), this process ends. If the switching control unit 201 determines that no termination instruction has been given (NO in S1628), the process returns to S1620 and the above process is repeated.

[0097] As described above, according to the fourth embodiment, the imaging device can acquire and display training images from the cloud system that possess representative features and special features as information representing the characteristics of the ML dictionary to be displayed. Therefore, by checking these, the user of the imaging device can grasp the general characteristics of the dictionary in advance without burdening the imaging device, and then download and apply the ML dictionary.

[0098] In the above, the configuration for generating the linking information in S1612 and acquiring the characteristic representation image in S1614 (S1622) is based on the configuration of the first or second embodiment, but it is not limited to these. The configuration of the third embodiment may also be used for generating the linking information in S1612 and acquiring the characteristic representation image in S1614 (S1622). In this case, the linking information generation process by S1612 is the same as in the third embodiment (Figures 10 and 12), except that it is executed by each part of the cloud system 1403. Also, the characteristic representation image of the ML dictionary, which is generated by the cloud system 1403 in S1614, received via the communication unit 161 of the imaging device 1402, and displayed in S1622, is the same as the example shown in Figure 13 in the third embodiment.

[0099] According to the fourth embodiment described above, icon images are displayed by the cloud system, and by checking them, the user can understand the overall dictionary characteristics in advance without placing a load on the imaging device, and then download and apply the ML dictionary.

[0100] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0101] The disclosures herein include the following image processing apparatus and control methods thereof, imaging apparatus, and programs. (Item 1) An acquisition means for acquiring a dictionary obtained by machine learning and a plurality of training images used in the machine learning of the dictionary, A selection means for selecting one or more learning images from the aforementioned plurality of learning images to be used as information representing the characteristics of the dictionary, An image processing apparatus comprising: a generation means for generating association information that links one or more learning images selected by the selection means with the dictionary. (Item 2) The image processing apparatus according to item 1, wherein the selection means selects one or more training images based on a plurality of feature vectors obtained from the plurality of training images. (Item 3) The system further includes a detection means for detecting an object from an image using one of several dictionaries obtained through machine learning. The image processing apparatus according to item 2, wherein the selection means obtains intermediate data obtained from the detection means as the plurality of feature vectors by inputting the plurality of training images into the detection means. (Item 4) The image processing apparatus according to item 2 or 3, wherein the selection means selects a training image of the feature vector that has the smallest distance from the mean vector of the plurality of feature vectors. (Item 5) The image processing apparatus according to any one of items 2 to 4, wherein the selection means selects a predetermined number of feature vectors from the plurality of feature vectors in descending order of their distance from the mean vector of the plurality of feature vectors, and selects a predetermined number of training images corresponding to the predetermined number of feature vectors. (Item 6) The image processing apparatus according to item 4, wherein the selection means selects a predetermined number of feature vectors from the plurality of feature vectors in order of the magnitude of the difference with the feature vector having the smallest distance from the mean vector, and selects a predetermined number of training images corresponding to the predetermined number of feature vectors. (Item 7) The image processing apparatus according to item 3, wherein the selection means selects one or more learning images from among learning images that satisfy the condition that the detection means can detect a subject from the learning images using the dictionary. (Item 8) The image processing apparatus according to any one of items 2 to 7, wherein the selection means selects one or more learning images from among the learning images that satisfy the condition that other learning images having a feature vector whose distance from the acquired feature vector is less than a predetermined value exist in a predetermined proportion or more among the plurality of learning images. (Item 9) The image processing apparatus according to any one of items 1 to 8, further comprising a display information generation means for generating display information for associating the dictionary with one or more learning images based on the aforementioned association information. (Item 10) The system further comprises a classification means for classifying the aforementioned multiple learning images into multiple classes. The image processing apparatus according to item 1, characterized in that the selection means selects a training image based on the classification result obtained by inputting the plurality of training images into the classification means. (Item 11) The image processing apparatus according to item 10, wherein the selection means selects a learning image from among the plurality of classes that belongs to the class in which the most learning images were classified in the classification result. (Item 12) The image processing apparatus according to item 10 or 11, wherein the selection means selects a predetermined number of learning images from a predetermined number of classes selected in order from the plurality of classes in order of the smallest number of learning images classified in the classification result. (Item 13) The image processing apparatus according to any one of items 10 to 12, wherein the selection means selects one or more learning images from learning images belonging to a class in which the ratio of the number of classified learning images to the number of the plurality of learning images is equal to or greater than a predetermined value. (Item 14) The selection means further selects one or more icon images from the multiple icon images corresponding to the multiple classes, each of the one or more training images corresponding to the class to which it belongs. The image processing apparatus according to any one of items 10 to 13, characterized in that the generation means includes in the linking information information that links the dictionary and one or more icon images. (Item 15) The image processing apparatus according to item 14, further comprising a display information generation means for generating display information for associating the dictionary with one or more learning images or one or more icon images based on the aforementioned association information. (Item 16) An acquisition means for acquiring a dictionary obtained by machine learning and a plurality of training images used in the machine learning of the dictionary, A classification means for classifying the aforementioned multiple training images into multiple classes, A selection means that selects one or more icon images from a plurality of icon images corresponding to the plurality of classes, based on the classification results of the plurality of learning images by the classification means, to be used as information representing the characteristics of the dictionary, An image processing apparatus comprising: a generation means for generating association information that links one or more icon images selected by the selection means with the dictionary. (Item 17) A receiving means for receiving the dictionary and the plurality of learning images from an external device, The image processing apparatus according to any one of items 1 to 16, further comprising a transmission means for transmitting the dictionary and the association information to an external device. (Item 18) Imaging means, An image processing device described in any one of items 1 through 16, A switching means for switching the dictionary applied to the detection means for detecting a subject from an image captured by the imaging means, according to user operation, An imaging device comprising: a display means for displaying the characteristics of a dictionary that can be switched by the switching means based on the association information. (Item 19) Imaging means, A communication means for communicating with the image processing device described in item 17, A switching means for switching the dictionary applied to the detection means for detecting a subject from an image captured by the imaging means, according to user operation, An imaging device comprising: a display means for displaying the characteristics of a dictionary to be switched by the switching means based on the association information received from the image processing device by the communication means. (Item 20) An acquisition step of acquiring a dictionary obtained by machine learning and a plurality of training images used in the machine learning of the dictionary, A selection step of selecting one or more learning images from the aforementioned plurality of learning images to be used as information representing the characteristics of the dictionary, A control method for an image processing apparatus, comprising: a generation step of generating association information that links the one or more learning images selected in the selection step with the dictionary. (Item 21) An acquisition step of acquiring a dictionary obtained by machine learning and a plurality of training images used in the machine learning of the dictionary, A classification step of classifying the aforementioned multiple training images into multiple classes, A selection step in which one or more icon images are selected from a plurality of icon images corresponding to the plurality of classes, based on the classification results of the plurality of learning images by the classification step, to be used as information representing the characteristics of the dictionary, A control method for an image processing apparatus, characterized by comprising: a generation step of generating association information that links the one or more icon images selected in the selection step with the dictionary. (Item 22) A program for causing an image processing device to function as one of the means of an image processing device described in any one of items 1 to 17.

[0102] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0103] 100: Imaging device, 150: Display, 151: CPU, 152: Image processing unit, 153: Image compression / decompression unit, 154: RAM, 155: Flash memory, 156: Operation switch, 157: Image recording medium, 158: Power management unit, 159: Battery, 160: Bus, 161: Communication unit, 162: Subject detection unit, 163: Subject classification unit, 164: Dictionary switching unit

Claims

1. A dictionary that outputs information about the subject in an input image, obtained by machine learning, and an acquisition means for acquiring a plurality of training images used in the machine learning of the dictionary, A selection means for selecting one or more learning images from the aforementioned plurality of learning images to be used as information representing the characteristics of the dictionary, A generation means for generating linking information that links one or more learning images selected by the selection means with the dictionary, The system includes a detection means for detecting a subject from an image using one of several dictionaries obtained through machine learning, The selection means selects one or more training images based on a plurality of feature vectors obtained from the plurality of training images. The selection means is characterized by inputting the plurality of training images into the detection means to obtain intermediate data obtained from the detection means as the plurality of feature vectors.

2. The image processing apparatus according to claim 1, wherein the selection means selects a training image of the feature vector that has the smallest distance from the mean vector of the plurality of feature vectors.

3. The image processing apparatus according to claim 1, wherein the selection means selects a predetermined number of feature vectors from the plurality of feature vectors in descending order of their distance from the mean vector of the plurality of feature vectors, and selects a predetermined number of training images corresponding to the predetermined number of feature vectors.

4. The image processing apparatus according to claim 2, wherein the selection means selects a predetermined number of feature vectors from the plurality of feature vectors in order of the magnitude of the difference with the feature vector having the smallest distance from the mean vector, and selects a predetermined number of training images corresponding to the predetermined number of feature vectors.

5. The image processing apparatus according to claim 1, characterized in that the selection means selects one or more learning images from among learning images that satisfy the condition that the detection means can detect a subject from the learning images using the dictionary.

6. The image processing apparatus according to claim 1, wherein the selection means selects one or more learning images from among the learning images that satisfy the condition that other learning images having a feature vector whose distance from the acquired feature vector is less than a predetermined value exist in a predetermined proportion or more among the plurality of learning images.

7. The image processing apparatus according to claim 1, further comprising a display information generation means for generating display information for associating the dictionary with one or more learning images based on the aforementioned linking information.

8. Imaging means, An image processing apparatus according to any one of claims 1 to 7, A switching means for switching the dictionary applied to the detection means for detecting a subject from an image captured by the imaging means, according to user operation, An imaging device comprising: a display means for displaying the characteristics of a dictionary that can be switched by the switching means based on the association information.

9. A dictionary that outputs information about the subject in the input image, obtained by machine learning, and an acquisition step of acquiring a plurality of training images used in the machine learning of the dictionary, A selection step of selecting one or more learning images from the plurality of learning images to be used as information representing the characteristics of the dictionary, A generation step that generates linking information that links the one or more learning images selected in the selection step with the dictionary, The system includes a detection step of detecting a subject from an image using one of several dictionaries obtained by machine learning, In the selection step, one or more training images are selected based on a plurality of feature vectors obtained from the plurality of training images. A control method for an image processing apparatus, characterized in that, in the selection step, the intermediate data obtained in the detection step is acquired as the multiple feature vectors by executing the detection step with the multiple learning images as input.

10. A program for causing an image processing device to function as one of the means of an image processing device described in any one of claims 1 to 7.