Image processing device and image processing method

JP2023171053A5Active Publication Date: 2025-05-19CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022083265
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-19
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Existing learning techniques for facial recognition models suffer from accuracy bias due to imbalanced data sampling frequencies in mini-batch learning, leading to reduced learning proficiency for certain attributes.

Method used

A method that adjusts the sampling probability of images based on the learning proficiency level of the learning target, ensuring that images from individuals with lower proficiency are sampled more frequently to improve learning efficiency and reduce bias.

Benefits of technology

This approach enables more balanced learning by increasing the sampling of images from individuals with lower proficiency levels, thereby reducing attribute bias and enhancing the accuracy of facial recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a technique for enabling learning of a learning model based on sampling images sampled on the basis of sampling probabilities of images of learning targets according to learning proficiency levels of the learning targets.SOLUTION: An image processing device is configured to: obtain a learning proficiency level of a learning target in a learning model; obtain a sampling probability of each image of the plurality of learning targets on the basis of the proficiency levels of the plurality of learning targets; and learn the learning model on the basis of sampling images sampled from images of the plurality of learning targets on the basis of the obtained sampling probabilities.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning technique for a learning model. [Background technology]

[0002] In recent years, accuracy bias in face recognition due to attributes such as race and gender has become a problem. Accuracy bias is strongly influenced by imbalances in training data. Mini-batch learning, which is the primary training method for learning models, reduces the sampling frequency of attribute data with small amounts of data, making it difficult to improve the learning proficiency of such attributes. As a result, accuracy bias occurs depending on the amount of data for each attribute. The technology disclosed in Non-Patent Document 1 reduces accuracy bias by adjusting the loss function according to the learning proficiency for each data. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Xingkun Xu et al.Consistent Instance False Positive Improves Fairness in Face Recognition.In CVPR 2021. Summary of the Invention [Problem to be solved by the invention]

[0004] However, the technology disclosed in Non-Patent Document 1 cannot resolve the difference in sampling frequency in mini-batch learning that is caused by accuracy bias. The present invention provides a technology that enables learning of a learning model based on sampled images sampled based on the sampling probability of images of the learning subject according to the learning proficiency of the learning subject. [Means for solving the problem]

[0005] One aspect of the present invention is characterized by comprising a first acquisition means for acquiring the learning proficiency of a learning subject in a learning model, a second acquisition means for acquiring a sampling probability of each image of a plurality of learning subjects based on the proficiency of the plurality of learning subjects, and a learning means for learning the learning model based on sampled images sampled from the images of the plurality of learning subjects based on the sampling probability acquired by the second acquisition means. [Effects of the Invention]

[0006] According to the present invention, it is possible to enable learning of a learning model based on sampled images sampled based on the sampling probability of images of the learning subject according to the learning proficiency of the learning subject. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 2 is a block diagram showing an example of the hardware configuration of the system. [Figure 2] FIG. 2 is a block diagram showing an example of the functional configuration of a learning device. [Figure 3] 10 is a flowchart of a learning process for a learning model. [Figure 4] 10 is a detailed flowchart of the process in step S102. [Figure 5] 10 is a detailed flowchart of the process in step S103. [Figure 6] FIG. 2 is a block diagram showing an example of the functional configuration of an inference device. [Figure 7] 10 is a flowchart of a face authentication process. [Figure 8] FIG. 2 is a block diagram showing an example of the functional configuration of a learning device. [Figure 9] 10 is a flowchart of a learning process for a learning model. [Figure 10] 10 is a flowchart showing details of the process in step S503. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0009] [First embodiment] In this embodiment, an example of an image processing device will be described, which acquires the learning proficiency of a learning object in a learning model, acquires a sampling probability of each image of the plurality of learning objects based on the proficiency of the plurality of learning objects, and performs learning of the learning model based on sampled images sampled from the images of the plurality of learning objects based on the sampling probability of each image of the plurality of learning objects. Also, in this embodiment, a case will be described in which the learning object is a human face.

[0010] First, an example of the hardware configuration of a system according to this embodiment will be described using the block diagram of Fig. 1. As shown in Fig. 1, the system according to this embodiment includes an image processing device 100 and an imaging device 112, and is configured so that the image processing device 100 and the imaging device 112 can communicate data with each other via a wired and / or wireless network 111. In addition, an input device 109 and a monitor 110 are connected to the image processing device 100.

[0011] First, the imaging device 112 will be described. The imaging device 112 captures a moving image within a range corresponding to its own imaging direction and angle of view, and transmits the images of each frame in the moving image as captured images to the image processing device 100 via the network 111. Note that the imaging device 112 may also capture still images periodically or irregularly, and transmit the captured still images as captured images to the image processing device 100 via the network 111.

[0012] Next, a description will be given of the image processing device 100. The image processing device 100 can be applied to a computer device such as a PC (personal computer), a smartphone, or a tablet terminal device.

[0013] The CPU 101 executes various processes using computer programs and data stored in the ROM 102 and the RAM 103. As a result, the CPU 101 controls the overall operation of the image processing device 100, and also executes or controls various processes that will be described as being performed by the image processing device 100.

[0014] The ROM 102 stores setting data for the image processing device 100, computer programs and data related to the startup of the image processing device 100, computer programs and data related to the basic operations of the image processing device 100, and the like.

[0015] The RAM 103 has an area for storing computer programs and data loaded from the ROM 102 or the external storage device 104, and an area for storing captured images received from the imaging device 112 via the communication I / F 107. The RAM 103 also has a work area used by the CPU 101 when executing various processes. In this way, the RAM 103 can provide various areas as needed.

[0016] The external storage device 104 is a large-capacity information storage device (non-volatile memory) such as a hard disk drive. The external storage device 104 stores an OS (operating system), computer programs and data for causing the CPU 101 to execute or control various processes described as being performed by the image processing device 100. The computer programs and data stored in the external storage device 104 are loaded into the RAM 103 as appropriate under the control of the CPU 101, and become targets for processing by the CPU 101.

[0017] The external storage device 104 may include an optical disk such as a flexible disk (FD) or compact disk (CD) that is detachable from the image processing device 100, a magnetic or optical card, an IC card, a memory card, or the like.

[0018] An input device 109 is connected to the I / F 105. The input device 109 is a user interface such as a keyboard, a mouse, or a touch panel, and can be operated by a user to input various instructions to the CPU 101.

[0019] A monitor 110 is connected to the I / F 106. The monitor 110 has a liquid crystal screen or a touch panel screen, and can display the processing results of the CPU 101 as images, text, etc. Furthermore, if the monitor 110 has a touch panel screen, it can accept operational inputs such as touches and swipes from the user. The operational inputs are notified to the CPU 101.

[0020] The communication I / F 107 is an interface for connecting the image processing device 100 to a network 111 , and the image processing device 100 performs data communication with an image capturing device 112 on the network 111 via the communication I / F 107 .

[0021] The CPU 101, ROM 102, RAM 103, external storage device 104, I / F 105, I / F 106, and communication I / F 107 are all connected to a system bus 108. Note that the configuration shown in Fig. 1 is an example of a configuration applicable to the system according to this embodiment, and can be changed / modified as appropriate.

[0022] An example of the functional configuration of a learning device that learns a learning model is shown in the block diagram of FIG. 2. In this embodiment, a case will be described in which an image processing device 100 is applied to such a learning device. In this embodiment, a case will be described in which each functional unit shown in FIG. 2 is implemented as software (computer program). In the following, the functional units shown in FIG. 2 will be described as the main processing units, but in reality, the functions of the functional units are realized by the CPU 101 executing the computer program corresponding to the functional units. Note that one or more of the functional units shown in FIG. 2 may be implemented as hardware.

[0023] The learning process of a learning model by a learning device will be described with reference to the flowchart in FIG. 3. In this embodiment, a case where Convolutional Neural Networks (CNN) is used as the learning model will be described, but the following description can be similarly applied to other learning models. In addition, in this embodiment, the learning process of a learning model for face recognition using the <representative vector method>, which is well known in the literature "Jiankang Deng, et. al, ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In CVPR, 2019", etc., will be described as an example. The <representative vector method> is a face recognition learning method that sets feature vectors of the face images (face images) of each person included in the learning data and uses these in combination to improve learning efficiency. For example, when the learning data includes face images of n people (n is a natural number equal to or greater than 2), the representative vector is a fully connected layer adjacent to the output layer of the learning model.

[0024]

number

[0025] is a vector that constitutes the representative vector

[0026]

number

[0027] is a representative vector corresponding to person ID=j. Person ID is identification information unique to each person. d is the number of dimensions of the representative vector. The feature vector x obtained (generated) by inputting the face image of person ID=i (image of interest for learning) into the learning model and running the learning model is i teeth

[0028]

number

[0029] In this embodiment, in the learning of the learning model, the feature vector x i and the representative vector W j Distance between vectors based on cosine similarity with

[0030]

number

[0031] For example, the representative vector W i and the feature vector x i The distance between the vectors (i.e., the representative vector W of the same person (person ID = i) i and the feature vector x i (vector distance between

[0032]

number

[0033] takes a smaller value, and the feature vector x of person ID=i i and the representative vector W of person ID=j j The sum of the distances between vectors j

[0034]

number

[0035] The learning model is trained by updating the parameters (weighting coefficients, etc.) of the learning model so that σ takes a larger value.

[0036]

number

[0037] In this embodiment, the learning proficiency of the learning model for a person with person ID=i is calculated based on the number of representative vectors (other than the representative vector of the person with person ID=i) that have a small inter-vector distance with the feature vector obtained from a learning model to which the face image of the person with person ID=i is input.The sampling frequency of each person's face image is then controlled so that the sampling frequency of face images of people with low proficiency is high and the sampling frequency of face images of people with high proficiency is low.The learning model is then trained using the face images of each person sampled according to the controlled sampling frequency.

[0038] In this embodiment, the learning loop of steps S101 to S104 is repeated until a learning termination condition is met. Examples of the learning termination condition include the number of learning attempts reaching a specified number, a specified time having elapsed since the start of learning, the learning error being equal to or less than a specified value, or the amount of change in the learning error being equal to or less than a specified amount. The learning termination condition may be a combination of two or more conditions.

[0039] In step S102, a face image of each person to be used for training the learning model is acquired. In step S103, the learning model is trained using the face image of each person acquired in step S102.

[0040] Details of the process in step S102 will be described with reference to the flowchart in Fig. 4. In step S201, the acquisition unit 202 acquires the proficiency level calculated by the calculation unit 206 in the previous step S302.

[0041] In step S202, the acquisition unit 202 acquires the probability (sampling probability) of sampling a face image of each of the n people from the learning data from the managed "proficiency levels of each of the n people." For example, the acquisition unit 202 acquires the proficiency level of a person with person ID=i as N i Then, the sampling probability of person ID=i is P i is calculated according to the following formula (2).

[0042]

number

[0043] In addition, proficiency level N i The sampling probability P of a person with a higher i (sampling frequency) becomes higher, and proficiency level N i The sampling probability P of a person with a lower i If (sampling frequency) is lower, the sampling probability P i The formula for calculating is not limited to the above formula (2).

[0044] In the first step S201, the acquisition unit 202 may acquire and manage preset values ​​as the proficiency levels of the n people. n Alternatively, a predetermined value (a real number between 0 and 1) may be acquired as .times. ...

[0045] Then, the processes of steps S203 to S206 are repeated a specified number of times (for example, a predetermined batch size number). In step S204, the acquiring unit 202 acquires the sampling probabilities P1 to P n Based on this, one of the person IDs 1 to n is selected. For example, suppose n=3 and the sampling probabilities of P1 to P3 are P1=0.3 (30%), P2=0.2 (20%), and P3=0.5 (50%). In this case, if the person ID=i is (the number of elements of the array x P), as in the array [1, 1, 1, 2, 2, 3, 3, 3, 3, 3], i) array and randomly select one from the array, thereby enabling selection of a person ID based on sampling probability.

[0046] In step S205, the acquisition unit 202 selects one face image from the face image group of the person having the person ID selected in step S204 in the learning data (including face images of n people) read by the acquisition unit 201 from the external storage device 104.

[0047] Next, details of the processing in step S103 above will be described with reference to the flowchart in Fig. 5. In step S301, the generation unit 203 inputs each face image selected in step S205 (face image of the person with the person ID selected in step S204) into a learning model and operates the learning model (forward processing), thereby acquiring a feature vector of the face image. As described above, in this embodiment, a CNN is used as the learning model.

[0048] A CNN extracts abstract information from an input image by repeatedly performing a set of processes—convolution, activation, and pooling—on the input image. A processing unit consisting of convolution, activation, and pooling is often referred to as a layer. There are several well-known activation methods, such as the Rectified Linear Unit (ReLU). There are also several well-known pooling methods, such as max pooling. For example, ResNet, introduced in non-patent literature (K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016), may be used as a CNN structure.

[0049] In step S302, the calculation unit 206 calculates the proficiency level of the person with the person ID selected in step S204 using the feature vectors of each person acquired in step S301 and the representative vectors of each of the n people acquired by the acquisition unit 204 from the external storage device 104. For example, the calculation unit 206 calculates the proficiency level N of person ID=i. i is calculated according to the following formula (3).

[0050]

number

[0051] where T u is the feature vector x i and the representative vector W j is the threshold for determining whether or not s is a neighborhood. I(s) is a function that returns 1 if s is true and 0 if s is false. Note that the following inequality

[0052]

number

[0053] The more j's that satisfy i The smaller the value of and the fewer the number of j that satisfy the inequality, the smaller N i The calculation is not limited to the above formula (3), and any formula may be used as long as it increases the value of .

[0054] Furthermore, the feature vector x i The more representative vectors whose inter-vector distance is less than the threshold, the greater the number of N i The value of becomes smaller, and the feature vector x i The fewer the number of representative vectors whose inter-vector distance is less than the threshold, the smaller the number of N i If the formula is such that the value of N becomes large, what formula can be used? i You may also ask for:

[0055] Then, in the next step S201, the acquisition unit 202 acquires the "proficiency level of the person having the person ID selected in step S204" thus obtained, and updates the proficiency level of the person having the person ID in the proficiency levels of the n people being managed to the acquired proficiency level. This allows the acquisition unit 202 to manage the latest proficiency levels for each of the n people.

[0056] In step S303, the update unit 205 performs a learning process for the learning model by updating the parameters of the learning model using the feature vector acquired in step S301 and the representative vector of the person having the person ID selected in step S204 from the representative vectors of each of the n people acquired by the acquisition unit 204 from the external storage device 104, and also stores the representative vector obtained by the learning in the external storage device 104. As a result, the latest representative vectors for each of the n people are always stored in the external storage device 104.

[0057] As a method for updating the parameters of the learning model, for example, a method using a loss function such as ArcFace disclosed in the literature "Jiankang Deng, et. al, ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In CVPR, 2019" can be applied.

[0058] 5, the processing is described as being executed in the order of steps S302 and S303, but the processing order of steps S302 and S303 is not limited to this. For example, steps S303 and S302 may be executed in that order, or the processing of step S302 and the processing of step S303 may be executed in parallel.

[0059] Next, an example of the functional configuration of an inference device that performs inference of a main task using a learning model trained by a learning device is shown in the block diagram of Fig. 6. In this embodiment, the main task is face recognition, and therefore the inference of the main task involves determining whether a registered face image and a face in an input image are the faces of the same person.

[0060] The face authentication process performed by the inference device will be described with reference to the flowchart in FIG.

[0061] In step S401, the acquisition unit 301 acquires a facial image of the user registered in the external storage device 104 and a captured image captured by the imaging device 112. The method and configuration for acquiring the facial image of the user registered in the external storage device 104 and the captured image captured by the imaging device 112 are not limited to a specific method or configuration.

[0062] For example, the external storage device 104 stores a set of authentication information and a facial image (face image) of each user for each of multiple users. When a user operates the input device 109 to input their authentication information (such as a user ID or password), the acquisition unit 301 checks whether the input authentication information has been registered in the external storage device 104. If the check shows that the information has been registered, the acquisition unit 301 determines that the first authentication has been successful and acquires the facial image associated with the input authentication information from the external storage device 104. Meanwhile, the imaging device 112 is configured to capture an image of the face of the user operating the input device 109. If the first authentication has been successful, the acquisition unit 301 instructs the user to turn their face toward the imaging device 112. The instruction may be notified to the user by displaying a message on the monitor 110 or by voice. After a certain time has elapsed since the user operated the input device 109, the imaging device 112 captures an image, and the acquisition unit 301 acquires the captured image. The method of instructing the image capture device 112 to start capturing an image is not limited to a specific method.

[0063] In step S402, the generation unit 302 operates in the same manner as the generation unit 203, and acquires a feature vector of the face image obtained by inputting the user's face image acquired by the acquisition unit 301 in step S401 into a learning model trained by a learning device and operating the learning model. Similarly, the generation unit 302 acquires a feature vector of the captured image obtained by inputting the captured image acquired by the acquisition unit 301 in step S401 into a learning model trained by a learning device and operating the learning model.

[0064] In step S403, the determination unit 303 calculates the similarity between the feature vector of the face image acquired in step S402 and the feature vector of the captured image acquired in step S402. For example, the determination unit 303 calculates the reciprocal of the Euclidean distance between the feature vector of the face image acquired in step S402 and the feature vector of the captured image acquired in step S402 as the similarity.

[0065] If the similarity is equal to or greater than the threshold, the determination unit 303 determines that the user is the user corresponding to the authentication information, and as a result, determines that the second authentication is successful. On the other hand, if the similarity is less than the threshold, the determination unit 303 determines that the user is not the user corresponding to the authentication information, and as a result, determines that the second authentication is unsuccessful.

[0066] In step S404, the output unit 304 outputs the result of the second authentication in step S403. The output form of the result of the second authentication is not limited to a specific output form. For example, the output unit 304 may display the result of the second authentication on the monitor 110 using an image or text. For example, if the result of the second authentication is successful (the user is the user corresponding to the authentication information), the output unit 304 may display an image or text indicating "Login successful" on the monitor 110. For example, if the result of the second authentication is unsuccessful (the user is not the user corresponding to the authentication information), the output unit 304 may display an image or text indicating "Login failed" on the monitor 110.

[0067] Also, for example, the output unit 304 may be configured to store the result of the second authentication in the external storage device 104. In this case, if the result of the second authentication indicating failure is stored in the external storage device 104 a predetermined number of times in succession, the inference device may perform processing such as notifying the administrator user of the possibility of unauthorized access.

[0068] For example, the output unit 304 may output the result of the second authentication as a voice, or may transmit a message representing the result of the second authentication as an image or text to an external device via the network 111.

[0069] As described above, according to this embodiment, in mini-batch learning, the learning proficiency level is calculated, and the sampling probability of images of people with low proficiency is increased, thereby reducing the difference in proficiency level between people, thereby reducing attribute bias in authentication accuracy during inference.

[0070] <Modification> In the first embodiment, the proficiency level N i was calculated using the above formula (3). However, if the training data contains many face images that are not suitable for face recognition training, such as face images with strong blurring, it becomes impossible to calculate the proficiency level appropriately.

[0071] Considering these points, proficiency level N i Instead of using a feature vector to calculate the proficiency level N, the acquisition unit 204 uses the representative vector of each of the n people acquired from the external storage device 104 and the representative vector of the person with the person ID selected in step S204 from the representative vectors, as shown in the following formula (4). i It is also possible to obtain the following.

[0072]

number

[0073] In the first embodiment, the learning device and the inference device are described as separate devices, but the learning device and the inference device may be integrated into a single processing device, in which case the processing device will have both the functions of the learning device and the inference device.

[0074] [Second embodiment] In this embodiment, differences from the first embodiment will be described, and unless otherwise specified below, it will be assumed that the present embodiment is the same as the first embodiment. An example of the functional configuration of a learning device according to this embodiment is shown in the block diagram of Fig. 8. In Fig. 8, functional units similar to those shown in Fig. 2 are assigned the same reference numerals, and descriptions of these functional units will be omitted.

[0075] In the first embodiment, the sampling probability in mini-batch learning is calculated according to the proficiency level of each person. In this embodiment, the facial images of the people are further processed according to the proficiency level. When the facial images used for learning are sampled, the process of performing processing such as changing the color tone with a certain probability is known as data augmentation, and is used as a method to improve the versatility of the learning model. By increasing the probability of performing data augmentation on facial images of people with low proficiency, it is possible to improve versatility for various facial images of people with low proficiency, such as facial expressions and lighting conditions.

[0076] The learning process of the learning model by the learning device will be described with reference to the flowchart in Fig. 9. In Fig. 9, the same processing steps as those shown in Fig. 3 are assigned the same step numbers, and descriptions of those processing steps will be omitted. The flowchart in Fig. 9 is a flowchart in which step S503 is inserted between step S102 and step S103 of the flowchart in Fig. 3.

[0077] In step S503, for each person ID selected in step S204, the probability of performing data augmentation on the face image of the person with that person ID is calculated, and the face image is processed based on that probability. Details of the processing in step S503 will be described with reference to the flowchart in FIG.

[0078] In step S601, the processing unit 403 acquires the proficiency level calculated by the calculation unit 206 in the previous step S302.

[0079] In step S602, the processing unit 403 acquires the probability (processing application probability) of performing Data Augmentation on each face image selected in step S205 from the "proficiency levels of each of n people" that it manages. For example, the processing unit 403 acquires the processing application probability D i is calculated according to the following formula (5).

[0080]

number

[0081] Here, d is a constant defined by equation (6). Then, the processes of steps S603 to S605 are repeated a specified number of times. In step S604, the processing unit 403 selects one unselected face image from the face images selected in step S205 as a selected face image. Then, the processing unit 403 obtains the processing application probability calculated in step S602 for the person in the selected face image, and determines whether to apply Data Augmentation to the selected face image based on the processing application probability.

[0082] For example, suppose that the face image of a person with person ID = 1 is selected as the selected face image, and the probability of applying processing corresponding to the person with person ID = 1 is D1 = 0.3 (30%). In this case, the flag value "0" indicating that Data Augmentation is applied is (number of elements of the array x D), as in the array [1, 1, 1, 1, 1, 1, 1, 0, 0, 0]. i ) array, and randomly select one element (flag value) from the array. If the selected flag value is "0", it is determined that Data Augmentation is to be applied, and if the selected flag value is "1", it is determined that Data Augmentation is not to be applied. As a method for processing face images with Data Augmentation, Cutout, etc., as shown in the literature "Terrance DeVries, et al. Improved Regularization of Convolutional Neural Networks with Cutout. arXiv:1708.04552. 2017" is used.

[0083] If the selected flag value is "0", the processing unit 403 applies Data Augmentation to the selected face image, and if the selected flag value is "1", the processing unit 403 does not apply Data Augmentation to the selected face image.

[0084] In step S605, the processing unit 403 determines whether or not all of the face images selected in step S205 have been selected as selected face images. If all of the face images selected in step S205 have been selected as selected face images, the loop of steps S603 to S605 ends. On the other hand, if any face images selected in step S205 remain that have not been selected as selected face images, the process proceeds to step S603.

[0085] In this way, according to this embodiment, by adjusting the probability of applying Data Augmentation to the face image of each person according to the learning proficiency of that person, it is possible to improve versatility for various face images of people with low learning proficiency, thereby reducing attribute bias.

[0086] In the above embodiments and variants, the person ID is assumed to be identification information unique to each person, but this is not limited to this and may be identification information unique to various units such as country, gender, type, etc.

[0087] Furthermore, the numerical values, processing timing, processing order, processing subject, destination / source / storage location of data (information) and the like used in the above-described embodiments and variant examples are given as examples to provide a concrete explanation, and are not intended to be limited to such examples.

[0088] Furthermore, some or all of the above-described embodiments and modified examples may be used in appropriate combination, and some or all of the above-described embodiments and modified examples may be selectively used.

[0089] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0090] The disclosure of this specification includes the following image processing device, image processing method, and computer program.

[0091] (Item 1) a first acquisition means for acquiring a learning proficiency of a learning subject in a learning model; a second acquisition means for acquiring a sampling probability of each image of the plurality of learning objects based on the proficiency of the plurality of learning objects; a learning means for learning the learning model based on sampled images sampled from the plurality of learning target images based on the sampling probability acquired by the second acquisition means; and An image processing device comprising:

[0092] (Item 2) The first acquisition means The image processing device described in item 1 is characterized in that the more representative vectors of the multiple learning objects there are for which the distance between the sampled image of the learning object of interest and the feature vector obtained based on the learning model is less than a threshold, the lower the proficiency level acquired as the proficiency level of the learning object of interest.

[0093] (Item 3) The first acquisition means The image processing device described in item 1 is characterized in that the greater the number of representative vectors among the plurality of learning objects whose distance from the representative vector of the target learning object is less than a threshold, the lower the proficiency level acquired as the proficiency level of the target learning object.

[0094] (Item 4) 4. The image processing device according to any one of items 1 to 3, wherein the second acquisition means acquires a smaller sampling probability as the sampling probability of a learning object with a higher level of proficiency.

[0095] (Item 5) moreover, a third obtaining means for obtaining a probability of processing the sampled image based on the proficiency of the sampled image; a processing means for processing the sampled image in accordance with the probability of processing the sampled image; 5. The image processing device according to any one of items 1 to 4, comprising:

[0096] (Item 6) moreover, 6. The image processing device according to any one of items 1 to 5, further comprising means for performing second authentication of a person based on a facial image of the person authenticated in the first authentication, a captured image including the person, and a learning model trained by the learning means.

[0097] (Item 7) 7. The information processing device according to item 6, wherein the first authentication is authentication using authentication information.

[0098] (Item 8) An image processing method performed by an image processing device, a first acquisition step in which a first acquisition means of the image processing device acquires a learning proficiency of a learning subject in a learning model; a second acquisition step in which a second acquisition means of the image processing device acquires a sampling probability of each image of the plurality of learning objects based on the proficiency of the plurality of learning objects; a learning step in which a learning means of the image processing device learns the learning model based on sampled images sampled from the plurality of learning target images based on the sampling probability acquired in the second acquisition step; An image processing method comprising:

[0099] (Item 9) A computer program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 7.

[0100] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0101] 201: Acquisition unit 202: Acquisition unit 203: Generation unit 204: Acquisition unit 205: Update unit 206: Calculation unit

Claims

1. A first acquisition means for acquiring a learning proficiency of a learning subject in a learning model; A second acquisition means for acquiring a sampling probability of each image of the plurality of learning objects based on the proficiency of the plurality of learning objects; a learning means for learning the learning model based on sampled images sampled from the plurality of learning target images based on the sampling probability acquired by the second acquisition means; An image processing device comprising:

2. The first acquisition means is The image processing device described in claim 1, characterized in that the more representative vectors of the multiple learning objects there are for which the distance between the sampled image of the learning object of interest and the feature vector obtained based on the learning model is less than a threshold, the lower the proficiency level acquired as the proficiency level of the learning object of interest.

3. The first acquisition means is The image processing device according to claim 1, characterized in that the greater the number of representative vectors among the plurality of learning objects whose distance from the representative vector of a target learning object is less than a threshold, the lower the proficiency level acquired as the proficiency level of the target learning object.

4. The image processing apparatus according to claim 1 , wherein the second acquisition means acquires a higher sampling probability as the sampling probability of a learning object having a lower proficiency level.

5. moreover, A third acquisition means for acquiring a probability of performing processing on the sampled image based on the proficiency of the sampled image; a processing means for processing the sampled image in accordance with a probability of processing the sampled image; The image processing device according to claim 1 , further comprising:

6. The image processing device described in Claim 5, characterized in that the processing means generates an array consisting of flag values ​​for processing and flag values ​​for not processing based on the probability of performing processing on the sampled image, and determines whether to perform image processing on the sampled image depending on the flag value selected from the array.

7. moreover, 2. The image processing device according to claim 1, further comprising: a means for performing a second authentication of a person authenticated in a first authentication based on a facial image of the person, a captured image including the person, and a learning model learned by the learning means.

8. 8. The image processing apparatus according to claim 7, wherein the first authentication is authentication using authentication information indicating an individual input by a user.

9. The image processing device described in Claim 1, characterized in that the learning means generates an array with a number of elements corresponding to the sampling probability of the images of each of the multiple learning objects, and learns the learning model based on an image corresponding to a learning object selected from the array.

10. An image processing method performed by an image processing device, comprising: a first acquisition step in which a first acquisition means of the image processing device acquires a learning proficiency of a learning subject in a learning model; a second acquisition step in which a second acquisition means of the image processing device acquires a sampling probability of each image of the plurality of learning objects based on the proficiency of the plurality of learning objects; a learning step in which a learning means of the image processing device learns the learning model based on a sampling image sampled from the plurality of learning target images based on the sampling probability acquired in the second acquisition step; An image processing method comprising:

11. A computer program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 9.