Learning device and identification device

The learning device generates pseudo-images to enhance training data for machine learning models, improving the accuracy of distinguishing between living and non-living objects in impersonation detection.

JP2026072226APending Publication Date: 2026-05-01CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing machine learning models for impersonation detection face challenges in accurately distinguishing between living and non-living objects due to limited real-world training data and minute differences between the two, particularly in various shooting environments and impersonation materials.

Method used

A learning device that generates pseudo-living and pseudo-non-living images to augment training data, using convolutional neural networks to calculate feature vectors and similarities, and employs methods like ArcFace and triplet loss to train models for precise identification.

Benefits of technology

Enhances the accuracy of impersonation detection by effectively identifying real living beings versus artificial objects even in different shooting environments and methods, reducing misidentification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072226000001_ABST
    Figure 2026072226000001_ABST
Patent Text Reader

Abstract

The goal is to enable high-precision identification of whether an image contains a real living organism or a non-living, artificial object. [Solution] A learning device for learning a model to identify whether the subject of an image is a real living organism or a non-living artificial object, the device acquires living organism images including living organisms and non-living organism images including non-living organisms, processes the acquired living organism images or non-living organism images to generate at least one of pseudo-living organism images including pseudo-living organisms and pseudo-non-living organism images including pseudo-non-living organisms, and learns a model to identify the attributes of living organisms, non-living organisms, pseudo-living organisms, and pseudo-non-living organisms based on the acquired living organism images and non-living organism images and at least one of the generated pseudo-living organism images and pseudo-non-living organism images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates, in particular, to a learning device, an identification device, a control method for the learning device, and a program suitable for use in impersonation detection. [Background technology]

[0002] Conventionally, in biometric authentication, there is a known method for detecting impersonation by determining whether the subject in the input image is a real person (biological) or a non-biological object that does not exist in the image, such as a printed image of a person. As such an image-based impersonation detection method, a method for distinguishing between biological and non-biological objects using machine learning models has been proposed (see Non-Patent Document 1). On the other hand, when training a machine learning model, a technique has been proposed to improve the efficiency of training under conditions where the size of the training data is limited by generating pseudo-data similar to the training data using a generative model (see Patent Document 1). [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-99834 [Non-patent literature]

[0004] [Non-Patent Document 1] Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang, Shang-Hong Lai. PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 20281-20290. [Non-Patent Document 2] K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016 [Non-Patent Document 3] J. Deng, J. Guo, N. Xue, and S. Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019 [Non-Patent Document 4] F. Schroff, D. Kalenichenko, and J. Philbin. Facenet: A unified embedding for face recognition and clustering. In Proc. CVPR, 2015 [Overview of the project] [Problems that the invention aims to solve]

[0005] However, in order to perform impersonation detection using machine learning models, the amount of real-world training data is limited, and it is difficult to obtain a sufficient amount of data in advance for both living and non-living objects for each shooting environment, impersonation method, and impersonation material. Here, impersonation materials include, for example, printing paper, display models, and mask materials. Furthermore, in impersonation detection, the difference between living and non-living objects is minute, which presents a challenge in training machine learning models.

[0006] In view of the aforementioned problems, the present invention aims to enable high-precision identification of whether an object in an image is a real living organism or a non-living artificial object. [Means for solving the problem]

[0007] The learning device according to the present invention is a learning device that learns a model for identifying whether the subject in an image is a living body in which the subject exists or a non-living body that is an artificial object, and includes an acquisition unit that acquires a living body image including the living body and a non-living body image including the non-living body, and a generation unit that processes the living body image or the non-living body image acquired by the acquisition unit to generate at least one of a pseudo-living body image including a pseudo-living body and a pseudo-non-living body image including a pseudo-non-living body, and a learning unit that learns a model for identifying the attributes of the living body, the non-living body, the pseudo-living body, and the pseudo-non-living body based on at least one of the living body image and the non-living body image acquired by the acquisition unit and at least one of the pseudo-living body image and the pseudo-non-living body image generated by the generation unit.

Advantages of the Invention

[0008] According to the present invention, it is possible to accurately identify whether the subject in the image is a real living body or a non-living body that is an artificial object.

Brief Description of the Drawings

[0009] [Figure 1] It is a block diagram showing an example of the hardware configuration of a learning device and an identification device. [Figure 2] It is a block diagram showing an example of the functional configuration of a learning device and an identification device according to the first embodiment. [Figure 3] It is a flowchart showing an example of the procedure of the learning loop process in the learning device. [Figure 4] It is a flowchart showing an example of the detailed processing procedure of the learning process in the first embodiment. [Figure 5] It is a flowchart showing an example of the identification processing procedure by the identification device. [Figure 6] It is a diagram for explaining the structure of an identification device that identifies whether it is a living body or a non-living body in the first embodiment.​​​​​This is a diagram illustrating the structure of an identification device for identifying living organisms or non-living organisms in a third embodiment. [Figure 9] This is a block diagram showing an example of the functional configuration of a learning device according to the fourth embodiment. [Figure 10] This flowchart shows an example of a detailed processing procedure for the learning process in the fourth embodiment. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described below with reference to the drawings. However, the present invention is not limited to the embodiments described below, and various forms that do not depart from the spirit of the invention are also included. Furthermore, each embodiment described below is merely one embodiment of the present invention, and it is possible to combine each embodiment as appropriate.

[0011] (First embodiment) Figure 1 is a block diagram showing an example of the hardware configuration of the learning device 10 and the identification device 20, which will be described later. The CPU (Central Processing Unit) 101 controls the entire device. The ROM (Read Only Memory) 102 is a non-volatile memory for storing programs and parameters that do not need to be changed. The RAM (Random Access Memory) 103 is a memory for temporarily storing programs and data supplied from external devices.

[0012] The external storage device 104 is a storage device such as a hard disk drive or memory card that is fixedly installed inside the device. The external storage device 104 may also be a removable flexible disk, optical disk, magnetic card, optical card, IC card, memory card, etc.

[0013] The input device interface 105 is an interface with input devices 109 such as pointing devices and keyboards. The output device interface 106 is an interface with a monitor 110 for displaying data held by the device or supplied data.

[0014] The communication interface 107 is an interface for connecting to a network line 111, such as the Internet, and connects to the NW (network) camera 112 via the network line 111. The NW camera 112 is an imaging device that captures images. The system bus 108 is a transmission path that connects each unit in a communicative manner. The processes described later are performed by the CPU 101 executing programs stored in a computer-readable storage medium such as ROM 102.

[0015] [Overview of Identification Process] In this embodiment, the structure of an identification device that identifies whether the subject in an image is a living person or a non-living object such as a printed image of a person is explained with reference to Figure 6. As shown in Figure 6, the identification device takes an image containing living or non-living objects as input and outputs feature vectors for identification into living, non-living, pseudo-living, and pseudo-non-living. A pseudo-living object refers to a pseudo-living object included in an image that reproduces a different shooting environment from the training data, generated by processing a living image from existing training data. A pseudo-non-living object refers to a pseudo-non-living object included in an image that reproduces phenomena specific to non-living objects, such as reflections or vertical streaks in printed materials, generated by processing a living or non-living image from existing training data. The CNN (Convolutional Neural Networks) 601 is a network structure corresponding to the feature vector calculation unit 125 of the learning device 10, which will be described in detail later.

[0016] On the other hand, the classification head 602 is a network structure corresponding to the biological detection unit 126, non-biological detection unit 127, pseudo-non-biological detection unit 128, and pseudo-biological detection unit 129 of the learning device 10, which will be described later. The classification head 602 takes the feature vector output by CNN 601 as input and calculates the similarity for each attribute. In this case, each attribute is the shooting environment for biological organisms, the shooting environment and the impersonation method (material) for non-biological organisms, and the processing method for pseudo-biological organisms and pseudo-non-biological organisms.

[0017] The shooting environment includes shooting equipment, shooting location, and shooting time, and is not limited to these, as long as the conditions are determined at the time of shooting. Impersonation methods include holding up printed materials, holding up a display showing a biological image, wearing a mask, and holding up a model. Materials include printing paper, display model, mask material, model material, etc. In this embodiment, the above-mentioned impersonation methods and materials are used, but there are other impersonation methods and materials, and they are not limited to these.

[0018] Figure 6 illustrates a living organism photographed in either shooting environment A or shooting environment B, a pseudo-living organism generated by processing method A, a non-living organism photographed in either shooting environment A or shooting environment B as seen in a printed document or on a display, and a pseudo-non-living organism generated by processing method B. In this embodiment, the similarity to each attribute shown in Figure 6 is calculated. For example, shooting environment A is an environment photographed with a network camera, and shooting environment B is an environment photographed with a webcam. Processing method A is a method of generating a pseudo-living organism by applying a blur filter to a living organism, and processing method B is a method of generating a pseudo-non-living organism by superimposing another image onto a living organism.

[0019] The identification device determines whether the subject of the input image is a living or non-living organism based on the similarity output by the classification head 302. The method for determining whether it is a living or non-living organism is to calculate the spoofing score by summing the similarities with non-living organisms and pseudo-non-living organisms, and if the spoofing score exceeds a threshold, it is determined to be a non-living organism. The number of attribute types is N, and the i-th attribute is t i , input and attribute ti The similarity to sim(t i Assuming this is the case, the impersonation score is calculated using the following formula (1).

[0020]

number

[0021] Alternatively, another determination method is to use the attribute t with the highest similarity. max However, if it is from a living organism or a pseudo-living organism, it may be determined to be a living organism, and if it is from a non-living organism or a pseudo-non-living organism, it may be determined to be a non-living organism. In this case, the attribute t with the highest similarity is... max This is calculated using the following formula (2).

[0022]

number

[0023] Furthermore, the method is not limited to the above-described determination method, as it can determine whether the input is living or non-living based on its similarity to each attribute.

[0024] [Configuration of the learning device] Figure 2(a) is a block diagram showing an example of the functional configuration of the learning device 10 of this embodiment. The learning device 10 of this embodiment learns to classify living organisms according to their shooting environment, non-living organisms according to their shooting environment and impersonation method (material), and pseudo-living organisms and pseudo-non-living organisms according to their processing method.

[0025] The learning data management unit 121 manages the learning data stored in the external storage device 104. Here, the learning data includes biological images, non-biological images, and attribute labels. Biological images are images of real people. Non-biological images are images of artificial objects created based on biological images. Artificial objects may be printed materials, displays showing biological images, 3D masks, or models. They are not limited to these, as long as they are artificial objects that can be used for impersonation using a face. Attribute labels are labels that are uniquely assigned to a combination of biological (non-biological) objects, impersonation methods (materials), shooting environment information, and data processing methods. For example, attribute labels are assigned to each attribute shown in Figure 6. Here, the shooting environment information is information about the shooting equipment. Note that the shooting environment information may also be shooting time information or shooting location information. The shooting environment information is not limited to these, as long as it is a condition determined at the time of shooting.

[0026] The data acquisition unit 122 acquires bio-images, non-biological images, and attribute labels to be used for training from the training data management unit 121.

[0027] The pseudo-non-living image generation unit 123 generates a pseudo-non-living image by processing the data acquired by the data acquisition unit 122 to reproduce a non-living image. At this time, the processing is performed to reproduce phenomena specific to non-living images, such as reflections on paper or monitors, or vertical streaks in printed materials. The pseudo-non-living image generation unit 123 may generate a pseudo-non-living image by combining multiple images, or it may generate a pseudo-non-living image by performing image processing on a single input image. In this embodiment, as processing method B, a pseudo-non-living image is generated by superimposing another image on the data acquired by the data acquisition unit 122. The superimposed image may be held by the learning data management unit 121 and acquired by the data acquisition unit 122 when generating the pseudo-non-living image, or an image with random RGB values ​​may be generated as the superimposed image when generating the pseudo-non-living image.

[0028] Furthermore, the superimposed images are not limited to those mentioned above, as long as the subject does not contain living organisms. The superimposition position may be determined using feature point information of the input image. Feature point information refers to organ point information when the subject is a person's face. Specifically, in processing that reproduces reflections, reflections occur in the pupils and glasses even in living organisms, so organ point information is used to prevent misidentification between living organisms and the attributes of the reflection-reproducing processing by not performing the reflection-reproducing processing around the eyes. Organ point information can be acquired using known methods such as organ point detection, and may be stored in the learning data management unit 121 after prior organ point detection, or it may be performed when generating pseudo-non-living images. The process for generating pseudo-non-living images is not limited to those mentioned above, as long as it outputs a single image.

[0029] The pseudo-biological generation unit 124 generates pseudo-biological images by processing the data acquired by the data acquisition unit 122 to reproduce different shooting environments. Specifically, it generates images taken with a low-resolution camera or images taken with a camera with strong blur. In addition, the shooting environment to be reproduced is not limited to these conditions, as long as they are conditions determined at the time of shooting, such as the shooting location and shooting time. The pseudo-biological generation unit 124 may generate pseudo-biological images by combining multiple images, or it may generate pseudo-biological images by performing image processing on a single input image. In this embodiment, as processing method A, a blur filter is applied to the data acquired by the data acquisition unit 122 to generate pseudo-biological images. Alternatively, pseudo-biological images may be generated by degrading the image quality of the data acquired by the data acquisition unit 122 by repeatedly enlarging and reducing it. The process for generating pseudo-biological images is not limited to these methods as long as it outputs a single image.

[0030] In this embodiment, we describe an example of generating both pseudo-biological images and pseudo-non-biological images, but it is also possible to generate only one of them.

[0031] The feature vector calculation unit 125 calculates a feature vector from an image. Specifically, a CNN, which is a type of neural network, is used. A CNN extracts information abstracted from an input image by repeatedly performing a process composed of a convolution process, an activation process, and a pooling process on the input image multiple times. At this time, a processing unit composed of a convolution process, an activation process, and a pooling process is called a layer. In the activation process, a known method such as a method called ReLU (Rectified Linear Unit) is used. Also, in the pooling process, a known method such as a method called max pooling is used.

[0032] For example, as the structure of the CNN, ResNet or the like described in Non-Patent Document 2 or the like may be used. Note that, for the neural network used by the feature vector calculation unit 125, a Transformer described in U.S. Patent No. 10956819 or the like may be used. The neural network used by the feature vector calculation unit 125 is not limited to these as long as it takes an image as an input and outputs a feature vector.

[0033] The living body determination unit 126, the non-living body determination unit 127, the pseudo non-living body determination unit 128, and the pseudo living body determination unit 129 calculate the similarity for each attribute of a living body, a non-living body, a pseudo non-living body, and a pseudo living body, respectively, using the feature vector calculated by the feature vector calculation unit 125. As a method for calculating the similarity, a method using a representative vector described in Non-Patent Document 3 is used. In the method using a representative vector, a representative vector corresponding to the type of attribute label is set. The representative vector is a vector that constitutes a fully connected layer that takes a feature vector as an input when the number of types of attribute labels is n and the dimensionality of the feature vector is d. Here, let the feature vector calculated by the feature vector calculation unit 125 from the i-th training data be x i and the representative vector corresponding to the j-th type of attribute label be W j . In this case, the similarity between the feature vector x i and the representative vector W j is the cosine similarity cosθ WjxiThis is shown in equation (3) below.

[0034]

number

[0035] Furthermore, the method for calculating similarity is not limited to the method described above, as long as it allows the similarity between the feature vector and each attribute to be determined as a real number.

[0036] The loss calculation unit 130 calculates the impersonation detection loss using the feature vector calculated by the feature vector calculation unit 125, or the similarity to the attributes of living organisms, non-living organisms, pseudo-non-living organisms, and pseudo-living organisms. As the impersonation detection loss, a loss function such as ArcFace shown in Non-Patent Document 3 is used. Specifically, the aforementioned feature vector x i and representative vector W j The cosine similarity with cosθ Wjxi Using this, the loss is calculated by equation (4) below. ArcFace is a method that simultaneously learns feature vectors and representative vectors using the loss calculated by equation (4) and uses these to perform distance learning as a classification problem.

[0037]

number

[0038] In equation (4), N represents the batch size, s and m represent hyperparameters, and y i represents the attribute label of the i-th data. Note that, as the spoofing detection loss, a triplet loss as shown in Non-Patent Document 4 may also be used. The type of spoofing detection loss is not limited to these, as long as it can be used for multi-class classification.

[0039] The gradient calculation unit 131 calculates the gradient with respect to parameters such as the weights of the neural network based on the impersonation detection loss calculated by the loss calculation unit 130. The NN (Neural Network) update unit 132 updates the neural network based on the gradient calculated by the gradient calculation unit 131.

[0040] [Learning Loop Processing] Figure 3 is a flowchart showing an example of the procedure for the learning loop processing in the learning device 10 of this embodiment. S301 marks the beginning of an epoch loop. Here, one epoch is defined as the time when all the training data managed by the training data management unit 121 is used in the training process. In this embodiment, the number of epochs to be repeated is predetermined. To count the number of epoch repetitions, the variable i is initialized to 1. If the variable i, which indicates the number of repetitions, is less than or equal to the predetermined number of epochs, the process proceeds to S302. If the variable i exceeds the predetermined number of epochs, the loop is exited and the process ends.

[0041] S302 marks the beginning of the training data loop. The training data is assumed to be numbered sequentially starting from 1. To reference the training data to be processed using the variable j, the variable j is initially initialized to 1. If the value of variable j is less than or equal to the number of training data, proceed to S303. If the value of variable j exceeds the number of training data, exit the training data loop and proceed to S306.

[0042] In S303, the data acquisition unit 122 acquires training data from the training data management unit 121. Specifically, the data acquisition unit 122 acquires biological images, non-biological images, and attribute labels as training data from the training data management unit 121. Alternatively, the acquired images may be processed using image processing or other methods before being used as training data.

[0043] In S304, learning is performed using the acquired data. Details of the processing in S304 will be described later with reference to Figure 4. S305 marks the end of the training data loop, where 1 is added to the variable j, and the program returns to S302. S306 marks the end of the epoch loop, where 1 is added to the variable i, and the program returns to S301.

[0044] [Learning Process] Figure 4 is a flowchart showing an example of a detailed processing procedure for the learning process in S304 of Figure 3.

[0045] In S401, the data acquisition unit 122 determines whether or not to perform data processing on the biological and non-biological images acquired in S303. The decision on whether or not to perform data processing may be made based on a predetermined probability or based on attribute labels. The method for determining whether or not to perform data processing is not limited to these methods. If the result of this decision is to perform data processing, the process proceeds to S402; otherwise, it proceeds to S403.

[0046] In S402, data processing is performed on images that were determined to require data processing in S401. For example, the pseudo-biological generation unit 124 performs data processing on a biological image to generate a pseudo-biological image. The pseudo-biological generation unit 124 then assigns attribute labels to the pseudo-biological image according to the data processing method. Similarly, the pseudo-non-biological generation unit 123 performs data processing on a biological image or a non-biological image to generate a pseudo-non-biological image. The pseudo-non-biological generation unit 123 then assigns attribute labels to the pseudo-non-biological image according to the data processing method.

[0047] In S403, the feature vector calculation unit 125 inputs both biological and non-biological images into the neural network and calculates feature vectors for each. Furthermore, if data processing is performed in S402, the feature vector calculation unit 125 also inputs pseudo-biological and pseudo-non-biological images into the neural network and calculates feature vectors for each.

[0048] In S404, the biological detection unit 126, the non-biological detection unit 127, the pseudo-non-biological detection unit 128, and the pseudo-biological detection unit 129 each calculate the similarity to the attributes of a biological organism, a non-biological organism, a pseudo-non-biological organism, and a pseudo-biological organism, respectively, using feature vectors and attribute labels. In S405, the loss calculation unit 130 calculates the spoofing detection loss using the feature vector or the similarity to each attribute.

[0049] In S406, the gradient calculation unit 131 uses the impersonation detection loss calculated in S405 to calculate the gradient of parameters such as the weights of the neural network from the biological image, non-biological image, pseudo-biological image, and pseudo-non-biological image until estimated attributes are obtained. In S407, the NN update unit 132 updates the neural network model by updating parameters such as the weights of the neural network based on the gradient calculated in S406. The parameters are updated by applying an optimization algorithm such as the stochastic gradient method.

[0050] [Configuration of the identification device] Figure 2(b) is a block diagram showing an example of the functional configuration of the identification device 20 in this embodiment. In this embodiment, the identification device 20 uses a neural network model trained by the learning device 10 to perform identification for impersonation detection. The identification process for impersonation detection involves determining whether the subject in the input image is a living or non-living being.

[0051] The identification data acquisition unit 141 acquires the image to be judged. The image to be judged can be a biological image or a non-biological image. For example, consider a system where, when performing identity verification for login using the image to be judged acquired from a camera, identity verification is performed only if the image to be judged acquired from the camera is a biological image. In this case, the monitor 110 displays a login screen and instructs the user to look at the network camera 112. At this time, the identification device 20 acquires the image to be judged from the network camera 112 via the communication interface 107. Note that the method of acquiring the image to be judged and the system configuration are not limited to these.

[0052] The feature vector calculation unit 142 performs the same processing as the feature vector calculation unit 125 of the learning device 10. The discrimination device 20 holds parameters such as the weights of the neural network learned by the learning device 10 and calculates feature vectors from the image to be judged acquired by the discrimination data acquisition unit 141.

[0053] The biological detection unit 143, non-biological detection unit 144, pseudo-non-biological detection unit 145, and pseudo-biological detection unit 146 perform the same processing as the biological detection unit 126, non-biological detection unit 127, pseudo-non-biological detection unit 128, and pseudo-biological detection unit 129 of the learning device 10, respectively. Using the feature vectors obtained by the feature vector calculation unit 142, they calculate the similarity to each attribute (each element of the attribute label) of the target image for detection, for biological, non-biological, pseudo-non-biological, and pseudo-biological. In this embodiment, the similarity to each attribute is calculated based on parameters such as the weights of the neural network learned by the learning device 10.

[0054] The judgment integration unit 147 integrates the similarity scores for each element of the attribute labels calculated by the biological determination unit 143, non-biological determination unit 144, pseudo-non-biological determination unit 145, and pseudo-biological determination unit 146, respectively, to determine whether the subject of the image to be judged is a living or non-living being. The determination of whether it is a living or non-living being is performed by methods such as calculating a spoofing score, as described above.

[0055] The identification result output unit 148 outputs the identification result obtained by the judgment integration unit 147. Specifically, it outputs content corresponding to the judgment result to the monitor 110. For example, if the image to be judged is determined to be a living being, identity verification is performed. If it is determined to be non-living, a message indicating login failure is displayed. Furthermore, the identification result of whether it is a living being or non-living being may be stored in the external storage device 104. If the identification result of non-living being stored more than a predetermined number of times, processing such as notifying the administrator user of the possibility of unauthorized access may be performed. The method of outputting the identification result is not limited to outputting the determination result of whether it is a living being or non-living being.

[0056] [Identification process] Figure 5 is a flowchart showing an example of the identification processing procedure by the identification device 20 of this embodiment.

[0057] In S501, the identification data acquisition unit 141 acquires the target image that is subject to detection for impersonation. In S502, the feature vector calculation unit 142 calculates the feature vector of the image to be judged, which was acquired in S501.

[0058] In S503, the biological determination unit 143, non-biological determination unit 144, pseudo-non-biological determination unit 145, and pseudo-biological determination unit 146 use the feature vectors calculated in S502 to calculate similarity to the attributes of biological, non-biological, pseudo-non-biological, and pseudo-non-biological entities, respectively. Then, the determination integration unit 147 integrates the calculated similarity to determine whether the subject of the image to be determined is a biological or non-biological entity.

[0059] In S504, the identification result output unit 148 outputs the judgment result from S503 to the monitor 110, and also outputs the judgment result to the external storage device 104 as needed. This allows the user to be informed of the judgment result and the judgment result to be recorded.

[0060] [Effects of this embodiment] The effects of this embodiment will be explained using Figures 7(a) and 7(b). Figures 7(a) and 7(b) are diagrams showing examples of the distribution of representative vectors in the feature space, and represent a comparison of the representative vectors and feature vectors for each attribute when an unknown living organism 701 and an unknown non-living organism 702 are given as input during classification.

[0061] Figures 7(a) and 7(b) show the attributes to be identified as living organism and shooting environment A (703), living organism and shooting environment B (704), printed material and shooting environment A (705), display and shooting environment A (706), pseudo-living organism 707, and pseudo-non-living organism 708. However, the attributes to be identified are not limited to these. In Figures 7(a) and 7(b), the angle between the representative vector and the feature vector represents the degree of similarity, with a smaller angle indicating a higher degree of similarity between the representative vector and the feature vector. Therefore, the unknown living organism 701 and unknown non-living organism 702, which are inputs for identification, are identified by the attribute with a small angle.

[0062] Figure 7(a) shows an overview of a method for classifying the processed pseudo-living organism 707 and pseudo-non-living organism 708 as attributes of a real class. In the example in Figure 7(a), the living organism and shooting environment A (703) and the pseudo-living organism 707 are treated as having the same attribute, and the display and shooting environment A (706) and the pseudo-non-living organism 708 are treated as having the same attribute, and only one representative vector is set for each. As a result, the unknown living organism 701 may be mistakenly identified as a printed material and shooting environment A (705), which has a smaller vector angle. Similarly, the unknown non-living organism 702 may be mistakenly identified as a living organism and shooting environment B (704). Conventional methods have the potential to misidentify living and non-living organisms in this way. In particular, in spoofing detection for shooting environments or spoofing methods (materials) that differ from the training data, the discrepancy between the input feature vector and the representative vector is large, and the accuracy of spoofing detection tends to decrease.

[0063] Figure 7(b) shows an overview of the method for identifying pseudo-living organisms 707 and pseudo-non-living organisms 708 as having attributes different from those of the real class in this embodiment. In this embodiment, representative vectors can be set individually for pseudo-living organisms 707 and pseudo-non-living organisms 708, and these can be compared with feature vectors. Therefore, even in impersonation detection for different shooting environments or impersonation methods (materials) than the training data, an unknown living organism 701 can be identified as a pseudo-living organism 707, and an unknown non-living organism 702 can be identified as a pseudo-non-living organism 708, thereby correctly identifying whether something is living or non-living.

[0064] As described above, according to this embodiment, a pseudo-biological image and a pseudo-non-biological image are generated by processing a biological image (or non-biological image), and a neural network is trained to identify them as having attributes different from those of the real class. For example, a pseudo-non-biological image is generated by superimposing another image onto a biological image, and the processing method is identified. In this way, the accuracy of spoofing detection for different shooting environments and spoofing methods (materials) from the training data can be improved. In particular, when generating a pseudo-non-biological image, the accuracy of spoofing detection can be improved by identifying the processing method by reproducing phenomena specific to non-biological images, such as monitor reflections. Although the effects of this embodiment were shown using a representative vector method as an example, the method is not limited to the representative vector method, as long as the similarity between the input feature vector and each attribute can be calculated as a real number.

[0065] (Second embodiment) In the first embodiment, a recognition device and its learning method were described for identifying whether an object in an image is living or non-living by identifying whether it is living, non-living, pseudo-living, or pseudo-non-living. In this embodiment, a method for preventing the learning of features that are not important to identify is described by specifying a combination of attributes for which loss is not calculated. The configuration and processing content of the learning device and recognition device in this embodiment are basically the same as in the first embodiment. Below, only the differences from the first embodiment will be described.

[0066] [Configuration of the learning device] The configuration of the learning device in this embodiment is basically the same as the configuration shown in Figure 2(a). However, the processing contents of the learning data management unit 121 and the loss calculation unit 130 differ in part from those of the first embodiment.

[0067] The learning data management unit 121 manages the learning data stored in the external storage device 104. The learning data includes not only biological images, non-biological images, and attribute labels, but also information on attribute combinations for which loss is not calculated.

[0068] The loss calculation unit 130 calculates the impersonation detection loss in the same manner as in the first embodiment. However, it does not calculate the loss in cases of misidentification between combinations of biological and pseudo-biological entities with specific attributes specified in advance, or between specific non-biological entities and pseudo-non-biological entities. For example, it does not calculate the loss between a pseudo-non-biological entity generated by processing to reproduce reflection and a non-biological entity displayed on a display. The specified combinations may be, but are not limited to, combinations of biological and pseudo-biological entities, or combinations of non-biological entities and pseudo-non-biological entities.

[0069] Based on the above points, when calculating the spoofing detection loss, the loss function of equation (5) below, which is based on the loss function of ArcFace etc. shown in Non-Patent Document 3, is used to calculate the loss as the spoofing detection loss.

[0070]

number

[0071] In equation (5), N represents the batch size, s and m represent hyperparameters, and y i represents the attribute label of the i-th data. Also, C i is attribute label y i This is a set of attribute labels for attributes for which the loss is not calculated. For example, attribute label y i In the case of a pseudo-non-living organism generated by processing that reproduces reflection, set C i These are non-living entities displayed on various screens. Furthermore, the type of loss can be used for multi-class classification, and is not limited to this method as long as it does not calculate losses between specific attributes.

[0072] [Effects of this embodiment] According to this embodiment, losses are not calculated between combinations of living organisms and pseudo-living organisms with pre-specified attributes, or between non-living organisms and pseudo-non-living organisms. This allows the neural network to be trained while ignoring errors between attributes that are not important to distinguish, such as the combination of a pseudo-non-living organism generated by processing and a non-living organism displayed on a screen. Therefore, it is possible to prevent a decrease in discrimination accuracy while preventing the learning of features that are not important to distinguish.

[0073] (Third embodiment) In the second embodiment, unintended learning, such as learning features that are not important to identify, is prevented by not calculating the loss between specified combinations of attributes. This embodiment describes an example of training an identification device characterized by having multiple two-class classification heads. The configuration and processing of the learning device and identification device in this embodiment are basically the same as in the first embodiment. Only the differences from the first embodiment will be described below.

[0074] [Overview of Identification Process] In this embodiment, the structure of the identification device for distinguishing between living and non-living organisms will be explained with reference to Figure 8. Unlike the first embodiment, this embodiment uses three two-class classification heads: classification head A (802), classification head B (803), and classification head X (804). Each classification head divides the attributes extracted from all attributes of living organisms, non-living organisms, pseudo-non-living organisms, and pseudo-living organisms into two sets, and determines which attribute set the feature vector belongs to.

[0075] For example, classification head A (802) distinguishes between living organisms in shooting environment A and non-living printed materials in shooting environment A, while classification head B (803) distinguishes between living organisms in shooting environment B and pseudo-non-living materials processed using method B. Classification head X (804) then distinguishes between living organisms in shooting environments A and B and non-living materials on a display in shooting environment A and pseudo-non-living materials processed using method B. The number of classification heads does not need to be the same as in the example in Figure 8, and the identification performed by each classification head is not limited to those that can be specified by a set of attributes. By training an identification device with multiple two-class classification heads in this way, it is possible to focus the training on attributes that are important to identify.

[0076] [Configuration of the learning device] The configuration of the learning device in this embodiment is basically the same as the configuration shown in Figure 2(a). However, the processing contents of the biological determination unit 126, non-biological determination unit 127, pseudo-non-biological determination unit 128, pseudo-biological determination unit 129, and loss calculation unit 130 differ in part from those of the first embodiment.

[0077] The biological detection unit 126, non-biological detection unit 127, pseudo-non-biological detection unit 128, and pseudo-biological detection unit 129 use the feature vector calculated by the feature vector calculation unit 125 to calculate the similarity of the set of attributes to be compared. In this embodiment, corresponding to the multiple classification heads shown in Figure 8, the probability that a feature vector belongs to the set of attributes compared by each classification head is calculated as the similarity between the feature vector and the set of attributes compared by each classification head. At this time, parameters such as the weights of the classification heads are also learned. The calculation of the probability that a feature vector belongs to the set of attributes compared by each classification head is performed by applying an activation function to the output of the classification head. For example, the sigmoid function shown in equation (6) below is used as the activation function. Note that a function other than the sigmoid function may be used as the activation function.

[0078]

number

[0079] The loss calculation unit 130 calculates the loss using the probability that the feature vector belongs to the set of attributes compared by each classification head, which is calculated using, for example, equation (6) described above. The loss function for each classification head is the Binary Cross Entropy Loss shown in equation (7) below.

[0080]

number

[0081] In equation (7), N represents the batch size, and p(x i ) is the probability that the i-th feature vector belongs to the set of correct attributes, and q(x i ) is the probability that the i-th feature vector belongs to the attribute set. A loss is calculated for each classification head, and the losses obtained from the classification heads related to the input data are combined to calculate a spoofing detection loss for calculating the gradient.

[0082] [Operation of the identification device] The configuration of the identification device in this embodiment is basically the same as the configuration shown in Figure 2(b). However, the processing contents of the biological determination unit 143, non-biological determination unit 144, pseudo-non-biological determination unit 145, pseudo-biological determination unit 146, and determination integration unit 147 differ in part from those of the first embodiment.

[0083] The biological detection unit 143, non-biological detection unit 144, pseudo-non-biological detection unit 145, and pseudo-biological detection unit 146 perform the same processing as the biological detection unit 126, non-biological detection unit 127, pseudo-non-biological detection unit 128, and pseudo-biological detection unit 129 of the learning device 10. That is, the biological detection unit 143, non-biological detection unit 144, pseudo-non-biological detection unit 145, and pseudo-biological detection unit 146 use the feature vectors calculated by the feature vector calculation unit 142 to calculate the similarity of the set of attributes to be compared. The identification device 20 also holds parameters such as the classification head and neural network weights learned by the learning device 10. Then, the biological detection unit 143, non-biological detection unit 144, pseudo-non-biological detection unit 145, and pseudo-biological detection unit 146 calculate the probability that the feature vector belongs to the set of attributes to be compared by each classification head from the feature vectors obtained by the feature vector calculation unit 142.

[0084] The judgment integration unit 147 determines whether the subject of the image to be judged is a living organism or a non-living organism by integrating the probabilities that feature vectors belong to the set of attributes compared by each classification head.

[0085] [Effects of this embodiment] As described above, according to this embodiment, by training an identification device having multiple two-class classification heads, the elements to be trained can be narrowed down to the important elements.

[0086] (Fourth embodiment) In the second embodiment, features that are not important to identify are not learned by specifying combinations of attributes for which loss is not calculated. Furthermore, in the third embodiment, the learning is focused on features that are important to identify by training an identification device with multiple two-class classification heads. In this embodiment, a method for integrating the attributes of living organisms and pseudo-living organisms, and non-living organisms and pseudo-non-living organisms, which are not important to identify during learning, by calculating the similarity between each attribute is described. The configuration and processing content of the learning device and identification device in this embodiment are basically the same as in the first embodiment. Below, only the differences from the first embodiment will be described.

[0087] [Configuration of the learning device] Figure 9 is a block diagram showing an example of the functional configuration of the learning device 90 in this embodiment. Compared to the configuration of the learning device 10 shown in Figure 2(a), this embodiment differs in that it further includes an attribute integration unit 901.

[0088] The attribute integration unit 901 calculates the similarity between each attribute and, if there are combinations where the similarity between a living organism and a pseudo-living organism, or between a non-living organism and a pseudo-non-living organism exceeds a predetermined value, it integrates the pseudo-living organism into a living organism and the pseudo-non-living organism into a non-living organism. For example, the similarity between the attributes of a pseudo-non-living organism generated by a process that reproduces reflection and a non-living organism displayed on a display tends to be high. Therefore, it is not important to distinguish between a pseudo-non-living organism generated by a process that reproduces reflection and a non-living organism displayed on a display.

[0089] The method for calculating the similarity between each attribute uses the representative vector method described in Non-Patent Document 3. In the representative vector method, a representative vector is set corresponding to the type of attribute label. When the number of types of attribute labels is n and the number of dimensions of the feature vector is d, the representative vector is a vector that constitutes a fully connected layer that takes the feature vector as input. Here, the representative vector corresponding to the i-th type of attribute label is W. i W is the representative vector corresponding to the j-th attribute label. j Therefore, the representative vector W i and representative vector W j The degree of similarity is expressed as cosine similarity cosθ WiWj This is shown in equation (8) below.

[0090]

number

[0091] Alternatively, as a method for calculating the similarity between each attribute, the feature vectors calculated by the feature vector calculation unit 125 may be stored, and the similarity may be calculated using the inter-class variance between each attribute of the training data. The method for calculating the similarity between each attribute is not limited to these methods, as long as it is possible to calculate the similarity between attributes as a real number from the training data.

[0092] If the similarity is greater than or equal to a predetermined value, the attribute integration unit 901 generates the representative vector W i and representative vector W j The two vectors are merged. The merging method may involve retaining the representative vector corresponding to living or non-living entities, or averaging the two representative vectors and merging them. The merging method is not limited to these methods, as long as it takes two attributes as input and outputs one attribute. By merging the two representative vectors, the fully connected layer has one fewer attribute label (n). The attribute labels of the data corresponding to the attribute deleted due to the merging are assigned the attribute labels of the merged entity.

[0093] [Learning Process] Figure 10 is a flowchart showing an example of a detailed processing procedure for the learning process in S304 of Figure 3 in this embodiment. Steps S401 to S403 in Figure 10 are the same as steps S401 to S403 in Figure 4, so their explanation is omitted.

[0094] In S1001, the attribute integration unit 901 calculates the similarity between each attribute. In S1002, the attribute integration unit 901 calculates the similarity between each attribute and determines whether there are any combinations of attributes corresponding to living organisms and pseudo-living organisms, or non-living organisms and pseudo-non-living organisms, where the similarity is greater than or equal to a predetermined value. If, as a result of this determination, there are combinations where the similarity is greater than or equal to a predetermined value, the process proceeds to S1003; otherwise, attribute integration is unnecessary, and the process proceeds to S404.

[0095] In S1003, the attribute integration unit 901 integrates attributes for combinations where the similarity between attributes of living organisms and pseudo-living organisms, or non-living organisms and pseudo-non-living organisms, is greater than or equal to a predetermined value. For example, pseudo-living organisms are integrated with their corresponding living organisms, and pseudo-non-living organisms are integrated with their corresponding non-living organisms. Steps S404 to S407 in Figure 10 are the same as steps S404 to S407 in Figure 4, so their explanation is omitted.

[0096] [Effects of this embodiment] As described above, according to this embodiment, the similarity between each attribute is calculated, and if the similarity is greater than or equal to a predetermined value, the pseudo-living organism is integrated into the living organism, and the pseudo-non-living organism is integrated into the non-living organism. This prevents a decrease in identification accuracy while preventing the learning of features that are not important to identify, without having to specify in advance combinations of attributes for which loss will not be calculated or design classification heads to be identified in advance.

[0097] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0098] This embodiment includes the following configurations, methods, and programs.

[0099] (Composition 1) A learning device that trains a model to distinguish whether the subject of an image is a real living organism or a non-living artificial object, Acquisition means for acquiring a biological image including the said biological organism and a non-biological image including the said non-biological organism, A generation means that processes a biological image or non-biological image acquired by the acquisition means to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism, A learning means for learning a model that identifies the attributes of a living organism, a non-living organism, a pseudo-living organism, and a pseudo-non-living organism based on the living organism image and non-living organism image acquired by the acquisition means and at least one of the pseudo-living organism image and pseudo-non-living organism image generated by the generation means, A learning device characterized by having the following features.

[0100] (Configuration 2) The learning device according to configuration 1, characterized in that the subject is a person's face. (Composition 3) The learning device according to configuration 1 or 2, characterized in that the generation means performs image processing on the biological image or the non-biological image to generate the pseudo-non-biological image. (Composition 4) The learning device according to any one of configurations 1 to 3, characterized in that the generation means generates the pseudo-non-biological image by synthesizing a plurality of images. (Composition 5) The learning device according to any one of configurations 1 to 4, characterized in that the generation means generates the pseudo-non-living image by superimposing another image on the living image or the non-living image. (Composition 6) The learning device according to any one of configurations 1 to 5, characterized in that the generation means performs processing according to the feature point information in the biological image or the non-biological image to generate the pseudo-non-biological image.

[0101] (Composition 7) The learning device according to any one of configurations 1 to 6, characterized in that the generation means generates the pseudo-non-biological image by performing processing to reproduce reflection. (Composition 8) The learning device according to any one of configurations 1 to 7, characterized in that the generation means generates the pseudo-biological image from the biological image by changing the shooting environment. (Composition 9) The learning device according to any one of configurations 1 to 8, characterized in that the aforementioned artifact is a printed material containing a living organism or a display that shows a living organism. (Composition 10) The learning device according to any one of configurations 1 to 9, characterized in that the learning means performs learning by calculating a loss based on the similarity to the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism. (Composition 11) The learning device according to any one of configurations 1 to 10, characterized in that the attributes are the imaging environment for the living organism, the method or material using the artificial object for the non-living organism, and the processing method for the pseudo-living organism and the pseudo-non-living organism.

[0102] (Composition 12) The learning device according to any one of configurations 1 to 11, characterized in that the learning means learns not to calculate losses for mistakes between a specific living organism and a pseudo-living organism, or between a specific non-living organism and a pseudo-non-living organism. (Composition 13) The learning device according to any one of configurations 1 to 12, characterized in that the learning means performs learning corresponding to an identification device having multiple classification heads consisting of two classes. (Composition 14) A calculation means for calculating the similarity between the attributes of living organisms and pseudo-living organisms, and between the attributes of non-living organisms and pseudo-non-living organisms, The system further includes an integration means for integrating a pseudo-living organism into a living organism or a pseudo-non-living organism into a non-living organism when the similarity calculated by the calculation means is equal to or greater than a predetermined value. The learning device according to any one of configurations 1 to 13, characterized in that the learning means performs integration by the integration means to learn how to identify the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism.

[0103] (Composition 15) An identification device characterized by identifying whether a subject included in an input image is a living organism or a non-living organism using a model trained by a learning device described in any of configurations 1 to 14. (Composition 16) The identification device according to configuration 15, characterized in that it identifies whether a subject included in an input image is a living organism or a non-living organism by calculating the similarity of the subject included in the input image to the attributes of living organism, non-living organism, pseudo-living organism, and pseudo-non-living organism using the aforementioned model.

[0104] (method) A method for controlling a learning device that learns a model for identifying whether the subject of an image is a real living organism or a non-living artificial object, An acquisition step to acquire a biological image including the said living organism and a non-living image including the said non-living organism, A generation step involves processing the biological image or non-biological image acquired in the acquisition step to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism. A learning step involves learning a model that identifies the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism based on the living organism image and non-living organism image acquired in the acquisition step and at least one of the pseudo-living organism image and pseudo-non-living organism image generated in the generation step. A control method for a learning device, characterized by having the following features.

[0105] (program) A program for controlling a learning device that trains a model to distinguish whether the subject of an image is a real living organism or a non-living artificial object, An acquisition step to acquire a biological image including the said living organism and a non-living image including the said non-living organism, A generation step involves processing the biological image or non-biological image acquired in the acquisition step to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism. A learning step involves learning a model that identifies the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism based on the living organism image and non-living organism image acquired in the acquisition step and at least one of the pseudo-living organism image and pseudo-non-living organism image generated in the generation step. A program that causes a computer to execute something. [Explanation of symbols]

[0106] 122 Data acquisition unit, 123 Pseudo-non-living generation unit, 124 Pseudo-living generation unit, 125 Feature vector calculation unit, 126 Living organism determination unit, 127 Non-living organism determination unit, 128 Pseudo-non-living organism determination unit, 129 Pseudo-living organism determination unit, 130 Loss calculation unit, 131 Gradient calculation unit, 132 NN update unit

Claims

1. A learning device that trains a model to distinguish whether the subject of an image is a real living organism or a non-living artificial object, Acquisition means for acquiring a biological image including the said biological organism and a non-biological image including the said non-biological organism, A generation means that processes a biological image or non-biological image acquired by the acquisition means to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism, A learning means for learning a model that identifies the attributes of a living organism, a non-living organism, a pseudo-living organism, and a pseudo-non-living organism based on the living organism image and non-living organism image acquired by the acquisition means and at least one of the pseudo-living organism image and pseudo-non-living organism image generated by the generation means, A learning device characterized by having the following features.

2. The learning device according to claim 1, characterized in that the subject is a person's face.

3. The learning device according to claim 1, characterized in that the generation means performs image processing on the biological image or the non-biological image to generate the pseudo-non-biological image.

4. The learning device according to claim 1, characterized in that the generation means generates the pseudo-non-biological image by synthesizing a plurality of images.

5. The learning device according to claim 1, characterized in that the generation means generates the pseudo-non-living image by superimposing another image on the living image or the non-living image.

6. The learning device according to claim 1, characterized in that the generation means generates the pseudo-non-biological image by processing according to the feature point information in the biological image or the non-biological image.

7. The learning device according to claim 1, characterized in that the generation means generates the pseudo-non-biological image by performing processing to reproduce reflection.

8. The learning device according to claim 1, characterized in that the generation means generates the pseudo-biological image from the biological image by changing the shooting environment.

9. The learning device according to claim 1, characterized in that the aforementioned artifact is a printed material containing a living organism or a display that shows a living organism.

10. The learning device according to claim 1, characterized in that the learning means performs learning by calculating a loss based on the similarity to the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism.

11. The learning device according to claim 1, characterized in that the attributes are the imaging environment for the living organism, the method or material using the artificial object for the non-living organism, and the processing method for the pseudo-living organism and the pseudo-non-living organism.

12. The learning device according to claim 1, characterized in that the learning means learns not to calculate losses due to mistakes between a specific living organism and a pseudo-living organism, or between a specific non-living organism and a pseudo-non-living organism.

13. The learning device according to claim 1, characterized in that the learning means performs learning corresponding to an identification device having a plurality of classification heads consisting of two classes.

14. A calculation means for calculating the similarity between the attributes of living organisms and pseudo-living organisms, and between the attributes of non-living organisms and pseudo-non-living organisms, The system further includes an integration means for integrating a pseudo-living organism into a living organism or a pseudo-non-living organism into a non-living organism when the similarity calculated by the calculation means is equal to or greater than a predetermined value. The learning device according to claim 1, characterized in that the learning means performs integration by the integration means to learn how to identify the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism.

15. An identification device characterized by identifying whether a subject included in an input image is a living organism or a non-living organism using a model learned by the learning device described in claim 1.

16. The identification device according to claim 15, characterized in that it identifies whether a subject included in the input image is a living organism or a non-living organism by calculating the similarity of the subject included in the input image to the attributes of living organism, non-living organism, pseudo-living organism, and pseudo-non-living organism using the aforementioned model.

17. A method for controlling a learning device that learns a model for identifying whether the subject of an image is a real living organism or a non-living artificial object, An acquisition step to acquire a biological image including the said living organism and a non-living image including the said non-living organism, A generation step involves processing the biological image or non-biological image acquired in the acquisition step to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism. A learning step involves learning a model that identifies the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism based on the living organism image and non-living organism image acquired in the acquisition step and at least one of the pseudo-living organism image and pseudo-non-living organism image generated in the generation step. A control method for a learning device, characterized by having the following features.

18. A program for controlling a learning device that trains a model to distinguish whether the subject of an image is a real living organism or a non-living artificial object, An acquisition step to acquire a biological image including the said living organism and a non-living image including the said non-living organism, A generation step involves processing the biological image or non-biological image acquired in the acquisition step to generate at least one of a pseudo-biological image including a pseudo-biological organism and a pseudo-non-biological image including a pseudo-non-biological organism. A learning step involves learning a model that identifies the attributes of the living organism, the non-living organism, the pseudo-living organism, and the pseudo-non-living organism based on the living organism image and non-living organism image acquired in the acquisition step and at least one of the pseudo-living organism image and pseudo-non-living organism image generated in the generation step. A program that causes a computer to execute something.

Citation Information

Patent Citations

  • Classification device, classification method, and program

    JP2021099834A