Information processing device, information processing method, and computer program

By training detection mechanisms with both annotated and unannotated images, the method addresses the inefficiencies in existing technologies, achieving faster and more accurate feature detection with reduced labor and time costs.

JP7859592B2Active Publication Date: 2026-05-15NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-03-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing information processing technologies face challenges in accurately detecting feature information from images, particularly when limited annotated data is available, leading to inefficiencies and high labor and time costs.

Method used

A method involving the use of both annotated and unannotated images to train a detection mechanism, employing a combination of supervised and unsupervised learning techniques to enhance the detection of feature information, reducing the reliance on extensive annotated data.

Benefits of technology

This approach enables accurate and efficient detection of feature information with reduced human effort and time, utilizing computationally efficient algorithms and neural networks to improve detection speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859592000001
    Figure 0007859592000001
  • Figure 0007859592000002
    Figure 0007859592000002
  • Figure 0007859592000003
    Figure 0007859592000003
Patent Text Reader

Abstract

An information processing device 1 includes: an acquisition unit 11 that acquires a correct-answer-present image having correct answer feature information of a correct answer, and a correct-answer-absent image having no feature information of the correct answer; and a training unit 12 that causes a detection mechanism to learn a detection method for detecting feature information from an input image by using correct-answer-present detection information based on the correct answer feature information and correct-answer-present feature information detected from the correct-answer-present image by the detection mechanism, and correct-answer-absent detection information based on correct-answer-absent feature information detected from the correct-answer-absent image by the detection mechanism and extended feature information detected from an extended image obtained by extending the correct-answer-absent image by the detection mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the technical fields of information processing apparatuses, information processing methods, and recording media.

Background Art

[0002] Patent Document 1 stores information of an identifier that outputs part-likeness regarding a predetermined part of a detection target, obtains the part-likeness for an image region included in an input image based on the information of the identifier, determines the position of the predetermined part in the input image based on the obtained part-likeness, stores information regarding the reference position of the predetermined part, obtains difference information between the reference position of the predetermined part and the determined position of the predetermined part based on the information of the reference position, and obtains a reliability indicating the possibility that the input image is an image of the detection target based on the difference information. A technique is described. Patent Document 2 generates a face orientation conversion image in which the face orientation represented in an input image is converted into a predetermined orientation according to a hypothesized face orientation for each of a plurality of hypothesized face orientations, and for each hypothesized face orientation, generates a reversed face image by reversing the face represented in the face orientation conversion image, and whether to convert the face orientation represented in the reversed face image to the hypothesized face orientation or convert the face orientation represented in the input image to a predetermined orientation according to the hypothesized face orientation, and based on the conversion result, calculates an evaluation value representing the degree of difference between the face represented in the reversed face image and the face represented in the input image for each hypothesized face orientation, and based on the evaluation value, specifies one of the hypothesized face orientations as the face orientation represented in the input image. A technique is described. Patent Document 3 acquires data regarding a detection target, extracts a common feature amount common to a plurality of candidates of attributes possessed by the detection target from the data, detects feature information for each of the plurality of candidates based on the common feature amount, discriminates the attribute of the detection target based on the data, and outputs feature information corresponding to the discriminated attribute. A technique is described. Non-patent document 1 describes a technique in which a network is trained using a large number of annotated, faceless images to reconstruct faces using an hourglass-shaped autoencoder, and then the weights of the autoencoder are fixed, and a branching network is added to train a feature point detection network using a small number of annotated, face images. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] International Publication No. 2013 / 122009 [Patent Document 2] Japanese Patent Publication No. 2018-097573 [Patent Document 3] International Publication No. 2022 / 003982 [Non-patent literature]

[0004] [Non-Patent Document 1] Bjorn Browatzki and Christian Wallraven. "3fabrec: Fast few-shot face alignment by reconstruction." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020. [Overview of the project] [Problems that the invention aims to solve]

[0005] This disclosure aims to provide an information processing device, an information processing method, and a recording medium that improve upon the technologies described in prior art documents. [Means for solving the problem]

[0006] One aspect of an information processing device includes an acquisition means for acquiring a correct image having correct feature information and a correct image not having correct feature information, and a learning means for causing the detection mechanism to learn a detection method for detecting feature information from an input image using correct feature information detected by the detection mechanism from the correct image, correct detection information based on the correct feature information, correct non-feature information detected by the detection mechanism from the non-feature image, and correct non-detection information based on extended feature information detected by the detection mechanism from an extended image obtained by extending the non-feature image.

[0007] One aspect of the information processing method involves acquiring a ground truth image having correct ground truth feature information and a ground truth image not having correct ground truth feature information, and training the detection mechanism to learn a detection method for detecting feature information from an input image using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth detection information based on the ground truth feature information, the ground truth feature information detected by the detection mechanism from the ground truth image not having correct ground truth, and the ground truth detection information based on the extended feature information detected by the detection mechanism from an extended image obtained by extending the ground truth image not having correct ground truth.

[0008] One embodiment of a recording medium contains a computer program that causes a computer to execute an information processing method for training a detection mechanism to detect feature information from an input image. This method involves acquiring a ground truth image having ground truth feature information and a ground truth image not having ground truth feature information, and using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth detection information based on the ground truth feature information, the ground truth feature information detected by the detection mechanism from the ground truth image not having ground truth, and the ground truth detection information based on the extended feature information detected by the detection mechanism from an extended image obtained by extending the ground truth image not having ground truth. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a block diagram showing the configuration of the information processing device in the first embodiment. [Figure 2] Figure 2 is a block diagram showing the configuration of the learning device in the second embodiment. [Figure 3] Figure 3 is a flowchart showing the flow of the learning operation of the learning device in the second embodiment. [Figure 4] Figure 4 is a conceptual diagram showing the flow of the learning operation of the learning device in the second embodiment. [Figure 5] Figure 5 is a block diagram showing the configuration of the detection device in the third embodiment. [Figure 6] Figure 6 is a flowchart showing the detection operation flow of the detection device in the third embodiment. [Figure 7] Figure 7 is a block diagram showing the configuration of the authentication device in the fourth embodiment. [Figure 8] Figure 8 is a flowchart showing the authentication operation flow of the authentication device in the fourth embodiment. [Modes for carrying out the invention]

[0010] The following describes embodiments of the information processing device, information processing method, and recording medium with reference to the drawings. [1: First Embodiment]

[0011] A first embodiment of the information processing device, information processing method, and recording medium will be described below. The first embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 1 to which the first embodiment of the information processing device, information processing method, and recording medium is applied. [1-1: Configuration of Information Processing Device 1]

[0012] FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1 in the first embodiment. As shown in FIG. 1, the information processing apparatus 1 includes an acquisition unit 11 and a learning unit 12. The acquisition unit 11 acquires a correct-image having correct correct feature information and a non-correct-image having no correct feature information. The learning unit 12 uses the correct feature information detected by the detection mechanism from the correct-image, the correct detection information based on the correct feature information, the non-correct feature information detected by the detection mechanism from the non-correct-image, and the non-correct detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the non-correct-image, to cause the detection mechanism to learn a detection method for detecting feature information from the input image. [1-2: Technical effects of the information processing apparatus 1]

[0013] The information processing apparatus 1 in the first embodiment causes the detection mechanism to learn a detection method for detecting feature information from the input image, using a correct-image having correct correct feature information and a non-correct-image having no correct feature information. By using, in addition to the correct-image having correct correct feature information, a non-correct-image having no correct feature information which is relatively easy to prepare, the information processing apparatus 1 can cause the detection mechanism to learn a detection method capable of detecting feature information accurately and relatively easily. [2: Second embodiment]

[0014] Next, a second embodiment of the information processing apparatus, the information processing method, and the recording medium will be described. Hereinafter, the second embodiment of the information processing apparatus, the information processing method, and the recording medium will be described using a learning apparatus 2 to which the second embodiment of the information processing apparatus, the information processing method, and the recording medium is applied. [2-1: Configuration of the learning apparatus 2]

[0015] Figure 2 is a block diagram showing the configuration of the learning device 2 in the second embodiment. As shown in Figure 2, the learning device 2 comprises an arithmetic unit 21 and a storage device 22. Furthermore, the learning device 2 may also comprise a communication device 23, an input device 24, and an output device 25. However, the learning device 2 does not have to comprise at least one of the communication device 23, the input device 24, and the output device 25. The arithmetic unit 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.

[0016] The arithmetic unit 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic unit 21 reads a computer program. For example, the arithmetic unit 21 may read a computer program stored in the storage device 22. For example, the arithmetic unit 21 may read a computer program stored in a computer-readable and non-temporary recording medium using a recording medium reading device (not shown) provided by the learning device 2 (for example, an input device 24 described later). The arithmetic unit 21 may obtain a computer program from a device (not shown) located outside the learning device 2 via a communication device 23 (or other communication device) (i.e., it may download or read the program). The arithmetic unit 21 executes the read computer program. As a result, logical functional blocks for performing the operations that the learning device 2 should perform are realized within the arithmetic unit 21. In other words, the arithmetic unit 21 can function as a controller for realizing the logical functional blocks necessary to perform the actions (in other words, processes) that the learning device 2 should perform.

[0017] Figure 2 shows an example of a logical functional block implemented within the arithmetic unit 21 to perform information processing operations. As shown in Figure 3, the arithmetic unit 21 implements an acquisition unit 211, which is a specific example of the "acquisition means" described in the appendix later, and a learning unit 212, which is a specific example of the "learning means" described in the appendix later. The learning unit 212 may also have a loss function calculation unit 2121 and a parameter update unit 2122. Furthermore, the arithmetic unit 21 implements a detection unit 213, a conversion unit 214, and an inverse conversion unit 215. However, at least one of the detection unit 213, the conversion unit 214, and the inverse conversion unit 215 does not have to be implemented within the arithmetic unit 21. The details of the operation of each of the acquisition unit 211, learning unit 212, detection unit 213, conversion unit 214, and inverse conversion unit 215 will be explained later with reference to Figures 3 and 4.

[0018] The storage device 22 is capable of storing desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic unit 21. The storage device 22 may temporarily store data that the arithmetic unit 21 uses temporarily when it is executing a computer program. The storage device 22 may store data that the learning device 2 stores long-term. The storage device 22 may include at least one of the following: RAM (Random Access Memory), ROM (Read Only Memory), hard disk drive, magneto-optical disk drive, SSD (Solid State Drive), and disk array device. In other words, the storage device 22 may include a non-temporary recording medium. The storage device 22 may also contain a learning data storage unit 221 and a parameter storage unit 222.

[0019] The learning data storage unit 221 may store images with correct answers and images without correct answers as learning data. Images with correct answers have correct answer feature information. Images with correct answers may also be annotated images on which correct answer feature information has been annotated. Images without correct answers do not have correct answer feature information. Images without correct answers may also be unannotated images on which correct answer feature information has not been annotated.

[0020] The learning data storage unit 221 may store, for example, 1,000 images with correct answers and 100,000 images without correct answers as learning data. The numbers 1,000 and 100,000 are just examples, but the learning data storage unit 221 may store a vastly larger number of images without correct answers compared to the images with correct answers.

[0021] The communication device 23 can communicate with devices outside the learning device 2 via a communication network (not shown). The communication device 23 may have a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), or USB (Universal Serial Bus).

[0022] The input device 24 is a device that receives information input to the learning device 2 from outside the learning device 2. For example, the input device 24 may include an operating device (e.g., at least one of a keyboard, mouse, and touch panel) that can be operated by the operator of the learning device 2. For example, the input device 24 may include a reading device that can read information recorded as data on an external recording medium that can be attached to the learning device 2.

[0023] The output device 25 is a device that outputs information to the outside of the learning device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (so-called display) capable of displaying an image that shows the information to be output. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (so-called speaker) capable of outputting sound. For example, the output device 25 may output information onto paper. That is, the output device 25 may include a printing device (so-called printer) capable of printing the desired information onto paper. [2-2: Learning actions performed by learning device 2]

[0024] The flow of learning operations performed by the learning device 2 will be explained with reference to Figures 3 and 4. Figure 3 is a flowchart showing the flow of learning operations performed by the learning device 2. Figure 4 is a conceptual diagram of the learning operations performed by the learning device 2. [Detection mechanism DM]

[0025] The detection mechanism DM detects feature information from the input image. In the second embodiment, the detection mechanism DM is a model trained to output facial feature points when a face image is input. The detection mechanism DM detects points of facial organs such as eyes, nose, and mouth. The detection mechanism DM may be a model that outputs the positions of both eyes, the tip of the nose, and both corners of the mouth as feature points, for example, as illustrated in Figure 4.

[0026] The parameters necessary for the detection mechanism DM to function are stored in the parameter storage unit 222. For example, the parameter storage unit 222 stores the parameters necessary for the detection mechanism DM to constitute a neural network.

[0027] The detection mechanism DM includes a convolutional neural network and a fully connected layer. The convolutional neural network and the fully connected layer are connected in series. In the following explanation, the convolutional neural network and the fully connected layer will not be distinguished.

[0028] The detection mechanism DM used in Network 1, illustrated in Figure 4(a), and Network 2, illustrated in Figure 4(b), has the same structure and learns using common weight parameters. That is, the detection mechanism DM includes a neural network composed of common weight parameters. Both Network 1 and Network 2 use the detection mechanism DM with common weight parameters. [Network 1: Operation using annotated image Ia]

[0029] As shown in Figure 3, the learning unit 212 trains the detection mechanism DM using the annotated image Ia as the ground truth image (step S20). Figure 4(a) is a conceptual diagram of the training operation of the detection mechanism DM using the annotated image Ia.

[0030] As shown in Figure 4(a), the acquisition unit 211 acquires an annotated image Ia. The annotated image Ia may have the correct coordinate values ​​of the correct feature points FP1 annotated. The acquisition unit 211 may also acquire the annotated image Ia from the training data storage unit 221.

[0031] The detection unit 213 detects annotated feature points FP2 from the annotated image Ia. The detection unit 213 inputs the annotated image Ia to the detection mechanism DM and outputs the annotated feature points FP2 output by the detection mechanism DM as the detection result. If the feature points are defined as illustrated in Figure 4, the detection mechanism DM may output 10 numerical values ​​representing the x and y coordinates of the 5 feature points, respectively.

[0032] The learning unit 212 uses the ground truth feature point FP1 and the ground truth detection information based on the annotated feature point FP2 detected by the detection mechanism DM from the annotated image Ia to train the detection mechanism DM on how to detect feature points. The loss function calculation unit 2121-a calculates the loss La based on a loss function in which the loss increases as the ground truth feature point FP1 and the annotated feature point FP2 become less similar. This loss La is called the supervised feature point error loss La. The loss function calculation unit 2121-a may, for example, calculate the supervised feature point error loss La based on the 10 coordinate values ​​output when the annotated image Ia is input to the detection mechanism DM and the coordinate value of the ground truth feature point FP1. The supervised feature point error loss La is the regression learning loss, and the loss decreases as the information annotated on the input image matches the detection result from the input image. [Network 2: Operation using unannotated image Ib]

[0033] As shown in Figure 3, the learning unit 212 trains the detection mechanism DM using the unannotated image Ib as the ground truth image (step S21). Figure 4(b) is a conceptual diagram of the training operation of the detection mechanism DM using the unannotated image Ib.

[0034] As shown in Figure 4(b), the acquisition unit 211 acquires an unannotated image Ib. Hereinafter, the unannotated image Ib will be referred to as the original image Ib. The detection unit 213 detects the original feature points FP3 from the original image Ib. The detection unit 213 inputs the original image Ib to the detection mechanism DM and outputs the original feature points FP3 output by the detection mechanism DM as the detection result.

[0035] The conversion unit 214 performs a transformation on the original image Ib to generate an extended image Ic, which is an extension of the original image Ib. For example, as illustrated in Figure 4(b), the conversion unit 214 may generate the extended image Ic by flipping the original image Ib horizontally. In addition to the horizontal flipping of the input image as illustrated in Figure 4(b), other extensions of the input image by the conversion unit 214 include translation of the input image and rotation of the input image.

[0036] The detection unit 213 detects the extended feature point FP4 from the extended image Ic. The detection unit 213 inputs the extended image Ic to the detection mechanism DM and outputs the extended feature point FP4 output by the detection mechanism DM as the detection result.

[0037] The inverse transform unit 215 applies the inverse transform of the transformation performed by the transformation unit 214 on the original image Ib to the extended feature point FP4. For example, as illustrated in Figure 4(b), the inverse transform unit 215 may generate the inverse transformed feature point FP5 by flipping the left and right sides of the extended feature point FP4.

[0038] The learning unit 212 uses the original feature point FP3 and the inversely transformed feature point FP5 to train the detection mechanism DM on a feature point detection method. The loss function calculation unit 2121-b calculates the loss Lb based on a loss function in which the loss increases as the original feature point FP3 and the inversely transformed feature point FP5 become less similar. This loss Lb is called the unsupervised invariant loss Lb. The loss function calculation unit 2121-b may, for example, calculate the unsupervised invariant loss Lb based on the coordinate values ​​obtained by inversely transforming the 10 coordinate values ​​output when the original image Ib is input to the detection mechanism DM and the 10 coordinate values ​​output when the augmented image Ic is input to the detection mechanism DM. The unsupervised invariant loss Lb decreases as the detection result with and without data augmentation matches. Matching results with and without data augmentation can be said to be a necessary condition that both are correctly detecting feature points.

[0039] Furthermore, as illustrated in Figure 4(b), the system can learn equally well the cases where the face is facing right and the case where the face is facing left. This prevents bias, such as being good at detecting faces facing right or left, but poor at detecting them from the other side.

[0040] Note that the processing order of steps S20 and S21 may be reversed, or they may be executed simultaneously. [Update parameters for detection mechanism DM]

[0041] The parameter update unit 2122 updates the weight parameters of the detection mechanism so that the supervised feature point error loss La and the unsupervised invariance loss Lb are minimized. The parameter update unit 2122 may also update the weight parameters of the detection mechanism so that L, expressed by the following equation 1, is minimized. L = La + λLb … (Equation 1)

[0042] The algorithm used to determine the above parameters in order to minimize the supervised feature point error loss La and the unsupervised invariance loss Lb may be any learning algorithm used in machine learning, such as gradient descent or backpropagation.

[0043] The parameter update unit 2122 may perform a parameter update operation each time it performs a detection operation on a predetermined number of face images. The predetermined number of face images may include images with and without ground truth in the same proportion as all face images stored in the learning data storage unit 221. For example, if the ratio of images with and without ground truth is 1:10, the predetermined number of face images may include 10 times the number of images with ground truth as the number of images without ground truth. For example, if the ratio of images with and without ground truth is 1:10, λ in equation 1 above may be set to "0.1". That is, the loss L may be calculated so that the supervised feature point error loss La and the unsupervised invariance loss Lb contribute equally.

[0044] Alternatively, the learning unit 212 may perform training using the unannotated image Ib, and then perform training using the annotated image Ia. Conversely, the learning unit 212 may perform training using the annotated image Ia, and then perform training using the unannotated image Ib.

[0045] The learning unit determines whether or not the learning process is complete (step S22). For example, the learning unit 212 may determine that the learning process is complete when it has performed learning using all the face images stored in the learning data storage unit 221. In another example, the learning unit 212 may determine that the learning process is complete when L, expressed in the above formula 1, falls below a predetermined threshold. Alternatively, the learning unit 212 may determine that the learning process is complete when the losses calculated in each step of step S20 and step S21 fall below a predetermined threshold. In yet another example, the learning unit 212 may determine that the learning process is complete when steps S20 and S21 have been repeated a predetermined number of times.

[0046] If the learning unit 212 determines that learning is complete (step S22; Yes), it determines the parameters of the detection mechanism DM and determines the detection mechanism DM (step S23). The learning unit 212 may also transmit the determined parameters of the detection mechanism DM to the device that performs the detection operation using the detection mechanism DM. On the other hand, if the learning unit 212 determines that learning is not complete (step S22; No), it returns to step S20. [2-3: Variations]

[0047] In the second embodiment described above, the case where the image feature information is the feature points of the image was used as an example, but the feature information may be information about features other than feature points. For example, the image feature information may be information that indicates a feature region within the image. Information that indicates a feature region may be information used for semantic segmentation, etc. The learning operation in the second embodiment may be broadly applied to tasks that take images as input.

[0048] Furthermore, although the second embodiment was described using the example of a face image, the image may be something other than a face. For example, it may be a body image of an organism. The organism may be a mammal including a human, a bird, etc. Alternatively, it may be an image containing numbers, such as a car license plate.

[0049] For example, in image recognition, the ability to accurately detect feature points is crucial. Feature points detected from faces can be used, for example, in facial recognition and expression recognition. Feature points detected from bodies can be used, for example, in estimating human body poses, analyzing human behavior, and human body matching. Feature points detected from bodies may also include, for example, joint points. [2-4: Technical effects of learning device 2]

[0050] In the second embodiment, the learning device 2 learns a detection method so that the detected feature information is similar to the correct feature information, and also learns a detection method so that the feature information detected from the image is similar to the feature information detected from the augmented image, thereby enabling it to learn a detection method that accurately detects feature information. The learning device 2 learns a detection method so that the feature information obtained by inversely transforming the feature information detected from the augmented image (which is created by transforming the original image) is similar to the feature information detected from the original image, thereby enabling it to learn a detection method that accurately detects feature information.

[0051] To accurately detect feature points from images, it is common to prepare a large number of annotated images (images with feature points marked) and use them for training. Preparing a large number of annotated images is a time-consuming and labor-intensive task. If accurate feature point detection can be achieved by training with a small number of annotated images, then both labor costs and time can be reduced.

[0052] Learning device 2 can relatively easily train its detection mechanism to accurately detect feature information by using both annotated and unannotated images, which are relatively easy to prepare. Learning device 2 uses a relatively large number of unannotated images, which are relatively easy to prepare, thus reducing the human cost and time required to prepare training data. Learning device 2 can train its detection mechanism to accurately detect feature points using a small number of annotated images.

[0053] Non-patent document 1, mentioned above, describes a comparative example that utilizes a small number of annotated facial images and a large number of annotated unannotated facial images. The comparative example described in Non-patent document 1 uses an hourglass-shaped network that requires an encoder and a decoder. In other words, there are limitations to the network. Hourglass-shaped network processing is computationally intensive, computationally inefficient, and slow in both learning and inference speeds. Furthermore, the comparative example is relatively computationally intensive because it derives feature points from the peaks of the heatmap of each feature point output by the hourglass-shaped network.

[0054] In contrast, the detection mechanism DM in this embodiment directly outputs feature information without using a heatmap. Since the detection mechanism DM in this embodiment does not use an hourglass network, the amount of computation can be reduced and computation efficiency is good. Also, since the detection mechanism DM in this embodiment directly outputs feature information, the processing is lighter compared to the processing of the comparative example. [3: Third Embodiment]

[0055] Next, a third embodiment of the information processing device, information processing method, and recording medium will be described. In the following, the third embodiment of the information processing device, information processing method, and recording medium will be described using a detection device 3 to which the third embodiment of the information processing device, information processing method, and recording medium is applied. [3-1: Configuration of detection device 3]

[0056] The configuration of the detection device 3 in the third embodiment will be described with reference to Figure 5. Figure 5 is a block diagram showing the configuration of the detection device 3 in the third embodiment.

[0057] As shown in Figure 5, the detection device 3 in the third embodiment includes a processing unit 21 and a storage device 22, similar to the learning device 2 in the second embodiment. Furthermore, the detection device 3 in the third embodiment may also include a communication device 23, an input device 24, and an output device 25, similar to the learning device 2 in the second embodiment. However, the detection device 3 does not need to include at least one of the communication device 23, the input device 24, and the output device 25. The detection device 3 in the third embodiment differs from the learning device 2 in that it performs detection operations using the detection mechanism DM learned by the learning device 2 in the second embodiment. Other features of the detection device 3 may be the same as other features of the learning device 2. For this reason, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate. [3-2: Detection operations performed by detection device 3]

[0058] The flow of the detection operation performed by the detection device 3 will be explained with reference to Figure 6. Figure 6 is a flowchart showing the detection operation performed by the detection device 3. In the third embodiment, the case of detecting facial feature points from the face region of an image will be explained as an example. In the parameter storage unit 322 in the third embodiment, the parameters of the detection mechanism DM learned in the second embodiment are stored.

[0059] As shown in Figure 6, the acquisition unit 311 acquires an image (step S30). The face detection unit 316 detects the face region from the image (step S31). The feature detection unit 313 detects feature points from the face region using the detection mechanism DM (step S32). The feature points detected from the face region may be applied to face recognition, facial expression analysis, etc.

[0060] Furthermore, if the input image is a human body image, the detection mechanism DM may output features such as elbows, knees, and heels as characteristic points. [3-3: Technical effects of detection device 3]

[0061] The detection mechanism DM utilizes computationally efficient algorithms and convolutional neural networks, enabling high-speed computation. In the third embodiment, the detection device 3 uses the detection mechanism DM to detect feature information from an image, allowing for accurate and high-speed detection of feature information. [4: Fourth Embodiment]

[0062] Next, a fourth embodiment of the information processing device, information processing method, and recording medium will be described. In the following, the fourth embodiment of the information processing device, information processing method, and recording medium will be described using an authentication device 4 to which the fourth embodiment of the information processing device, information processing method, and recording medium is applied. [4-1: Configuration of Authentication Device 4]

[0063] The configuration of the authentication device 4 in the fourth embodiment will be described with reference to Figure 7. Figure 7 is a block diagram showing the configuration of the authentication device 4 in the fourth embodiment.

[0064] As shown in Figure 7, the authentication device 4 in the fourth embodiment includes a processing unit 21 and a storage device 22, similar to the learning device 2 in the second embodiment and the detection device 3 in the third embodiment. Furthermore, the authentication device 4 in the fourth embodiment may also include a communication device 23, an input device 24, and an output device 25, similar to the learning device 2 in the second embodiment and the detection device 3 in the third embodiment. However, the authentication device 4 does not have to include at least one of the communication device 23, the input device 24, and the output device 25. The authentication device 4 in the fourth embodiment differs from the learning device 2 in the second embodiment and the detection device 3 in the third embodiment in that it performs authentication operations using the detection mechanism DM learned by the learning device 2 in the second embodiment. Other features of the authentication device 4 may be the same as other features of at least one of the learning device 2 and the detection device 3. For this reason, the parts that differ from the embodiments already described will be explained in detail below, and other overlapping parts will be omitted as appropriate. [4-2: Detection operations performed by authentication device 4]

[0065] The flow of the detection operation performed by the authentication device 4 will be explained with reference to Figure 8. Figure 8 is a flowchart of the detection operation performed by the authentication device 4. In the fourth embodiment, the case of facial recognition using the face region of an image will be explained as an example. In the parameter storage unit 322 in the fourth embodiment, the parameters of the detection mechanism DM learned in the second embodiment are stored.

[0066] As shown in Figure 8, the acquisition unit 311 acquires an image (step S30). The face detection unit 316 detects the face region from the image (step S31). The feature detection unit 313 detects feature points from the face region using the detection mechanism DM (step S32). The normalization unit 417 normalizes the face image by aligning the orientation of the face to a predetermined orientation and aligning the size of the face to a predetermined size (step S40). The extraction unit 418 extracts feature quantities from the normalized face (step S41). The extraction unit 418 extracts feature quantities useful for face recognition from the face image. The authentication unit 419 compares and matches the feature quantities extracted in step S41 with the feature quantities registered in the registration information storage unit 423 to determine whether they are the same person or not (step S42). [4-3: Technical effects of authentication device 4]

[0067] The detection mechanism DM can detect feature points from a facial image accurately and quickly. Therefore, the authentication device 4 in the fourth embodiment performs facial authentication based on the feature points detected from the facial image using the detection mechanism DM, thus enabling accurate and fast facial authentication. [5: Addendum]

[0068] The following additional information is disclosed regarding the embodiments described above. [Note 1] An acquisition means for acquiring a correct image with correct feature information and a correct image without correct feature information, A learning means for causing the detection mechanism to learn a detection method for detecting feature information from an input image, using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth detection information based on the ground truth feature information, the ground truth feature information detected by the detection mechanism from the ground truth image, and the ground truth detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the ground truth image. An information processing device equipped with the following features. [Note 2] The aforementioned ground truth detection information is information based on a ground truth loss function, where the loss increases as the ground truth feature information becomes less similar to the ground truth feature information. The aforementioned ground truth no-detection information is information based on a ground truth lossless function, where the loss increases as the ground truth no-feature information and the extended feature information become less similar. An information processing device as described in Appendix 1, comprising: [Note 3] A conversion means that performs a conversion on the aforementioned no-ground image to generate the extended image obtained by expanding the said no-ground image, The detection mechanism provides an inverse transformation means for applying the inverse transformation of the transformation to the extended feature information detected from the extended image. An information processing device according to Appendix 1 or 2, comprising: [Note 4] The acquisition means acquires more images without a correct answer compared to the images with a correct answer. An information processing device as described in any one of the items 1 to 3 in the appendix. [Note 5] Detection means for detecting feature information from an image using the aforementioned detection mechanism An information processing device described in any one of the appendices 1 to 4, comprising the above. [Note 6] An extraction means for extracting feature quantities from the image using the aforementioned feature information, Authentication means for authenticating the image using the aforementioned features, An information processing device described in any one of the appendices 1 to 5, comprising the features described herein. [Note 7] The aforementioned image is a facial image. An information processing device as described in any one of the items 1 to 6 of the appendix. [Note 8] The aforementioned feature information is the feature points of the image. An information processing device as described in any one of the items 1 through 7 of the appendix. [Note 9] We obtain ground truth images that have correct feature information, and ground truth images that do not have correct feature information. The detection mechanism is trained to learn a detection method for detecting feature information from an input image, using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth feature detection information based on the ground truth feature information, the ground truth featureless information detected by the detection mechanism from the ground truth image, and the ground truth non-detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the ground truth image. Information processing methods. [Note 10] A recording medium storing a computer program that causes a computer to execute an information processing method for training a detection mechanism to detect feature information from an input image, using the correct image having correct feature information and the image not having correct feature information, the correct detection information detected by the detection mechanism from the image having correct feature information and the correct detection information based on the correct feature information, the correct non-feature information detected by the detection mechanism from the image not having correct feature information and the correct non-detection information based on the extended feature information detected by the detection mechanism from an extended image obtained by extending the image not having correct feature information.

[0069] This disclosure may be modified as appropriate, insofar as it does not contradict the technical idea that can be inferred from the claims and the entire specification. Information processing devices, information processing methods, and recording media that include such modifications are also included in the technical idea of ​​this disclosure. [Explanation of Symbols]

[0070] 1. Information Processing Device 11,211,311 Acquisition Department 12,212 Learning Department 2 Learning device 2121 Loss Function Calculation Unit 2122 Parameter update section DM detection mechanism 213 Detection unit 214 Conversion Unit 215 Inverse Transform Section 221 Learning Data Storage Unit 222,322 Parameter storage unit 3. Detection device 313 Feature detection unit 316 Face detection unit 4 Authentication device 417 Normalization section 418 Extraction part 419 Authentication Department 423 Registration Information Storage Unit

Claims

1. An acquisition means for acquiring a correct image with correct feature information and a correct image without correct feature information, A learning means for causing the detection mechanism to learn a detection method for detecting feature information from an input image, using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth detection information based on the ground truth feature information, the ground truth feature information detected by the detection mechanism from the ground truth image, and the ground truth detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the ground truth image. An information processing device equipped with the following features.

2. The aforementioned ground truth detection information is information based on a ground truth loss function, where the loss increases as the ground truth feature information becomes less similar to the ground truth feature information. The aforementioned ground truth no-detection information is information based on a ground truth lossless function, where the loss increases as the ground truth no-feature information and the extended feature information become less similar. The information processing apparatus according to claim 1, comprising:

3. A conversion means that performs a conversion on the aforementioned no-ground image to generate the extended image obtained by expanding the said no-ground image, The detection mechanism provides an inverse transformation means for applying the inverse transformation of the transformation to the extended feature information detected from the extended image. The information processing apparatus according to claim 1 or 2, comprising:

4. The acquisition means acquires more images without a correct answer compared to the images with a correct answer. The information processing apparatus according to claim 1 or 2.

5. Detection means for detecting feature information from an image using the aforementioned detection mechanism The information processing apparatus according to claim 1 or 2, comprising:

6. An extraction means for extracting feature quantities from the image using the aforementioned feature information, Authentication means for authenticating the image using the aforementioned features, The information processing apparatus according to claim 1 or 2, comprising:

7. The aforementioned image is a facial image. The information processing apparatus according to claim 1 or 2.

8. The aforementioned feature information is the feature points of the image. The information processing apparatus according to claim 1 or 2.

9. We obtain ground truth images that have correct feature information, and ground truth images that do not have correct feature information. The detection mechanism is trained to learn a detection method for detecting feature information from an input image, using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth feature detection information based on the ground truth feature information, the ground truth featureless information detected by the detection mechanism from the ground truth image, and the ground truth non-detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the ground truth image. The information processing method performed by computers.

10. On the computer, We obtain ground truth images that have correct feature information, and ground truth images that do not have correct feature information. The detection mechanism is trained to learn a detection method for detecting feature information from an input image, using the ground truth feature information detected by the detection mechanism from the ground truth image, the ground truth feature detection information based on the ground truth feature information, the ground truth featureless information detected by the detection mechanism from the ground truth image, and the ground truth non-detection information based on the extended feature information detected by the detection mechanism from the extended image obtained by extending the ground truth image. A computer program that executes an information processing method.