Information processing device, information processing method, and recording medium

The apparatus generates face images with controlled similarity and variation using machine-learned models, addressing data collection challenges and enhancing face recognition model training.

WO2025158567A1PCT designated stage Publication Date: 2025-07-31NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002014
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for generating face images struggle to produce images with desired levels of similarity to input faces while maintaining individuality and variation, making it difficult to train effective face recognition models due to privacy concerns and limitations in collecting diverse face data.

Method used

An information processing apparatus and method that uses machine-learned generation models, such as GANs and diffusion models, to generate new face images based on input scores and texts, allowing control over the degree of similarity and adding specific information, and a learning mechanism to adjust the generation model parameters.

Benefits of technology

Enables the generation of face images with controlled similarity and variation, facilitating effective training of face authentication and verification models by overcoming data collection limitations and improving model training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002014_31072025_PF_FP_ABST
    Figure JP2024002014_31072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises: an acquisition means for acquiring an arbitrary input score and an input facial image including a target face; and a generation means for generating, on the basis of the input facial image and the input score, a new facial image for which the value of a collation score indicating the degree of coincidence with the input facial image is the value of the input score. Such an information processing device can generate a new facial image that has a desired degree of coincidence with an input facial image. The generated new facial image can be used to, for example, train a face authentication model.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and recording medium

[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.

[0002] Known examples of this type of device include those that generate facial images of people using a generative model trained by machine learning. For example, Patent Literature 1 discloses training a facial image generator that generates facial images based on feedback from facial recognition of facial images.

[0003] Special table 2021-534526 publication

[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the techniques disclosed in prior art documents.

[0005] One aspect of the information processing device disclosed herein includes an acquisition means for acquiring an input face image including a target face and an arbitrary input score, and a generation means for generating a new face image based on the input face image and the input score, the new face image having a matching score value indicating the degree of match with the input face image equal to the value of the input score.

[0006] One aspect of the information processing method disclosed herein is to use at least one computer to obtain an input face image including a target face and an arbitrary input score, and to generate a new face image based on the input face image and the input score, in which the value of the matching score indicating the degree of match with the input face image is the value of the input score.

[0007] One aspect of the recording medium of this disclosure is a recording medium having a computer program recorded thereon that causes at least one computer to execute an information processing method, which acquires an input face image including a target face and an arbitrary input score, and generates a new face image based on the input face image and the input score, the new face image having a matching score value indicating the degree of match with the input face image equal to the value of the input score.

[0008] 1 is a block diagram showing the hardware configuration of a first information processing device. FIG. 2 is a block diagram showing the functional configuration of the first information processing device. FIG. 3 is a flowchart showing the flow of a generation operation by the first information processing device. FIG. 4 is a conceptual diagram showing an example of a generation operation by the first information processing device. FIG. 5 is a conceptual diagram showing a specific example of a matching score used in the first information processing device. FIG. 6 is a conceptual diagram showing an example of a generation operation by a second information processing device. FIG. 7 is a conceptual diagram showing an example of a generation operation by a third information processing device. FIG. 8 is a block diagram showing the functional configuration of a fourth information processing device. FIG. 9 is a flowchart showing the flow of a learning operation by the fourth information processing device. FIG. 10 is a conceptual diagram showing an example of a learning operation by the fourth information processing device.

[0009] Hereinafter, embodiments of an information processing device, an information processing method, and a recording medium will be described with reference to the drawings.

[0010] First Embodiment A first information processing apparatus will be described with reference to FIGS. 1 to 5. FIG.

[0011] (Hardware Configuration) First, the hardware configuration of the first information processing apparatus will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the first information processing apparatus.

[0012] 1, a first information processing device 10 includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The information processing device 10 may further include an input device 15 and an output device 16. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected to each other via a data bus 17. The data bus 17 may be an interface other than a data bus (for example, a LAN, a USB, etc.).

[0013] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the information processing device 10 via a network interface. The processor 11 controls the RAM 12, the storage device 14, the input device 15, and the output device 16 by executing the loaded computer program. In particular, in this embodiment, when the processor 11 executes the loaded computer program, functional blocks for executing processing related to the generation of a facial image are realized within the processor 11. In other words, the processor 11 may function as a controller that executes each control in the information processing device 10.

[0014] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a quantum processor. The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.

[0015] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that the processor 11 temporarily uses while it is executing the computer programs. The RAM 12 may be, for example, a dynamic random access memory (D-RAM) or a static random access memory (SRAM). Alternatively, other types of volatile memory may be used instead of the RAM 12.

[0016] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a programmable read-only memory (PROM) or an erasable read-only memory (EPROM). Alternatively, other types of non-volatile memory may be used instead of the ROM 13.

[0017] The storage device 14 stores data that the information processing device 10 stores for a long period of time. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device may store computer programs executed by the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.

[0018] The input device 15 is a device that receives input instructions from a user of the information processing device 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as part of a smartphone, a tablet terminal, an earphone-type terminal, a watch-type terminal, an HMD (Head Mounted Display) terminal, etc. The input device 15 may be, for example, a device that includes a microphone and is capable of voice input.

[0019] The output device 16 is a device that outputs information related to the information processing device 10 to the outside. For example, the output device 16 may be a display device (e.g., a display or digital signage) that can display information related to the information processing device 10. The output device 16 may also be a speaker or the like that can output information related to the information processing device 10 as audio.

[0020] 1 may be configured to be included in a device external to the first information processing device 10. For example, the first information processing device 10 may be configured to include a processor 11, a RAM 12, and a ROM 13, and the other devices, such as a storage device 14, an input device 15, and an output device 16, may be configured as external devices. That is, the first information processing device 10 may be configured as an information processing system including a plurality of different devices. Furthermore, some of the calculation functions of the first information processing device 10 may be realized by an external server, a cloud, or the like.

[0021] (Functional Configuration) Next, the functional configuration of the first information processing device 10 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the first information processing device.

[0022] 2, the first information processing device 10 is configured to include, as processing blocks for realizing its functions, a data acquisition unit 110 and an image generation unit 120. Note that each of the data acquisition unit 110 and the image generation unit 120 may be realized by, for example, the above-mentioned processor 11 (see FIG. 1).

[0023] The data acquisition unit 110 is configured to be able to acquire an input face image including the face of a target. The target is typically a human, but may also be an animal such as a cat or a dog. The data acquisition unit 110 may acquire the input face image, for example, from a database that stores a plurality of images. Alternatively, the data acquisition unit 110 may acquire the input face image by photographing the target with a camera. The data acquisition unit 110 is also configured to be able to acquire an arbitrary input score. The input score is a parameter used when generating a new face image in the image generation unit 120, which will be described later. The input score is an arbitrary value and may be input by, for example, a user operation. The input face image and input score acquired by the data acquisition unit 110 are each configured to be output to the image generation unit 120.

[0024] The image generation unit 120 is configured to be able to generate a new face image (i.e., a face image different from the input face image) based on the input face image and input score acquired by the data acquisition unit 110. Specifically, the image generation unit 120 is configured to generate a new face image (hereinafter, appropriately referred to as a "generated face image") in which the value of the matching score indicating the degree of match with the input face image is the value of the input score. The image generation unit 120 may be configured to generate the generated face image using an image generation model configured by, for example, a neural network. This image generation model may be one that has been previously learned by machine learning. The learning of the image generation model will be described in detail in another embodiment described later.

[0025] (Generation Operation) Next, the flow of the generation operation (i.e., a series of operations when generating a new face image) in the first information processing device 10 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of the generation operation by the first information processing device.

[0026] 3, when the generation operation by the first information processing device 10 is started, the data acquisition unit 110 first acquires an input face image and an input score (step S101). The input face image and input score acquired by the data acquisition unit 110 are output to the image generation unit 120.

[0027] Next, the image generation unit 120 generates a generated face image based on the input face image and the input score acquired by the data acquisition unit 110 (step S102). For example, the image generation unit 120 inputs the input face image and the input score to an image generation model, and obtains a generated face image as an output of the image generation model.

[0028] Note that image generation unit 120 may be configured to generate a plurality of different generated face images for one input face image. In this case, a random number may be input to image generation unit 120 in addition to the input face image and input score. By inputting the random number, even if the input core image and input score are the same, the generated face images are generated as different images according to the random number.

[0029] Next, the image generation unit 120 outputs the generated facial image (step S103). Note that the output destination of the generated facial image is not particularly limited. For example, the image generation unit 120 may output the generated facial image to a display or the like. Alternatively, the image generation unit 120 may output the generated facial image to a storage device or the like that stores images.

[0030] By repeatedly executing the above-described series of processes, generated face images according to different input face images and input scores may be generated. In this case, for example, the input face image may be kept fixed and only the input score may be changed to generate generated face images. In this way, generated face images according to different input scores can be generated from the same input face image. Alternatively, the input score may be kept fixed and only the input face image may be changed to generate generated face images. In this way, generated face images according to a common input score can be generated from different input face images.

[0031] (Specific Operation Example) Next, a specific operation example regarding the above-mentioned generation operation will be described with reference to Fig. 4 and Fig. 5. Fig. 4 is a conceptual diagram showing an example of a generation operation by the first information processing device. Fig. 5 is a conceptual diagram showing a specific example of a matching score used in the first information processing device.

[0032] As shown in FIG. 4 , the image generation unit 120 may input an input facial image and an input score S into the image generation model 125 to generate a generated facial image. The generated facial image is generated as an image whose matching score with the input facial image is the value of the input score S. Therefore, the generated facial image is generated as a facial image including an object that resembles an object included in the input facial image. For example, when the input score S is input as a relatively high value, the object included in the input facial image and the object included in the generated facial image are more likely to be recognized as the same person. Conversely, when the input score S is input as a relatively low value, the object included in the input facial image and the object included in the generated facial image are less likely to be recognized as the same person.

[0033] As shown in FIG. 5, the matching score S C may be a value calculated from the feature amount of the input face image and the feature amount of the generated face image. For example, when the input face image is input to the feature amount extractor 200, the feature amount F A The generated face image is input to the feature extractor 200 to obtain the feature F B In this case, the matching score S C may be a value calculated as the distance d between each feature. C = d(F A , F B ) The above-mentioned matching score S C The calculation method is merely an example and is not particularly limited.

[0034] (Technical Effects) Next, technical effects obtained by the first information processing device 10 will be described.

[0035] 1 to 5, the first information processing device 10 generates a new face image based on an input face image and an input score. In this way, it is possible to generate a new face image that matches the input face image to a desired degree of similarity. For example, it is possible to generate an image in which an object included in the input face image resembles an object included in the generated face image with a desired degree of similarity.

[0036] The generated facial images generated by the first information processing device 10 may be used, for example, to train a facial recognition model that performs facial recognition by matching facial images. Specifically, the generated facial images may be used as training data for training the facial recognition model. When training a facial recognition model, it is necessary to prepare a sufficient number of facial images as training data, but due to increasing awareness of privacy and personal information protection, it is becoming more difficult to collect facial images. However, the first information processing device 10 can generate a new facial image from an input facial image, making it easy to collect facial images as training data.

[0037] When attempting to generate new facial images using existing methods, the resulting images tend to all look similar. Such images have little intra-personal variation and are therefore unsuitable for training a facial recognition model. On the other hand, attempts to increase variation tend to result in facial images of different people. Such images also lose individuality and are therefore unsuitable for training a facial recognition model. However, the first information processing device 10 can adjust the degree of matching of facial images (in other words, the difficulty of learning) according to the input score value when generating the image. Therefore, it is possible to prepare facial images with a wide variety while preserving individuality, thereby enabling appropriate training of a facial recognition model.

[0038] Second Embodiment A second information processing device 10 will be described with reference to Fig. 6. The second information processing device 10 differs in some configurations and operations from the first information processing device 10 described above, but other parts may be similar to the first information processing device 10. Therefore, the following will describe in detail the parts that differ from the first embodiment, and will omit explanations of other overlapping parts as appropriate.

[0039] (Example of Generative Model) First, a generative model used in the second information processing device 10 will be described with reference to Fig. 6. Fig. 6 is a conceptual diagram showing an example of a generating operation by the second information processing device. Note that in Fig. 6, elements similar to those shown in Fig. 4 are assigned the same reference numerals.

[0040] In FIG. 6, an image generation unit 120 in the second information processing device 10 generates a generated face image from an input face image using an image generation model 125 configured as a Generative Adversarial Network (GAN).

[0041] For example, the image generation model 125 may be configured using StyleGAN. In this case, the image generation unit 120 first inversely calculates an input vector (latent variable) corresponding to an input facial image (GAN inversion). Then, the image generation unit 120 minutely changes a part of the calculated input vector. Thereafter, the image generation model 125 generates a generated facial image based on the minutely changed input vector.

[0042] The generated facial image output from the image generation model 125 is a facial image that is slightly different from the input facial image due to the slight change in the input vector. Therefore, when the image generation unit 120 makes a slight change in the input vector, it may change the input vector according to the input score S. For example, if the input score S has a high value, the amount of change in the input vector may be reduced. In this case, a generated facial image that is very close to the input facial image will be output. On the other hand, if the input score S has a low value, the amount of change in the input vector may be increased. In this case, a generated facial image that is slightly different from the input facial image will be output.

[0043] (Technical Effects) Next, technical effects obtained by the second information processing device 10 will be described.

[0044] 6, in the second information processing device 10, a generated facial image is generated by the image generation model 125 configured by a generative adversarial network. In this way, it is possible to appropriately generate a new facial image that has a desired degree of match when matched with an input facial image.

[0045] <Third embodiment> A third information processing device 10 will be described with reference to Fig. 7. The third information processing device 10 differs in some configurations and operations from the first and second information processing devices 10 described above, but other parts may be similar to the first and second information processing devices 10. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.

[0046] (Example of Generative Model) First, a generative model used in the third information processing device 10 will be described with reference to Fig. 7. Fig. 7 is a conceptual diagram showing an example of a generating operation by the third information processing device. Note that in Fig. 7, the same elements as those shown in Figs. 4 and 6 are denoted by the same reference numerals.

[0047] In FIG. 7, the image generation unit 120 in the third information processing apparatus 10 generates a generated facial image from an input facial image using an image generation model 125 configured as a diffusion model that generates a facial image based on an input text.

[0048] For example, the image generation model 125 may be configured using stable diffusion. In this case, input text is input to the image generation unit 120 in addition to the input face image and the input score S. The input text may be acquired by, for example, the data acquisition unit 110. The input text may be any text input by, for example, a user.

[0049] The generated facial image output from the image generation model 125 is one to which information from the input text has been added. In the example shown in FIG. 7 , the word "glasses" is input as input text. Therefore, glasses corresponding to the input text are added to the target face in the generated facial image. In this way, the input text may be information regarding accessories to be added to the target face. For example, when the input text "cap" is input, a facial image wearing a hat may be generated. Also, when the input text "mask" is input, a facial image wearing a mask may be generated.

[0050] Note that the input text may include information other than the above-described attachment, as long as it is information about the target face. For example, the input text may be information specifying the direction of the face. For example, if the input text "right" is input, a face image of a profile face facing to the right may be generated. Also, if the input text "up" is input, a face image of a face facing up may be generated.

[0051] (Technical Effects) Next, technical effects obtained by the third information processing apparatus 10 will be described.

[0052] As described in FIG. 7 , in the third information processing device 10, a new facial image is generated based on the input text in addition to the input facial image and input score. In this way, it is possible to add information corresponding to the input text to the new facial image to be generated. Therefore, a new facial image can be generated more appropriately compared to when the input text is not used. Note that if the data acquisition unit 110 is configured to acquire the input text in addition to the input image and input score, it is possible to generate a facial image corresponding to the input text more efficiently.

[0053] <Fourth embodiment> A fourth information processing device 10 will be described with reference to Figures 8 to 10. The third information processing device 10 differs in some configurations and operations from the first to third information processing devices 10 described above, but other parts may be similar to the first to third information processing devices 10. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.

[0054] (Functional Configuration) First, the functional configuration of the fourth information processing device 10 will be described with reference to Fig. 8. Fig. 8 is a block diagram showing the functional configuration of the fourth information processing device. Note that in Fig. 8, the same elements as those described in Fig. 2 are denoted by the same reference numerals.

[0055] 8, the fourth information processing device 10 is configured to include, as processing blocks for realizing its functions, a data acquisition unit 110, an image generation unit 120, and a learning unit 130. That is, the fourth information processing device 10 further includes the learning unit 130 in addition to the configuration described in the first embodiment (see FIG. 2). Note that the learning unit 130 may be realized by, for example, the above-mentioned processor 11 (see FIG. 1).

[0056] The learning unit 130 is configured to train the image generation unit 120. Specifically, the learning unit 130 may be configured to perform machine learning of the image generation model 125 used by the image generation unit 120. The learning unit 130 uses a first facial image, a second facial image, and a learning score as learning data. The first facial image and the second facial image are facial images including a face of a common target. The learning score is a matching score obtained when the first facial image and the second facial image are matched. The learning score may be obtained by matching the first facial image and the second facial image in advance. Note that if the image generation model 125 is a diffusion model using input text described in the third embodiment (see FIG. 7 ), the learning data may include input text in addition to the first facial image, the second facial image, and the learning score. In this case, the input text may indicate information added to the second facial image when the first facial image and the second facial image are compared.

[0057] The learning unit 130 includes a loss calculation unit 131 and a parameter update unit 132. The loss calculation unit 131 is configured to calculate a loss (in other words, a difference between the generated face image and the second face image) by comparing a generated face image generated from the first face image and the learning score with the second face image. The parameter update unit 132 is configured to update parameters of the image generation model 125 so as to reduce the loss calculated by the loss calculation unit 132. For example, the parameter update unit 132 calculates the gradient of a loss function indicating the loss calculated by the loss calculation unit 131, and adjusts the parameters in a direction that reduces the loss function. In other words, the learning unit 130 may learn the image generation model 125 using an error backpropagation algorithm.

[0058] (Learning Operation) Next, the flow of the learning operation (i.e., a series of operations when learning the generative model 125) in the fourth information processing device 10 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of the learning operation by the fourth information processing device.

[0059] As shown in FIG. 9, when the learning operation by the fourth information processing device 10 is started, a learning data set including a plurality of sets of a first face image, a second face image, and a learning score is first acquired (step S301).

[0060] Next, the data acquisition unit 110 acquires a set of first face images and training scores from the training data set (step S302), and the image generation unit 120 generates a new face image based on the first face images and training scores acquired by the data acquisition unit 110 (step S303).

[0061] Next, the loss calculation unit 131 calculates a loss by comparing the generated face image generated by the image generation unit 120 with the second face image (specifically, the first face image used to generate the image and the second face image paired with the learning score) (step S304). Then, the parameter update unit 132 updates the parameters of the image generation model 125 so as to reduce the loss calculated by the loss calculation unit 131 (step S305).

[0062] Next, the learning unit 130 determines whether learning has ended (step S306). For example, the learning unit 130 may determine that learning has ended when all of the learning data included in the learning data set acquired in step S301 has been used. If it is determined that learning has not ended (step S306: NO), the processing from step S302 onwards is repeated. On the other hand, if it is determined that learning has ended (step S306: YES), the series of operations ends.

[0063] (Specific Operation Example) Next, a specific operation example regarding the above-mentioned learning operation will be described with reference to Fig. 10. Fig. 10 is a conceptual diagram showing an example of a learning operation by the fourth information processing device. Note that in Fig. 10, the same elements as those shown in Figs. 4, 6, and 7 are assigned the same reference numerals.

[0064] As shown in FIG. 10, in the learning operation, the first face image and the learning score S T is input to the image generation model 125. The image generation model 125 then receives the first face image and the learning score S T The generated face image is output based on the generated face image.

[0065] The learning unit 130 learns the image generation model 125 so as to minimize the difference between the generated face image generated by the image generation model 125 and the second face image. Here, the second face image is a face image that includes an object that is common to the first face image. In other words, the first face image and the second face image are face images of an object that is the same person. The second face image is compared with the first face image based on a matching score S C is the learning score S T Therefore, if the image generation model 125 is trained so as to minimize the difference between the generated face image and the second face image, the generated face image will contain a common object with the input face image and have a matching score S C It is possible to create a model that generates a face image such that the input score S is the same value as the input score S.

[0066] The difference between the generated facial image and the second facial image that is minimized when the learning unit 130 learns may be, for example, the average difference in pixel values ​​(brightness values) between the two images, the maximum difference in pixel values, or the distance between vectors obtained by GAN inversion.

[0067] (Example) The fourth information processing device 10 can be used to create a face image dataset to be used for machine learning, for example. For example, by using the fourth information processing device 10, it is possible to create a face image dataset to be used when training a face matching model.

[0068] The facial image dataset suitable for training a face matching model varies depending on the characteristics of the model (such as model size and model structure). For example, it is effective to train a small-sized model with a dataset that allows for relatively easy score separation (in other words, a low learning difficulty). The fourth information processing device 10 can generate facial image data according to a desired score S, thereby creating a facial image dataset suitable for the characteristics of the face matching model.

[0069] Facial image data included in a facial image dataset is collected, for example, via the Internet or by capturing images with a camera. However, with such collection methods, it is difficult to adjust the learning difficulty for a facial matching model. Specifically, since the learning difficulty from a human perspective differs from the learning difficulty from a facial matching model perspective, it is difficult to collect facial images with an appropriate learning difficulty. However, the fourth information processing device 10 can generate facial image data by directly adjusting the learning difficulty from a facial matching model perspective (i.e., the score S). Therefore, facial image data that is lacking in an existing facial image dataset (e.g., facial image data in a low score range) can be generated. In other words, it is possible to appropriately supplement an existing facial image dataset.

[0070] (Technical Effects) Next, technical effects obtained by the fourth information processing apparatus 10 will be described.

[0071] 8 to 10 , in the fourth information processing device 10, the image generation unit 120 is trained by the learning unit 130. In this way, the image generation unit 120 can be trained to be able to generate an appropriate facial image based on the input image and the input score S. In other words, it is possible to train the image generation unit 120 to be able to generate a new facial image that has a desired degree of match when matched with the input facial image.

[0072] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.

[0073] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processes by themselves, but also includes programs that execute processes by operating on an OS in conjunction with other software or expansion board functions. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal. The program may be provided to the user in, for example, a SaaS (Software as a Service) format.

[0074] <Supplementary Notes> The above-described embodiment may be further described as in the following supplementary notes, but is not limited to the following.

[0075] (Supplementary Note 1) The information processing device described in Supplementary Note 1 is an information processing device that includes an acquisition means that acquires an input face image including a target face and an arbitrary input score, and a generation means that generates a new face image based on the input face image and the input score, the new face image having a matching score value that indicates the degree of match with the input face image as the value of the input score.

[0076] (Supplementary Note 2) The information processing device according to Supplementary Note 2 is the information processing device according to Supplementary Note 1, wherein the generating means generates the new face image using a generative adversarial network.

[0077] (Supplementary Note 3) The information processing device according to Supplementary Note 3 is the information processing device according to Supplementary Note 1, wherein the generating means generates the new facial image using a diffusion model that generates an image based on input text.

[0078] (Supplementary Note 4) The information processing device described in Supplementary Note 4 is the information processing device described in Supplementary Note 3, wherein the acquisition means acquires the input text in addition to the input facial image and the input score, and the generation means generates the new facial image based on the input facial image, the input score, and the input text.

[0079] (Supplementary Note 5) The information processing device described in Supplementary Note 5 is the information processing device described in any one of Supplements 1 to 4, further including a learning means that learns the generating means by using a first face image and a second face image including a face of a common target and a learning score that is a matching score of the first face image and the second face image as learning data.

[0080] (Appendix 6) The information processing device described in Appendix 6 is the information processing device described in Appendix 5, in which the learning means trains the generation means so as to minimize the difference between the new facial image generated by the generation means based on the first image and the learning score and the second image.

[0081] (Supplementary Note 7) The information processing method described in Supplementary Note 7 is an information processing method in which an input face image including a target face and an arbitrary input score are acquired by at least one computer, and a new face image is generated based on the input face image and the input score, in which the value of the matching score indicating the degree of match with the input face image is the value of the input score.

[0082] (Appendix 8) The recording medium described in Appendix 8 is a recording medium having recorded thereon a computer program for causing at least one computer to execute an information processing method, which acquires an input face image including a target face and an arbitrary input score, and generates, based on the input face image and the input score, a new face image whose matching score value, which indicates the degree of match with the input face image, is the value of the input score.

[0083] (Supplementary Note 9) The computer program described in Supplementary Note 9 is a computer program that causes at least one computer to execute an information processing method of acquiring an input face image including a target face and an arbitrary input score, and generating a new face image based on the input face image and the input score, the new face image having a matching score value that indicates the degree of match with the input face image equal to the value of the input score.

[0084] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of ​​the invention that can be read from the claims and the entire specification, and information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical idea of ​​this disclosure.

[0085] REFERENCE SIGNS LIST 10 Information processing device 11 Processor 12 RAM 13 ROM 14 Storage device 15 Input device 16 Output device 110 Data acquisition unit 120 Image generation unit 125 Image generation model 130 Learning unit 131 Loss calculation unit 132 Parameter update unit S Input score S C Matching score S T Study score

Claims

1. An information processing apparatus comprising: an acquisition unit that acquires an input face image including a target face and an arbitrary input score; and a generation unit that generates a new face image in which the value of a collation score indicating the degree of coincidence with the input face image becomes the value of the input score based on the input face image and the input score.

2. The information processing apparatus according to claim 1, wherein the generation unit generates the new face image using an adversarial generation network.

3. The information processing apparatus according to claim 1, wherein the generation unit generates the new face image using a diffusion model that generates an image based on input text.

4. The information processing apparatus according to claim 3, wherein the acquisition unit acquires the input text in addition to the input face image and the input score, and the generation unit generates the new face image based on the input face image, the input score, and the input text.

5. The information processing apparatus according to any one of claims 1 to 4, further comprising a learning unit that learns the generation unit by using, as learning data, a first face image and a second face image including a common target face and a learning score that is a collation score between the first face image and the second face image.

6. The information processing apparatus according to claim 5, wherein the learning unit learns the generation unit so as to minimize the difference between the new face image generated by the generation unit based on the first image and the learning score and the second image.

7. An information processing method in which at least one computer acquires an input face image including a target face and an arbitrary input score, and generates a new face image in which the value of a collation score indicating the degree of coincidence with the input face image becomes the value of the input score based on the input face image and the input score.

8. A recording medium on which a computer program for causing at least one computer to execute an information processing method is recorded, the information processing method including acquiring an input face image including a target face and an arbitrary input score, and generating a new face image in which the value of a collation score indicating the degree of coincidence with the input face image becomes the value of the input score based on the input face image and the input score.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and computer readable storage medium

    CN116580127A

  • Program, image data generating device, and image data generating method

    JP7095935B2