Information processing device, information processing method, and program

The information processing device uses a physiognomy judgment network to adjust super-resolution network generation power based on facial resemblance, addressing discrepancies in facial features and enhancing image fidelity.

JP7835220B2Active Publication Date: 2026-03-25SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Super-resolution networks using GANs generate high-frequency components not present in the input signal, leading to discrepancies and changes in facial features, particularly in human faces, which existing methods struggle to control effectively.

Method used

An information processing device comprising a physiognomy judgment network to calculate facial resemblance before and after super-resolution processing, and a super-resolution network that adjusts generation power based on this resemblance to minimize changes in facial features.

Benefits of technology

The solution effectively suppresses changes in facial features during super-resolution processing by selecting generators with appropriate generation power levels, ensuring higher fidelity in image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835220000001
    Figure 0007835220000001
  • Figure 0007835220000002
    Figure 0007835220000002
  • Figure 0007835220000003
    Figure 0007835220000003
Patent Text Reader

Abstract

This information processing device (IP) has a physiognomy assessment network (PN) and a super-resolution network (SRN). The physiognomy assessment network (PN) calculates a degree of physiognomy matching between an input image (IMI) prior to undergoing super-resolution processing and an input image (IMI) after undergoing super-resolution processing. The super-resolution network (SRN) adjusts the generation force of the super-resolution processing on the basis of the degree of physiognomy matching.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Super-resolution techniques that enhance the resolution of input images are well-known. More recently, super-resolution networks have been proposed that use an image generation method called Generative Adversarial Systems (GANs) to reproduce even fine details that are difficult to discern from the input image. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 10-240920 [Non-patent literature]

[0004] [Non-Patent Document 1] [online], Few-shot Video-to-Video Synthesis, [Retrieved June 4, 2021], Internet<URL:https: / / nvlabs.github.io / few-shot-vid2vid / main.pdf> [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] In a super-resolution network using GANs, high-frequency components not present in the input signal are newly generated based on the learning results. The higher the super-resolution network's ability to generate signals (generation power), the higher the resolution image it can produce. However, the addition of signals not present in the input signal can sometimes cause a discrepancy between the input image and the generated image. For example, when dealing with a human face, subtle shifts in the shape of the eyes or mouth can alter the facial features.

[0006] Therefore, this disclosure proposes an information processing device, an information processing method, and a program capable of suppressing changes in facial features caused by super-resolution processing. [Means for solving the problem]

[0007] According to this disclosure, an information processing device is provided, comprising: a physiognomy judgment network that calculates the degree of facial resemblance between an input image before super-resolution processing and the input image after super-resolution processing; and a super-resolution network that adjusts the generation power of the super-resolution processing based on the degree of facial resemblance. Furthermore, according to this disclosure, an information processing method is provided in which the information processing of the information processing device is performed by a computer, and a program is provided in which the computer implements the information processing of the information processing device. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example of image processing using super-resolution technology. [Figure 2] This figure shows the changes in facial features caused by super-resolution processing. [Figure 3] This figure shows the changes in facial features caused by super-resolution processing. [Figure 4] This figure shows an example of a conventional super-resolution processing system. [Figure 5] This figure shows an example of a conventional super-resolution processing system. [Figure 6] This is a diagram showing the configuration of the information processing device according to the first embodiment. [Figure 7] This figure shows an example of the relationship between the degree of facial resemblance and the generative power control value. [Figure 8] A flowchart illustrating an example of information processing in an information processing device. [Figure 9] This figure shows an example of a training method for a super-resolution network. [Figure 10] This figure shows an example of a weight combination corresponding to the generative power level. [Figure 11] This figure shows the configuration of the information processing device according to the second embodiment. [Figure 12]This figure shows an example of a method for comparing facial posture, size, and position. [Figure 13] A flowchart illustrating an example of information processing in an information processing device. [Figure 14] This figure shows an example of the hardware configuration of an information processing device. [Modes for carrying out the invention]

[0009] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0010] The explanation will proceed in the following order. [1. Background] [1-1. Super-resolution technology] [1-2. Changes in facial features caused by super-resolution processing] [2. First Embodiment] [2-1. Configuration of Information Processing Devices] [2-2. Information Processing Methods] [2-3. Learning Methods] [2-4. Effects] [3. Second Embodiment] [3-1. Configuration of Information Processing Equipment] [3-2. Information Processing Methods] [3-3. Effects] [4. Hardware Configuration Examples]

[0011] [1. Background] [1-1. Super-resolution technology] Figure 1 shows an example of image processing using super-resolution technology (super-resolution processing).

[0012] The image in the upper left of Figure 1 is the original image (high-resolution image) IM O This is the generated image IM. G1 ~IM G7 This is the original image IM, which has been reduced in resolution due to compression, etc. O This is a reconstruction using super-resolution processing. Generated image IM G1 Image generated from IMG7 The generative power of super-resolution processing is increasing towards [it]. Note that the generative power means the ability to newly generate signals of high-frequency components that are not present in the input signal. The stronger the generative power, the higher-resolution the image obtained.

[0013] In super-resolution processing with weak generative power, information (such as patterns) lost in the input signal is not sufficiently restored. However, since the difference from the input signal is small, an image deviating from the original image IM O is difficult to generate. In super-resolution processing with strong generative power, since information lost in the input signal is also generated, an image close to the original image IM O is obtained. However, if the signal is not correctly generated, there is a possibility that an image deviating from the original image IM O will be generated.

[0014] For example, in the example of FIG. 1, an image of the beard of a mandrill is shown. In the original image IM O a large number of fine beards are shown. The blurring of the beard decreases from the generated image IMG1 to the generated image IMG7, and the generated image IMG7 has a resolution comparable to that of the original image IM O However, in the generated image IMG7, the shape of each beard is slightly different, resulting in an image with a slightly different atmosphere from the original image IM O Such subtle changes in the generated image appear as changes in the facial features when processing a human face.

[0015] ] [1-2. Changes in facial features due to super-resolution processing] FIG. 2 and FIG. 3 are diagrams showing changes in facial features due to super-resolution processing.

[0016] In the example of FIG. 2, a male face is the processing target. The input image IM I is the original image IM OThe image is generated by reducing its resolution. This reduction causes some information to be lost, such as the contours of facial features like the eyes, nose, and mouth, as well as the texture of the skin. In super-resolution processing, the lost information is restored (generated) based on the results of machine learning. However, if there is a discrepancy between the restored information and the original information, the facial features will change.

[0017] In the example in Figure 2, the original image IM O Compared to the original image, the generated image (IM) has subtle differences in eye size and shape, beard and hair density, and skin texture and wrinkles. G This is what is output. The shape of the eyes greatly influences a person's appearance, so even a slight change in the size or shape of the eyes can make a person's appearance seem to have changed.

[0018] In the example in Figure 3, the original image IM O Compared to the original, the generated image (IM) has subtle differences in eye size and shape, nose shape, hair texture, lip shape, and the degree to which the corners of the mouth are turned up. G The output shows that the appearance has changed significantly due to the alteration of the shape of facial features such as the eyes, mouth, and nose.

[0019] Figures 4 and 5 show an example of a conventional super-resolution processing system.

[0020] Figure 4 shows a typical super-resolution network (SRN) using GANs. A This is shown. Super-resolution network (SRN) A So, with powerful generation capabilities, generated image IM G While the resolution of the generated images can be increased, it is difficult to control unexpected generation results. This is because it is difficult to clarify the input-output dependencies obtained by machine learning, and the learning process is complex, making it difficult to generate images as intended. G This is because it is practically impossible to correct the error. Furthermore, because the learning process cannot be controlled, even if the processing result for a particular input is incorrect, it is difficult to correct only that specific input.

[0021] Figure 5 shows a reference image IM of the same person's face.R Super-resolution network (SRN) used as B This is shown. This type of super-resolution network (SRN) B This is disclosed in Non-Patent Document 1. Super-resolution network (SRN) B This is a reference image IM R Using the feature information of the reference image IM, some of the parameters used in super-resolution processing are dynamically adjusted. R An image with a similar facial appearance is generated. However, the reference image IM R The causal relationship between the input and the output is obtained through deep learning, so it is not possible to generate a perfectly matching facial feature in all cases. Therefore, the super-resolution network SRN B Even with these methods, it is not possible to completely suppress changes in facial features.

[0022] Therefore, this disclosure proposes a new method to solve the above-mentioned problems. The information processing device IP of this disclosure calculates the degree of facial resemblance before and after super-resolution processing, and adjusts the generation power of the super-resolution network SRN based on the calculated degree of facial resemblance. With this configuration, the generated image IM G The facial features are fed back into the super-resolution processing. Therefore, changes in facial features caused by the super-resolution processing are less likely to occur.

[0023] The information processing device IP can be used for enhancing the image quality of older video materials (such as movies and photographs) and for highly efficient video compression and transmission systems (video conferencing, online meetings, live video broadcasting, and online distribution of video content). When enhancing the image quality of movies and photographs, high fidelity to the faces of subjects is required, so the method disclosed herein is suitably adopted. In video compression and transmission systems, the information of the original video is significantly reduced, which can easily lead to changes in facial features during restoration. By using the method disclosed herein, such problems can be avoided.

[0024] The following describes in detail an embodiment of the information processing device IP.

[0025] [2. First Embodiment] [2-1. Configuration of Information Processing Devices] Figure 6 shows the configuration of the information processing device IP1 according to the first embodiment.

[0026] The information processing device IP1 uses super-resolution technology to process the input image IM. I High-resolution generated image IM G This is a device for restoring [something]. The information processing device IP1 has a super-resolution network SRN1, a physiognomy judgment network PN, and a generation power control value calculation unit GCU.

[0027] The super-resolution network SRN1 uses the input image IM I Super-resolution processing is used to generate the image IM G The super-resolution network SRN1 can change the generative power of the super-resolution process in multiple stages. For example, the super-resolution network SRN1 includes generator GEs of multiple GANs with different generative power levels (LV). In the example in Figure 6, four generator GEs (generative power levels LV=0~3) are held in the trained database, but the number of generator GEs is not limited to four. The number of generator GEs can be two or more.

[0028] Multiple generators (GEs) are generated using the same neural network. However, each generator uses different parameters to optimize the neural network. These differences in optimization parameters result in variations in the generative power level (LV) of each generator (GE).

[0029] The super-resolution network SRN1 uses the input image IM I The facial image of the same person as the subject is a facial morphological image (IM). PR It may be obtained as follows. The super-resolution network SRN1 uses a facial reference image IM. PR Using feature information of the input image IM I It can perform super-resolution processing on facial features reference images (IM). PR This is a reference image IM for adjusting facial features. R It is used as such. For example, the super-resolution network SRN1 uses a facial reference image IM. PRUsing the characteristic information, some of the parameters used in super-resolution processing are dynamically adjusted. This allows for the creation of a facial reference image (IM). PR Image generation of facial features similar to IM G This can be obtained. (Image based on facial features) PR As for methods of adjusting facial features using this method, known methods described in Non-Patent Document 1, etc., are used.

[0030] The facial recognition network PN uses the input image IM before super-resolution processing. I and the input image IM after super-resolution processing I The facial resemblance score DC is calculated. The facial recognition network PN is a neural network that performs face recognition. The facial recognition network PN calculates the facial resemblance score DC as the similarity between the face of a person included in the generated image and the face of the same person included in the facial recognition reference image. The similarity is calculated using known face recognition techniques such as feature point matching.

[0031] The super-resolution network SRN1 adjusts the generation power of super-resolution processing based on the facial resemblance score DC. For example, the super-resolution network SRN1 selects and uses a generator GE from multiple generators GE with different generation power levels LV, specifically one whose facial resemblance score DC meets an acceptable standard. The super-resolution network SRN1 determines whether the facial resemblance score DC meets the acceptable standard for each generator GE, starting with the one with the highest generation power level LV. The super-resolution network SRN1 then selects and uses the generator GE that is initially determined to meet the acceptable standard.

[0032] The Generative Power Control Value Calculation Unit (GCU) calculates the Generative Power Control Value (CV) based on the facial resemblance degree (DC). The Generative Power Control Value (CV) indicates the reduction from the current Generative Power Level (LV). The reduction is larger the lower the facial resemblance degree (DC). The Super-Resolution Network (SRN1) calculates the Generative Power Level (LV) based on the Generative Power Control Value (CV). The Super-Resolution Network (SRN1) performs super-resolution processing using the generator (GE) corresponding to the calculated Generative Power Level (LV).

[0033] Figure 7 shows an example of the relationship between the degree of facial resemblance DC and the generative power control value CV.

[0034] In the example in Figure 7, the acceptable criterion is the threshold T. A , threshold T B and threshold T C (threshold T) A <Threshold T B <Threshold T C ) is set. For example, the facial resemblance degree DC is set to threshold T. A If it is smaller than (-), the generative power control value CV is set to (-3). The facial resemblance DC is below the threshold T. A The above and threshold T B If it is smaller than (-), the generative power control value CV is set to (-2). Facial Conformity D C The threshold T B The above and threshold T C If it is smaller than the threshold T, the generative power control value CV is set to (-1). C If the above conditions are met, the generative power control value CV is set to 0. The reduction in the generative power level LV is set in stages according to the facial matching degree DC, allowing for the rapid detection of an appropriate generator GE.

[0035] [2-2. Information Processing Methods] Figure 8 is a flowchart showing an example of information processing by the information processing device IP1.

[0036] In step ST1, the super-resolution network SRN1 selects the generator GE with the highest generation power level LV. In step ST2, the super-resolution network SRN1 performs super-resolution processing using the selected generator GE.

[0037] In step ST3, the super-resolution network SRN1 determines whether the generation power level LV of the currently selected generator GE is at its minimum. If it is determined in step ST3 that the generation power level LV is at its minimum (step ST3: yes), the super-resolution network SRN1 continues to use the currently selected generator GE.

[0038] If it is determined in step ST3 that the generative power level LV is not the minimum (step ST3: no), proceed to step ST4. In step ST4, the physiognomy network PN generates the image IM. G and facial features reference image IM PR The degree of facial resemblance DC is calculated using this method, and facial physiognomy is used to make a judgment.

[0039] In step ST5, the generation power control value calculation unit GCU calculates the facial matching degree DC to the threshold T. C Determine whether the facial resemblance DC is equal to or greater than the threshold T. C If it is determined that the above is true (Step ST5: yes), the generation force control value calculation unit GCU sets the generation force control value CV to 0. The super-resolution network SRN1 continues to use the currently selected generator GE.

[0040] In step ST5, the facial resemblance degree DC is at threshold T C If it is determined to be smaller than (step ST5: no), the process proceeds to step ST6. In step ST6, the generative force control value calculation unit GCU calculates a generative force control value CV corresponding to the facial matching degree DC. In step ST7, the super-resolution network SRN1 selects a generator GE with a generative force level LV identified by the generative force control value CV. Then, returning to step ST2, the super-resolution network SRN1 performs super-resolution processing using the generator GE with the modified generative force level LV. After that, the above process is repeated.

[0041] [2-3. Learning Methods] Figure 9 shows an example of a training method for the super-resolution network SRN1.

[0042] The super-resolution network SRN1 uses student image IM S and generated image IM G Includes a generator GE of multiple GANs that have been machine-learned using and . Student image IM S This is a training image IM T This is low-resolution input data for machine learning. (Generated image IM)G This is student image IM S This is the output data after super-resolution processing of the training image IM. T Various facial images of different people are used.

[0043] In the GAN generator GE, the generated image IM G and teacher image IM T Machine learning is performed to minimize the difference with the target image. In the discriminator DI of GANs, the training image IM is used. T When this is entered, the identification value becomes 0, student image IM S Machine learning is performed so that the identification value when input is 1. Generated image IM G and teacher image IM T From each, the object recognition network ORN extracts the feature quantity C. The object recognition network ORN is a pre-trained neural network that extracts the feature quantity C from an image. In the generator GE, the generated image IM G Feature vector C and training image IM T Machine learning is performed in a way that minimizes the difference between the feature C and the target feature.

[0044] For example, training image IM T and generated image IM G Let D1 be the pixel-by-pixel difference value from the original image. Let D2 be the discriminator DI's discrimination value. (Training image IM) T and generated image IM G Let D3 be the difference between feature C and the original feature. Let w1 be the weight of the difference D1. Let w2 be the weight of the discriminant value D2. Let w3 be the weight of the difference D3. In each GAN, machine learning is performed so that the weighted sum of the difference D1, discriminant value D2, and difference D3 (w1 × D1 + w2 × D2 + w3 × D3) is minimized. The ratios of weights w1, w2, and w3 differ for each GAN.

[0045] GANs are a widely known Convolutional Neural Network (CNN), and they learn by minimizing the weighted sum of the three values ​​mentioned above (difference value D1, discriminant value D2, and difference value D3). The optimal values ​​for the three weights w1, w2, and w3 change depending on the CNN and training dataset used for training. Normally, one optimal set of values ​​is used to obtain the greatest generative power, but in this disclosure, by changing the three weights w1, w2, and w3, it is possible to obtain learning results with progressively different generative powers using the same CNN.

[0046] Figure 10 shows an example of a combination of weights w1, w2, and w3 corresponding to the generation power level LV.

[0047] ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) is a well-known CNN for super-resolution processing that uses GANs. ESRGAN is described below [1].

[0048] [1]Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, Chen Change Loy, “ESRGAN:Enhanced Super-Resolution Generative Adversarial Networks”, Published in ECCV Workshops 2018

[0049] For example, in this disclosure, the generator GE of ESRGAN is applied to the super-resolution network SRN1. Generator GEs with higher generative power levels (LV) have higher ratios of weights w2 and w3 to weight w1. Generator GEs with lower generative power levels (LV) have lower ratios of weights w2 and w3 to weight w1.

[0050] In the example in Figure 10, a generator GE with a generation power level of 0 is obtained when w1=1.0, w2=0, and w3=0. A generator GE with a generation power level of 1 is obtained when w1=0.1, w2=0.05, and w3=0.1. A generator GE with a generation power level of 2 is obtained when w1=0.01, w2=0.05, and w3=0.1. A generator GE with a generation power level of 3 is obtained when w1=0.01, w2=0.05, and w3=1.0.

[0051] Note that the values ​​of weights w1, w2, and w3 can change depending on the neural network configuration, the number of images in the training dataset, the content of the images, and other conditions such as the learning rate of the CNN. Even with different combinations of weight values, the learning results may converge to the optimal value under the same conditions.

[0052] [2-4. Effects] The information processing device IP1 includes a physiognomy judgment network PN and a super-resolution network SRN1. The physiognomy judgment network PN processes the input image IM before super-resolution processing. I and the input image IM after super-resolution processing I The facial resemblance score DC is calculated. The super-resolution network SRN1 adjusts the generation power of the super-resolution processing based on the facial resemblance score DC. In the information processing method of this disclosure, the processing of the information processing device IP1 is executed by the computer 1000 (see Figure 14). The program of this disclosure (program data 1450: see Figure 14) causes the computer 1000 to implement the processing of the information processing device IP1.

[0053] In this configuration, the generation power of the super-resolution network SRN1 is adjusted based on the changes in facial features before and after super-resolution processing. Therefore, changes in facial features caused by super-resolution processing are suppressed.

[0054] The super-resolution network SRN1 selects and uses a generator GE from multiple generators GE with different generation power levels LV, the generator GE whose facial resemblance degree DC meets an acceptable standard.

[0055] According to this configuration, the generation power of the super-resolution network SRN1 is adjusted by the selection of the generator GE.

[0056] The super-resolution network SRN1 includes a plurality of GAN generators GE that are machine-learned using the student image IM T obtained by downscaling the teacher image IM S and the generated image IM S obtained by super-resolving the student image IM G Let the pixel-by-pixel difference value between the teacher image IM T and the generated image IM G be D1, the discrimination value of the discriminator DI of the GAN be D2, and the difference value of the feature amount C between the teacher image IM T and the generated image IM G be D3. Let the weight of the difference value D1 be w1, the weight of the discrimination value D2 be w2, and the weight of the difference value D3 be w3. In each GAN, machine learning is performed so that the weighted sum (w1×D1 + w2×D2 + w3×D3) of the difference value D1, the discrimination value D2, and the difference value D3 is minimized. The ratios of the weights w1, w2, and w3 are different for each GAN.

[0057] [ According to this configuration, the neural networks of each generator GE can be shared. Also, the generation power of each generator GE can be easily controlled by the ratios of the weights w1, w2, and w3.

[0058] The super-resolution network SRN1 determines, in order from the generator GE with a high generation power level LV, whether the human likeness meets the tolerance standard. The super-resolution network SRN1 selects and uses the generator GE that is first determined to meet the tolerance standard.

[0059] According to this configuration, the generator GE with the maximum allowable generation power is selected.

[0060] The information processing device IP1 has a generation power control value calculation unit GCU. The generation power control value calculation unit GCU calculates a generation power control value CV indicating the reduction width from the current generation power level LV based on the human consistency degree DC. The reduction width is larger as the human consistency degree DC is lower.

[0061] According to this configuration, an appropriate generator GE can be detected quickly.

[0062] The super-resolution network SRN1 performs super-resolution processing on the input image IM using the feature information of the human face reference image IM PR I

[0063] According to this configuration, the human consistency degree DC before and after the super-resolution processing is increased.

[0064] Note that the effects described in this specification are merely examples and are not limiting, and there may be other effects.

[0065] [3. Second Embodiment] [3-1. Configuration of Information Processing Device] FIG. 11 is a diagram showing the configuration of the information processing device IP2 according to the second embodiment.

[0066] The difference from the first embodiment in this embodiment is that the generation power of the super-resolution network SRN2 is adjusted by switching the human face reference image IM PR

[0067] In the first embodiment, a plurality of generators GE are switched and used based on the human consistency degree DC. However, in this embodiment, only one generator GE is used. The super-resolution network SRN2 performs super-resolution processing on the input image IM using the feature information of the human face reference image IM PR I R From among the plurality of reference images IM included in the reference image group RG, a reference image IM whose human consistency degree DC satisfies the allowable standard RIllustrative image IM PR Select it as such.

[0068] The reference image group RG is obtained from image data located inside or outside the information processing device IP2. For example, the input image IM I If the person in the photo is a celebrity, multiple reference images that allow for identification of the person's features should be obtained from the internet or other sources. R (Reference image group RG) is obtained. Input image IM I If the image is from a scene in past footage (such as a movie), a reference image IM can be found from a close-up of a face in a different scene within the same footage. R A set of images that could potentially be the result is extracted. Input image IM I If the person in the image is a user of the information processing device IP2, and the information processing device IP2 is a device with a camera function such as a smartphone, a reference image IM will be taken from the photo data stored in the information processing device IP2. R A set of images that could potentially be this is extracted.

[0069] From the reference image group RRG, the reference image IM is suitable for physiognomy. R In order, physiognomy reference images IM PR It is selected as such. The super-resolution network SRN2 uses multiple reference images IM R Prioritize the reference images (IMs) according to their respective priorities. R Illustrative image IM PR For example, the super-resolution network SRN2 uses the pose, size, and position of the subject's face as input to the image IM. I Similar reference image IM R The first step is to determine whether the facial resemblance score DC meets the acceptable criteria. The super-resolution network SRN2 then uses the reference image IM, which is the first to be determined to meet the acceptable criteria. R Illustrative image IM PR This is selected as the setting. This ensures that super-resolution processing is performed with the maximum possible generative power.

[0070] Figure 12 shows an example of a method for comparing facial posture, size, and position.

[0071] In the Super Resolution Network SRN2, the left and right eyes, eyebrows, nose, upper and lower lips, and lower jaw are pre-set as facial features to be compared. The Super Resolution Network SRN2 uses an input image IM I and reference image IM R The coordinates of each point on the contour line of each facial feature are extracted from this data. Facial feature detection is performed using, for example, a known facial recognition technique as shown in [2] below.

[0072] [2]Kazemi, V., & Josephine, S. “One Millisecond Face Alignment with an Ensemble of Regression Trees. Computer Vision and Pattern Recognition (CVPR)”, 2014

[0073] The super-resolution network SRN2 uses techniques such as correspondence point matching to process the input image IM I And reference image IM R It extracts corresponding points (corresponding points) from the input image IM. The super-resolution network SRN2 uses the input image IM. I and reference image IM R Reference image IM shows a small sum of the absolute values ​​of the differences in the coordinates of corresponding points. R Prioritize this as much as possible. This will allow for appropriate facial features reference images (IM). PR It is quickly detected. In the example in Figure 12, the reference image IM RA The reference image is IM RB The posture of the facial features is more important than the input image IM I It is similar to the reference image IM. RA The priority is shown in the reference image IM. RB It will be set higher than that.

[0074] [3-2. Information Processing Methods] Figure 13 is a flowchart showing an example of information processing by the information processing device IP2.

[0075] In step ST11, the super-resolution network SRN2 selects one reference image IM from the reference image group RG according to priority. R Illustrative image IM PR Select as follows. In step ST12, the super-resolution network SRN2 selects the reference image IM. R Super-resolution processing is performed using the characteristic information.

[0076] In step ST13, the super-resolution network SRN2 uses the facial reference image IM. PR The currently selected reference image IM R The last reference image IM according to priority R Determine whether or not this is the case. In step ST13, the current reference image IM R This is the last reference image IM R If it is determined that (Step ST13: yes), the super-resolution network SRN2 will use the currently selected reference image IM. R Illustrative image IM PR Continue using it as such.

[0077] In step ST13, the current reference image IM R This is the last reference image IM R If it is determined that this is not the case (step ST13: no), proceed to step ST14. In step ST14, the super-resolution network SRN2 generates the image IM G The currently selected reference image IM R The degree of facial resemblance DC is calculated using this method, and facial physiognomy is used to make a judgment.

[0078] In step ST15, the super-resolution network SRN2 determines the facial resemblance degree DC to the threshold T. C Determine whether the facial resemblance degree DC is equal to or greater than the threshold T. C If it is determined that the above is true (Step ST15: yes), the super-resolution network SRN2 will use the currently selected reference image IM. R Illustrative image IM PR Continue using it as such.

[0079] In step ST15, the facial resemblance DC is at threshold T C If it is determined to be smaller than (step ST15: no), proceed to step ST16. In step ST16, the super-resolution network SRN2 selects the reference image IM that has not yet been selected according to priority. R Illustrative image IM PR Select as such. Then, return to step ST12, and the super-resolution network SRN2 will select the newly selected reference image IM. R Super-resolution processing is performed using [the specified method]. The process described above is then repeated.

[0080] [3-3. Effects] The super-resolution network SRN2 of this embodiment uses multiple reference image IMs. R Therefore, the reference image IM meets the acceptable standard for facial resemblance (DC). R Illustrative image IM PR Select as such. According to this configuration, the facial reference image IM PR The generation power of the super-resolution network SRN2 is adjusted according to the selection. As a result, changes in facial features caused by super-resolution processing are suppressed.

[0081] [4. Hardware Configuration Examples] Figure 14 shows an example of the hardware configuration of an information processing device IP. For example, the information processing device IP is implemented by a computer 1000. The computer 1000 has a CPU 1100, RAM 1200, ROM (Read Only Memory) 1300, HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The various parts of the computer 1000 are connected by a bus 1050.

[0082] The CPU 1100 operates based on programs stored in the ROM 1300 or HDD 1400, and controls various parts. For example, the CPU 1100 loads the programs stored in the ROM 1300 or HDD 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0083] ROM1300 stores boot programs such as the BIOS (Basic Input Output System) executed by CPU1100 when computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0084] HDD1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU1100 and data used by such programs. Specifically, HDD1400 is a recording medium that records an information processing program related to this disclosure, which is an example of program data 1450.

[0085] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (e.g., the Internet). For example, the CPU 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.

[0086] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard or mouse via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, speaker, or printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.

[0087] For example, when computer 1000 functions as an information processing device IP, the CPU 1100 of computer 1000 implements various functions for super-resolution processing by executing a program loaded onto RAM 1200. The HDD 1400 stores a program that enables the computer to function as an information processing device IP. While the CPU 1100 reads and executes program data 1450 from HDD 1400, another example is that these programs may be obtained from other devices via an external network 1550.

[0088] [Note] Furthermore, this technology can also be configured as follows. (1) A facial recognition network that calculates the degree of facial resemblance between an input image before super-resolution processing and the input image after super-resolution processing, A super-resolution network that adjusts the generation power of the super-resolution processing based on the degree of facial resemblance, An information processing device having (2) The super-resolution network selects and uses a generator from a plurality of generators with different generative power levels that satisfies the acceptable standard for facial resemblance. The information processing device described in (1) above. (3) The super-resolution network includes a generator of multiple GANs that have been machine-trained using student images obtained by reducing the resolution of training images and generated images obtained by super-resolution processing of the student images. Let D1 be the pixel-wise difference between the training image and the generated image, D2 be the discriminator's discrimination value of the GAN, D3 be the difference in features between the training image and the generated image, w1 be the weight of the difference value D1, w2 be the weight of the discriminator D2, and w3 be the weight of the difference value D3. In each GAN, machine learning is performed so as to minimize the weighted sum (w1×D1+w2×D2+w3×D3) of the difference value D1, the discrimination value D2, and the difference value D3. The ratios of weights w1, w2, and w3 differ for each GAN. The information processing device described in (2) above. (4) The super-resolution network determines whether the degree of facial resemblance meets the acceptable criteria, starting with the generators with the highest generation power levels, and selects and uses the generator that is determined to meet the acceptable criteria first. The information processing device described in (2) or (3) above. (5) The system includes a generation force control value calculation unit that calculates a generation force control value indicating the amount of reduction from the current generation force level based on the degree of facial resemblance, The aforementioned reduction is greater the lower the degree of facial resemblance. An information processing device as described in any one of (2) through (4) above. (6) The super-resolution network performs super-resolution processing on the input image using feature information from a physiognomy reference image. An information processing device as described in any one of (2) through (5) above. (7) The super-resolution network performs super-resolution processing on the input image using feature information from a physiognomy reference image. The super-resolution network selects a reference image from a plurality of reference images that satisfies the acceptable standard for facial resemblance, as the facial reference image. The information processing device described in (1) above. (8) The super-resolution network determines whether the degree of facial resemblance meets the acceptable criteria, starting with reference images whose facial posture, size, and position are closest to the input image, and selects the first reference image that is determined to meet the acceptable criteria as the facial reference image. The information processing device described in (7) above. (9) The super-resolution network extracts the coordinates of each point on the contour line of the facial features from the input image and the reference image, and prioritizes the reference image for which the sum of the absolute values ​​of the differences in the coordinates of corresponding points in the input image and the reference image is small. The information processing device described in (8) above. (10) The degree of facial resemblance between the input image before super-resolution processing and the input image after super-resolution processing is calculated. The generation power of the super-resolution processing is adjusted based on the degree of facial resemblance. An information processing method performed by a computer, which includes the ability to perform the following actions. (11) The degree of facial resemblance between the input image before super-resolution processing and the input image after super-resolution processing is calculated. The generation power of the super-resolution processing is adjusted based on the degree of facial resemblance. A program that allows a computer to accomplish something. [Explanation of symbols]

[0089] C features CV generation power control value D1, D3 difference value D2 Identifier DC physiognomy consistency DI Discriminator GCU Generation Power Control Value Calculation Unit GE Generator IM G Generated image IM I Input image IM PR physiognomy standard images IM R Reference image IM S Student images IM T Teacher image IP, IP1, IP2 Information Processing Devices LV Generation Power Level PN (Psychological Judgment Network) SRN, SRN1, SRN2 Super-Resolution Networks w1, w2, w3 weights

Claims

1. A physiognomy judgment network that calculates the degree of facial resemblance between an input image before super-resolution processing and a generated image obtained by performing the super-resolution processing on the input image, based on a facial reference image corresponding to the person included in the input image, A super-resolution network that adjusts the generation power of the super-resolution processing based on the degree of facial resemblance, An information processing device having

2. The super-resolution network selects and uses a generator from a plurality of generators with different generative power levels that satisfies the acceptable standard for facial resemblance. The information processing apparatus according to claim 1.

3. The super-resolution network includes a generator of multiple GANs that have been machine-trained using student images obtained by reducing the resolution of training images and training generated images obtained by super-resolution processing of the student images. Let D1 be the pixel-by-pixel difference between the training image and the generated training image, D2 be the discriminator's discrimination value of the GAN, D3 be the difference in features between the training image and the generated training image, w1 be the weight of the difference value D1, w2 be the weight of the discrimination value D2, and w3 be the weight of the difference value D3. In each GAN, machine learning is performed so as to minimize the weighted sum (w1 × D1 + w2 × D2 + w3 × D3) of the difference value D1, the identification value D2, and the difference value D3. The ratios of weights w1, w2, and w3 differ for each GAN. The information processing apparatus according to claim 2.

4. The super-resolution network determines whether the degree of facial resemblance meets the acceptable criteria, starting with the generators with the highest generation power levels, and selects and uses the generator that is determined to meet the acceptable criteria first. The information processing apparatus according to claim 2.

5. The system includes a generation force control value calculation unit that calculates a generation force control value indicating the amount of reduction from the current generation force level based on the degree of facial resemblance, The aforementioned reduction is greater the lower the degree of facial resemblance. The information processing apparatus according to claim 2.

6. The super-resolution network performs super-resolution processing on the input image using the feature information of the physiognomy reference image. The information processing apparatus according to claim 2.

7. The super-resolution network uses the feature information of the physiognomy reference image to perform super-resolution processing on the input image. The super-resolution network selects a reference image from a plurality of reference images that satisfies the acceptable standard for facial resemblance, as the facial reference image. The information processing apparatus according to claim 1.

8. The super-resolution network determines whether the degree of facial resemblance meets the acceptable criteria, starting with reference images whose facial posture, size, and position are closest to the input image, and selects the first reference image that is determined to meet the acceptable criteria as the facial reference image. The information processing apparatus according to claim 7.

9. The super-resolution network extracts the coordinates of each point on the contour line of the facial features from the input image and the reference image, and prioritizes the reference image whose sum of the absolute values ​​of the differences in the coordinates of corresponding points in the input image and the reference image is small. The information processing apparatus according to claim 8.

10. The degree of facial resemblance between the input image before super-resolution processing and the generated image obtained by performing the super-resolution processing on the input image is calculated based on a facial reference image corresponding to the person included in the input image. The generation power of the super-resolution processing is adjusted based on the degree of facial resemblance. An information processing method performed by a computer, which includes the ability to perform the following actions.

11. The degree of facial resemblance between the input image before super-resolution processing and the generated image obtained by performing the super-resolution processing on the input image is calculated based on a facial reference image corresponding to the person included in the input image. The generation power of the super-resolution processing is adjusted based on the degree of facial resemblance. A program that allows a computer to accomplish something.

Citation Information

Patent Citations

  • Face super-resolution implementation method and device, electronic equipment and storage medium

    CN111709878A

  • Image forming device, its method and image forming program storage medium

    JP1998240920A

  • Method, device and program for enhancing face image resolution

    JP2010286959A

  • Facial image processing method and device, and storage medium

    JP2019504386A