Image conversion device, method, and program

The image conversion device seamlessly converts facial expressions by correcting deformation amounts based on face angle and hidden area ratios, addressing issues of unnatural conversions due to angle changes or obstructions.

JP7700951B2Active Publication Date: 2025-07-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024502365
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-07-01
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing image conversion methods fail to seamlessly convert facial expressions when the face angle changes or parts of the face are hidden, resulting in unnatural conversions.

Method used

An image conversion device and method that recognizes feature points, calculates the face angle and hidden area ratio, and corrects the deformation amount of these points to ensure seamless expression conversion.

Benefits of technology

Enables seamless and natural conversion of facial expressions even when face angles change or parts are obscured, preventing unnatural conversions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700951000002
    Figure 0007700951000002
  • Figure 0007700951000003
    Figure 0007700951000003
  • Figure 0007700951000004
    Figure 0007700951000004
Patent Text Reader

Abstract

An image transformation device according to an embodiment comprises: a feature point recognition unit that recognizes feature points of facial parts recognized from an image including a human face; a change amount correction unit that, when transforming the expression on the recognized face in the image into a target transformed facial expression, corrects an amount of change representing the amount of deformation for each feature point of the facial parts that corresponds to the transformed facial expression, on the basis of at least one of the ratio of the angle of the face in the image as measured from the front to a limit angle of the face at which the face can be recognized from the front, and the ratio of the area of the face not hidden by any objects to the entire area of the face; and a facial expression transformation unit that deforms the feature points according to the corrected amount of change and thereby obtains a transformed image in which the expression on the human face has been transformed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an image conversion apparatus, method, and program.

Background Art

[0002] Non-Patent Document 1 discloses the possibility of manipulating emotional experiences through real-time facial deformation (expression conversion) feedback. In Non-Patent Document 1, the face of a subject is tracked in real time and natural facial deformation processing is performed. In Non-Patent Document 1, the Rigid MLS (Moving Least Squares) method is used as an image conversion method to deform the expression in a face image. The Rigid MLS method is a technique of distorting an image by recognizing feature points in the image recognized from the image and moving them. Such a technique is also disclosed in Non-Patent Document 2. Note that a face image is an image obtained by photographing the face of a subject, an image obtained by extracting the face of an avatar generated by a computer, or the like.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, if the face angle of the subject changes or a part of the face is hidden, and the above-mentioned feature points cannot be recognized, the expression conversion will stop at an unnatural timing, and only face images with unnatural conversions can be obtained. That is, the expression appearing in the face image cannot be seamlessly converted.

[0005] This invention is made by paying attention to the above circumstances, and its object is to provide an image conversion device, method, and program capable of seamlessly converting the expression appearing in a face image.

Means for Solving the Problems

[0006] In order to solve the above problems, an image conversion device according to an aspect of this invention includes a feature point recognition unit that recognizes feature points of face parts recognized from an image including a human face, and the ratio of the angle of the face from the front to the limit angle at which the face in the image cannot be recognized from the front, and the ratio of the area of the face excluding the area hidden by an object to the entire area of the face to Based on this, a change amount correction unit that corrects the change amount representing the deformation amount of each of the feature points of the face parts corresponding to the conversion expression when converting the recognized face expression into a conversion expression to be converted, and an expression conversion unit that obtains a converted image in which the expression of the human face is converted by deforming the feature points with the corrected change amount.

[0007] In order to solve the above problems, an image conversion method according to this aspect is performed by an image conversion device that converts an expression in an image of a human face. waA method comprising: recognizing feature points of facial parts recognized from an image including a human face by a feature point recognition unit of the image conversion device; and calculating, by a change amount correction unit of the image conversion device, a ratio of the angle of the face from the front to the limit angle at which the face in the image cannot be recognized from the front, and a ratio of the area of the face excluding the area hidden by an object to the entire area of the face ni Based on these, correcting a change amount representing a deformation amount for each of the feature points of the facial parts corresponding to the converted expression when converting to a converted expression for which the expression of the recognized face is to be converted; and obtaining, by an expression conversion unit of the image conversion device, a converted image in which the expression of the human face is converted by deforming the feature points according to the corrected change amount.

Advantages of the Invention

[0008] According to the present invention, the expression shown in the face image can be seamlessly converted.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0010] [One Embodiment] Hereinafter, with reference to the drawings, one embodiment of the present invention will be described. (Configuration Example) FIG. 1 is a block diagram showing an example of the configuration of an image conversion apparatus according to one embodiment of the present invention. In the example shown in FIG. 1, an image conversion apparatus 100 according to one embodiment of the present invention includes an image acquisition unit 11, a feature point recognition unit 12, a face angle calculation unit 13, a display ratio calculation unit 14, a converted expression input unit 15, a change amount storage unit 16, a change amount correction unit 17, an expression conversion unit 18, and an image output unit 19.

[0011] The image acquisition unit 11 acquires a face image of a user from, for example, an image captured by a web camera or an avatar. The image acquisition unit 11 outputs the acquired face image to the feature point recognition unit 12, the display ratio calculation unit 14, and the expression conversion unit 18.

[0012] The feature point recognition unit 12 takes as input the face image acquired by the image acquisition unit 11 and recognizes the feature points of the face parts recognized from the face image. The method of recognizing the feature points in the feature point recognition unit 12 will be described later. The feature point recognition unit 12 outputs the recognized feature points to the face angle calculation unit 13 and the change amount correction unit 17.

[0013] The face angle calculation unit 13 takes as input the feature points recognized by the feature point recognition unit 12 and calculates the angle of the face in the face image, for example, the angle between the current position of the center of the face and the position when the face is facing forward (sometimes referred to as the angle of the face from the front), and outputs the data of the calculated angle to the change amount correction unit 17.

[0014] The display ratio calculation unit 14 takes the face image acquired by the image acquisition unit 11 as input, calculates the ratio of the hidden part in the whole face for the face image, and outputs the data of the calculated ratio to the change amount correction unit 17.

[0015] The converted expression input unit 15 acquires a converted expression (which may be referred to as the converted expression to be converted), such as a smiling expression, specified and input by the user from a user interface such as a keyboard. The converted expression input unit 15 outputs the acquired converted expression to the change amount correction unit 17.

[0016] In the change amount storage unit 16, change amounts representing the amount of deformation (the amount of movement of coordinate values) for each feature point are stored (memorized) in advance for each expression to be converted. The change amount is information indicating how much the coordinate values of each feature point should be moved according to the expression to be converted. The change amount can be obtained in advance, for example, by the user adjusting while applying an expression deformation process from a neutral face to a natural expression for a specific face image.

[0017] The change amount correction unit 17 inputs the feature points recognized by the feature point recognition unit 12, the face angle calculated by the face angle calculation unit 13, and the display ratio calculated by the display ratio calculation unit 14. Also, the change amount correction unit 17 reads out from the change amount storage unit 16 the change amount corresponding to the expression to be converted indicated by the converted expression input from the converted expression input unit 15. The change amount correction unit 17 calculates a corrected change amount of the change amount in the expression to be converted according to a formula described later based on these input feature points, face angle, and display ratio, and outputs the data of the calculated change amount to the expression conversion unit 18.

[0018] The facial expression conversion unit 18 takes as input the amount of change corrected by the amount of change correction unit 17. Based on the corrected amount of change, that is, the amount of change representing the amount of deformation corresponding to the conversion facial expression to be converted, the facial expression conversion unit 18 moves each feature point in the input face image based on the input movement amount, which is the corrected amount of change of the feature point, to obtain a face image with the facial expression of the face image converted. The facial expression conversion unit 18 outputs the converted face image to the image output unit 19.

[0019] The image output unit 19 takes as input the face image after conversion from the facial expression conversion unit 18 and outputs the input face image. Here, output includes, for example, storing in a storage medium, displaying on a display, transmitting to other devices via a communication network, and the like.

[0020] FIG. 2 is a diagram showing an example of the hardware configuration of the image conversion apparatus 100. The image conversion apparatus 100 is constituted by a computer such as, for example, a personal computer, a smart phone, a server computer, and the like. As shown in FIG. 2, the image conversion apparatus 100 has a hardware processor (sometimes simply referred to as a processor) 111A such as a CPU (Central Processing Unit). Note that the CPU can execute a plurality of information processes simultaneously by using a multi-core and multi-thread one. Also, the processor 111A may include a plurality of CPUs. In the image conversion apparatus 100, a program memory 111B, a data memory 112, a communication interface 114, and an input / output interface 113 are connected to the processor 111A via a bus 115.

[0021] The communication interface 114 can include, for example, one or more wired or wireless communication modules. The communication interface 114 can communicate with other computers, web cameras, etc. connected via a cable, a local area network (LAN), or a network (NW) such as the Internet.

[0022] An input / output interface 113 has an input device 200 and an output device 300 connected thereto. The input device 200 includes input devices such as a keyboard, a pointing device such as a mouse, and sensor devices such as a camera. Also, the output device 300 is a display device such as a liquid crystal display or a cathode ray tube (CRT) display. The input device 200 and the output device 300 can also use a so-called tablet type input / display device. This type of input / display device is configured by arranging an input detection sheet using an electrostatic method or a pressure method on the display screen of a display device using, for example, liquid crystal or organic electro luminescence (EL). The input / output interface 113 inputs the operation information input in the input device 200 to the processor 111A and causes the output device 300 to display the display information generated by the processor 111A.

[0023] Note that the input device 200 and the output device 300 do not necessarily have to be connected to the input / output interface 113. The input device 200 and the output device 300 can be provided with a communication unit for directly connecting or connecting via a network to the communication interface 114, so as to be able to exchange information with the processor 111A.

[0024] In addition, the input / output interface 113 may have a read / write function for a recording medium such as a semiconductor memory like a flash memory, or alternatively, may have a connection function with a reader writer having a read / write function for such a recording medium. Further, the input / output interface 113 may have a connection function with other devices.

[0025] The program memory 111B is a non-transitory tangible computer-readable storage medium, which is a combination of a non-volatile memory that can be written to and read from at any time and a non-volatile memory that can only be read from at any time. The non-volatile memory that can be written to and read from at any time is, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. The non-volatile memory that can only be read from at any time is, for example, a ROM (Read Only Memory). In this program memory 111B, programs necessary for the processor 111A to execute various control processes according to an embodiment, such as an image conversion program, are stored. That is, the processing function units in each of the above-described image acquisition unit 11, feature point recognition unit 12, face angle calculation unit 13, display ratio calculation unit 14, converted expression input unit 15, change amount correction unit 17, expression conversion unit 18, and image output unit 19 can all be realized by causing the processor 111A to read out and execute the image conversion program stored in the program memory 111B. Note that some or all of these processing function units may be realized in other various forms including integrated circuits such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0026] The data memory 112 is a tangible computer-readable storage medium, which is, for example, a combination of the above-mentioned non-volatile memory and a volatile memory such as RAM (Random Access Memory). This data memory 112 is used to store various data acquired and created during the process of performing various processes. That is, in the data memory 112, an area for storing various data is appropriately secured during the process of performing various processes.

[0027] Figure 3 is a diagram showing an example of facial feature points. The star marks in Figure 3 are the feature points recognized by the processor 111A, and the numbers attached to the side of each feature point are unique feature point IDs (Identifiers) for identifying each feature point. The number of feature point IDs and the facial parts corresponding to each feature point ID are determined by the adopted feature point recognition method. For example, the feature point with the feature point ID "18" is predetermined to be the left end of the left eyebrow facing forward.

[0028] Figure 4 is a diagram showing an example of the storage form of feature points. As shown in Figure 4, in the data memory 112, the x-coordinates and y-coordinates of the feature points in the facial image are stored in a table format in association with the feature point IDs. The coordinate values are in pixels. Therefore, in the data memory 112, in the example of Figure 3, the xy coordinates of the feature points related to the feature point IDs "1" to "68" are stored.

[0029] The data memory 112 stores the conversion expression specified by the user, which is acquired when the processor 111A operates as the above-mentioned converted expression input unit 15. The data memory 112 may store the conversion amount stored in the above-mentioned change amount storage unit 16.

[0030] FIG. 5 is a diagram showing an example of a storage form of a change amount. As shown in FIG. 5, in the data memory 112, for each converted expression, the change amount of the x coordinate and the change amount of the y coordinate of the feature point are stored in a table format as change amounts independent of the person who is the subject, in association with the feature point ID. The value of the change amount is in pixels. The change amount is represented by the moving direction and the moving amount of the feature point. For example, the moving amount of “+1” represents moving 1 pixel in the positive direction.

[0031] The data memory 112 may store the face image converted when the processor 111A operates as the above-described expression conversion unit 18. In addition, the data memory 112 may store various intermediate data generated during the operation of the processor 111A.

[0032] (Operation) Next, the operation of the image conversion apparatus 100 will be described. FIG. 6 is a flowchart showing an example of an image conversion processing operation by the image conversion apparatus 100. The processor 111A of the image conversion apparatus 100 starts the operation as the image conversion apparatus 100 shown in this flowchart by reading and executing the image conversion program stored in the program memory 111B. The execution of the image conversion program by the processor 111A is started by being instructed to perform image conversion from the input device 200 via the input / output interface 113 or via the communication interface 114.

[0033] The processor 111A operates as the conversion expression input unit 15 and waits for a user to specify and input a conversion expression, which is the expression to be converted such as a smiling face (step S1). For example, the processor 111A determines whether an input signal from the input device 200 via the input / output interface 113 or the communication interface 114 includes a specification input of the conversion expression. If there is a specification input of the conversion expression, the processor 111A proceeds to the process of step S2.

[0034] The processor 111A stores the specified conversion expression in the data memory 112 (step S2).

[0035] Processor 111A operates as an image acquisition unit 11 to acquire a face image (step S3). For example, processor 111A acquires a captured image of the face of a subject by a camera of input device 200 via input / output interface 113. Alternatively, processor 111A acquires a face image captured by a web camera connected to a network or the face of an avatar generated by another computer via communication interface 114. Processor 111A stores the acquired face image in data memory 112.

[0036] Processor 111A operates as a feature point recognition unit 12 to recognize feature points from the face image stored in data memory 112 (step S4). Processor 111A recognizes feature points for the face image, for example, by using the face_landmark_detection function of dlib (see, for example, http: / / dlib.net / face_landmark_detection.py.html). Specifically, processor 111A extracts the distribution of the gradient directions of luminance called HOG (Histogram of Oriented Gradients) features for the input face image. A model that has been learned based on data associating HOG features with the positions of facial feature points is generally provided. Therefore, processor 111A inputs the extracted HOG features into this learning model to obtain the positions of the facial feature points. Processor 111A stores the obtained positions of the feature points in data memory 112.

[0037] Processor 111A operates as a face angle calculation unit 13 to calculate, for example, using opencv or the like, the angle of the face in the face image (step S5). Specifically, processor 111A measures in advance the three-dimensional position (P_3d) of the feature points of the facial parts when the face is facing forward and holds this in data memory 112. Processor 111A obtains the two-dimensional position (P’_2d) of the current feature points of the facial parts of the face image. Processor 111A calculates the two-dimensional position (P_2d) of the feature points of the face parts when the three-dimensional position (P_3d) is rotated or moved. Processor 111A calculates each of the above two-dimensional positions by using, for example, the ProjectPoints2 function of opencv (see, for example, http: / / opencv.jp / opencv-2svn / py / camera_calibration_and_3d_reconstruction.html#projectpoints2).

[0038] Processor 111A calculates the sum of squares (D) of the distances between the two-dimensional position (P_2d) and the two-dimensional position (P'_2d). Processor 111A obtains, by global optimization, the angle (and the amount of movement) that minimizes this sum of squares D.

[0039] Processor 111A calculates, as the angle (and the amount of movement) that minimizes the above, the angle (a) of the face from the front by using, for example, the solvePnP function of opencv (see, for example, http: / / opencv.jp / opencv-2svn / cpp / camera_calibration_and_3d_reconstruction.html#cv-solvepnp).

[0040] While starting the face recognition tool and moving the face, Processor 111A obtains the positions of the feature points when recognition becomes impossible, and thus calculates in advance, as an angle independent of the subject person, the limit angle (A) of the face at which recognition is possible, and stores this in the data memory 112.

[0041] Next, Processor 111A operates as the display ratio calculation unit 14 to calculate the display ratio of the face, which is the ratio of the area hidden by objects other than the face in the entire area of the face in the face image (step S6). For example, if 10% of the entire face is hidden by objects other than the face, the above display ratio of the face is 10%.

[0042] Here, an example of the calculation by the display ratio calculation unit 14 will be described with reference to FIGS. 7 and 8.

[0043] FIG. 7 is a diagram showing an example of a neural network used by the display ratio calculation unit. FIG. 8 is a diagram showing an example of a grid cell processed by the display ratio calculation unit. Here, an example related to an input image including animals and various objects will be described, but the same applies when these are a human face and an object hiding the face, such as a hand or other object.

[0044] In the examples shown in FIGS. 7 and 8, a known YOLO (You Only Look Once) (a general object detection method by deep learning) can be used. This method is disclosed, for example, in the following materials. “Joseph Redmon, et al., “YOLOv3: An Incremental Improvement”, arXiv preprint, arXiv:1804.02767, 2018.”

[0045] In this method, the processor 111A resizes the face image to a square and inputs it into a CNN (Convolutional Neural Network), which is a neural network widely used in the field of image processing as shown in FIG. 7. The processor 111A extracts features from the face image through 24 convolutional layers (Conv. Layer) and 4 pooling layers (see reference numeral a in FIG. 7) in the CNN shown in FIG. 7, and can estimate the Bounding Box of the object in the image and the probability of the type of the object in 2 fully connected layers (see reference numeral b in FIG. 7). The final output size of 7×7 of the convolutional layer matches the number of divisions of the grid cell.

[0046] The above input image is divided into S×S grid cells as shown in FIG. 8 (see (a) in FIG. 8). Processor 111A estimates the bounding boxes of B objects for each of the above - divided grid cells. For each bounding box, Processor 111A outputs a total of five values, namely the coordinate values, width, and height (x, y, w, h) of the bounding box and the confidence score indicating that the bounding box is an object (see (b) of FIG. 8).

[0047] The x and y of the coordinate values are the center coordinates of the bounding box based on the boundaries of the grid cell. The width w and height h are relative values with respect to the size of the entire image, and the confidence score represents the probability that the bounding box is an object or a background. This probability is "1" if it is an object and "0" if it is a background.

[0048] As an index for measuring the estimation accuracy of the object region, there is IoU (Intersection over Union) which represents the degree of coincidence between the ground - truth bounding box and the estimated bounding box. In the above - mentioned YOLO, the confidence score of the bounding box represents IoU.

[0049] Processor 111A estimates the probability of the type of object for each grid cell. For example, Processor 111A estimates the probability, that is, the conditional probability, of which class a grid cell belongs to when it is an object in C classification classes (see (c) of FIG. 8).

[0050] Processor 111A integrates the class probabilities estimated here with the above - mentioned bounding boxes to obtain a plurality of bounding boxes indicating what the object is (see (d) of FIG. 8).

[0051] Processor 111A selects these Bounding Boxes, including overlapping regions, using the NMS (Non-Maximum Suppression) method based on the Bounding Box with the highest confidence score (see (e) of FIG. 8). NMS suppresses regions with large IoU values (high overlap) using a threshold. As a result, the detection result of the object region is obtained.

[0052] When there is a face region and an object region overlapping this region, Processor 111A can calculate the above-mentioned face display ratio by dividing the area of the overlapping region by the area of the face region.

[0053] Next, Processor 111A operates as the amount-of-change correction unit 17, reads out the amount of change corresponding to the target expression to be transformed from the amount-of-change storage unit 16, and calculates an amount of change corrected according to the target expression to be transformed based on the feature points recognized in S4, the face angle calculated in S5, and the display ratio calculated in S6 (step S7).

[0054] Specifically, Processor 111A acquires the face angle, that is, the face angle a from the front and the limit face angle A that can be recognized, and the ratio H of the area where the face is hidden to the entire face area, and accordingly attenuates the amount of change in the facial expression transformation, that is, corrects the amount of change, and holds the corrected result in the data memory 112 according to the following formula (1). ΔP new =ΔP·(1 - H)·a / A … Formula (1) The left side ΔP of Formula (1) new is the attenuated, that is, the corrected amount of change in the facial expression transformation, and the right side ΔP is the amount of change before the correction of the facial expression transformation.

[0055] That is, in the above example, (1) based on the ratio a / A between the face angle a from the front and the limit face angle A that can be recognized, and (2) the ratio H of the area where the face is hidden to the entire face area, the corrected amount of change is calculated. Note that, not limited to this example, for example, within the range of allowable accuracy, (1) the ratio a / A between the angle a of the face from the front and the limit angle A of the face that can be recognized, and (2) the ratio H of the area where the face is hidden to the area of the entire face, the corrected change amount may be calculated based on one of them.

[0056] By correcting the change amount in this way, even if the feature points cannot be recognized due to a change in the face angle or a part of the face being hidden, the expression conversion will not stop at an unnatural timing, and the expression of the face image can be naturally converted.

[0057] The processor 111A operates as an expression conversion unit 18 to convert the expression of the face image stored in the data memory 112 (step S8). That is, the processor 111A converts the face image based on the result in which the change amount corresponding to the conversion expression stored in the data memory 112 is corrected. For example, the processor 111A utilizes the implementation of MLS (see, for example, https: / / github.com / Jarvis73 / Moving-Least-Squares).

[0058] Specifically, for each feature point, the processor 111A moves it by the corrected change amount corresponding to the conversion expression stored in the data memory 112. For example, when converting the expression to a smiling face, for the control point of the feature point ID "1", since the xy coordinates before conversion are (23, 45) (see FIG. 4), the processor 111A performs a conversion such that the pixels of the feature point are moved to (24, 47) by setting the x coordinate to "+1" and the y coordinate to "+2" (see FIG. 5).

[0059] Then, for the feature points, the processor 111A applies an Affine transformation (including Helmert transformation = similarity transformation and rigid deformation = rigid body deformation) shown in the following formula (2).

[0060]

Equation

[0061] However, x and y in the above formula (2) are the coordinates of the neighboring feature points, x' and y' are the coordinates obtained by adding the variation amount to the coordinates of the feature points, a, b, c, and d are parameters, and t x , t y is the translation parameter. The processor 111A calculates the least square means of the coordinates x, y of the feature points and the coordinates x', y' obtained by adding the variation amount, and the parameters a, b, c, d, t x , t y that minimize this are obtained by global optimization. Then, taking the coordinates of the target points to be transformed by the processor 111A as x and y, the coordinates after transformation are obtained using these obtained parameters. The processor 111A uses the parameters a, b, c, d, t x , t y to obtain the coordinates after transformation from the feature points by the above affine transformation.

[0062] The processor 111A stores the face image after such transformation in the data memory 112 as a transformed image.

[0063] The processor 111A operates as the image output unit 19 and outputs the transformed image stored in the data memory 112 (step S9). For example, the processor 111A causes the output device 300 to display the face image via the input / output interface 113. Alternatively, the processor 111A transmits it over the network via the communication interface 114 and causes it to be displayed on a display device connected to the network or on the display unit of another computer connected to the network.

[0064] The processor 111A determines whether to end the operation as the image conversion device 100 shown in the flowchart of FIG. 6 (step S10). For example, the processor 111A checks whether the user has instructed to end the image conversion from the input device 200, via the input / output interface 113, or via the communication interface 114. Here, if the above operation is to be ended (YES in step S10), the processor 111A ends the operation shown in the flowchart of FIG. 6.

[0065] On the other hand, if the above operation has not been ended yet (NO in step S10), the processor 111A operates as the conversion expression input unit 15 and determines whether there is an input for designating a change in the conversion expression by the user (step S11). If there is no input for designating a change in the conversion expression (NO in step S11), the processor 111A proceeds to the process of step S3. Also, if there is an input for designating a change in the conversion expression (YES in step S10), the processor 111A proceeds to the process of step S2.

[0066] The image conversion device 100 according to the above-described embodiment includes a face angle calculation unit 13, a display ratio calculation unit 14, a change amount correction unit 17, and an expression conversion unit 18. The expression conversion unit 18 obtains a converted image in which the expression of a human face is converted by converting feature points according to the amount of deformation corresponding to the conversion expression to be converted. Therefore, even if the feature points cannot be recognized due to a change in the face angle or a part of the face being hidden, the image conversion device 100 according to an embodiment does not stop the expression conversion at an unnatural timing, and can naturally convert the expression of the face image.

[0067] [Other Embodiments] Note that the present invention is not limited to the above-described embodiment. For example, the flow of each process described above is not limited to the described procedure, and the order of some steps may be changed, or some steps may be performed in parallel.

[0068] In addition, the flow of each process described above was for the case of converting the expression of a face image acquired in real time in real time. However, it can be similarly applied to applications for converting the expression of a saved face image, rather than real-time processing.

[0069] In addition, the methods described in each embodiment can be stored as a program (software means) to be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy (registered trademark) disk, a hard disk, etc.), an optical disc (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), and can also be transmitted and distributed through a communication medium. Note that the program stored on the medium includes a setting program for configuring software means (including not only an execution program but also tables and data structures) to be executed by a computer in the computer. The computer that realizes this device reads the program recorded on the recording medium, and in some cases, constructs software means using the setting program, and executes the above-described processing by being controlled by this software means. Note that the recording medium referred to in this specification includes not only a medium for distribution but also storage media such as a magnetic disk and a semiconductor memory provided inside a computer or in a device connected via a network.

[0070] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, the embodiments may be combined as appropriate, and in that case, the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from a plurality of disclosed constituent elements. For example, even if some constituent elements are deleted from all the constituent elements shown in the embodiments, if the problem can be solved and the effect can be obtained, the configuration from which these constituent elements are deleted can be extracted as an invention.

Explanation of Reference Numerals

[0071] 100... Image conversion device 11… Image acquisition unit 12… Feature point recognition unit 13… Face angle calculation unit 14… Display ratio calculation unit 15… Transformed expression input unit 16… Change amount storage unit 17… Change amount correction unit 18… Expression conversion unit 19… Image output unit 111A… Processor 111B… Program memory 112… Data memory 113… Input / output interface 114… Communication interface 115… Bus 200… Input device 300… Output device

Claims

1. A feature point recognition unit that recognizes feature points of facial parts recognized from an image including a human face, Based on the ratio of the angle of the face from the front to the limit angle at which the face in the image cannot be recognized from the front, and the ratio of the area excluding the area where the face is hidden by an object to the entire area of the face, when converting to a conversion expression to be converted to the expression of the recognized face, a change amount correction unit that corrects the change amount representing the amount of deformation for each of the feature points of the facial parts corresponding to the conversion expression, An expression conversion unit that obtains a converted image in which the expression of the human face is converted by deforming the feature points with the corrected change amount, An image conversion device comprising:

2. The change amount correction unit, By multiplying the ratio of the angle of the face from the front to the limit angle at which the face in the image cannot be recognized from the front, and the ratio of the area excluding the area where the face is hidden by an object to the entire area of the face, by a predetermined change amount for each of the feature points of the facial parts, the change amount is corrected. The image conversion device according to claim 1.

3. Calculating the two-dimensional position of the feature point of the facial part when the three-dimensional position of the feature point of the facial part when the face is facing forward is rotated or moved, and calculating the angle at which the sum of the squares of the distances between the calculated two-dimensional position and the current two-dimensional position of the feature point of the facial part is minimized as the angle of the face from the front. The image conversion device according to claim 1.

4. A storage device that stores in advance a change amount representing the amount of deformation for each of the feature points for each conversion expression to be converted, A conversion expression input unit that inputs the conversion expression to be converted, Further comprising, The change amount correction unit, Reads out the change amount corresponding to the input conversion expression from the storage device and corrects the read change amount. The image conversion device according to any one of claims 1 to 3.

5. A method performed by an image conversion device that converts an expression in an image of a human face, Recognizing, by the feature point recognition unit of the image conversion device, feature points of facial parts recognized from an image including a human face, Based on the ratio of the angle of the face from the front to the angle at which the face in the image can no longer be recognized from the front by the change amount correction unit of the image conversion device, and the ratio of the area of the face excluding the area hidden by the object to the entire area of the face, correct the change amount representing the amount of deformation for each of the feature points of the face parts corresponding to the conversion expression when converting the recognized facial expression to the conversion expression to be converted. Obtain a converted image in which the facial expression of the person's face is converted by deforming the feature points by the corrected change amount by the facial expression conversion unit of the image conversion device. An image conversion method comprising the above.

6. An image conversion processing program that causes a processor to function as each part of the image conversion device according to any one of Claims 1 to 4.

Citation Information

Patent Citations

  • Micro expression fitting method and system based on displacement compensation

    CN112766063A

  • Method, device and program for image processing

    JP2005215763A

  • Image processing apparatus

    JP2011060038A

  • Image conversion device and method, and computer-readable recording medium

    JP2021077376A