Image conversion device, image conversion method, and image conversion program

The image conversion device addresses unnatural expressions by adding missing feature points to facilitate natural-looking facial expression conversion using Rigid MLS, ensuring accurate and natural expression representation.

JP7704288B2Active Publication Date: 2025-07-08NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024502364
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-07-08
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing facial expression conversion methods using Rigid MLS result in unnatural expressions when only one side of facial parts is recognized, as they fail to account for the missing feature points on the other side.

Method used

An image conversion device that adds a second feature point on the unrecognized side based on the recognized first feature point, using these points as control points for expression conversion.

Benefits of technology

Enables conversion to a natural-looking facial expression even when only one side of the facial part is recognized, by incorporating additional feature points to guide the deformation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704288000002
    Figure 0007704288000002
  • Figure 0007704288000003
    Figure 0007704288000003
  • Figure 0007704288000004
    Figure 0007704288000004
Patent Text Reader

Abstract

An image converting device according to one embodiment comprises a control point generating unit and an expression converting unit. On the basis of a first feature point that is on one side of a face part recognized from an image of a person's face, the control point generating unit adds a second feature point that is on the other side that is not recognized, and sets the first and second feature points as control points. The expression converting unit transforms the control points with a transformation amount that corresponds to a conversion expression to be converted, thereby obtaining a converted image in which the expression of the person's face is converted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an image conversion apparatus, an image conversion method, and an image conversion program.

Background Art

[0002] Non-Patent Document 1 discloses the possibility of operating emotional experiences through real-time facial expression deformation feedback. In Non-Patent Document 1, the face of a subject is tracked in real time and natural facial expression deformation processing is performed. In Non-Patent Document 1, the Rigid MLS (Moving Least Squares) method is used as an image conversion method to deform the facial expression in a face image. The Rigid MLS method is a technique of distorting an image by moving each control point using the feature points recognized in the image recognized from the image as control points. Note that the face image is an image obtained by photographing the face of a subject, an image obtained by extracting the face of an avatar generated by a computer, or the like.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When recognizing facial parts, there is a feature point recognition method that recognizes only one side of the facial parts. When trying to move the facial parts recognized by such a feature point recognition method using the Rigid MLS method, only a face image with an unnatural expression can be obtained.

[0005] For example, regarding the eyebrows, there is a feature point recognition method in which only the upper side, which is one side of the eyebrows, is recognized, and the lower side of the other side is not recognized. When the eyebrows recognized only on the upper side are image-converted to move upward by this method, the obtained face image will have thick eyebrows and only an unnatural-expression face image can be obtained. In addition, regarding the width of the double eyelids, the shadow of the contour, etc., if only one side is recognized, when the image is deformed, similarly, an unnatural-expression face image will be obtained.

[0006] This invention aims to provide an image conversion technology that enables conversion to an image with a natural expression even when only one side of a facial part is recognized.

Means for Solving the Problem

[0007] In order to solve the above problem, an image conversion device according to an aspect of this invention includes a control point generation unit and an expression conversion unit. The control point generation unit adds a second feature point, which is the feature point on the unrecognized other side, based on the first feature point, which is the feature point on one side of the facial part recognized from an image of a person's face, and uses the first and second feature points as control points. The expression conversion unit obtains a converted image in which the expression of a person's face is converted by deforming the control points by a deformation amount corresponding to the conversion expression to be converted.

Effect of the Invention

[0008] According to an aspect of this invention, since an image is converted by adding the feature point on the other side based on the feature point on one side of the facial part, it is possible to provide an image conversion technology that enables conversion to a face image with a natural expression even when only one side of the facial part is recognized.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0010] [One Embodiment] Hereinafter, with reference to the drawings, one embodiment of the present invention will be described.

[0011] (Configuration Example) FIG. 1 is a block diagram showing an example of the configuration of an image conversion device 1 according to one embodiment of the present invention. The image conversion device 1 includes an image acquisition unit 11, a feature point recognition unit 12, a control point generation unit 13, a conversion expression input unit 14, a change amount storage unit 15, an expression conversion unit 16, and an image output unit 17.

[0012] The image acquisition unit 11 acquires a face image from a web camera, an avatar, or the like. The image acquisition unit 11 outputs the acquired face image to the feature point recognition unit 12 and the expression conversion unit 16.

[0013] The feature point recognition unit 12 takes the face image acquired by the image acquisition unit 11 as an input and recognizes feature points from the face image. The method of recognizing feature points in the feature point recognition unit 12 will be described later. The feature point recognition unit 12 outputs the recognized feature points to the control point generation unit 13.

[0014] The control point generation unit 13 takes as input the first feature points, which are the feature points recognized by the feature point recognition unit 12, and adds second feature points, which are unrecognized feature points, based on the input first feature points. For example, the control point generation unit 13 calculates the distance between the feature points of the eyebrows, which are the first feature points, and the feature points of the eyes, and adds second feature points downward by half of the distance obtained from each feature point of the eyebrows. The method of adding the second feature points will be described in detail later. The control point generation unit 13 outputs the first and second feature points as control points to the expression conversion unit 16. Which of the first feature points the second feature points are added based on, and the number of second feature points to be added are determined in advance. Therefore, the number of control points is also determined in advance.

[0015] The conversion expression input unit 14 acquires a conversion expression, which is an expression to be converted, such as a smiling face, specified and input by the user from a user interface such as a keyboard. The conversion expression input unit 14 outputs the acquired conversion expression to the expression conversion unit 16.

[0016] The change amount storage unit 15 stores in advance the change amount for each control point for each expression to be converted. The change amount is information indicating how much the control point should be moved. The change amount can be obtained in advance, for example, by the user adjusting while applying an expression deformation process from a neutral face to a natural expression for a specific face image.

[0017] The expression conversion unit 16 takes as input the face image acquired by the image acquisition unit 11, the control points output by the control point generation unit 13, and the conversion expression acquired by the conversion expression input unit 14. Also, the expression conversion unit 16 reads from the change amount storage unit 15 the change amount in the expression to be converted indicated by the conversion expression input from the conversion expression input unit 14. The change amount storage unit 15 moves each control point in the input face image based on the read movement amount of the control point to obtain a face image with the expression of the face image converted. The expression conversion unit 16 outputs the converted face image to the image output unit 17.

[0018] The image output unit 17 takes the converted face image from the expression conversion unit 16 as input and outputs the input face image. Here, output includes, for example, storing in a storage medium, displaying on a display, transmitting to other devices via a communication network, and so on.

[0019] FIG. 2 is a diagram showing an example of the hardware configuration of the image conversion apparatus 1.

[0020] The image conversion apparatus 1 is constituted by, for example, a computer such as a personal computer, a smartphone, or a server computer. As shown in FIG. 2, the image conversion apparatus 1 has a hardware processor 100 such as a CPU (Central Processing Unit). Note that the CPU can execute a plurality of information processes simultaneously by using a multi-core and multi-thread one. Also, the processor 100 may include a plurality of CPUs. In the image conversion apparatus 1, a program memory 200, a data memory 300, a communication interface 400, and an input / output interface (denoted as input / output IF in FIG. 2) 500 are connected to the processor 100 via a bus 600.

[0021] The communication interface 400 can include, for example, one or more wired or wireless communication modules. The communication interface 400 can communicate with other computers, web cameras, etc. connected via a cable or a network such as a LAN (Local Area Network) or the Internet.

[0022] An input / output interface 500 has an input unit 700 and a display unit 800 connected thereto. The input unit 700 includes input devices such as a keyboard, a pointing device such as a mouse, and sensor devices such as a camera. Further, the display unit 800 is a display device such as a liquid crystal display or a CRT (Cathode Ray Tube) display. The input unit 700 and the display unit 800 may use what is called a tablet-type input / display device. This type of input / display device is configured by arranging an input detection sheet adopting an electrostatic method or a pressure method on the display screen of a display device using, for example, liquid crystal or organic EL (Electro Luminescence). The input / output interface 500 inputs the operation information input in the input unit 700 to the processor 100 and causes the display unit 800 to display the display information generated by the processor 100.

[0023] Note that the input unit 700 and the display unit 800 do not necessarily have to be connected to the input / output interface 500. The input unit 700 and the display unit 800 can exchange information with the processor 100 by including a communication unit for directly connecting or connecting via a network to the communication interface 400.

[0024] Further, the input / output interface 500 may have a read / write function for a recording medium such as a semiconductor memory such as a flash memory, or may have a connection function with a reader / writer having a read / write function for such a recording medium. Furthermore, the input / output interface 500 may have a connection function with other devices.

[0025] The program memory 200 is a non-transitory tangible computer-readable storage medium, which is a combination of a non-volatile memory that can be written and read at any time and a non-volatile memory that can only be read at any time. The non-volatile memory that can be written and read at any time is, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. The non-volatile memory that can only be read at any time is, for example, a ROM, etc. In this program memory 200, programs necessary for the processor 100 to execute various control processes according to an embodiment, such as an image conversion program, are stored. That is, each processing function unit in each of the above-described image acquisition unit 11, feature point recognition unit 12, control point generation unit 13, converted expression input unit 14, change amount storage unit 15, expression conversion unit 16, and image output unit 17 can be realized by causing the processor 100 to read and execute the image conversion program stored in the program memory 200. Note that some or all of these processing function units may be realized in other various forms including an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0026] The data memory 300 is a tangible computer-readable storage medium, which is, for example, a combination of the above-described non-volatile memory and a volatile memory such as a RAM (Random Access Memory). This data memory 300 is used to store various data acquired and created during the process of performing various processes. That is, in the data memory 300, areas for appropriately storing various data are secured during the process of performing various processes. As such areas, for example, an acquired image storage unit 301, a feature point storage unit 302, a control point storage unit 303, a converted expression designation storage unit 304, a change amount storage unit 305, a converted image storage unit 306, and a temporary storage unit 307 can be provided in the data memory 300.

[0027] The acquired image storage unit 301 is used to store the face image acquired when the processor 100 operates as the above-described image acquisition unit 11.

[0028] The feature point storage unit 302 is used to store the feature points acquired when the processor 100 operates as the above-described feature point recognition unit 12.

[0029] FIG. 3 is a diagram showing an example of face feature points. The star marks in FIG. 3 are the feature points recognized by the processor 100, and the numbers attached to the sides of the respective feature points are unique feature point IDs for identifying the respective feature points. The number of feature point IDs and the face portions corresponding to the respective feature point IDs are determined by the feature point recognition method employed. For example, the feature point with the feature point ID "18" is predetermined to be the left end of the left eyebrow facing forward.

[0030] FIG. 4 is a diagram showing an example of the storage form of feature points in the feature point storage unit 302. As shown in FIG. 4, the feature point storage unit 302 stores the x coordinate and y coordinate in the face image in association with the feature point ID in a table format. The values of the coordinates are in pixels. Therefore, in the example of FIG. 3, the feature point storage unit 302 stores the xy coordinates of the feature points with the feature point IDs "1" to "68".

[0031] The control point storage unit 303 is used to store the control points generated when the processor 100 operates as the above-described control point generation unit 13. The storage form of the control points in this control point storage unit 303 is, for example, the same as the storage form of the feature points in the feature point storage unit 302 shown in FIG. 4. That is, the control point storage unit 303 can store the x coordinate and y coordinate in the face image in association with the control point ID in a table format. The control point storage unit 303 stores the xy coordinates of the feature point IDs "1" to "68" in association with each other, with the feature point IDs "1" to "68" assigned to the feature points shown in FIG. 3 remaining as the control point IDs "1" to "68". Further, the processor 100 stores the xy coordinates of each of the second feature points, which are the added feature points, in association with the control point IDs "69" to.

[0032] The conversion expression specification memory unit 304 is used to store the conversion expression specified by the user, which is obtained when the processor 100 operates as the above-mentioned conversion expression input unit 14.

[0033] The change amount memory unit 305 corresponds to the above-mentioned change amount storage unit 15.

[0034] FIG. 5 is a diagram showing an example of the storage form of the change amount in the change amount memory unit 305. As shown in FIG. 5, the change amount memory unit 305 can be in a table format that stores the change amount of the x coordinate and the change amount of the y coordinate in association with the control point ID for each conversion expression. The value of the change amount is in pixels. The change amount is represented by the moving direction and the moving amount of the control point. For example, the moving amount of “+1” represents moving 1 pixel in the positive direction.

[0035] The converted image memory unit 306 is used to store the face image converted when the processor 100 operates as the above-mentioned expression conversion unit 16.

[0036] The temporary memory unit 307 is used to store various intermediate data that occur during the operation of the processor 100 and are not stored in the above-mentioned acquired image memory unit 301, feature point memory unit 302, control point memory unit 303, conversion expression specification memory unit 304, change amount memory unit 305, and converted image memory unit 306.

[0037] (Operation) Next, the operation of the image conversion apparatus 1 having the image conversion apparatus 1 will be described.

[0038] FIG. 6 is a flowchart showing an example of the image conversion processing operation by the image conversion apparatus 1. The processor 100 of the image conversion apparatus 1 starts the operation as the image conversion apparatus 1 shown in this flowchart by reading and executing the image conversion program stored in the program memory 200. The execution of the image conversion program by the processor 100 is started by being instructed to perform image conversion from the input unit 700 via the input / output interface 500 or via the communication interface 400.

[0039] The processor 100 operates as a converted expression input unit 14 and waits for a user to specify and input a converted expression, which is an expression such as a smiling face that the user wants to convert (step S1). For example, the processor 100 determines whether an input signal from the input unit 700 via the input / output interface 500 or the communication interface 400 includes a specified input of a converted expression. If there is a specified input of a converted expression, the processor 100 proceeds to the process of step S2.

[0040] The processor 100 stores the specified converted expression in the converted expression specification storage unit 304 of the data memory 300 (step S2).

[0041] The processor 100 operates as an image acquisition unit 11 and acquires a face image (step S3). For example, the processor 100 acquires an image of a subject's face taken by the camera of the input unit 700 via the input / output interface 500. Alternatively, the processor 100 acquires a face image taken by a web camera connected to the network or the face of an avatar generated by another computer via the communication interface 400. The processor 100 stores the acquired face image in the acquired image storage unit 301 of the data memory 300.

[0042] The processor 100 operates as the feature point recognition unit 12 to recognize the first feature points from the face images stored in the acquired image storage unit 301 (step S4). For example, the processor 100 uses the face_landmark_detection function of dlib (see, for example, http: / / dlib.net / face_landmark_detection.py.html) to recognize the feature points for the face image. Specifically, the processor 100 extracts the distribution of the gradient directions of luminance, called HOG (Histogram of Oriented Gradients) features, for the input face image. A model learned based on the data associating the HOG features with the positions of the facial feature points is generally provided. Therefore, the processor 100 inputs the extracted HOG features into this learning model to obtain the positions of the facial feature points. The processor 100 stores the positions of the acquired first feature points in the feature point storage unit 302 of the data memory 300.

[0043] The processor 100 operates as the control point generation unit 13 to generate control points (step S5). Specifically, the processor 100 stores the recognized first feature points as control points in the control point storage unit 303 of the data memory 300. Further, for the facial parts where feature points are recognized only on one side, the processor 100 adds the second feature points, which are the feature points on the other side. Then the processor 100 also stores the added second feature points as additional control points in the control point storage unit 303.

[0044] Examples of facial parts where feature points are recognized only on one side include eyebrows, eyelids, and contours. For the eyebrows, since only the upper feature points are recognized, the processor 100 adds the lower feature points of the eyebrows. For the eyelids, since only the upper feature points of the eyes, which are the inner and lower feature points of the double eyelids, are recognized, the processor 100 adds the upper feature points. For the contours, since the parts without shadows are recognized as feature points, the processor 100 adds the feature points of the shadow parts.

[0045] For example, the processor 100 adds the feature points of the eyebrows as follows. FIG. 7 is a schematic diagram for explaining the relationship between the feature points of the eyebrows and the feature points above the eyes. The processor 100 drops perpendicular lines from each feature point of the eyebrows (the feature points with feature point IDs "18" to "22"), and obtains the feature point among the feature points above the eyes (the feature points with feature point IDs "37" to "40") that is closest to the perpendicular line. That is, if the feature point is the feature point with feature point ID "18", the feature point with feature point ID "37"; if the feature point is the feature point with feature point ID "19", the feature point with feature point ID "37"; if the feature point is the feature point with feature point ID "20", the feature point with feature point ID "38"; if the feature point is the feature point with feature point ID "21", the feature point with feature point ID "39"; if the feature point is the feature point with feature point ID "22", the feature point with feature point ID "40";..., if the feature point is the feature point with feature point ID "27", the feature point with feature point ID "46", is obtained. At this time, as the distance (difference in the vertical coordinates) between each feature point, d18, d19, d20, d21, d22,..., d27, is obtained.

[0046] FIG. 8 is a schematic diagram for explaining a method of adding the lower feature points that are the second feature points of the eyebrows. The processor 100 calculates the average distance da of the distances d18 to d27 between the above-mentioned feature points. At this time, the average distance da may be obtained without distinguishing between the right eye and the left eye. Generally, since there are slight differences between the left and right eyes, the average distance da may also be obtained separately. Here, it is assumed that the average distance is obtained by distinguishing between the left and right. That is, the processor 100 calculates the average distance da as da = (d18 + d19 + d20 + d21 + d22) / 5 for the left eyebrow facing, and da = (d23 + d24 + d25 + d26 + d27) / 5 for the right eyebrow facing, respectively.

[0047] The processor 100 adds a second feature point at a distance d below each of the first feature points of the eyebrows, where d is half of the average distance da thus calculated, i.e., da / 2. That is, the processor 100 adds a second feature point with the feature point ID "69" at a distance d below the first feature point with the feature point ID "18". Similarly, the processor 100 adds a second feature point with the feature point ID "70" below the first feature point with the feature point ID "19", a second feature point with the feature point ID "71" below the first feature point with the feature point ID "20", …, and a second feature point with the feature point ID "78" below the first feature point with the feature point ID "27".

[0048] In addition, the processor 100 adds feature points of the eyelids as follows, for example. FIG. 9 is a schematic diagram for explaining a method of adding a second feature point of the eyelid. The recognized feature points on one side of the eyelid are the feature points above the eye. Therefore, when adding the second feature point of the eyelid, the average distance da used when adding the second feature point of the eyebrows can also be used. The processor 100 sets the feature point addition distance d to 1 / 4 of this average distance da for each side, i.e., da / 4. The processor 100 adds a second feature point at a distance d above each of the feature points above the eyes (the first feature points with the feature point IDs "37" to "40"). That is, the processor 100 adds a second feature point with the feature point ID "79" at a distance d above the first feature point with the feature point ID "37". Similarly, the processor 100 adds a second feature point with the feature point ID "80" above the first feature point with the feature point ID "38", a second feature point with the feature point ID "81" above the first feature point with the feature point ID "39", …, and a second feature point with the feature point ID "86" above the first feature point with the feature point ID "46".

[0049] Regarding the size of the shadow of the contour, since the individual differences are not particularly large, the processor 100 adds a second feature point of the shadow of the contour as follows, for example. For example, the processor 100 adds a second feature point at a position with a predetermined direction and distance for each of the feature points of the contour (the first feature points with the feature point IDs "1" to "17").

[0050] The processor 100 operates as an expression conversion unit 16 to convert the expression of the face image stored in the acquired image storage unit 301 (step S6). That is, the processor 100 converts the face image based on the control points stored in the control point storage unit 303 and the change amount corresponding to the converted expression stored in the change amount storage unit 305 according to the converted expression stored in the converted expression specification storage unit 304. For example, the processor 100 utilizes the implementation of MLS (for example, refer to https: / / github.com / Jarvis73 / Moving-Least-Squares). Specifically, for each control point, the processor 100 moves it by the amount of change corresponding to the converted expression stored in the converted expression specification storage unit 304. For example, when changing the expression to a smiling face, for the control point with control point ID "1", since the xy coordinates before conversion are (23, 45) (refer to FIG. 4), the processor 100 performs a conversion such that the pixels of the control point are moved to (24, 47) by adding " + 1" to the x coordinate and " + 2" to the y coordinate (refer to FIG. 5).

[0051] Then, for points other than the control points, the processor 100 applies the following affine transformation (including Hermite transformation = similarity transformation and rigid deformation = rigid body deformation).

[0052]

Equation

[0053] However, x and y are the coordinates of neighboring control points, x' and y' are the coordinates obtained by adding the change amount to the coordinates of the control points, a, b, c, d are parameters, and t x , t y is a translation parameter. The processor 100 calculates the least squares mean of the coordinates x, y of the control points and the coordinates x', y' obtained by adding the change amount, and the parameters a, b, c, d, t x , t yIt is obtained by global optimization. Then, taking the coordinates of the target points to be transformed as x and y, the coordinates after transformation are obtained using these obtained parameters. The processor 100 uses the parameters a, b, c, d, t x ,t y to obtain the coordinates after transformation by the above affine transformation from the added control points.

[0054] The processor 100 stores the face image thus transformed in the transformed image storage unit 306 of the data memory 300 as a transformed image.

[0055] The processor 100 operates as an image output unit 17 and outputs the transformed image stored in the transformed image storage unit 306 (step S7). For example, the processor 100 causes the face image to be displayed on the display unit 800 via the input / output interface 500. Alternatively, the processor 100 transmits it via the communication interface 400 onto the network and causes it to be displayed on a display device connected to the network or on the display unit of another computer connected to the network.

[0056] The processor 100 determines whether to end the operation as the image conversion device 1 shown in this flowchart (step S8). For example, the processor 100 checks whether it has been instructed by the user to end the image conversion from the input unit 700, via the input / output interface 500, or via the communication interface 400. Here, if the operation is to end, the processor 100 ends the operation shown in this flowchart.

[0057] On the other hand, if the operation has not yet ended, the processor 100 operates as a transformed expression input unit 14 and determines whether there has been an input for specifying a change in the transformed expression by the user (step S9). If there is no input for specifying a change in the transformed expression, the processor 100 proceeds to the process of step S3. Also, if there has been an input for specifying a change in the transformed expression, the processor 100 proceeds to the process of step S2.

[0058] The image conversion device 1 according to the above-described embodiment includes a control point generation unit 13 and an expression conversion unit 16. The control point generation unit 13 adds a second feature point, which is a feature point on the unrecognized other side, based on a first feature point that is a feature point on one side of a face part recognized from an image of a person's face, and uses the first and second feature points as control points. The expression conversion unit 16 obtains a converted image in which the expression of a person's face is converted by deforming the control points by a deformation amount corresponding to the conversion expression to be converted. Therefore, since the image conversion device 1 according to the embodiment adds a feature point on the other side based on a feature point on one side of the face part and converts the image, it is possible to provide an image conversion technique that enables conversion into a face image with a natural expression even when only one side of the face part is recognized.

[0059] Furthermore, in the image conversion device 1 according to the embodiment, the face part includes at least eyebrows or eyelids. In this way, even if only the feature points above the eyebrows and / or below the eyelids can be recognized, it is possible to add the feature points below the eyebrows and / or above the eyelids and convert them into a face image with a natural expression.

[0060] Here, the control point generation unit 13 calculates a feature point addition distance d based on the distance between the feature point above the eyebrows, which is the first feature point, and the feature point above the eyes, and adds a point below the eyebrows that is at a distance d from the feature point above the eyebrows, which is the first feature point, as the second feature point. Therefore, it is possible to easily add the second feature point on the unrecognized lower side of the eyebrows.

[0061] Alternatively, the control point generation unit 13 calculates a feature point addition distance d based on the distance between the feature point above the eyebrows and the feature point below the eyelids, which is the first feature point, and adds a point above the eyelids that is at a distance d from the feature point below the eyelids, which is the first feature point, as the second feature point. Therefore, it is possible to easily add the second feature point on the unrecognized upper side of the eyelids.

[0062] Also, in the image conversion device 1 according to one embodiment, a change amount storage unit 15 that stores in advance a change amount representing a deformation amount for each control point for each conversion expression to be converted, and a conversion expression input unit 14 that inputs a conversion expression to be converted are provided. Then, the expression conversion unit 16 reads out the change amount corresponding to the input conversion expression from the change amount storage unit, and obtains a converted image using the read-out change amount. In this way, by storing in advance the change amount also for the control point corresponding to the second feature point to be added, it becomes possible to convert the face image with a natural expression using also the control point corresponding to the second feature point.

[0063] [Other Embodiments] Note that the present invention is not limited to the above-described embodiment. For example, the flow of each process described above is not limited to the described procedure, and the order of some steps may be interchanged, or some steps may be performed in parallel.

[0064] Also, although the flow of each process described above was for the case of converting the expression of the face image acquired in real time in real time, it can be similarly applied to the use of converting the expression of the stored face image instead of real-time processing.

[0065] When adding the second feature point, although a fixed value of 1 / 2 or 1 / 4 is used for the average distance da, the user may be able to specify an arbitrary value.

[0066] Also, the user may be able to select to which part of the face the second feature point, which is the feature point on the other side, is added.

[0067] In addition, the method described in the embodiment can be stored as a program (software means) to be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy (registered trademark) disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), and can also be transmitted and distributed via a communication medium. Note that the program stored on the medium side includes a setting program for configuring software means (including not only an execution program but also tables and data structures) to be executed by a computer in the computer. The computer that realizes this apparatus reads the program recorded on the recording medium, and in some cases, constructs software means using the setting program, and executes the above-described processing by being controlled by this software means. Note that the recording medium referred to in this specification includes not only a medium for distribution but also a storage medium such as a magnetic disk or a semiconductor memory provided inside a computer or in a device connected via a network.

[0068] In short, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, each embodiment may be implemented in an appropriate combination as much as possible, and in that case, the combined effects can be obtained. Furthermore, the above-described embodiments include inventions at various stages, and various inventions can be extracted by appropriate combinations of a plurality of disclosed constituent elements.

Explanation of Reference Numerals

[0069] 1... Image conversion device 11... Image acquisition unit 12... Feature point recognition unit 13... Control point generation unit 14... Transformed expression input unit 15... Variation amount storage unit 16... Expression conversion unit 17... Image output unit 100... Processor 200... Program memory 300... Data memory 301... Acquired image storage unit 302… Feature point memory unit 303… Control point memory unit 304… Transformed expression specification memory unit 305… Variation amount memory unit 306… Transformed image memory unit 307… Temporary memory unit 400… Communication interface 500… Input / output interface 600… Bus 700… Input unit 800… Display unit

Claims

1. Based on a first feature point which is a feature point on one side of a facial part recognized from an image of a human face, a second feature point which is a feature point on the other unrecognized side is added, and a control point generation unit using the first and second feature points as control points; An expression conversion unit that obtains a converted image in which the expression of the human face is converted by deforming the control points by a deformation amount corresponding to a conversion expression to be converted; Comprising: The facial part includes at least eyebrows or eyelids; The control point generation unit: Calculates an additional feature point distance based on the distance between a feature point above the eyebrow which is the first feature point and a feature point above the eye; An image conversion device that adds, as the second feature point, a point below the eyebrow that is separated from the feature point above the eyebrow which is the first feature point by the additional feature point distance.

2. Based on a first feature point which is a feature point on one side of a facial part recognized from an image of a human face, a second feature point which is a feature point on the other unrecognized side is added, and a control point generation unit using the first and second feature points as control points; An expression conversion unit that obtains a converted image in which the expression of the human face is converted by deforming the control points by a deformation amount corresponding to a conversion expression to be converted; Comprising: The facial part includes at least eyebrows or eyelids; The control point generation unit: Calculates an additional feature point distance based on the distance between a feature point above the eyebrow and a feature point below the eyelid which is the first feature point; An image conversion device that adds, as the second feature point, a point above the eyelid that is separated from the feature point below the eyelid which is the first feature point by the additional feature point distance.

3. A change amount storage unit that stores in advance a change amount representing the deformation amount for each of the control points for each conversion expression to be converted; A conversion expression input unit that inputs the conversion expression to be converted; Further comprising: The expression conversion unit reads out the change amount corresponding to the input conversion expression from the change amount storage unit, and obtains the converted image using the read change amount. The image conversion device according to claim 1 or 2.

4. An image conversion method in an image conversion device having a processor for converting an expression in an image of a human face, the method comprising: By the processor, based on a first feature point which is a feature point on one side of a facial part recognized from the image of the human face, a second feature point which is a feature point on the other unrecognized side is added, and the first and second feature points are used as control points; The facial part includes at least eyebrows or eyelids; Calculate a feature point additional distance based on the distance between the feature point above the eyebrow, which is the first feature point, and the feature point above the eye. Add, as the second feature point, a point below the eyebrow that is separated from the feature point above the eyebrow, which is the first feature point, by the feature point additional distance. Obtain a transformed image in which the expression of the person's face is transformed by deforming the control points by an amount of deformation corresponding to the transformed expression to be transformed by the processor. Image conversion method.

5. An image conversion method in an image conversion apparatus having a processor for converting an expression in an image of a person's face, the method comprising: Based on a first feature point, which is a feature point on one side of a face part recognized from the image of the person's face, the processor adds a second feature point, which is a feature point on the other side that has not been recognized, and uses the first and second feature points as control points. The face part includes at least an eyebrow or an eyelid. Calculate a feature point additional distance based on the distance between the feature point above the eyebrow and the feature point below the eyelid, which is the first feature point. Add, as the second feature point, a point above the eyelid that is separated from the feature point below the eyelid, which is the first feature point, by the feature point additional distance. Obtain a transformed image in which the expression of the person's face is transformed by deforming the control points by an amount of deformation corresponding to the transformed expression to be transformed by the processor. Image conversion method.

6. An image conversion program for causing a processor to function as each part of the image conversion apparatus according to any one of Claims 1 to 3.

Citation Information

Patent Citations

  • Method and system for improving portrait image processed in batch mode

    JP2004265406A

  • Method, device and program for image processing

    JP2005215763A

  • Model creation apparatus

    JP2007026088A

  • Face direction detecting program, face direction detecting method, and face direction detecting unit

    JP2009294999A

  • Method for aging appearance simulation

    JP2020515952A