Image generation method and image generation program
The image generation method corrects facial features based on the shooting direction to address the issue of unnatural expressions in output face images, ensuring a more natural and realistic appearance.
Patent Information
- Application Number
- JP2024089022
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-05-31
AI Technical Summary
Conventional face tracking techniques struggle to maintain natural expressions of output face images when the shooting direction and orientation of the face change.
The proposed method involves an image generation method and program that corrects the shape or position of facial features based on the shooting direction to prevent unnatural expressions in output face images.
By dynamically adjusting facial features according to the shooting direction, the method effectively prevents unnatural expressions and maintains a more realistic appearance in output face images.
Smart Images

Figure 0007696478000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation method and an image generation program for generating a facial image of an expression according to an input facial image.
Background Art
[0002] Conventionally, for example, a technique has been proposed in which by performing face tracking on a face photographed with a smartphone, the expression of the photographed face can be reflected in the expression of an avatar such as a character created by CG or the like (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the conventional technique of reflecting the expression of a face obtained by face tracking in the expression of an avatar, when outputting a face image and an image associated therewith, there is a problem that the expression of the output face image becomes unnatural depending on the shooting direction and the orientation of the face to be tracked.
[0005] An object of the present invention is to provide an image generation method and an image generation program that prevent the expression of a face image output by face tracking from becoming unnatural.
Means for Solving the Problems
[0006] The image generation method of Means 1 is a face image input step of inputting a photographed face image, a face image generation step of generating a face image of an expression according to the face image input in the face image input step, An image output step of outputting an output image including the face image generated in the face image generation step; In an image generation method using a computer including: In the face image generation step, correcting the shape or position of the parts constituting the face image according to the shooting direction It is characterized by this. According to this feature, by correcting the shape or position of the parts constituting the face image according to the shooting direction, it is possible to prevent the expression of the face image from becoming unnatural.
[0007] The image generation method of means 2 is the image generation method described in means 1, In the face image generation step, correcting the contour of the face image according to the shooting direction It is characterized by this. According to this feature, it is possible to prevent the contour of the face image from becoming unnatural even when the shooting direction changes.
[0008] The image generation method of means 3 is the image generation method described in means 1 or 2, In the face image generation step, correcting the position or shape of the mouth constituting the face image according to the shooting direction It is characterized by this. According to this feature, it is possible to prevent the position or shape of the mouth constituting the face image from becoming unnatural even when the shooting direction changes.
[0009] The image generation method of means 4 is the image generation method described in any one of means 1 to 3, In the face image generation step, correcting the position or shape of the eyes constituting the face image according to the shooting direction It is characterized by this. According to this feature, it is possible to prevent the position or shape of the eyes constituting the face image from becoming unnatural even when the shooting direction changes.
[0010] The image generation method of means 5 is the image generation method described in any one of means 1 to 4, In the face image generation step, the correction amount of the shape or position of the parts constituting the face image is gradually increased according to the shooting direction. It is characterized by this. According to this feature, since the correction amount of the shape or position of the parts constituting the face image gradually increases according to the change in the shooting direction, the shape does not change suddenly.
[0011] The image generation program of means 6 A face image input step for inputting a captured face image, A face image generation step for generating a face image with an expression corresponding to the face image input in the face image input step, An image output step for outputting an output image including the face image generated in the face image generation step, In an image generation program that causes a computer to execute, In the face image generation step, the shape or position of the parts constituting the face image is corrected according to the shooting direction. It is characterized by this. According to this feature, by correcting the shape or position of the parts constituting the face image according to the shooting direction, it is possible to prevent the expression of the face image from becoming unnatural.
[0012] Note that the present invention may have only the invention specific matters described in the claims of the present invention, or may have configurations other than the invention specific matters together with the invention specific matters described in the claims of the present invention.
Brief Description of Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Embodiment for Carrying Out the Invention
[0014] [Embodiment 1] The image generation method of Embodiment 1-1 includes a face image input step of inputting a captured face image, a parameter detection step of detecting parameters indicating the states of a plurality of parts in the face image input by the face image input step, a face image generation step of generating a face image with an expression according to a plurality of parameters, an image output step of outputting an output image including the face image generated in the face image generation step, In an image generation method using a computer including an information input step of inputting information other than the face image, a parameter correction step of correcting the parameters detected in the parameter detection step based on the information other than the face image input in the information input step, further includes In the face image generation step, a face image with an expression according to the parameters corrected by the parameter correction step is generated which is characterized by According to this feature, since a face image with an expression according to the parameters corrected based on information other than the captured face image is generated, the face image detected by face tracking can be output as a richly variable image.
[0015] The image generation method of Embodiment 1-2 is the image generation method described in Embodiment 1-1, and further includes an emotion value identification step of identifying an emotion value based on the information other than the face image, In the parameter correction step, the parameters detected in the parameter detection step are corrected based on the emotion value identified in the emotion value identification step which is characterized by According to this feature, a facial image of an expression based on an emotion value specified from information other than the facial image can be generated.
[0016] The image generation method of Form 1-3 is the image generation method described in Form 1-2, In the emotion value specification step, the emotion value is specified based on information other than the facial image and parameters of some parts detected in the parameter detection step. It is characterized by this. According to this feature, the emotion value can be specified more accurately.
[0017] The image generation method of Form 1-4 is the image generation method described in Form 1-3, In the parameter correction step, when the emotion value specified in the emotion value specification step exceeds a threshold value, the parameters detected in the parameter detection step are corrected. It is characterized by this. According to this feature, a sense of rhythm can be added to the change of the expression.
[0018] The image generation method of Form 1-5 is the image generation method described in any one of Forms 1-1 to 1-4, The information other than the facial image includes information related to at least any one of voice, vital signs, and operation input. It is characterized by this. According to this feature, a facial image of an expression based on information related to at least any one of voice, vital signs, and operation input can be generated.
[0019] The image generation method of Form 1-6 is the image generation method described in any one of Forms 1-1 to 1-5, It further includes a set value reception step of receiving a set value related to the change amount of the plurality of parts, The information other than the facial image includes the set value set in the set value reception step. It is characterized by this. According to this feature, a face image with an expression corresponding to a parameter corrected based on set values related to the amounts of change of a plurality of parts can be generated.
[0020] The correction content setting method of Form 1-7 is a correction content setting method using a computer for setting the correction content of the parameter used in the image generation method described in any of Forms 1-1 to 1-6, a first node setting step of setting a first node for designating a parameter to be corrected; a second node setting step of setting a second node for designating a parameter to be used for correction; a third node setting step of setting a third node for setting correction content; a connection step of connecting the first node and the second node to the third node; a correction content setting step of setting, in the third node, the correction content of the parameter of the first node by the parameter of the second node; a step of creating setting data based on the set content; including and is characterized by this. According to this feature, the correction content of the parameter can be easily set only by connecting the nodes to set the correction content.
[0021] The image generation program of Form 1-8 a face image input step of inputting a photographed face image; a parameter detection step of detecting a parameter indicating the states of a plurality of parts in the face image input by the face image input step; a face image generation step of generating a face image with an expression corresponding to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program for causing a computer to execute, an information input step of inputting information other than the face image; A parameter correction step for correcting the parameters detected in the parameter detection step based on information other than the face image input in the information input step; further execute In the face image generation step, a face image with an expression corresponding to the parameters corrected by the parameter correction step is generated which is characterized by. According to this feature, since a face image with an expression corresponding to the parameters corrected based on information other than the photographed face image is generated, the face image detected by face tracking can be output as a richly changing image.
[0022] The correction content setting program of Form 1-9 is a correction content setting program for setting, using a computer, the correction content of the parameters used in the image generation program described in Form 1-8, A first node setting step for setting a first node that designates a parameter to be corrected; A second node setting step for setting a second node that designates a parameter to be used for correction; A third node setting step for setting a third node that sets correction content; A connection step of connecting the first node, the second node, and the third node; A correction content setting step of setting, in the third node, the correction content of the parameter of the first node by the parameter of the second node; A step of creating setting data based on the setting content; execute which is characterized by. According to this feature, the correction content of the parameters can be easily set only by connecting the nodes and setting the correction content.
[0023] [Form 2] The image generation method of Form 2-1 is A face image input step for inputting a photographed face image; A parameter detection step of detecting parameters indicating the states of a plurality of parts among the face images input in the face image input step; A face image generation step of generating a face image of an expression according to a plurality of parameters; An image output step of outputting an output image including the face image generated in the face image generation step; In an image generation method using a computer including: Further including a parameter correction step of correcting the parameters of other parts based on the parameters of some parts detected in the parameter detection step; In the face image generation step, a face image of an expression according to the parameters corrected by the parameter correction step is generated It is characterized by this. According to this feature, based on the parameters of some parts detected from the captured face image, the parameters of other parts are corrected, and a face image of an expression according to the corrected parameters is generated. Therefore, the face image detected by face tracking can be output as a richly changing image.
[0024] The image generation method of Form 2-2 is the image generation method described in Form 2-1, Further including an emotion value identification step of identifying an emotion value based on the parameters of the some parts; In the parameter correction step, the parameters of the other parts detected in the parameter detection step are corrected based on the emotion value identified in the emotion value identification step It is characterized by this. According to this feature, a face image of an expression based on the emotion value identified based on the parameters of some parts can be generated.
[0025] The image generation method of Form 2-3 is the image generation method described in Form 2-2, In the emotional value determination step, a first value indicating joy and a second value indicating anger are determined based on the parameters of the partial parts, and when the first value is greater than the second value, an emotional value indicating joy is determined, and when the second value is greater than the first value, an emotional value indicating anger is determined. It is characterized by this. According to this feature, a face image of an expression based on an emotional value reflecting both joy and anger specified from a face image can be generated.
[0026] The image generation method of Form 2-4 is the image generation method described in Form 2-3, and in the emotional value determination step, the first value and the second value are determined based on a parameter indicating the degree of movement of the corners of the mouth. It is characterized by this. According to this feature, emotions of joy and anger can be determined from the degree of movement of the corners of the mouth where emotions are easily reflected.
[0027] The image generation method of Form 2-5 is the image generation method described in any one of Forms 2-2 to 2-4, and in the parameter correction step, when the emotional value determined in the emotional value determination step exceeds a threshold value, the parameters of the other parts are corrected. It is characterized by this. According to this feature, sharpness can be added to the change of the expression.
[0028] The correction content setting method of Form 2-6 is a correction content setting method using a computer for setting the correction content of the parameters used in the image generation method described in any one of Forms 2-1 to 2-5, and a first node setting step of setting a first node for designating a parameter to be corrected, a second node setting step of setting a second node for designating a parameter to be used for correction, a third node setting step of setting a third node for setting correction content, and a connection step of connecting the first node, the second node, and the third node. A correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; A step of creating setting data based on the set content; including is characterized in that. According to this feature, it is possible to easily set the correction content of the parameters only by connecting the nodes and setting the correction content.
[0029] The image generation program of Form 2-7 is A face image input step of inputting a captured face image; A parameter detection step of detecting parameters indicating the states of a plurality of parts in the face image input by the face image input step; A face image generation step of generating a face image with an expression according to a plurality of parameters; An image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program for causing a computer to execute, further execute a parameter correction step of correcting the parameters of other parts based on the parameters of some parts detected in the parameter detection step; In the face image generation step, generate a face image with an expression according to the parameters corrected by the parameter correction step is characterized in that. According to this feature, since the parameters of other parts are corrected based on the parameters of some parts detected from the captured face image, and a face image with an expression according to the corrected parameters is generated, it is possible to output the face image detected by face tracking as an image rich in changes.
[0030] The correction content setting program of Form 2-8 is a correction content setting program for setting, using a computer, the correction content of the parameters used in the image generation program described in Form 2-7, A first node setting step of setting a first node that designates a parameter to be corrected; A second node setting step of setting a second node that designates a parameter to be used for correction; A third node setting step of setting a third node that sets correction content; A connection step of connecting the first node, the second node, and the third node; A correction content setting step of setting, in the third node, correction content of the parameter of the first node based on the parameter of the second node; A step of creating setting data based on the set content; which is characterized by the above. According to this feature, the correction content of the parameter can be easily set only by connecting the nodes and setting the correction content.
[0031] [Mode 3] The image generation method of Mode 3-1 is A face image input step of inputting a captured face image; A parameter detection step of detecting parameters indicating the states of a plurality of parts in the face image input by the face image input step; A face image generation step of generating a face image with an expression corresponding to a plurality of parameters; An image output step of outputting an output image including the face image generated in the face image generation step; In an image generation method using a computer including further includes a decorative image generation step of generating a decorative image other than the face image based on the parameters detected in the parameter detection step, and in the image output step, the decorative image generated in the decorative image generation step is output together with the face image which is characterized by the above. According to this feature, since a decorative image other than the face image is generated based on the parameters detected from the captured face image, the face image detected by face tracking can be output as a richly changing image.
[0032] The image generation method of Form 3-2 is the image generation method described in Form 3-1, further including an emotional value identification step of identifying an emotional value based on parameters of at least some parts, In the decorative image generation step, the decorative image is generated based on the emotional value identified in the emotional value identification step. It is characterized by this. According to this feature, a decorative image can be generated according to the parameters corrected based on the emotional value identified based on the parameters of at least some parts.
[0033] The image generation method of Form 3-3 is the image generation method described in Form 3-2, In the emotional value identification step, based on the parameters of the some parts, a first value indicating joy and a second value indicating anger are identified. When the first value is greater than the second value, an emotional value indicating joy is identified. When the second value is greater than the first value, an emotional value indicating anger is identified. It is characterized by this. According to this feature, a decorative image based on the emotional value reflecting both joy and anger identified from the face image can be generated.
[0034] The image generation method of Form 3-4 is the image generation method described in Form 3-3, In the emotional value identification step, the first value and the second value are identified based on the parameter indicating the degree of movement of the corners of the mouth. It is characterized by this. According to this feature, the emotions of joy and anger can be identified from the degree of movement of the corners of the mouth where emotions are easily reflected.
[0035] The image generation method of Form 3-5 is the image generation method described in any one of Forms 3-2 to 3-4, In the decorative image generation step, when the emotional value identified in the emotional value identification step exceeds a threshold value, the decorative image is generated. It is characterized by the following. According to this feature, a decorative image can be generated at an appropriate frequency.
[0036] The image generation program of Forms 3-6 A face image input step for inputting a captured face image, A parameter detection step for detecting parameters indicating the states of a plurality of parts in the face image input by the face image input step, A face image generation step for generating a face image with an expression according to a plurality of parameters, An image output step for outputting an output image including the face image generated in the face image generation step, In an image generation program for causing a computer to execute, Further execute a decorative image generation step for generating a decorative image other than the face image based on the parameters detected in the parameter detection step, In the image output step, output the decorative image generated in the decorative image generation step together with the face image It is characterized by the following. According to this feature, since a decorative image other than the face image is generated based on the parameters detected from the captured face image, the face image detected by face tracking can be output as a richly changing image.
[0037] [Form 4] The image generation method of Form 4-1 A face image input step for inputting a captured face image, A face image generation step for generating a face image with an expression according to the face image input by the face image input step, An image output step for outputting an output image including the face image generated in the face image generation step, In an image generation method using a computer including Further include a specific image generation step for generating a specific image displayed around the face image, In the image output step, the specific image generated in the specific image generation step is output together with the face image. In the specific image generation step, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the specific image is deformed into a shape where the expression of the face image can be visually recognized. It is characterized by this. According to this feature, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the specific image is deformed into a shape where the expression of the face image can be visually recognized, so that it is possible to prevent the expression of the face image from being hidden by the specific image.
[0038] The image generation method of Form 4-2 is the image generation method described in Form 4-1, In the specific image generation step, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position where the expression of the face image cannot be visually recognized, the specific image is deformed toward a shape where the expression of the face image can be visually recognized. It is characterized by this. According to this feature, the specific image can be deformed toward a shape where the expression of the face image can be visually recognized before the expression of the face image is hidden.
[0039] The image generation method of Form 4-3 is the image generation method described in Form 4-2, In the specific image generation step, the specific image is gradually deformed toward a shape where the expression of the face image can be visually recognized according to the distance to the position where the expression of the face image cannot be visually recognized. It is characterized by this. According to this feature, the specific image can be naturally deformed toward a shape where the expression of the face image can be visually recognized.
[0040] The image generation method of Form 4-4 is a face image input step of inputting a captured face image, and a face image generation step of generating a face image with an expression corresponding to the face image input in the face image input step. An image output step of outputting an output image including the face image generated in the face image generation step; In an image generation method using a computer including further including a specific image generation step of generating a specific image displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generation step, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the position of the specific image is moved to a position where the expression of the face image can be visually recognized It is characterized by this. According to this feature, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the position of the specific image is moved to a position where the expression of the face image can be visually recognized, so that it is possible to prevent the expression of the face image from being hidden by the specific image.
[0041] The image generation method of Form 4-5 is the image generation method described in Form 4-4, In the specific image generation step, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position where the expression of the face image cannot be visually recognized, the position of the specific image is moved toward a position where the expression of the face image can be visually recognized It is characterized by this. According to this feature, the position of the specific image can be moved toward a position where the expression of the face image can be visually recognized before the expression of the face image is hidden.
[0042] The image generation method of Form 4-6 is the image generation method described in Form 4-5, In the specific image generation step, the position of the specific image is gradually moved toward a position where the expression of the face image can be visually recognized according to the distance to a position where the expression of the face image cannot be visually recognized It is characterized by this. According to this feature, the position of the specific image can be naturally moved toward a position where the expression of the face image can be visually recognized.
[0043] The image generation programs of Forms 4-7 include a face image input step of inputting a captured face image, a face image generation step of generating a face image with an expression corresponding to the face image input in the face image input step, an image output step of outputting an output image including the face image generated in the face image generation step, In an image generation program for causing a computer to execute, further execute a specific image generation step of generating a specific image to be displayed around the face image, In the image output step, output the specific image generated in the specific image generation step together with the face image, In the specific image generation step, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, deform the specific image into a shape where the expression of the face image can be visually recognized It is characterized by this. According to this feature, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the specific image is deformed into a shape where the expression of the face image can be visually recognized, so that it is possible to prevent the expression of the face image from being hidden by the specific image.
[0044] The image generation program of Form 4-8 includes a face image input step of inputting a captured face image, a face image generation step of generating a face image with an expression corresponding to the face image input in the face image input step, an image output step of outputting an output image including the face image generated in the face image generation step, In an image generation program for causing a computer to execute, further execute a specific image generation step of generating a specific image to be displayed around the face image, In the image output step, output the specific image generated in the specific image generation step together with the face image, In the specific image generation step, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, move the position of the specific image to a position where the expression of the face image can be visually recognized It is characterized by this According to this feature, when the position of the specific image becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction, the position of the specific image is moved to a position where the expression of the face image can be visually recognized, so that it is possible to prevent the expression of the face image from being hidden by the specific image
[0045] [Form 5] The image generation method of Form 5-1 is A face image input step of inputting a captured face image, A face image generation step of generating a face image with an expression corresponding to the face image input in the face image input step, An image output step of outputting an output image including the face image generated in the face image generation step, In an image generation method using a computer including In the face image generation step, correct the shape or position of the parts constituting the face image according to the shooting direction It is characterized by this According to this feature, by correcting the shape or position of the parts constituting the face image according to the shooting direction, it is possible to prevent the expression of the face image from becoming unnatural
[0046] The image generation method of Form 5-2 is the image generation method described in Form 5-1, and In the face image generation step, correct the contour of the face image according to the shooting direction It is characterized by this According to this feature, it is possible to prevent the contour of the face image from becoming unnatural even when the shooting direction changes
[0047] The image generation method of Form 5-3 is the image generation method described in Form 5-1 or 5-2, and In the face image generation step, correct the position or shape of the mouth that constitutes the face image according to the shooting direction. It is characterized by this. According to this feature, it is possible to prevent the position or shape of the mouth that constitutes the face image from becoming unnatural even when the shooting direction changes.
[0048] The image generation method of Form 5-4 is the image generation method described in any one of Forms 5-1 to 5-3, In the face image generation step, correct the position or shape of the eyes that constitute the face image according to the shooting direction. It is characterized by this. According to this feature, it is possible to prevent the position or shape of the eyes that constitute the face image from becoming unnatural even when the shooting direction changes.
[0049] The image generation method of Form 5-5 is the image generation method described in any one of Forms 5-1 to 5-4, In the face image generation step, gradually increase the correction amount of the shape or position of the parts that constitute the face image according to the shooting direction. It is characterized by this. According to this feature, since the correction amount of the shape or position of the parts that constitute the face image gradually increases according to the change in the shooting direction, the shape and position do not change suddenly.
[0050] The image generation program of Form 5-6 is A face image input step for inputting a captured face image, A face image generation step for generating a face image with an expression corresponding to the face image input by the face image input step, An image output step for outputting an output image including the face image generated in the face image generation step, In an image generation program that causes a computer to execute, In the face image generation step, correct the shape or position of the parts that constitute the face image according to the shooting direction. It is characterized by this. According to this feature, by correcting the shape or position of the parts that make up the face image according to the shooting direction, it is possible to prevent the expression of the face image from becoming unnatural.
[0051] The embodiments for carrying out the present invention will be described below with reference to the drawings based on examples.
Example
[0052] [Computer terminal] FIG. 1 is a block diagram showing a configuration example of a computer terminal 1 used in an embodiment of the present invention. The computer terminal 1 shown in FIG. 1 may use an input device such as a display device, a keyboard, and a mouse, such as a separate desktop PC or workstation, a notebook PC with an integrated display device and input device, a tablet terminal or a mobile terminal such as a smartphone with a touch panel or the like mounted on the display device. Here, however, a mobile terminal such as a tablet terminal or a smartphone will be assumed and described.
[0053] As shown in FIG. 1, the computer terminal 1 is equipped with a processor 101, a memory 102, and a storage 103 such as a hard disk or SSD. These processor 101, memory, and storage 103 are connected via a data bus 111, and the processor 101 can execute various processes according to the programs stored in the storage 103.
[0054] Also, a display device 105 and a speaker 106 are connected to the data bus 111 via an output interface (not shown), and it is possible to output images and sounds based on the processes executed by the processor 101.
[0055] Also, input devices such as an input device 107, a camera 108, a microphone 109, and a vital measurement device 110 are connected to the data bus 111 via an input interface (not shown), and the information input by these input devices can be input. The input device 107 includes a touch panel integrated with the display device 105, and can input a command input operation using the touch panel. Note that a configuration in which commands input by an input device such as a keyboard or a mouse can be input may also be used. The camera 108 is a 3D camera and can input data including depth data as photographed image data. Further, the camera 108 may be configured to include both an out-camera provided on the opposite side of the display device 105 and an in-camera provided on the display device 105 side, or may include only one of them. The microphone 109 may be a stereo microphone or a monaural microphone, and can input ambient sound including the voice of the subject. The vital measurement device 110 is a wearable device such as a smart band and can input vital data (heart rate, blood pressure, body temperature, etc.).
[0056] Also, a communication interface 104 is connected to the data bus 111, and is configured to be able to communicate with other computer terminals and server computers via a public network such as a local area network or the Internet, either wired or wirelessly.
[0057] In this embodiment, a configuration in which the image generation method and the correction content setting method of the present invention are implemented using one computer terminal is illustrated, but a configuration in which the image generation method and the correction content setting method of the present invention are implemented using a plurality of computer terminals may also be used. Further, in the computer terminal 1 of this embodiment, the display device 105 and the speaker 106 are provided as output devices, but a configuration in which these output devices are not provided and images and sounds are output from the output devices of another computer terminal may also be used. Further, in the computer terminal 1 of this embodiment, an input device 107 such as a touch panel, and input devices such as a camera 108 and a microphone 109 are provided, but a configuration including at least an input device such as a touch panel, a keyboard, a mouse, or the like, and a camera capable of photographing images may be used.
[0058] In the storage 103 of the computer terminal 1, in addition to an operation system (OS) (not shown), various programs executed by the computer terminal 1 are stored as shown in FIG. 2. Specifically, a face tracking program, a motion tracking program, a voice detection program, an input detection program, a vital detection program, a parameter correction program, an image generation program, a correction content setting program, and a distribution program are stored. Although not particularly shown, setting data used by the parameter correction program, configuration setting values, material data used by the image generation program, etc. are also stored.
[0059] As shown in FIG. 3, the computer terminal 1 creates avatar image data of an expression based on the expression parameters and a form based on the motion parameters from the image data output from the camera 108 by the image generation program, and the distribution program can distribute a video using the avatar image data created by the image generation program. Also, the expression parameters output by the face tracking program are not directly output to the image generation program, but are output to the image generation program via the parameter correction program. The parameter correction program corrects at least some of the expression parameters using some of the expression parameters and parameters other than the expression parameters (voice parameters, input parameters, vital parameters, configuration parameters), and outputs the corrected expression parameters to the image generation program.
[0060] Next, the programs executed by the computer terminal 1 will be described.
[0061] As shown in FIG. 3, the face tracking program is a program that uses image data including depth data input from the camera 108 to detect the states of a plurality of parts constituting the face image from the face image of the subject, and outputs expression parameters obtained by quantifying the degree of movement of each part from the detected states. The expression parameters are parameters that can identify the expression of the subject. In this embodiment, as shown in FIG. 4, the expression parameters are parameters obtained by quantifying the degrees of six types of movements related to the movement of the left eye, the degrees of six types of movements related to the movement of the right eye, the degrees of 27 types of movements related to the movement of the mouth and jaw, the degrees of 10 types of movements related to the movement of the eyebrows, cheeks, and nose, and the degree of one type of movement related to the movement of the tongue, each in the range of 0.0 to 1.0. Note that the present invention is not limited to the illustrated expression parameters, and more finely divided and many types of expression parameters may be used, or fewer types of expression parameters may be used.
[0062] As shown in FIG. 3, the motion tracking program is a program that uses image data including depth data input from the camera 108 to detect the movement of the body of the subject, and outputs motion parameters that specify the positions and directions of the parts constituting the body. The motion parameters are parameters that can identify the movement of the body of the subject. In this embodiment, as shown in FIG. 4, the motion parameters include a head pose indicating the position coordinates and direction of the head, all finger joints indicating the position coordinates and direction of all finger joints, and a hand pose indicating the position coordinates and direction of the hand. Note that the motion parameters of this embodiment do not include parameters regarding the movement of the entire body, but may be configured to include parameters regarding the movement of the entire body.
[0063] As shown in FIG. 2, the voice detection program is a program that uses voice data input from the microphone 109 to detect the voice of the subject, and outputs voice parameters obtained by quantifying the components of the detected voice. The voice parameters are parameters that can identify the voice of the subject. In this embodiment, as shown in FIG. 4, the voice parameters are parameters obtained by quantifying the voice volume and the voice pitch, each in the range of 0.0 to 1.0.
[0064] As shown in FIG. 3, the input detection program is a program that converts input data from the input device 107 into input parameters and outputs them. The input parameters are parameters that can specify the input status from the input device 107. In this embodiment, as shown in FIG. 4, the input parameters are parameters indicating the type of command input by a touch operation corresponding to the command displayed on the display device 105. The command input includes, for example, command inputs for gradually specifying emotions such as joy, anger, sorrow, and pleasure, and command inputs for specifying decorative images.
[0065] As shown in FIG. 3, the vital detection program is a program that converts vital data (such as heart rate, blood pressure, body temperature, etc.) input from the vital measurement device 110 and outputs vital parameters. The vital parameters are parameters that can specify the vital status of the subject. In this embodiment, as shown in FIG. 4, the vital parameters are parameters obtained by converting the heart rate, blood pressure, and body temperature into numerical values in the range of 0.0 to 1.0 respectively.
[0066] [Parameter Correction Program] As shown in FIG. 3, the parameter correction program is a program that mainly corrects the expression parameters created by the face tracking program. The parameter correction program includes a parameter correction process for correcting at least some of the expression parameters created by the face tracking program based on some of the expression parameters, voice parameters from the voice detection program, input parameters from the input detection program, vital parameters from the vital detection program, and config parameters (set values) set in advance, and a parameter creation process for newly creating expression parameters based on the input parameters. The parameter correction program is executed for each frame at intervals corresponding to the frame rate of the video distributed by the distribution program. For example, when the frame rate is 60 fps, the parameter correction program is executed 60 times per second.
[0067] The config (set value) is a value that can set the amount of change in the facial expression parameters. It is possible to display a config setting screen on the display device 105 and set it using the input device 107. The config (set value) can be set for each type of facial expression parameter. In this embodiment, as shown in FIG. 4, for each of the facial expression parameters related to eye movement, the facial expression parameters related to mouth and jaw movement, the facial expression parameters related to eyebrow movement, the facial expression parameters related to cheek and nose movement, and the facial expression parameters related to tongue movement, numerical values from 0.0 to 1.0 can be individually set, and the set numerical values are used as config parameters. In this embodiment, the larger the numerical value set as the config, the larger the amount of change in the corresponding facial expression parameter, and the larger the movement of the corresponding part.
[0068] FIG. 5 is a flowchart showing the control content of the parameter correction program.
[0069] As shown in FIG. 5, in the parameter correction program, first, various parameters are fetched (Sa1). In the step of Sa1, the facial expression parameters output from the face tracking program, the voice parameters output from the voice detection program, the input parameters output by the input detection program, the vital parameters output by the vital detection program, and the config parameters based on the pre-set config (set value) are fetched.
[0070] Next, referring to the setting data, an emotion value is calculated based on the various parameters fetched in the step of Sa1 (Sa2). The setting data is data set by the correction content setting program, and is data that defines the type of input parameter, the correction content (calculation formula, output condition, etc.) based on the input parameter, and the type of output parameter. The emotion value, as will be described in detail later, identifies a Joy value indicating the degree of joy and an Anger value indicating the degree of anger from among some of the facial expression parameters, and is a value calculated from the AG value calculated from the Joy value and the Anger value, the voice parameters, and the vital parameters, and is a parameter indicating the degree of joy and anger of the subject.
[0071] Next, referring to the setting data, the expression parameters are corrected (Sa3) using the various parameters captured in the step of Sa1 and the emotion value calculated in the step of Sa2. In the step of Sa3, it is sufficient that at least a part of the expression parameters output from the face tracking program are corrected. Also, the expression parameters may be configured to be corrected by parameters including the emotion value, or may be configured to be corrected by parameters not including the emotion value. Further, as the parameters used for correcting the expression parameters, a configuration using only a part of the expression parameters, a configuration using only parameters other than the expression parameters, or a configuration using both a part of the expression parameters and parameters other than the expression parameters may be used. Also, the parameters used for correcting the expression parameters are not limited to the parameters output from the face tracking program or the like, and a configuration may be used in which the parameters corrected by a plurality of parameters are further used for correcting the expression parameters. Regarding the expression parameters to be corrected, it is not limited to the expression parameters output from the face tracking program, and a configuration may be used in which the expression parameters once corrected are further corrected using other parameters.
[0072]
[0073] Next, referring to the setting data, new expression parameters other than the expression parameters output from the face tracking program are created (Sa4) using the various parameters captured in the step of Sa1 and the emotion value calculated in the step of Sa2. The parameters created in the step of Sa4 may be added and output separately from the expression parameters output from the face tracking program, or the expression parameters output from the face tracking program may be switched to the newly created parameters.Next, the facial expression parameters corrected in the step of Sa3 and the facial expression parameters created in the step of Sa4 are output to the image generation program (Sa5). In the step of Sa5, the facial expression parameters output from the face tracking program and used without correction are also output. Also, in the step of Sa5, not only the facial expression parameters, but also parameters such as the emotion value created in the process of correcting or newly creating the facial expression parameters can be output to the image generation program as extended parameters. In this embodiment, the emotion value is output as an extended parameter. The facial expression parameters and extended parameters output to the image generation program in the step of Sa5 can be specified by the correction content setting process.
[0074] [Emotion value calculation process] Next, an example of a method for calculating the emotion value by the parameter correction program will be described. As shown in FIG. 6, in the emotion value calculation process for calculating the emotion value, among the facial expression parameters output from the face tracking program, BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthFunnel (MFu), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), MouthSmileRight (MSmR) are used, the volume (Vo) and pitch (Pi) are used as the voice parameters output from the voice detection program, and the pulse rate (PR) is used as the vital parameter output from the vital detection program.
[0075] As shown in Fig. 12(a), BrowDownLeft (BDoL) is a parameter representing the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter representing the downward movement of the right eyebrow. BrowDownLeft (BDoL) and BrowDownRight (BDoR) are parameters with higher numerical values in the case of an angry expression. Also, as shown in Fig. 12(b), CheekSquintLeft (CSqL) is a parameter representing the upward movement around the left eye and the lower cheek, and CheekSquintRight (CSqR) is a parameter representing the upward movement around the right eye and the lower cheek. CheekSquintLeft (CSqL) and CheekSquintRight (CSqR) are parameters with higher numerical values in the case of a happy expression. Also, as shown in Fig. 12(c), MouthFunnel (MFu) is a parameter representing the contraction of both lips into an open shape. MouthFunnel (MFu) is a parameter with higher numerical values in the case of an angry expression. Also, as shown in Fig. 12(d), MouthStretchLeft (MStL) is a parameter representing the leftward movement of the left mouth corner, and MouthStretchRight (MStR) is a parameter representing the rightward movement of the right mouth corner. MouthStretchLeft (MStL) and MouthStretchRight (MStR) are parameters with higher numerical values in the case of an angry expression. Also, as shown in Fig. 12(e), MouthSmileLeft (MSmL) is a parameter representing the upward movement of the left mouth corner, and MouthSmileRight (MSmR) is a parameter representing the upward movement of the right mouth corner. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters with higher numerical values in the case of a happy expression.
[0076] In the emotion value calculation process, first, (BDoL + BDoR) / 2, (CSqL + CSqR) / 2, (MStL + MStR) / 2, and (MSmL + MSmR) / 2 are performed, and BDo, CSq, MSt, and MSm are calculated as the average values of these expression parameters divided into left and right.
[0077] Next, when the expression is an angry expression, using BDo, MFu, and MSt whose numerical values increase, perform ((BDo - CSq) × 2 + MSt + MFu) / 4 × (-1) to calculate the Anger value indicating the degree of anger. At this time, not only BDo, MFu, and MSt whose numerical values increase when the expression is an angry expression, but also CSq whose numerical value increases when the expression is a happy expression is subtracted from BDo, so that the degree of anger can be calculated more accurately. The Anger value is a numerical value in the range of 0 to -1, and it is shown that the closer the numerical value is to -1, the greater the degree of anger.
[0078] Next, using CSq and MSm whose numerical values increase when the expression is a happy expression, perform (MSm + CSq) / 2 to calculate the Joy value indicating the degree of joy. The Joy value is a numerical value in the range of 0 to +1, and it is shown that the closer the numerical value is to +1, the greater the degree of joy.
[0079] Next, using the Anger value and the Joy value, perform Anger + Joy to calculate the emotion value AG. The emotion value AG is a numerical value in the range of 1 to -1.
[0080] Next, convert the volume (Vo), pitch (Pi) which are voice parameters, and the pulse rate (PR) which is a vital parameter into Vo1, Pi1, PR1 indicating values of 1 or 0. The volume (Vo), pitch (Pi), and pulse rate (PR) are all numerical values in the range of 0 to 1. If the numerical values of the volume (Vo) and pitch (Pi) are 0.8 or more, they are converted to 1, and if they are less than 0.8, they are converted to 0. Also, if the numerical value of the pulse rate (PR) is 0.6 or more, it is converted to 1, and if it is less than 0.6, it is converted to 0.
[0081] Next, when the emotion value AG exceeds 0, that is, when the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, perform AG×0.8+(Vo1+Pi1+PR1) / 3×0.2, and output the calculated result as the EmoJ indicating the degree of joy as the emotion value. Thus, when the volume (Vo), pitch (Pi), and pulse rate (PR) are above a certain value and the emotional excitement can be read, a larger value, that is, an EmoJ with a greater degree of joy, will be output than when the volume (Vo), pitch (Pi), and pulse rate (PR) are below the certain value.
[0082] On the other hand, when the emotion value AG is less than 0, that is, when the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, perform AG×0.8-(Vo1+Pi1+PR1) / 3×0.2, and output the calculated result as the EmoA indicating the degree of anger as the emotion value. Thus, when the volume (Vo), pitch (Pi), and pulse rate (PR) are above a certain value and the emotional excitement can be read, a smaller value, that is, an EmoA with a greater degree of anger, will be output than when the volume (Vo), pitch (Pi), and pulse rate (PR) are below the certain value.
[0083] In this way, in the emotion value calculation process, the Anger value indicating the degree of anger is identified using the expression parameter with a higher value in the case of an angry expression, the Joy value indicating the degree of joy is identified using the parameter with a higher value in the case of a happy expression, when the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, the emotion value EmoJ indicating the degree of joy is identified, and when the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, the emotion value EmoA indicating the degree of anger is identified.
[0084] Also, when the volume (Vo), pitch (Pi), and pulse rate (PR) are above a certain value, the emotion value EmoJ indicating the degree of joy and the emotion value EmoA indicating the degree of anger are such that the emotion value indicating a greater degree is identified.
[0085] In addition, in this embodiment, in addition to the expression parameters, voice parameters and vital parameters are used to identify the emotion value. However, a configuration that identifies the emotion value only with the expression parameters may be used, or a configuration that identifies the emotion value only with parameters other than the expression parameters such as voice parameters and vital parameters may be used. Further, a configuration may be adopted in which parameters other than the expression parameters, voice parameters, and vital parameters are used together with these parameters to identify the emotion value, or a configuration may be adopted in which the emotion value is identified only with parameters other than the expression parameters, voice parameters, and vital parameters. For example, when a command input for gradually specifying emotions such as joy, anger, sorrow, and pleasure is given by operating the input device 107, a configuration may be adopted in which the emotion value is identified using an input parameter that identifies the corresponding command input.
[0086] [Parameter correction process] Next, an example of a method for correcting expression parameters by a parameter correction program will be described.
[0087] FIG. 7 is a diagram showing the content of the MouthSmile correction process for correcting the expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) based on the emotion value.
[0088] In the MouthSmile correction process, the emotion value EmoJ indicating the degree of joy is used to correct MouthSmileLeft (MSmL) and MouthSmileRight (MSmR). Specifically, MSmL + EmoJ × 0.5 and MSmR + EmoJ × 0.5 are respectively performed, and the corrected expression parameters MSmLN and MSmRN corresponding to MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are respectively output. Note that when the calculated result exceeds 1.0, 1.0 is output.
[0089] MSmLN and MSmRN have a correction amount that becomes larger as the emotion value EmoJ indicating the degree of joy approaches +1, and a correction amount that becomes smaller as the emotion value EmoJ indicating the degree of joy approaches 0. MouthSmileLeft (MSmL) is a parameter representing the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter representing the upward movement of the right corner of the mouth. As the emotion value EmoJ approaches +1, the corners of the mouth become more upward. Therefore, as shown in Fig. 13(a), compared with the case of the emotion value 0, for example, when the emotion value EmoJ is +1, the corners of the mouth are upward, and as the emotion value EmoJ approaches +1, the expression of the output face image can be made closer to a happy expression.
[0090] Fig. 8 is a diagram showing the content of the MouthFrown correction process for correcting the expression parameters MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) based on the emotion value.
[0091] In the MouthFrown correction process, the emotion value EmoA indicating the degree of anger is used to correct MouthFrownLeft (MFrL) and MouthFrownRight (MFrR). Specifically, MFrL + EmoA × (-1) × 0.5 and MFrR + EmoA × (-1) × 0.5 are performed respectively, and the corrected expression parameters MFrLN and MFrRN corresponding to MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) are output respectively. When the calculated result exceeds 1.0, 1.0 is output.
[0092] MFrLN and MFrRN become correction amounts that are larger as the emotion value EmoA indicating the degree of anger approaches -1, and become smaller as the emotion value EmoA indicating the degree of anger approaches 0. MouthFrownLeft (MFrL) is a parameter representing the downward movement of the left corner of the mouth, and MouthFrownRight (MFrR) is a parameter representing the downward movement of the right corner of the mouth. As the emotion value EmoA approaches -1, the corners of the mouth turn downward. Therefore, as shown in Fig. 13(a), compared with the case of the emotion value 0, for example, when the emotion value EmoA is -1, the corners of the mouth turn downward, and as the emotion value EmoA approaches -1, the expression of the output face image can be made closer to an angry expression.
[0093] Fig. 9 is a diagram showing the content of the BrowDown correction process for correcting the expression parameters BrowDownLeft (BDoL) and BrowDownRight (BDoR) based on the emotion value.
[0094] In the BrowDown correction process, the emotion value EmoA indicating the degree of anger is used to correct BrowDownLeft (BDoL) and BrowDownRight (BDoR). Specifically, BDoL + EmoA × (-1) × 0.5 and BDoR + EmoA × (-1) × 0.5 are respectively performed, and the corrected expression parameters BDoLN and BDoRN corresponding to BrowDownLeft (BDoL) and BrowDownRight (BDoR) are respectively output. If the calculated result exceeds 1.0, 1.0 is output.
[0095] BDoLN and BDoRN become correction amounts that are larger as the emotion value EmoA indicating the degree of anger approaches -1, and smaller as the emotion value EmoA indicating the degree of anger approaches 0. BrowDownLeft (BDoL) is a parameter representing the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter representing the downward movement of the right eyebrow. As the emotion value EmoA approaches -1, the eyebrows move downward. Therefore, as shown in Fig. 13(b), compared with the case of the emotion value 0, for example, when the emotion value EmoA is -1, the eyebrows are in a lower position, and as the emotion value EmoA approaches -1, the expression of the output face image can be made closer to an angry expression.
[0096] Fig. 10 is a diagram showing the content of the EyeBlink switching process that switches the expression parameter corresponding to the expression parameter EyeBlinkLeft (EBL) to one of the expression parameters EyeBlinkLeft (EBL) and EyeSmileLeft (ESL) based on the expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR).
[0097] EyeBlinkLeft (EBL) is a parameter representing the degree of closing of the left eyelid, and EyeSmileLeft (ESL) is a parameter representing the degree of closing of the left eyelid with the outer corner of the eye lowered. It is an expression parameter used in place of EyeBlinkLeft (EBL).
[0098] In the EyeBlink switching process, MSm is calculated as the average value of these left and right parameters based on MouthSmileLeft (MSmL) and MouthSmileRight (MSmR). If MSm is a value of 0.5 or more, the value of EyeBlinkLeft (EBL) is output as the value of EyeSmileLeft (ESL) instead of EyeBlinkLeft (EBL). On the other hand, if MSm is a value less than 0.5, the value of EyeBlinkLeft (EBL) is output as the value of EyeBlinkLeft (EBL) as it is.
[0099] As shown in FIG. 12(e), MouthSmileLeft (MSmL) is a parameter representing the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter representing the upward movement of the right corner of the mouth. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters with higher numerical values in the case of a smiling expression. Therefore, in the EyeBlink switching process, when MSm is less than a certain value, as shown in FIG. 14(a), it becomes an expression of the normal closing state of the eyelids by EyeBlinkLeft (EBL), while when MSm is greater than or equal to a certain value, as shown in FIG. 14(b), it becomes an expression of the closing state of the eyelids with the outer corners of the eyes lowered by EyeSmileLeft (ESL).
[0100] FIG. 11 is a diagram showing the content of the config correction process for correcting expression parameters based on config parameters.
[0101] The config (set value), as shown in FIG. 4, is set for each expression parameter related to eye movement, expression parameter related to mouth and jaw movement, expression parameter related to eyebrow movement, expression parameter related to cheek and nose movement, and expression parameter related to tongue movement. In the config correction process, the numerical values of the corresponding config parameters are corrected using each config parameter. Specifically, (expression parameter + corresponding config parameter) / 2 is performed, and the corrected expression parameter N is output.
[0102] The config parameter is a numerical value between 0.0 and 1.0. The larger the numerical value of the config parameter, the larger the numerical value of the corresponding expression parameter N, and the smaller the numerical value of the config parameter, the smaller the numerical value of the corresponding expression parameter N. Therefore, the amount of change in the part corresponding to the expression parameter can be adjusted according to the preset numerical value of the config parameter.
[0103] Thus, in this embodiment, other expression parameters are corrected based on some of the expression parameters detected from the face image data captured by the camera 108, and a face image with an expression corresponding to the corrected expression parameter N is generated. Therefore, the face image detected by face tracking can be output as a richly varying image.
[0104] Further, in this embodiment, in addition to some of the expression parameters detected from the face image data captured by the camera 108, a face image with an expression corresponding to the expression parameter N corrected based on parameters other than the expression parameters is generated. Therefore, the face image detected by face tracking can be output as an even more richly varying image.
[0105] Note that in this embodiment, the expression parameters are corrected by both some of the expression parameters and parameters other than the expression parameters. However, even when configured to generate a face image with an expression corresponding to the expression parameter N corrected based only on some of the expression parameters, or even when configured to generate a face image with an expression corresponding to the expression parameter N corrected based only on parameters other than the expression parameters, in such a configuration as well, the face image detected by face tracking can be output as a richly varying image.
[0106] In this embodiment, as parameters other than the expression parameters, it includes voice parameters based on the voice input from the microphone 109. For example, a face image with an expression corresponding to the volume, pitch, etc. of the voice input from the microphone 109 can be generated.
[0107] Further, in this embodiment, as parameters other than the expression parameters, it includes input parameters based on the command operation input by the input device 107. For example, a face image with an expression corresponding to the command input, etc. input by the input device 107 can be generated.
[0108] In addition, in this embodiment, as parameters other than the expression parameters, it includes vital parameters based on vital data input by the vital measurement device 110. For example, a face image of an expression corresponding to the pulse, blood pressure, body temperature, etc. input by the vital measurement device 110 can be generated.
[0109] In addition, in this embodiment, as parameters other than the expression parameters, it includes configuration parameters based on a configuration (set value) set in advance by a configuration setting screen. For example, a face image of an expression can be generated with the amount of change adjusted by the configuration.
[0110] Note that in this embodiment, as parameters other than the expression parameters, it has a configuration including voice parameters, input parameters, vital parameters, and configuration parameters. However, it may also have a configuration including only some of these, or a configuration including other parameters, for example, parameters specified from the background image of the subject, the background image of the avatar to be distributed simultaneously, the temperature and humidity of the shooting location, the brightness and darkness of the shooting location, etc.
[0111] In this embodiment, an emotion value is specified based on some of the expression parameters detected from the face image data captured by the camera 108, and the expression parameters are corrected based on the specified emotion value. Since a face image of an expression corresponding to the corrected expression parameter N is generated, the emotion specified from the expression of the face image detected by face tracking can be reflected in the generated face image.
[0112] In addition, in this embodiment, in addition to some of the expression parameters detected from the face image data captured by the camera 108, an emotion value is specified based on parameters other than the expression parameters, and the emotion to be reflected in the generated face image can be specified more accurately.
[0113] In this embodiment, the emotional value is specified by both some of the expression parameters and the parameters other than the expression parameters. However, even when the emotional value is specified based only on some of the expression parameters or only on the parameters other than the expression parameters, in such a configuration, the generated emotion can be reflected in the face image.
[0114] Also, in this embodiment, based on some of the expression parameters, a Joy value indicating joy and an Anger value indicating anger are specified. When the Joy value is greater than the Anger value, an emotional value EmoJ indicating joy is specified, and when the Anger value is greater than the Joy value, an emotional value EmoA indicating anger is specified. Based on the emotional value reflecting both joy and anger specified from some of the expression parameters and the parameters other than the expression parameters, a face image of an expression can be generated.
[0115] Also, in this embodiment, MouthStretch and MouthSmile, which indicate the degree of movement of the corners of the mouth among the expression parameters, are used to specify the Joy value and the Anger value, and joy and anger emotions can be specified from the degree of movement of the corners of the mouth where emotions are likely to be reflected.
[0116] In this embodiment, a Joy value indicating joy and an Anger value indicating anger are specified and reflected in the expression parameters. However, it is also possible to specify a value indicating sadness and a value indicating fun, and configure to reflect the emotional values of joy, anger, sadness, and fun in the expression parameters respectively, or to specify only some of these emotional values and reflect the specified emotional values in the expression parameters.
[0117] Also, in this embodiment, regardless of the magnitude of the emotional value, the corresponding expression parameters are corrected. However, it is also possible to configure to correct the corresponding expression parameters when the emotional value exceeds a certain threshold. By adopting such a configuration, it is possible to add a sense of rhythm to the change in the expression of the generated face image.
[0118] In this embodiment, the emotional value is specified by some of the expression parameters and the parameters other than the expression parameters, and the expression parameters are corrected based on the specified emotional value. However, it is also possible to directly correct the expression parameters using some of the expression parameters that change according to emotions and the parameters other than the expression parameters. Even in such a configuration, when the parameter used for correction exceeds the threshold value, the corresponding expression parameter is corrected, so that the change in the expression of the generated face image can be given a sense of rhythm.
[0119] [Correction Content Setting Program] The correction content setting program is a program that performs a process of setting the setting data used by the parameter correction program. In addition, the correction content setting program also performs a process of designating the expression parameters and the extension parameters output to the image generation program in addition to the process of setting the setting data. Note that a detailed description of the process of designating the expression parameters and the extension parameters output to the image generation program is omitted. In this embodiment, the computer terminal 1 that executes the parameter correction program and the image generation program is configured to be able to execute the correction content setting program. However, the function of executing the correction content setting program may be installed in a computer terminal different from the computer terminal 1, and the setting data set in the different computer terminal may be applied as the setting data of the parameter correction program of the computer terminal 1.
[0120] FIG. 15 is a flowchart showing the control content of the correction content setting program.
[0121] As shown in FIG. 15, in the correction content setting program, first, the parameter to be corrected is designated (Sb1). In the step Sb1, for example, an input node is arranged on the setting screen, and the type of the parameter to be corrected is set in the input node.
[0122] Next, specify the parameters to be used for correction (Sb2). In the step of Sb2, for example, in the same manner as in the step of Sb1, input nodes are arranged on the setting screen, and the types of parameters to be used for correction are set in the input nodes.
[0123] Next, set the correction details such as calculation formulas and output conditions using multiple parameters (Sb3). In the step of Sb3, for example, setting nodes and output nodes are arranged on the setting screen, and the input nodes in which the types of parameters to be corrected are set and the input nodes in which the types of parameters to be used for correction are set are connected to the input part of the setting node, and the output node in which the types of parameters after correction are set is connected to the output part of the setting node. Further, the correction details such as calculation formulas and output conditions using the parameters input to the setting node are set.
[0124] Next, based on the setting contents of the input nodes, setting nodes, and output nodes, create setting data that defines the types of parameters to be input, the correction details (calculation formulas, output conditions, etc.) based on the input parameters, and the types of parameters to be output, and output it to the parameter correction program (Sb4).
[0125] Next, the specific setting procedure of the setting data by the correction detail setting program will be described. Here, the setting procedures of the setting data used for the emotion value setting process, the MouthSmile correction process, the MouthFrown correction process, and the BrowDown correction process that are corrected using the emotion value will be described.
[0126] Figs. 16 to 21 are diagrams for explaining the setting procedure of the setting data used for the emotion value setting process.
[0127] The setting data used for the emotional value setting process consists of a plurality of setting data. First, setting data for calculating BDo, CSq, MSt, and MSm, which are the averages of the left and right values of BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR), is set.
[0128] For example, when setting the setting data for calculating BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR), as shown in Fig. 16, two input nodes, one setting node, and one output node are arranged on the setting screen. The number and positions of these nodes can be arbitrarily specified.
[0129] Next, BDoL and BDoR are input to the two input nodes respectively. Also, two input parts A and B are set in the setting node, the input node of BDoL is connected to the input part A of the setting node, and the input node of BDoR is connected to the input part B of the setting node. Also, the BDo output to the output node is input, and the output part of the setting node and the output node are connected. Further, (A + B) / 2 is input as the calculation formula for taking the average of the parameters input from the input parts A and B in the setting node. By performing the determination operation in this state, the setting data for calculating BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR), is created.
[0130] The procedure for creating the setting data for calculating CSq, MSt, and MSm, which are the averages of the left and right values of CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR), is the same.
[0131] Next, set the setting data for calculating the Anger value indicating the degree of anger using BDo, CSq, MFu, and MSt.
[0132] Here, as shown in FIG. 17, arrange four input nodes, one setting node, and one output node on the setting screen.
[0133] Next, input BDo, CSq, MFu, and MSt into the four input nodes respectively. Also, set four input parts A to D in the setting node, connect the input node of BDo to the input part A of the setting node, connect the input node of CSq to the input part B of the setting node, connect the input node of MSt to the input part C of the setting node, and connect the input node of MFu to the input part D of the setting node. Further, input the Anger output to the output node and connect the output part of the setting node and the output node. Furthermore, input ((A - B) * 2 + C + D) / 4 * (-1) as the calculation formula for calculating the Anger value based on the parameters input from the input parts A to D to the setting node. By performing the determination operation in this state, the setting data for calculating the Anger value indicating the degree of anger is created.
[0134] Next, set the setting data for calculating the Joy value indicating the degree of joy using MSm and CSq.
[0135] Here, as shown in FIG. 18, arrange two input nodes, one setting node, and one output node on the setting screen.
[0136] Next, input MSm and CSq into the two input nodes respectively. Also, set two input parts A and B in the setting node, connect the input node of MSm to the input part A of the setting node, and connect the input node of CSq to the input part B of the setting node. Further, input Joy output to the output node, and connect the output part of the setting node and the output node. Additionally, input (A + B) / 2 as the calculation formula for calculating the Joy value based on the parameters input from the input parts A and B into the setting node. By performing the determination operation in this state, setting data for calculating the Joy value indicating the degree of joy is created.
[0137] Next, set the setting data for calculating the emotion value AG using the Anger value and the Joy value, outputting AGJ where the calculated AG exceeds 0, and AGA where the calculated AG is less than 0.
[0138] Here, as shown in FIG. 19, arrange two input nodes, two setting nodes, and two output nodes on the setting screen.
[0139] Next, input Anger and Joy into the two input nodes respectively. Also, set two input parts A and B in one of the setting nodes, connect the input node of Anger to the input part A of one of the setting nodes, connect the input node of Joy to the input part B of one of the setting nodes, and connect the output part to the input part of the other setting node. Further, input A + B as the calculation formula for calculating the AG value based on the parameters input from A and B into the input part of one of the setting nodes. Next, set two output parts A and B in the other setting node, input AGJ and AGA into the two output nodes respectively, connect the output part A of the other setting node and the output node of AGJ, and connect the output part B of the other setting node and the output node of AGA. Additionally, input In>0→A In<0→B as the conditional formula for distributing the parameters according to the AG value input from the input part into the other setting node. By performing the determination operation in this state, setting data for outputting AGJ where AG exceeds 0 and AGA where AG is less than 0 is created.
[0140] Next, set the setting data for converting the volume (Vo), pitch (Pi), and pulse rate (PR) into Vo1, Pi1, and PR1 that indicate 1 or 0.
[0141] For example, when converting the volume (Vo) to Vo1, as shown in FIG. 20, one input node, one setting node, and one output node are arranged on the setting screen.
[0142] Next, input Vo to the input node. Also, connect the input node of Vo to the input part of the setting node. Further, input Vo1 output to the output node and connect the output part of the setting node and the output node. Furthermore, input In≧0.8→1 In<0.8→0 as a conditional expression for converting parameters according to the Vo value input from the input part to the setting node. By performing the determination operation in this state, the setting data for converting the volume (Vo) to Vo1 is created.
[0143] The procedure for creating the setting data for converting the pitch (Pi) and pulse rate (PR) is the same. Note that for the conditional expression for converting the pulse rate (PR) to PR1, input In≧0.6→1 In<0.6→0.
[0144] Next, set the setting data for calculating the emotion values EmoJ and EmoA using AGJ, AGA, Vo1, Pi1, and PR1.
[0145] Here, as shown in FIG. 21, five input nodes, two setting nodes, and two output nodes are arranged on the setting screen.
[0146] Next, input AGJ, AGA, Vo1, Pi1, and PR1 to the five input nodes respectively. Also, set four input parts A to D to one setting node respectively, connect the input node of AGJ to input part A of one setting node, connect the input node of Vo1 to input part B of one setting node, connect the input node of Pi1 to input part C of one setting node, and connect the input node of PR1 to input part D of one setting node. Further, set four input parts A to D to the other setting node respectively, connect the input node of AGA to input part A of the other setting node, connect the input node of Vo1 to input part B of the other setting node, connect the input node of Pi1 to input part C of the other setting node, and connect the input node of PR1 to input part D of the other setting node. Also, set EmoJ output to one output node, connect the output part of one setting node and one output node, set EmoA output to the other output node, and connect the output part of the other setting node and the other output node. Furthermore, input A*0.8+(B+C+D) / 3*0.2 as the calculation formula for correcting the AGJ value with the parameters input from input parts A to D to one setting node, and input A*0.8-(B+C+D) / 3*0.2 as the calculation formula for correcting the AGA value with the parameters input from input parts A to D to the other setting node. By performing the determination operation in this state, setting data for calculating the emotional values EmoJ and EmoA using AGJ, AGA, Vo1, Pi1, and PR1 is created.
[0147] In this way, by setting all the setting data shown in FIGS. 16 to 21 by the correction content setting program, the parameter correction program can refer to these setting data and calculate the emotional values EmoJ and EmoA.
[0148] FIG. 22 is a diagram for explaining the setting procedure of the setting data used for the MouthSmile correction process.
[0149] When creating the setting data used for the MouthSmile correction process, as shown in FIG. 22, three input nodes, two setting nodes, and two output nodes are arranged on the setting screen.
[0150] Next, input the EmoJ used for correction, the MSmL to be corrected, and MSmR into the three input nodes respectively. Also, set two input parts A and B in one of the setting nodes, connect the input node of EmoJ to the input part A of one of the setting nodes, and connect the input node of MSmL to the input part B of one of the setting nodes. Also, set two input parts A and B in the other setting node, connect the input node of EmoJ to the input part A of the other setting node, and connect the input node of MSmR to the input part B of the other setting node. Also, set MSmLN output to one of the output nodes, connect the output part of one of the setting nodes and one of the output nodes, set MSmRN output to the other output node, and connect the output part of the other setting node and the other output node. Further, input B + A * 0.5 as the calculation formula for correcting MSmL input from the input part B by EmoJ input from the input part A in one of the setting nodes, and input B + A * 0.5 as the calculation formula for correcting MSmR input from the input part B by EmoJ input from the input part A in the other setting node. By performing the determination operation in this state, the setting data for correcting the expression parameters MSmL and MSmR using the emotion value EmoJ is created.
[0151] FIG. 23 is a diagram for explaining the setting procedure of the setting data used for the MouthFrown correction process and the BrowDown correction process.
[0152] When creating the setting data used for the MouthFrown correction process and the BrowDown correction process, as shown in FIG. 23, five input nodes, four setting nodes, and four output nodes are arranged on the setting screen.
[0153] Next, input EmoA used for correction, MFrL to be corrected, MFrR, BDoL, and BDoR to the five input nodes respectively. Also, set two input parts A and B to the first setting node, connect the input node of EmoA to input part A of the first setting node, and connect the input node of MFrL to input part B of the first setting node. Also, set two input parts A and B to the second setting node, connect the input node of EmoA to input part A of the second setting node, and connect the input node of MFrR to input part B of the second setting node. Also, set two input parts A and B to the third setting node, connect the input node of EmoA to input part A of the first setting node, and connect the input node of BDoL to input part B of the third setting node. Also, set two input parts A and B to the fourth setting node, connect the input node of EmoA to input part A of the fourth setting node, and connect the input node of BDoR to input part B of the fourth setting node. Also, set MFrLN output to the first output node, connect the output part of the first setting node and the first output node, set MFrRN output to the second output node, and connect the output part of the second setting node and the second output node. Also, set BDoLN output to the third output node, connect the output part of the third setting node and the third output node, set BDoRN output to the fourth output node, and connect the output part of the fourth setting node and the fourth output node. Further, input B + A * (-1) * 0.5 as the calculation formula for correcting MFrL input from input part B by EmoA input from input part A to the first setting node, input B + A * (-1) * 0.5 as the calculation formula for correcting MFrR input from input part B by EmoA input from input part A to the second setting node, input B + A * (-1) * 0.5 as the calculation formula for correcting BDoL input from input part B by EmoA input from input part A to the third setting node, and input B + A * (-1) * 0.5 as the calculation formula for correcting BDoR input from input part B by EmoA input from input part A to the fourth setting node. By performing the determination operation in this state, setting data for correcting the expression parameters MFrL, MFrR, BDoL, and BDoR using the emotion value EmoA is created.
[0154] In this way, in the correction content setting program, input nodes, setting nodes, and output nodes are arranged on the setting screen. The parameters to be corrected and the types of parameters used for correction are input to the input nodes and connected to the setting nodes. Also, the types of parameters to be output to the output nodes are input and connected to the setting nodes. The correction content is set by using calculation formulas or conditional expressions with the parameters input from the input nodes at the setting nodes, so that the setting data used in the parameter correction program can be easily set.
[0155] Also, in the correction content setting program, it is possible to read the created setting data. By reading the setting data, the input nodes, setting nodes, output nodes based on the setting data, and their setting contents are displayed. It is possible to change the connection between the nodes, change the types of parameters set in the input nodes and output nodes, change the calculation formulas and conditional expressions set in the setting nodes, and perform overwriting and saving or creating new setting data. It is possible to easily edit the existing setting data and easily create new setting data based on the existing setting data.
[0156] [Image Generation Program] As shown in FIG. 3, the image generation program is a program that executes a process of generating an avatar image based on the expression parameters N and extension parameters from the parameter correction program, the motion parameters from the motion tracking program, the voice parameters from the voice parameter, and the input parameters from the input detection program, and outputs the generated avatar image to the distribution program. It includes a face image generation process for generating a face image, a face image correction process for correcting the face image, a clothing image generation process for generating a clothing image, and a decoration image generation process for generating a decoration image. Similar to the parameter correction program, the image generation program is executed every frame at intervals corresponding to the frame rate of the video distributed by the distribution program. For example, when the frame rate is 60 fps, the image generation program is executed 60 times per second.
[0157] Figure 24 is a flowchart showing the control content of the image generation program.
[0158] As shown in Figure 24, in the image generation program, first, various parameters are captured (Sc1). In the step of Sc1, the facial expression parameter N and the extended parameter output from the parameter correction program, the motion parameter from the motion tracking program, the voice parameter from the voice parameter, and the input parameter from the input detection program are captured.
[0159] Next, a 3D model of the face is created based on the facial expression parameter N captured in the step of Sc1 (Sc2). In the step of Sc2, it is created by blending (blend shape) using a plurality of 3D models (face data) preset corresponding to the facial expression parameter N according to the value of the corresponding facial expression parameter N. Also, for some parts, it is created by setting the position and angle of the joints between a plurality of bones according to the value of the facial expression parameter N. Further, the 3D model created by the blend shape and the 3D model created by the bones may be blended and complemented.
[0160] Next, based on the motion parameters captured in the step of Sc1, determine how many degrees the shooting direction of camera 108 is inclined horizontally and vertically with respect to the front of the face, and correct the outline, part positions, and shapes of the 3D face model created in the step of Sc2 with correction amounts corresponding to the determined shooting directions. As a result, for example, as shown in FIGS. 25(a) and 25(b), instead of using the same 3D model whether shooting from the front of the face or from an oblique direction, a 3D model with corrected face outline, positions and shapes of the mouth, nose, and eyes according to the shooting direction will be used. The correction amounts corresponding to the shooting directions are calculated based on the maximum correction amounts in the left-right and up-down directions predetermined for each outline and part to be corrected, and the horizontal and vertical angles of the shooting direction. As shown in FIG. 26(a), as the horizontal or vertical angle of the shooting direction increases, the correction amounts in the left-right or up-down directions also gradually increase. Also, intermediate correction amounts corresponding to intermediate angles smaller than the maximum angle may be set. In this case, as shown in FIG. 26(b), by gradually increasing the correction amounts in the left-right or up-down directions so that they form a curve passing through the intermediate correction amounts at the intermediate angles, the correction amounts will not change abruptly due to changes in the shooting angle.
[0161] Next, based on the motion parameters captured in the step of Sc1, create a 3D model of the body, hands (fingers) (Sc4). In the step of Sc4, using a plurality of 3D models (body data) preset corresponding to the motion parameters, create a 3D model of the body, hands (fingers) in the corresponding pose.
[0162] Next, based on the motion parameters captured in the step of Sc1, create a 3D model of the clothing (including accessories such as hairstyles and ornaments) (Sc5). In the step of Sc5, using a plurality of 3D models (clothing data) preset corresponding to the motion parameters, create a 3D model of the clothing that fits the body, hands (fingers) in the corresponding pose.
[0163] Next, it is determined whether the 3D model of the clothing created in the step of Sc5 overlaps with the prohibited area (Sc6). The prohibited area is an area including the overlapping area where the clothing and the face overlap and the facial expression cannot be visually recognized, and also includes the peripheral area thereof, and includes areas other than the overlapping area between the clothing and the face.
[0164] If it is determined in the step of Sc6 that the 3D model of the clothing does not overlap with the prohibited area, the process proceeds to the step of Sc9. If it is determined that the 3D model of the clothing overlaps with the prohibited range, the amount of deformation or movement of the clothing is calculated according to the distance from the 3D model of the clothing to the overlapping area between the clothing and the face (Sc7), and the shape or position of the corresponding part of the clothing is corrected according to the amount of deformation or movement calculated in the step of Sc7 (Sc8).
[0165] Thereby, for example, as shown in FIG. 27(a), when the shooting direction is the front of the face, even if the clothing and the face do not overlap, as shown in FIG. 27(b), when the shooting direction changes and a part of the clothing overlaps with the face and the expression is hidden, a part of the clothing is deformed and the expression is not hidden. Also, as shown in FIG. 28(a), when the shooting direction is the front of the face, even if the accessory and the face do not overlap, as shown in FIG. 28(b), when the shooting direction changes and the accessory overlaps with the face and the expression is hidden, the accessory moves and the expression is not hidden.
[0166] Further, the amount of deformation or movement of the clothing after the 3D model of the clothing overlaps with the prohibited range is set to gradually increase as it approaches the overlapping range of the face and the clothing, as shown in FIG. 29. Furthermore, the closer it is to the overlapping range of the face and the clothing, the greater the increase amount of the amount of deformation or movement of the clothing is set. For this reason, even when the 3D model of the clothing overlaps with the prohibited range, when it is far from the overlapping range of the face and the clothing, the amount of deformation or movement is suppressed to be small, and as it approaches the overlapping range of the face and the clothing, the amount of deformation or movement of the clothing is increased to prevent the overlapping of the face and the clothing.
[0167] Next, it is determined whether or not an additional condition for adding a decorative image is satisfied based on the extended parameter (emotion value) and voice parameter captured in the step of Sc1 (Sc9). If it is determined in the step of Sc9 that the additional condition is not satisfied, the process proceeds to the step of Sc11. If it is determined that the additional condition is satisfied, a decorative image corresponding to the satisfied additional condition is set (Sc10). Note that the decorative image set in the step of Sc10 is retained for a certain period and cleared after the elapse of the certain period. For this reason, after the additional condition is satisfied, the decorative image is maintained and displayed for a certain period. If a new additional condition is satisfied, it will be overwritten.
[0168] The additional conditions and the decorative images corresponding to the additional conditions are set in advance. For example, as shown in FIG. 30, the additional condition is satisfied when the magnitude of the emotion value, a command input for designating a decorative image by the input device 107, the volume, or the pitch is equal to or greater than a specified value. In this embodiment, when the emotion value Emo is 0.7 or greater or the "joy (small)" command is input, a decorative image of a cracker is set. When the emotion value Emo is 0.9 or greater or the "joy (large)" command is input, a decorative image of a heart mark in addition to the cracker is set. When the emotion value Emo is -0.8 or less or the "anger" command is input, a decorative image of an anger mark is set. When the volume (Vo) or the pitch (Pi) is 0.8 or greater, a decorative image of a speaker mark is set.
[0169] Therefore, for example, when the emotion value (Emo) is 0.9 or more, or when a "big joy" command is input and a happy situation is identified, as shown in FIG. 31, decorative images of crackers and heart marks are displayed around the avatar image for a certain period of time. When the emotion value (Emo) is less than -0.8 or when an anger command is input and an angry situation is identified, as shown in FIG. 32, a decorative image of an anger mark is displayed around the avatar image for a certain period of time. Also, when the volume (Vo) or pitch (Pi) is 0.8 or more and a situation where the user is speaking in a loud or high voice is identified, as shown in FIG. 32, a decorative image of a speaker mark is displayed.
[0170] Next, the 3D model of the face created in step Sc2 and corrected in step Sc3, the 3D model of the body and hands (fingers) created in step Sc4, the 3D model of the clothing created in step Sc5, the decorative images set in Sc10, etc. are arranged to generate avatar image data, which is then output to the distribution program (Sc11).
[0171] Thus, in this embodiment, decorative images other than the face image are generated based on some of the expression parameters detected from the face image data captured by the camera 108, and the face image detected by face tracking can be output as a richly changing image.
[0172] Also, in this embodiment, an emotion value is identified based on some of the expression parameters detected from the face image data captured by the camera 108, and decorative images other than the face image are generated based on the identified emotion value, so that the decorative images can be displayed reflecting the emotion identified from the expression of the face image detected by face tracking.
[0173] Furthermore, in this embodiment, in addition to some of the expression parameters detected from the face image data captured by the camera 108, an emotion value is identified based on parameters other than the expression parameters, so that the decorative images can be displayed reflecting a more accurate emotion.
[0174] In this embodiment, the emotional value is specified by both some of the facial expression parameters and the parameters other than the facial expression parameters. However, even when the emotional value is specified based on only some of the facial expression parameters, or even when the emotional value is specified based on only the parameters other than the facial expression parameters, in such a configuration, the decorative image can be displayed reflecting the specified emotion.
[0175] Also, in this embodiment, when the emotional value exceeds a certain threshold, the corresponding decorative image is generated. By adopting such a configuration, the decorative image can be displayed at an appropriate frequency.
[0176] In this embodiment, the decorative image other than the face image is generated based on some of the facial expression parameters detected from the face image data captured by the camera 108. However, it may also be configured to change the design and color of the clothing, the composition and color tone of the background image, the shape, size, and color of the accessories, etc. based on some of the parameters detected from the face image data and the emotional value specified from these parameters. In such a configuration, changes can be made not only to the face image but also to other parts by reflecting the parameters specified from the expression of the face image detected by face tracking.
[0177] In this embodiment, when a part of the clothing becomes a position where the expression of the face image cannot be visually recognized according to the shooting direction of the camera 108, a part of the clothing is changed to a shape where the expression of the face image can be visually recognized, and it is possible to prevent the expression of the face image from being hidden by a part of the clothing when the shooting direction of the subject by the camera 108 changes.
[0178] Also, in this embodiment, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position (prohibited area) where the expression of the face image cannot be visually recognized, a part of the clothing is deformed toward a shape where the expression of the face image can be visually recognized, and a part of the clothing can be deformed toward a shape where the expression of the face image can be visually recognized before the expression of the face image is hidden.
[0179] In addition, in this embodiment, according to the distance to a position where the expression of the face image cannot be visually recognized, a part of the clothing is gradually deformed so as to be oriented in a shape where the expression of the face image can be visually recognized, and a part of the clothing can be naturally deformed so as to be oriented in a shape where the expression of the face image can be visually recognized.
[0180] Note that, in this embodiment, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position where the expression of the face image cannot be visually recognized (prohibited area), a part of the clothing is configured to be deformed so as to be oriented in a shape where the expression of the face image can be visually recognized. However, it may be configured such that a part of the clothing is deformed into a shape where the expression of the face image can be visually recognized at the timing when the position where the expression of the face image cannot be visually recognized is reached. Even with such a configuration, it is possible to prevent the photographing direction of the subject by the camera 108 from changing and the expression of the face image from being hidden by a part of the clothing.
[0181] In addition, in this embodiment, according to the distance to a position where the expression of the face image cannot be visually recognized, a part of the clothing is gradually deformed so as to be oriented in a shape where the expression of the face image can be visually recognized. However, it may also be configured such that a part of the clothing is deformed so as to be oriented in a shape where the expression of the face image can be visually recognized at the timing when the position where the expression of the face image cannot be visually recognized (prohibited area) is reached.
[0182] In this embodiment, when the accessory becomes a position where the expression of the face image cannot be visually recognized according to the photographing direction of the camera 108, the accessory is moved to a position where the expression of the face image can be visually recognized, and it is possible to prevent the photographing direction of the subject by the camera 108 from changing and the expression of the face image from being hidden by the accessory.
[0183] In addition, in this embodiment, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position where the expression of the face image cannot be visually recognized (prohibited area), the accessory is moved so as to be oriented in a position where the expression of the face image can be visually recognized, and the accessory can be moved so as to be oriented in a position where the expression of the face image can be visually recognized before the expression of the face image is hidden.
[0184] In addition, in this embodiment, the accessory is gradually moved toward a position where the expression of the face image can be visually recognized according to the distance to a position where the expression of the face image cannot be visually recognized, and the accessory can be naturally moved toward a position where the expression of the face image can be visually recognized.
[0185] Note that, in this embodiment, even when it is not a position where the expression of the face image cannot be visually recognized, when approaching a position where the expression of the face image cannot be visually recognized (prohibited area), the accessory is configured to move toward a position where the expression of the face image can be visually recognized. However, the accessory may be moved toward a position where the expression of the face image can be visually recognized at the timing when the position where the expression of the face image cannot be visually recognized is reached. Even with such a configuration, it is possible to prevent the photographing direction of the subject by the camera 108 from changing and the expression of the face image from being hidden by the accessory.
[0186] In addition, in this embodiment, the accessory is configured to be gradually moved toward a position where the expression of the face image can be visually recognized according to the distance to a position where the expression of the face image cannot be visually recognized. However, it may be configured such that the accessory is moved toward a position where the expression of the face image can be visually recognized at the timing when the position where the expression of the face image cannot be visually recognized (prohibited area) is reached.
[0187] In this embodiment, the shape or position of the parts constituting the face image is corrected according to the photographing direction of the camera 108, and it is possible to prevent the expression of the generated face image from becoming unnatural even when the photographing direction changes.
[0188] In addition, in this embodiment, the contour of the face image is corrected according to the photographing direction of the camera 108, and it is possible to prevent the contour of the generated face image from becoming unnatural even when the photographing direction changes.
[0189] In addition, in this embodiment, the position and shape of the mouth constituting the face image are corrected according to the photographing direction of the camera 108, and it is possible to prevent the position and shape of the mouth constituting the generated face image from becoming unnatural even when the photographing direction changes.
[0190] In addition, in this embodiment, the position and shape of the nose that make up the face image are corrected according to the shooting direction of the camera 108, and it is possible to prevent the position and shape of the nose that make up the generated face image from becoming unnatural even when the shooting direction changes.
[0191] In addition, in this embodiment, the position and shape of the eyes that make up the face image are corrected according to the shooting direction of the camera 108, and it is possible to prevent the position and shape of the eyes that make up the generated face image from becoming unnatural even when the shooting direction changes.
[0192] In addition, in this embodiment, the correction amount of the shape or position of the parts that make up the face image is gradually increased according to the shooting direction of the camera 108, so that the shape and position of the parts that make up the face image do not change suddenly.
[0193] As described above, the embodiments of the present invention have been described with reference to the drawings. However, the present invention is not limited to these embodiments, and it goes without saying that the present invention includes changes and additions within the scope that does not depart from the gist of the present invention.
[0194] For example, in the above embodiment, an example in which the avatar image generated by the image generation program is used for video distribution has been described. However, the use of the avatar image generated by the image generation program is arbitrary, and the avatar image may be recorded as an archive, or may be used for the production of animation or the like.
[0195] In addition, in the above embodiment, an example has been described in which the expression parameters output from the face tracking program are corrected by the parameter correction program, and the image generation program generates an avatar image using the corrected expression parameters. However, a configuration may be adopted in which the image generation program directly uses the parameters output from the face tracking program without using the parameter correction program to generate an image.
Explanation of Signs
[0196] 1 Computer terminal 101 Processor 102 Memory 103 Storage 104 Communication Interface 105 Display Device 106 Speaker 107 Input Device 108 Camera 109 Microphone 110 Vital Sign Monitor 111 Data Bus
Claims
1. A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generating step; A method for generating an image using a computer, comprising: an emotion value identification step of identifying an emotion value that varies based on the parameters of at least a portion of the parts detected in the parameter detection step; a parameter correction step of correcting the parameters detected in the parameter detection step based on the emotion value identified in the emotion value identification step, In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated, and the shape or position of features constituting the face image is corrected according to the photographing direction.
13. An image generating method comprising:
2. In the face image generating step, the contour of the face image is corrected according to the photographing direction. The image generating method according to claim 1 .
3. In the face image generating step, the position or shape of the mouth constituting the face image is corrected according to the photographing direction. The image generating method according to claim 1 .
4. In the face image generating step, the position or shape of the eyes constituting the face image is corrected according to the photographing direction. The image generating method according to claim 1 .
5. In the face image generating step, the correction amount of the shape or position of the features constituting the face image is gradually increased according to the photographing direction. The image generating method according to any one of claims 1 to 4.
6. A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generating step; In an image generating program for causing a computer to execute an emotion value identification step of identifying an emotion value that varies based on the parameters of at least a portion of the parts detected in the parameter detection step; a parameter correction step of correcting the parameters detected in the parameter detection step based on the emotion value identified in the emotion value identification step, In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated, and the shape or position of features constituting the face image is corrected according to the photographing direction.
13. An image generating program comprising:
Citation Information
Patent Citations
Picture generation system and information storage medium
JP2001084402A
Image generation device, image generation method, and program
JP2007156773A
System and method for processing image
JP2011039828A
Content distribution server, content distribution system, content distribution method, and program
JP2019186797A
Program for generating image, method, image generation device, and game system
JP2023165541A