IMAGE GENERATION METHOD AND IMAGE GENERATION PROGRAM
By detecting the state parameters of multiple parts in the facial image and generating facial images and decorative images based on these parameters, the problem of the output image of traditional facial tracking technology being too monotonous is solved, and the generation and output of diverse images are achieved.
Patent Information
- Application Number
- JP2024089020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2044-05-31
AI Technical Summary
When traditional facial tracking technology reflects facial expressions to virtual characters, the output image is too monotonous and lacks diversity.
By detecting the state parameters of multiple parts in the facial image and generating facial images and decorative images based on these parameters, and combining methods of facial image generation and decorative image generation, diverse images are output.
It realizes the output of diverse images through facial tracking technology, which enhances the fun and transformability of the images.
Smart Images

Figure 0007673295000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image generating method and an image generating program for generating a face image having an expression corresponding to an input face image. [Background technology]
[0002] Conventionally, a technology has been proposed that, for example, uses face tracking to capture a face photographed with a smartphone, allowing the facial expression of the captured face to be reflected in the facial expression of an avatar, such as a character created using CG, etc. (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2019-186797 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional technology for reflecting face-tracked facial expressions on an avatar's facial expressions simply reflected the state of the tracked facial features onto the corresponding features of the avatar, resulting in a monotonous output image.
[0005] An object of the present invention is to provide an image generating method and an image generating program that are capable of outputting a variety of images through face tracking. [Means for solving the problem]
[0006] The image generation method of the first means is as follows: A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: The method further includes a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step, In the image output step, the decorative image generated in the decorative image generating step is output together with the face image. It is characterized by the following. According to this feature, a decorative image other than a facial image is generated based on parameters detected from a captured facial image, so that the facial image detected by face tracking can be output as a varied image.
[0007] The image generating method according to the second aspect is the image generating method according to the first aspect, The method further includes the step of determining an emotion value based on the parameters of at least some of the parts; In the decorative image generating step, the decorative image is generated based on the emotion value identified in the emotion value identifying step. It is characterized by the following. According to this feature, a decorative image can be generated according to parameters corrected based on emotion values determined based on parameters of at least some of the features.
[0008] The image generating method of the third aspect is the image generating method according to the second aspect, In the emotion value identification step, a first value indicating happiness and a second value indicating anger are identified based on parameters of the partial part, and the emotion value indicating happiness is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. It is characterized by the following. According to this feature, it is possible to generate a decorative image based on an emotion value that reflects both joy and anger identified from a face image.
[0009] The image generating method according to the fourth aspect is the image generating method according to the third aspect, In the emotion value identification step, the first value and the second value are identified based on a parameter indicating a degree of movement of the corners of the mouth. It is characterized by the following. According to this feature, emotions such as joy and anger can be identified from the degree of movement of the corners of the mouth, which is likely to reflect emotions.
[0010] The image generating method of means 5 is the image generating method according to any one of means 2 to 4, In the decorative image generating step, when the emotion value specified in the emotion value specifying step exceeds a threshold, the decorative image is generated. It is characterized by the following. According to this feature, decorative images can be generated with an appropriate frequency.
[0011] The image generating program of the means 6 includes: A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step; In the image output step, the decorative image generated in the decorative image generating step is output together with the face image. It is characterized by the following. According to this feature, a decorative image other than a facial image is generated based on parameters detected from a captured facial image, so that the facial image detected by face tracking can be output as a varied image.
[0012] Furthermore, the present invention may have only the invention-specific matters described in the claims of the present invention, or may have the invention-specific matters described in the claims of the present invention as well as configurations other than the invention-specific matters. [Brief description of the drawings]
[0013] [Figure 1] FIG. 2 is a block diagram showing an example of the configuration of a computer terminal used in an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a configuration of a program installed in a computer terminal used in an embodiment of the present invention. [Diagram 3] FIG. 2 is a diagram showing the relationship between programs installed in a computer terminal in an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram showing a configuration of parameters used for image generation in an embodiment of the present invention. [Diagram 5] 10 is a flowchart showing the control process of a parameter correction program executed by a computer terminal in the embodiment of the present invention. [Figure 6] FIG. 13 is a diagram showing the contents of emotion value calculation processing performed by a computer terminal in an embodiment of the present invention. [Figure 7] FIG. 11 is a diagram showing the contents of a MouthSmile correction process performed by a computer terminal in an embodiment of the present invention. [Figure 8] FIG. 11 is a diagram showing the contents of a MouthFrown correction process performed by a computer terminal in an embodiment of the present invention. [Figure 9] 11A and 11B are diagrams illustrating the content of a BrowDown correction process performed by a computer terminal in an embodiment of the present invention. [Figure 10] FIG. 11 is a diagram showing the contents of EyeBlink switching processing performed by a computer terminal in an embodiment of the present invention. [Figure 11] 11A and 11B are diagrams illustrating the contents of a configuration correction process performed by a computer terminal in an embodiment of the present invention. [Figure 12] 5A to 5C are diagrams illustrating examples of facial expressions detected by a face tracking program in an embodiment of the present invention. [Figure 13] 11A to 11C are diagrams illustrating examples of changes in facial expression according to emotion values in an embodiment of the present invention. [Figure 14] 11A to 11C are diagrams showing examples of changes in facial expression in response to switching of facial expression parameters in an embodiment of the present invention. [Figure 15] 10 is a flowchart showing the control content of a correction content setting process performed by a computer terminal in the embodiment of the present invention. [Figure 16] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 17] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 18] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 19] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 20] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 21] FIG. 13 is a diagram illustrating a procedure for setting setting data used in the emotion value setting process in an embodiment of the present invention. [Figure 22] 11A to 11C are diagrams for explaining a procedure for setting setting data used in the MouthSmile correction process in an embodiment of the present invention. [Figure 23] 11A and 11B are diagrams for explaining a procedure for setting setting data used in MouthFrown correction processing and BrowDown correction processing in an embodiment of the present invention. [Figure 24] 5 is a flowchart showing the control of an image generating program executed by a computer terminal in the embodiment of the present invention. [Diagram 25] 11A to 11C are diagrams illustrating examples of changes in facial expressions accompanying correction according to the shooting direction in an embodiment of the present invention. [Figure 26] 11A and 11B are diagrams illustrating changes in correction amount depending on the shooting direction in an embodiment of the present invention. [Figure 27] 13A to 13C are diagrams illustrating examples of how clothing changes depending on the shooting direction in an embodiment of the present invention. [Figure 28] 11A to 11C are diagrams illustrating an example of how an accessory changes depending on the shooting direction in an embodiment of the present invention. [Figure 29] 11A and 11B are diagrams showing the amount of change in clothing or the amount of movement of accessories depending on the shooting direction in an embodiment of the present invention. [Diagram 30] 11A and 11B are diagrams illustrating a relationship between additional conditions and decorative images corresponding to the additional conditions in an embodiment of the present invention. [Diagram 31] 11A and 11B are diagrams illustrating examples of display of decorative images in an embodiment of the present invention. [Diagram 32] 11A and 11B are diagrams illustrating examples of display of decorative images in an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] [Form 1] The image generating method of form 1-1 is as follows: A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: an information input step of inputting information other than a face image; a parameter correction step of correcting the parameters detected in the parameter detection step based on information other than the face image input in the information input step; Further comprising: In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, a facial image with an expression corresponding to parameters corrected based on information other than the photographed facial image is generated, so that the facial image detected by face tracking can be output as a varied image.
[0015] The image generating method of embodiment 1-2 is the image generating method according to embodiment 1-1, further comprising an emotion value determination step of determining an emotion value based on information other than the face image; In the parameter correction step, the parameter detected in the parameter detection step is corrected based on the emotion value identified in the emotion value identification step. It is characterized by the following. According to this feature, it is possible to generate a facial image with an expression based on an emotion value specified from information other than the facial image.
[0016] The image generating method of embodiment 1-3 is the image generating method according to embodiment 1-2, In the emotion value determination step, the emotion value is determined based on information other than the face image and parameters of some of the features detected in the parameter detection step. It is characterized by the following. This feature allows for more accurate identification of emotion values.
[0017] The image generating method of embodiment 1-4 is the image generating method according to embodiment 1-3, In the parameter correction step, when the emotion value identified in the emotion value identification step exceeds a threshold value, the parameter detected in the parameter detection step is corrected. It is characterized by the following. This feature allows for more pronounced changes in facial expression.
[0018] The image generating method of embodiment 1-5 is the image generating method according to any one of embodiments 1-1 to 1-4, The information other than the face image includes information related to at least one of voice, vital signs, and operation input. It is characterized by the following. According to this feature, a facial image with an expression based on information regarding at least one of voice, vital signs, and operational input can be generated.
[0019] The image generating method of embodiment 1-6 is the image generating method according to any one of embodiments 1-1 to 1-5, The method further includes a setting value receiving step of receiving a setting value related to the amount of change of the plurality of parts, The information other than the face image includes the setting value set in the setting value receiving step. It is characterized by the following. According to this feature, it is possible to generate a face image with an expression corresponding to parameters corrected based on set values relating to the amounts of change in a plurality of facial features.
[0020] A correction content setting method of embodiment 1-7 is a correction content setting method using a computer to set correction content of the parameters used in the image generating method according to any one of embodiments 1-1 to 1-6, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting the correction content; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting contents; Includes It is characterized by the following. According to this feature, the correction contents of the parameters can be easily set by simply connecting the nodes and setting the correction contents.
[0021] The image generation program of form 1-8 is A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute an information input step of inputting information other than a face image; a parameter correction step of correcting the parameters detected in the parameter detection step based on information other than the face image input in the information input step; Then, In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, a facial image with an expression corresponding to parameters corrected based on information other than the photographed facial image is generated, so that the facial image detected by face tracking can be output as a varied image.
[0022] A correction content setting program of aspect 1-9 is a correction content setting program for setting correction content of the parameters used in the image generating program according to aspect 1-8 using a computer, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting the correction content; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting contents; Run It is characterized by the following. According to this feature, the correction contents of the parameters can be easily set by simply connecting the nodes and setting the correction contents.
[0023] [Form 2] The image generating method of form 2-1 is as follows: A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: The method further includes a parameter correction step of correcting parameters of other parts based on the parameters of the part of the parts detected in the parameter detection step, In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, parameters of some features are corrected based on the parameters of other features detected from the captured face image, and a face image with an expression according to the corrected parameters is generated, so that the face image detected by face tracking can be output as an image with a rich variety.
[0024] The image generating method of embodiment 2-2 is the image generating method according to embodiment 2-1, The method further includes an emotion value determination step of determining an emotion value based on the parameters of the part of the parts; In the parameter correction step, the parameters of the other parts detected in the parameter detection step are corrected based on the emotion value identified in the emotion value identification step. It is characterized by the following. According to this feature, it is possible to generate a face image with an expression based on an emotion value that is specified based on parameters of some facial features.
[0025] The image generating method of embodiment 2-3 is the image generating method according to embodiment 2-2, In the emotion value identification step, a first value indicating happiness and a second value indicating anger are identified based on parameters of the partial part, and the emotion value indicating happiness is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. It is characterized by the following. According to this feature, it is possible to generate a face image with an expression based on an emotion value that reflects both joy and anger identified from the face image.
[0026] The image generating method of embodiment 2-4 is the image generating method according to embodiment 2-3, In the emotion value identification step, the first value and the second value are identified based on a parameter indicating a degree of movement of the corners of the mouth. It is characterized by the following. According to this feature, emotions such as joy and anger can be identified from the degree of movement of the corners of the mouth, which is likely to reflect emotions.
[0027] The image generating method of embodiment 2-5 is the image generating method according to any one of embodiments 2-2 to 2-4, In the parameter correction step, when the emotion value specified in the emotion value specifying step exceeds a threshold value, the parameters of the other parts are corrected. It is characterized by the following. This feature allows for more pronounced changes in facial expression.
[0028] A correction content setting method of embodiment 2-6 is a correction content setting method using a computer to set correction content of the parameters used in the image generating method according to any one of embodiments 2-1 to 2-5, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting the correction content; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting contents; Includes It is characterized by the following. According to this feature, the correction contents of the parameters can be easily set by simply connecting the nodes and setting the correction contents.
[0029] The image generation program of form 2-7 is A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute a parameter correction step of correcting parameters of other parts based on the parameters of the part of the parts detected in the parameter detection step is further performed; In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, parameters of some features are corrected based on the parameters of other features detected from the captured face image, and a face image with an expression according to the corrected parameters is generated, so that the face image detected by face tracking can be output as an image with a rich variety.
[0030] A correction content setting program of aspect 2-8 is a correction content setting program for setting correction content of the parameters used in the image generating program according to aspect 2-7 using a computer, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting the correction content; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting contents; It is characterized by the following. According to this feature, the correction contents of the parameters can be easily set by simply connecting the nodes and setting the correction contents.
[0031] [Form 3] The image generating method of form 3-1 is as follows: A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: The method further includes a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step, In the image output step, the decorative image generated in the decorative image generating step is output together with the face image. It is characterized by the following. According to this feature, a decorative image other than a facial image is generated based on parameters detected from a captured facial image, so that the facial image detected by face tracking can be output as a varied image.
[0032] The image generating method of embodiment 3-2 is the image generating method according to embodiment 3-1, The method further includes the step of determining an emotion value based on the parameters of at least some of the parts; In the decorative image generating step, the decorative image is generated based on the emotion value identified in the emotion value identifying step. It is characterized by the following. According to this feature, a decorative image can be generated according to parameters corrected based on emotion values determined based on parameters of at least some of the features.
[0033] The image generating method of embodiment 3-3 is the image generating method according to embodiment 3-2, In the emotion value identification step, a first value indicating happiness and a second value indicating anger are identified based on parameters of the partial part, and the emotion value indicating happiness is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. It is characterized by the following. According to this feature, it is possible to generate a decorative image based on an emotion value that reflects both joy and anger identified from a face image.
[0034] The image generating method of embodiment 3-4 is the image generating method according to embodiment 3-3, In the emotion value identification step, the first value and the second value are identified based on a parameter indicating a degree of movement of the corners of the mouth. It is characterized by the following. According to this feature, emotions such as joy and anger can be identified from the degree of movement of the corners of the mouth, which is likely to reflect emotions.
[0035] The image generating method of embodiment 3-5 is the image generating method according to any one of embodiments 3-2 to 3-4, In the decorative image generating step, when the emotion value specified in the emotion value specifying step exceeds a threshold, the decorative image is generated. It is characterized by the following. According to this feature, decorative images can be generated with an appropriate frequency.
[0036] The image generation program of form 3-6 is A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step; In the image output step, the decorative image generated in the decorative image generating step is output together with the face image. It is characterized by the following. According to this feature, a decorative image other than a facial image is generated based on parameters detected from a captured facial image, so that the facial image detected by face tracking can be output as a varied image.
[0037] [Form 4] The image generating method of form 4-1 is as follows: A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: A specific image generating step of generating a specific image to be displayed around the face image, In the image output step, the specific image generated in the specific image generating step is output together with the face image, In the specific image generating step, when the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that allows the facial expression of the facial image to be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0038] The image generating method of embodiment 4-2 is the image generating method according to embodiment 4-1, In the specific image generating step, even if the specific image is not in a position where the facial expression of the facial image is not visually recognizable, when the specific image approaches a position where the facial expression of the facial image is not visually recognizable, the specific image is deformed into a shape where the facial expression of the facial image is visually recognizable. It is characterized by the following. According to this feature, the specific image can be deformed into a shape that allows the facial expression of the facial image to be visually recognized even before the facial expression of the facial image is hidden.
[0039] The image generating method of embodiment 4-3 is the image generating method according to embodiment 4-2, In the specific image generating step, the specific image is gradually deformed toward a shape in which the facial expression of the facial image is visible, depending on a distance to a position in which the facial expression of the facial image is not visible. It is characterized by the following. According to this feature, the specific image can be naturally transformed into a shape that allows the facial expression of the facial image to be visually recognized.
[0040] The image generating method of form 4-4 is as follows: A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: A specific image generating step of generating a specific image to be displayed around the face image, In the image output step, the specific image generated in the specific image generating step is output together with the face image, In the specific image generating step, when the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image becomes a position where the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0041] The image generating method of embodiment 4-5 is the image generating method according to embodiment 4-4, In the specific image generating step, even if the facial expression of the facial image is not visually recognizable, when the specific image approaches a position where the facial expression of the facial image is visually recognizable, the position of the specific image is moved toward a position where the facial expression of the facial image is visually recognizable. It is characterized by the following. According to this feature, the position of the specific image can be moved toward a position where the expression of the facial image can be visually recognized before the expression of the facial image is hidden.
[0042] The image generating method of embodiment 4-6 is the image generating method according to embodiment 4-5, In the specific image generating step, the position of the specific image is gradually moved toward a position where the facial expression of the facial image can be visually recognized, depending on the distance to the position where the facial expression of the facial image cannot be visually recognized. It is characterized by the following. According to this feature, the position of the specific image can be naturally moved toward a position where the expression of the facial image can be visually recognized.
[0043] The image generation program of form 4-7 is A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generating step is output together with the face image, In the specific image generating step, when the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that allows the facial expression of the facial image to be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0044] The image generation program of form 4-8 is A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generating step is output together with the face image, In the specific image generating step, when the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image becomes a position where the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0045] [Form 5] The image generating method of form 5-1 is as follows: A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: In the face image generating step, the shape or position of features constituting the face image is corrected according to the photographing direction. It is characterized by the following. According to this feature, by correcting the shape or position of the features constituting the face image in accordance with the shooting direction, it is possible to prevent the facial expression of the face image from becoming unnatural.
[0046] The image generating method of embodiment 5-2 is the image generating method according to embodiment 5-1, In the face image generating step, the contour of the face image is corrected according to the photographing direction. It is characterized by the following. According to this feature, even if the photographing direction changes, it is possible to prevent the contour of the face image from becoming unnatural.
[0047] The image generating method of embodiment 5-3 is the image generating method according to embodiment 5-1 or 5-2, In the face image generating step, the position or shape of the mouth constituting the face image is corrected according to the photographing direction. It is characterized by the following. According to this feature, even if the photographing direction changes, it is possible to prevent the position or shape of the mouth constituting the face image from becoming unnatural.
[0048] The image generating method of embodiment 5-4 is the image generating method according to any one of embodiments 5-1 to 5-3, In the face image generating step, the position or shape of the eyes constituting the face image is corrected according to the photographing direction. It is characterized by the following. According to this feature, even if the shooting direction changes, it is possible to prevent the position or shape of the eyes constituting the face image from becoming unnatural.
[0049] The image generating method of embodiment 5-5 is an image generating method according to any one of embodiments 5-1 to 5-4, In the face image generating step, the correction amount of the shape or position of the features constituting the face image is gradually increased according to the photographing direction. It is characterized by the following. According to this feature, the amount of correction for the shape or position of features constituting the face image increases gradually in response to changes in the shooting direction, so that the shape or position does not change suddenly.
[0050] The image generation program of form 5-6 is A face image input step of inputting a photographed face image; a face image generating step of generating a face image having an expression corresponding to the face image inputted in the face image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute In the face image generating step, the shape or position of features constituting the face image is corrected according to the photographing direction. It is characterized by the following. According to this feature, by correcting the shape or position of the features constituting the face image in accordance with the shooting direction, it is possible to prevent the facial expression of the face image from becoming unnatural.
[0051] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings. EXAMPLES
[0052] [Computer Terminal] Fig. 1 is a block diagram showing an example of the configuration of a computer terminal 1 used in an embodiment of the present invention. The computer terminal 1 shown in Fig. 1 may be a desktop PC or a workstation having separate display devices and input devices such as a keyboard and a mouse, a notebook PC having an integrated display device and input device, or a mobile terminal such as a tablet terminal or a smartphone having a touch panel or the like mounted on the display device, but the following description will be given assuming a mobile terminal such as a tablet terminal or a smartphone.
[0053] As shown in FIG. 1, the computer terminal 1 is equipped with a processor 101, a memory 102, and a storage 103 such as a hard disk or SSD. The processor 101, the memory, and the storage 103 are connected via a data bus 111, and the processor 101 can execute various processes according to programs stored in the storage 103.
[0054] In addition, a display device 105 and a speaker 106 are connected to the data bus 111 via an output interface (not shown), and images and sounds can be output based on the processing executed by the processor 101 .
[0055] In addition, input devices such as an input device 107, a camera 108, a microphone 109, and a vital sign measuring instrument 110 are connected to the data bus 111 via an input interface (not shown), and information input by these input devices can be input. The input device 107 includes a touch panel integrated with the display device 105, and can input a command input operation by the touch panel. Note that a configuration in which commands can be input by an input device such as a keyboard or a mouse can be input. The camera 108 is a 3D camera, and can input data including depth data as captured image data. In addition, the camera 108 may be configured to include both an out-camera provided on the opposite side of the display device 105 and an in-camera provided on the display device 105 side, or only one of them. The microphone 109 may be a stereo microphone or a monaural microphone, and can input surrounding sounds including the voice of the subject. The vital sign measuring instrument 110 is a wearable device such as a smart band, and can input vital sign data (heart rate, blood pressure, body temperature, etc.).
[0056] Furthermore, a communication interface 104 is connected to the data bus 111, and is configured to be able to communicate with other computer terminals and server computers via a local area network, the Internet, or other public network, either wired or wirelessly.
[0057] In this embodiment, a configuration in which the image generating method and the correction content setting method of the present invention are implemented using one computer terminal is illustrated, but a configuration in which the image generating method and the correction content setting method of the present invention are implemented using a plurality of computer terminals is also acceptable. In addition, the computer terminal 1 of this embodiment is configured to include a display device 105 and a speaker 106 as output devices, but these output devices may not be included and images and sounds may be output from an output device of another computer terminal. In addition, the computer terminal 1 of this embodiment is configured to include input devices such as an input device 107 such as a touch panel, a camera 108, and a microphone 109, but it is sufficient that the computer terminal 1 is configured to include at least an input device such as a touch panel, a keyboard, and a mouse, and a camera capable of capturing images.
[0058] In addition to an operation system (OS) (not shown), various programs executed by the computer terminal 1 are stored in the storage 103 of the computer terminal 1 as shown in Fig. 2. Specifically, a face tracking program, a motion tracking program, a voice detection program, an input detection program, a vital detection program, a parameter correction program, an image generation program, a correction content setting program, and a distribution program are stored, and, although not particularly shown, setting data used by the parameter correction program, configuration setting values, material data used by the image generation program, and the like are also stored.
[0059] 3, computer terminal 1 is configured such that an image generation program creates avatar image data having an expression based on the expression parameters and a form based on the motion parameters, based on the expression parameters and the motion parameters based on the image data output from camera 108, and a distribution program distributes a video using the avatar image data created by the image generation program. Furthermore, the expression parameters output by the face tracking program are not output to the image generation program as is, but are output to the image generation program via a parameter correction program, and the parameter correction program corrects at least some of the expression parameters using some of the expression parameters and parameters other than the expression parameters (voice parameters, input parameters, vital parameters, configuration parameters), and outputs the corrected expression parameters to the image generation program.
[0060] Next, the programs executed by the computer terminal 1 will be described.
[0061] The face tracking program is a program that detects the state of a plurality of parts constituting a face image of a subject from the face image using image data including depth data input from the camera 108, as shown in Fig. 3, and outputs facial expression parameters that quantify the degree of movement of each part from the detected state. The facial expression parameters are parameters that can identify the facial expression of the subject, and in this embodiment, as shown in Fig. 4, are parameters that quantify, in the range of 0.0 to 1.0, six types of movement degrees related to the movement of the left eye, six types of movement degrees related to the movement of the right eye, 27 types of movement degrees related to the movement of the mouth and jaw, 10 types of movement degrees related to the movement of the eyebrows, cheeks, and nose, and one type of movement degree related to the movement of the tongue. Note that the facial expression parameters are not limited to these exemplified ones, and many types of more detailed facial expression parameters may be used, or fewer types of facial expression parameters may be used.
[0062] The motion tracking program is a program that detects the body movement of the subject using image data including depth data input from the camera 108, as shown in Fig. 3, and outputs motion parameters that specify the positions and directions of parts that constitute the body. The motion parameters are parameters that can specify the body movement of the subject, and in this embodiment, as shown in Fig. 4, include a head pose that indicates the position coordinates and direction of the head, all finger joints that indicate the position coordinates and directions of all finger joints, and a hand pose that indicates the position coordinates and directions of the hand. Note that the motion parameters in this embodiment are configured not to include parameters regarding the movement of the entire body, but may be configured to include parameters regarding the movement of the entire body.
[0063] The voice detection program is a program that detects the voice of a target person using voice data input from microphone 109 and outputs voice parameters that digitize components of the detected voice, as shown in Fig. 2. The voice parameters are parameters that can identify the voice of a target person, and in this embodiment, as shown in Fig. 4, are parameters that digitize the voice volume and voice pitch in the range of 0.0 to 1.0.
[0064] The input detection program is a program that converts input data from the input device 107 into input parameters and outputs them, as shown in Fig. 3. The input parameters are parameters that can specify the input status from the input device 107, and in this embodiment, as shown in Fig. 4, are parameters that indicate the type of command input by a touch operation corresponding to a command displayed on the display device 105. The command input includes, for example, a command input that specifies emotions such as joy, anger, sadness, and happiness in stages, a command input that specifies a decorative image, and the like.
[0065] The vital sign detection program is a program that converts vital data (heart rate, blood pressure, body temperature, etc.) input from the vital sign measuring device 110 and outputs vital parameters, as shown in Fig. 3. The vital parameters are parameters that can identify the subject's vital signs, and in this embodiment, as shown in Fig. 4, the heart rate, blood pressure, and body temperature are each converted into a numerical value between 0.0 and 1.0.
[0066] [Parameter correction program] As shown in Fig. 3, the parameter correction program is a program that mainly corrects facial expression parameters created by the face tracking program, and includes a parameter correction process that corrects at least some of the facial expression parameters created by the face tracking program using some of the facial expression parameters, voice parameters from the voice detection program, input parameters from the input detection program, vital parameters from the vital detection program, and configuration parameters based on a preset configuration (setting value), and a parameter creation process that creates new facial expression parameters based on input parameters. The parameter correction program is executed for each frame at intervals according to the frame rate of the video distributed by the distribution program. For example, when the frame rate is 60 fps, the parameter correction program is executed 60 times per second.
[0067] The config (setting value) is a value by which the amount of change in the facial expression parameter can be set, and can be set by displaying a config setting screen on the display device 105 and using the input device 107. The config (setting value) can be set for each type of facial expression parameter, and in this embodiment, as shown in FIG. 4, a numerical value between 0.0 and 1.0 can be set individually for each of the facial expression parameters related to eye movement, the facial expression parameters related to mouth and jaw movement, the facial expression parameters related to eyebrow movement, the facial expression parameters related to cheek and nose movement, and the facial expression parameters related to tongue movement, and the set numerical values are used as the config parameters. In this embodiment, the larger the numerical value set as the config, the larger the amount of change in the corresponding facial expression parameter and the larger the movement of the corresponding part.
[0068] FIG. 5 is a flowchart showing the control contents of the parameter correction program.
[0069] 5, in the parameter correction program, first, various parameters are taken in (Sa1). In step Sa1, facial expression parameters output from the face tracking program, voice parameters output from the voice detection program, input parameters output from the input detection program, vital parameters output from the vital detection program, and configuration parameters based on a preset configuration (setting value) are taken in.
[0070] Next, the setting data is referred to and an emotion value is calculated based on the various parameters taken in step Sa1 (Sa2). The setting data is data set by a correction content setting program, and defines the type of input parameters, the correction content (calculation formula, output conditions, etc.) based on the input parameters, and the type of output parameters. The emotion value, which will be described in detail later, is a value calculated by specifying a Joy value indicating the degree of joy and an Anger value indicating the degree of anger from some parameters of the facial expression parameters, and calculating an AG value from the Joy value and Anger value, as well as voice parameters and vital parameters, and is a parameter indicating the degree of joy or anger of the subject.
[0071] Next, by referring to the setting data, the facial expression parameters are corrected by the various parameters acquired in step Sa1 and the emotion value calculated in step Sa2 (Sa3). In step Sa3, it is sufficient that at least some of the facial expression parameters output from the face tracking program are corrected. The facial expression parameters may be corrected by parameters including emotion values, or may be corrected by parameters not including emotion values. As parameters used to correct the facial expression parameters, only some of the facial expression parameters may be used, only parameters other than the facial expression parameters may be used, or both some of the facial expression parameters and parameters other than the facial expression parameters may be used. The parameters used to correct the facial expression parameters are not limited to parameters output from a face tracking program or the like, and may be parameters corrected by a plurality of parameters and further used to correct the facial expression parameters. The facial expression parameters to be corrected are not limited to those output from a face tracking program, and may be further corrected by using other parameters on the facial expression parameters that have been corrected once.
[0072] Next, by referring to the setting data, the various parameters taken in step Sa1 and the emotion value calculated in step Sa2 are used to newly create facial expression parameters other than those output from the face tracking program (Sa4). The parameters created in step Sa4 may be configured to be output separately from the facial expression parameters output from the face tracking program, or the facial expression parameters output from the face tracking program may be switched to the newly created parameters.
[0073] Next, the facial expression parameters corrected in step Sa3 and the facial expression parameters created in step Sa4 are output to the image generation program (Sa5). Note that in step Sa5, facial expression parameters output from the face tracking program and used without correction are also output. In step Sa5, it is possible to output not only the facial expression parameters, but also parameters such as emotion values created in the process of correcting or newly creating the facial expression parameters, as extended parameters to the image generation program. In this embodiment, the emotion values are output as extended parameters. The facial expression parameters and extended parameters output to the image generation program in step Sa5 can be specified by a correction content setting process.
[0074] [Emotion value calculation process] Next, an example of a method for calculating emotion values using a parameter correction program will be described. As shown in Fig. 6, the emotion value calculation process for calculating emotion values uses BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthFunnel (MFu), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR) from among the facial expression parameters output from the face tracking program, volume (Vo) and pitch (Pi) are used as voice parameters output from the voice detection program, and pulse (PR) is used as a vital parameter output from the vital detection program.
[0075] As shown in FIG. 12(a), BrowDownLeft (BDoL) is a parameter representing the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter representing the downward movement of the right eyebrow. BrowDownLeft (BDoL) and BrowDownRight (BDoR) are parameters whose numerical value increases in the case of an angry facial expression. Also, as shown in FIG. 12(b), CheekSquintLeft (CSqL) is a parameter representing the upward movement of the cheek around the left eye and the lower cheek, and CheekSquintRight (CSqR) is a parameter representing the upward movement of the cheek around the right eye and the lower cheek. CheekSquintLeft (CSqL) and CheekSquintRight (CSqR) are parameters whose numerical value increases in the case of a happy facial expression. Also, as shown in FIG. 12(c), MouthFunnel (MFu) is a parameter representing the contraction of both lips into an open shape. MouthFunnel (MFu) is a parameter whose numerical value increases in the case of an angry facial expression. As shown in FIG. 12(d), MouthStretchLeft (MStL) is a parameter that represents the left movement of the left corner of the mouth, and MouthStretchRight (MStR) is a parameter that represents the right movement of the right corner of the mouth. MouthStretchLeft (MStL) and MouthStretchRight (MStR) are parameters that have a higher value in the case of an angry facial expression. As shown in FIG. 12(e), MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters that have a higher value in the case of a happy facial expression.
[0076] The emotion value calculation process first calculates (BDoL+BDoR) / 2, (CSqL+CSqR) / 2, (MStL+MStR) / 2, and (MSmL+MSmR) / 2, and then calculates BDo, CSq, MSt, and MSm as the average values of the left and right facial expression parameters.
[0077] Next, using BDo, MFu, and MSt, which have higher numerical values for angry facial expressions, ((BDo-CSq) x 2 + MSt + MFu) / 4 x (-1) is performed to calculate the Anger value, which indicates the level of anger. At this time, a more accurate level of anger can be calculated by subtracting not only BDo, MFu, and MSt, which have higher numerical values for angry facial expressions, but also CSq, which has higher numerical values for happy facial expressions, from BDo. The Anger value is a numerical value ranging from 0 to -1, and the closer the numerical value is to -1, the greater the level of anger.
[0078] Next, using CSq and MSm, which have higher values when the facial expression is happy, we calculate the Joy value, which indicates the degree of joy, by calculating (MSm+CSq) / 2. The Joy value is a value in the range of 0 to +1, and the closer the value is to +1, the greater the degree of joy.
[0079] Next, the Anger value and the Joy value are used to perform Anger+Joy to calculate an emotion value AG. The emotion value AG is a numerical value ranging from 1 to -1.
[0080] Next, the voice parameters volume (Vo) and pitch (Pi), and the vital parameter pulse (PR) are converted to Vo1, Pi1, and PR1, which indicate values of 1 or 0. Volume (Vo), pitch (Pi), and pulse (PR) are all numerical values between 0 and 1. If the numerical values of volume (Vo) and pitch (Pi) are 0.8 or greater, they are converted to 1, and if they are less than 0.8, they are converted to 0. In addition, if the numerical value of pulse (PR) is 0.6 or greater, it is converted to 1, and if it is less than 0.6, it is converted to 0.
[0081] Next, if the emotion value AG exceeds 0, i.e., if the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, AG x 0.8 + (Vo1 + Pi1 + PR1) / 3 x 0.2 is calculated, and the calculated result is output as the emotion value EmoJ indicating the degree of joy. As a result, if the volume (Vo), pitch (Pi), and pulse rate (PR) are equal to or greater than a certain value, and emotional excitement can be read, a larger value, i.e., EmoJ indicating a greater degree of joy, is output than when the volume (Vo), pitch (Pi), and pulse rate (PR) are less than a certain value.
[0082] On the other hand, if the emotion value AG is less than 0, i.e., if the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, AG x 0.8 - (Vo1 + Pi1 + PR1) / 3 x 0.2 is calculated, and the calculated result is output as the emotion value EmoA indicating the degree of anger. As a result, if the volume (Vo), pitch (Pi), and pulse rate (PR) are equal to or greater than a certain value, and the emotional intensity can be read, a smaller value is output than when the volume (Vo), pitch (Pi), and pulse rate (PR) are less than a certain value, i.e., EmoA indicating a greater degree of anger is output.
[0083] In this way, in the emotion value calculation process, the Anger value indicating the degree of anger is determined using emotion parameters that are higher in the case of an angry expression, and the Joy value indicating the degree of joy is determined using parameters that are higher in the case of a happy expression. If the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, the emotion value EmoJ indicating the degree of joy is determined, and if the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, the emotion value EmoA indicating the degree of anger is determined.
[0084] In addition, when the volume (Vo), pitch (Pi), and pulse rate (PR) are equal to or greater than a certain value, an emotional value indicating a greater degree is specified, such as an emotional value EmoJ indicating a degree of joy or an emotional value EmoA indicating a degree of anger.
[0085] In this embodiment, the emotional value is specified using voice parameters and vital parameters in addition to facial expression parameters, but the emotional value may be specified only using facial expression parameters, or may be specified only using parameters other than facial expression parameters such as voice parameters and vital parameters.Furthermore, the emotional value may be specified using parameters other than facial expression parameters, voice parameters and vital parameters together with these parameters, or may be specified only using parameters other than facial expression parameters, voice parameters and vital parameters.For example, when a command input is made by operating the input device 107 to specify emotions such as joy, anger, sadness and happiness in stages, the emotional value may be specified using input parameters that specify the corresponding command input.
[0086] [Parameter correction process] Next, an example of a method for correcting facial expression parameters using a parameter correction program will be described.
[0087] FIG. 7 is a diagram showing the contents of MouthSmile correction processing for correcting the facial expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) based on emotion values.
[0088] In the MouthSmile correction process, MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are corrected using the emotion value EmoJ, which indicates the degree of joy. In detail, MSmL+EmoJ×0.5 and MSmR+EmoJ×0.5 are performed, and the corrected facial expression parameters MSmLN and MSmRN corresponding to MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are output. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0089] The correction amounts of MSmLN and MSmRN become larger as the emotion value EmoJ, which indicates the degree of joy, approaches +1, and become smaller as the emotion value EmoJ, which indicates the degree of joy, approaches 0. MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth, and the closer the emotion value EmoJ approaches +1, the more the mouth corners turn upward. For this reason, as shown in FIG. 13(a), compared to when the emotion value is 0, for example, when the emotion value EmoJ is +1, the corners of the mouth turn upward, and the closer the emotion value EmoJ approaches +1, the more the facial expression of the output facial image can be made to look like a happier expression.
[0090] FIG. 8 is a diagram showing the contents of MouthFrown correction processing for correcting the facial expression parameters MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) based on emotion values.
[0091] In the MouthFrown correction process, MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) are corrected using the emotion value EmoA, which indicates the degree of anger. In detail, MFrL+EmoA×(-1)×0.5 and MFrR+EmoA×(-1)×0.5 are performed, respectively, and the corrected facial expression parameters MFrLN and MFrRN corresponding to MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) are output, respectively. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0092] The correction amounts of MFrLN and MFrRN become larger as the emotion value EmoA indicating the degree of anger approaches -1, and become smaller as the emotion value EmoA indicating the degree of anger approaches 0. MouthFrownLeft (MFrL) is a parameter representing the downward movement of the left corner of the mouth, and MouthFrownRight (MFrR) is a parameter representing the downward movement of the right corner of the mouth, and the closer the emotion value EmoA approaches -1, the more the corners of the mouth turn downward. For this reason, as shown in Fig. 13(a), compared to the case where the emotion value is 0, for example, when the emotion value EmoA is -1, the corners of the mouth turn downward, and the closer the emotion value EmoA approaches -1, the more the facial expression of the output facial image can be made to look more angry.
[0093] FIG. 9 is a diagram showing the contents of the BrowDown correction process for correcting the facial expression parameters BrowDownLeft (BDoL) and BrowDownRight (BDoR) based on an emotion value.
[0094] In the BrowDown correction process, BrowDownLeft (BDoL) and BrowDownRight (BDoR) are corrected using the emotion value EmoA, which indicates the degree of anger. In detail, BDoL+EmoA×(-1)×0.5 and BDoR+EmoA×(-1)×0.5 are performed, and the corrected facial expression parameters BDoLN and BDoRN corresponding to BrowDownLeft (BDoL) and BrowDownRight (BDoR) are output. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0095] BDoLN and BDoRN become larger correction amounts as the emotion value EmoA indicating the degree of anger approaches -1, and become smaller correction amounts as the emotion value EmoA indicating the degree of anger approaches 0. BrowDownLeft (BDoL) is a parameter representing the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter representing the downward movement of the right eyebrow, and the closer the emotion value EmoA approaches -1, the more the eyebrows are downward. Therefore, as shown in FIG. 13(b), compared to the emotion value 0, for example, when the emotion value EmoA is -1, the eyebrows are positioned lower, and the closer the emotion value EmoA approaches -1, the more the facial expression of the output facial image can be made to look more angry.
[0096] FIG. 10 is a diagram showing the contents of an EyeBlink switching process in which the facial expression parameter corresponding to the facial expression parameter EyeBlinkLeft (EBL) is switched to one of the facial expression parameter EyeBlinkLeft (EBL) and the facial expression parameter EyeSmileLeft (ESL) based on the facial expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR).
[0097] EyeBlinkLeft (EBL) is a parameter that represents the degree of closure of the left eyelid, and EyeSmileLeft (ESL) is a parameter that represents the degree of closure of the left eyelid when the corner of the eye is drooping, and is an expression parameter that is used in place of EyeBlinkLeft (EBL).
[0098] In the EyeBlink switching process, MSm is calculated as the average value of the left and right parameters based on MouthSmileLeft (MSmL) and MouthSmileRight (MSmR). If MSm is a value of 0.5 or more, the value of EyeBlinkLeft (EBL) is output as the value of EyeSmileLeft (ESL) instead of EyeBlinkLeft (EBL). On the other hand, if MSm is a value less than 0.5, the value of EyeBlinkLeft (EBL) is output as it is as the value of EyeBlinkLeft (EBL).
[0099] As shown in FIG. 12(e), MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters whose numerical values become higher when the facial expression is smiling. Therefore, in the EyeBlink switching process, when MSm is less than a certain value, the facial expression is a normal facial expression with the eyelids closed by EyeBlinkLeft (EBL) as shown in FIG. 14(a), whereas when MSm is equal to or greater than a certain value, the facial expression is a facial expression with the eyelids closed with the corners of the eyes drooping as shown in FIG. 14(b) by EyeSmileLeft (ESL).
[0100] FIG. 11 is a diagram showing the contents of the configuration correction process for correcting facial expression parameters based on configuration parameters.
[0101] As shown in Figure 4, configurations (setting values) are set for each of the facial expression parameters related to eye movement, mouth and jaw movement, eyebrow movement, cheek and nose movement, and tongue movement, and the configuration correction process uses each configuration parameter to correct the value of the corresponding configuration parameter. Specifically, (facial expression parameter + corresponding configuration parameter) / 2 is calculated and the corrected facial expression parameter N is output.
[0102] A configuration parameter is a value between 0.0 and 1.0, and the larger the value of the configuration parameter, the larger the value of the corresponding facial expression parameter N, and the smaller the value of the configuration parameter, the smaller the value of the corresponding facial expression parameter N. For this reason, the amount of change in the part corresponding to the facial expression parameter can be adjusted according to the preset configuration parameter value.
[0103] In this manner, in this embodiment, other facial expression parameters are corrected based on some of the facial expression parameters detected from the facial image data captured by the camera 108, and a facial image with an expression corresponding to the corrected facial expression parameter N is generated, so that the facial image detected by face tracking can be output as an image with a variety of features.
[0104] Furthermore, in this embodiment, in addition to some of the facial expression parameters detected from the facial image data captured by the camera 108, a facial image with an expression corresponding to the facial expression parameter N corrected based on parameters other than the facial expression parameters is generated, so that the facial image detected by face tracking can be output as an image with even more variety.
[0105] In this embodiment, the facial parameters are corrected using both some of the facial parameters and parameters other than the facial parameters, but it is also possible to generate a facial image with an expression corresponding to the facial parameter N corrected based on only some of the facial parameters, or to generate a facial image with an expression corresponding to the facial parameter N corrected based on only parameters other than the facial parameters. Even with such a configuration, the facial image detected by face tracking can be output as a varied image.
[0106] In this embodiment, parameters other than facial expression parameters include voice parameters based on the voice input from the microphone 109, and it is possible to generate a facial image with an expression that corresponds to, for example, the volume, pitch, etc. of the voice input from the microphone 109.
[0107] In addition, in this embodiment, parameters other than facial expression parameters include input parameters based on the operational input of commands via the input device 107, and it is possible to generate a facial image with an expression corresponding to, for example, a command input via the input device 107.
[0108] In addition, in this embodiment, parameters other than facial expression parameters include vital parameters based on vital data input by the vital sign measuring device 110, and it is possible to generate a facial image with an expression corresponding to, for example, the pulse, blood pressure, body temperature, etc. input by the vital sign measuring device 110.
[0109] In addition, in this embodiment, parameters other than facial expression parameters include configuration parameters based on configurations (setting values) that are set in advance on a configuration setting screen, and for example, a facial image of an expression can be generated with an amount of change adjusted by the configuration.
[0110] In this embodiment, the parameters other than facial expression parameters include voice parameters, input parameters, vital parameters, and configuration parameters, but it may be a configuration that includes only some of these, or a configuration that includes other parameters, such as parameters identified from the background image of the subject, the background image of the avatar to be distributed at the same time, the temperature, humidity, and brightness of the shooting location.
[0111] In this embodiment, an emotion value is identified based on some of the emotion parameters detected from the facial image data captured by the camera 108, and the emotion parameters are corrected based on the identified emotion value. Since a facial image with an emotion corresponding to the corrected emotion parameter N is generated, the emotion identified from the emotion of the facial image detected by face tracking can be reflected in the generated facial image.
[0112] Furthermore, in this embodiment, in addition to some of the facial expression parameters detected from the face image data captured by the camera 108, the emotion value is determined based on parameters other than the facial expression parameters, so that the emotion to be reflected in the generated face image can be determined more accurately.
[0113] In this embodiment, the emotion value is specified based on both some of the facial expression parameters and parameters other than the facial expression parameters, but it is also possible to specify an emotion value based only on some of the facial expression parameters, or to specify an emotion value based only on parameters other than the facial expression parameters, and even in such a configuration, the specified emotion can be reflected in the generated facial image.
[0114] In addition, in this embodiment, a Joy value indicating joy and an Anger value indicating anger are identified based on some facial expression parameters, and when the Joy value is greater than the Anger value, an emotional value EmoJ indicating joy is identified, and when the Anger value is greater than the Joy value, an emotional value EmoA indicating anger is identified. This makes it possible to generate a facial image with an expression based on emotional values reflecting both joy and anger identified from some facial expression parameters or parameters other than the facial expression parameters.
[0115] In addition, in this embodiment, the Joy value and the Anger value are determined using MouthStretch and MouthSmile, which are facial expression parameters that indicate the degree of movement of the corners of the mouth, and emotions such as joy and anger can be determined from the degree of movement of the corners of the mouth, which tend to reflect emotions.
[0116] In this embodiment, a Joy value indicating happiness and an Anger value indicating anger are specified and reflected in the facial expression parameters, but a value indicating sadness and a value indicating enjoyment may also be specified and the emotional values of joy, anger, sadness, and happiness may each be reflected in the facial expression parameters, or only a portion of these emotional values may be specified and the specified emotional values may be reflected in the facial expression parameters.
[0117] Furthermore, in this embodiment, the configuration is such that the corresponding facial expression parameter is corrected regardless of the magnitude of the emotion value, but it may also be such that the corresponding facial expression parameter is corrected when the emotion value exceeds a certain threshold value. By using such a configuration, it is possible to add emphasis to the changes in facial expression in the generated facial image.
[0118] Furthermore, in this embodiment, an emotion value is specified by some of the emotion parameters and parameters other than the emotion parameters, and the emotion parameters are corrected by the specified emotion value. However, an emotion parameter may be directly corrected by some of the emotion parameters that change depending on the emotion or parameters other than the emotion parameters. Even in such a configuration, if a parameter used for correction exceeds a threshold value, the corresponding emotion parameter is corrected, thereby making it possible to add sharpness to the changes in the emotion of the generated facial image.
[0119] [Correction content setting program] The correction content setting program is a program that performs a process of setting setting data used by the parameter correction program. In addition to the process of setting setting data, the correction content setting program also performs a process of designating facial expression parameters and extended parameters to be output to the image generation program. A detailed description of the process of designating facial expression parameters and extended parameters to be output to the image generation program will be omitted. In this embodiment, the computer terminal 1 that executes the parameter correction program and the image generation program is configured to be able to execute the correction content setting program, but a function for executing the correction content setting program may be installed in a computer terminal other than the computer terminal 1, and the setting data set in the other computer terminal may be applied as the setting data for the parameter correction program of the computer terminal 1.
[0120] FIG. 15 is a flowchart showing the control contents of the correction content setting program.
[0121] 15, in the correction content setting program, first, a parameter to be corrected is specified (Sb1). In step Sb1, for example, an input node is placed on a setting screen, and the type of parameter to be corrected is set in the input node.
[0122] Next, parameters to be used for correction are designated (Sb2). In step Sb2, for example, like step Sb1, an input node is placed on the setting screen, and the type of parameter to be used for correction is set in the input node.
[0123] Next, correction contents such as a formula using a plurality of parameters and output conditions are set (Sb3). In step Sb3, for example, a setting node and an output node are arranged on a setting screen, an input node in which the type of parameter to be corrected is set and an input node in which the type of parameter to be used for correction is set are connected to an input section of the setting node, an output node in which the type of parameter after correction is set is connected to an output section of the setting node, and further correction contents such as a formula using the parameters input to the setting node and output conditions are set.
[0124] Next, based on the settings of the input node, setting node, and output node, setting data is created that defines the types of parameters to be input, the correction contents (calculation formulas, output conditions, etc.) based on the input parameters, and the types of parameters to be output, and is output to the parameter correction program (Sb4).
[0125] Next, a specific procedure for setting the setting data by the correction content setting program will be described. Here, the procedure for setting the setting data used in the emotion value setting process, and the MouthSmile correction process, MouthFrown correction process, and BrowDown correction process, which are corrections made using emotion values, will be described.
[0126] 16 to 21 are diagrams illustrating the procedure for setting the setting data used in the emotion value setting process.
[0127] The setting data used in the emotion value setting process consists of multiple setting data. First, setting data is set for calculating the left and right averages BDo, CSq, MSt, and MSm of BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR).
[0128] For example, when setting data for calculating BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR), two input nodes, one setting node, and one output node are placed on the setting screen as shown in Fig. 16. The number and positions of these nodes can be arbitrarily specified.
[0129] Next, BDoL and BDoR are input to the two input nodes, respectively. In addition, two input parts, A and B, are set to the setting node, and the input node for BDoL is connected to input part A of the setting node, and the input node for BDoR is connected to input part B of the setting node. In addition, BDo, which is output to the output node, is input, and the output part of the setting node is connected to the output node. Furthermore, (A+B) / 2 is input to the setting node as a calculation formula for taking the average of the parameters input from input parts A and B. By performing a decision operation in this state, setting data is created that calculates BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR).
[0130] The procedure for creating setting data for calculating CSq, MSt, and MSm, which are the left and right averages of CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR), is also similar.
[0131] Next, setting data is set for calculating an Anger value indicating the degree of anger using BDo, CSq, MFu, and MSt.
[0132] Here, as shown in FIG. 17, four input nodes, one setting node, and one output node are placed on the setting screen.
[0133] Next, BDo, CSq, MFu, and MSt are input to the four input nodes, respectively. Furthermore, four input parts A to D are set to the setting node, and the input node of BDo is connected to input part A of the setting node, the input node of CSq is connected to input part B of the setting node, the input node of MSt is connected to input part C of the setting node, and the input node of MFu is connected to input part D of the setting node. Furthermore, Anger to be output to the output node is input, and the output part of the setting node and the output node are connected. Furthermore, ((AB)*2+C+D) / 4*(-1) is input to the setting node as a formula for calculating the Anger value from the parameters input from input parts A to D. By performing a decision operation in this state, setting data for calculating the Anger value indicating the degree of anger is created.
[0134] Next, setting data is set to calculate a Joy value indicating the degree of joy using MSm and CSq.
[0135] Here, as shown in FIG. 18, two input nodes, one setting node, and one output node are placed on the setting screen.
[0136] Next, MSm and CSq are input to the two input nodes, respectively. Furthermore, two input parts A and B are set to the setting node, the input node of MSm is connected to input part A of the setting node, and the input node of CSq is connected to input part B of the setting node. Furthermore, Joy, which is output to the output node, is input, and the output part of the setting node is connected to the output node. Furthermore, (A+B) / 2 is input to the setting node as a formula for calculating a Joy value using the parameters input from input parts A and B. By performing a decision operation in this state, setting data for calculating a Joy value indicating the degree of joy is created.
[0137] Next, the Anger value and the Joy value are used to calculate an emotion value AG, and setting data is set to output AGJ when the calculated AG exceeds 0, and AGA when the calculated AG is less than 0.
[0138] Here, as shown in FIG. 19, two input nodes, two setting nodes, and two output nodes are placed on the setting screen.
[0139] Next, Anger and Joy are input to the two input nodes, respectively. In addition, two input parts A and B are set to one of the setting nodes, the input node of Anger is connected to the input part A of one of the setting nodes, the input node of Joy is connected to the input part B of one of the setting nodes, and the output part is connected to the input part of the other setting node. Furthermore, A+B is input to the input part of one of the setting nodes as a calculation formula for calculating the AG value from the parameters input from A and B. Next, two output parts A and B are set to the other setting node, AGJ and AGA are input to the two output nodes, respectively, the output part A of the other setting node is connected to the output node of AGJ, and the output part B of the other setting node is connected to the output node of AGA. Furthermore, In>0→A In<0→B is input to the other setting node as a conditional formula for distributing parameters according to the AG value input from the input part. By performing a decision operation in this state, setting data is created that outputs AGJ where AG is greater than 0 and AGA where AG is less than 0.
[0140] Next, setting data is set to convert volume (Vo), pitch (Pi), and pulse rate (PR) into Vo1, Pi1, and PR1, which indicate 1 or 0.
[0141] For example, when converting the volume (Vo) to Vo1, one input node, one setting node, and one output node are placed on the setting screen as shown in FIG.
[0142] Next, Vo is input to the input node. The Vo input node is also linked to the input section of the setting node. Vo1, which is output to the output node, is also input, and the output section of the setting node is linked to the output node. Furthermore, In≧0.8→1 In<0.8→0 is input to the setting node as a conditional expression for converting parameters according to the Vo value input from the input section. By performing a decision operation in this state, setting data for converting volume (Vo) to Vo1 is created.
[0143] The procedure for creating the setting data for converting the pitch (Pi) and pulse rate (PR) is similar. Note that the condition formula for converting the pulse rate (PR) to PR1 is In≧0.6→1 and In<0.6→0.
[0144] Next, setting data for calculating emotion values EmoJ and EmoA is set using AGJ, AGA, Vo1, Pi1, and PR1.
[0145] Here, as shown in FIG. 21, five input nodes, two setting nodes, and two output nodes are placed on the setting screen.
[0146] Next, AGJ, AGA, Vo1, Pi1, and PR1 are input to the five input nodes. Also, four input parts A to D are set on one of the setting nodes, the input node of AGJ is connected to the input part A of the setting node, the input node of Vo1 is connected to the input part B of the setting node, the input node of Pi1 is connected to the input part C of the setting node, and the input node of PR1 is connected to the input part D of the setting node. Also, four input parts A to D are set on the other setting node, the input node of AGA is connected to the input part A of the other setting node, the input node of Vo1 is connected to the input part B of the other setting node, the input node of Pi1 is connected to the input part C of the other setting node, and the input node of PR1 is connected to the input part D of the other setting node. Also, EmoJ is set to be output to one output node, the output part of one setting node is connected to one output node, and EmoA is set to be output to the other output node, and the output part of the other setting node is connected to the other output node. Furthermore, A*0.8+(B+C+D) / 3*0.2 is input to one setting node as a formula for correcting the AGJ value using parameters input from input units A to D, and A*0.8-(B+C+D) / 3*0.2 is input to the other setting node as a formula for correcting the AGA value using parameters input from input units A to D. By performing a decision operation in this state, setting data is created that calculates emotion values EmoJ and EmoA using AGJ, AGA, Vo1, Pi1, and PR1.
[0147] In this way, by setting all the setting data shown in Figs. 16 to 21 using the correction content setting program, the parameter correction program can refer to this setting data and calculate the emotion values EmoJ and EmoA.
[0148] FIG. 22 is a diagram for explaining a procedure for setting the setting data used in the MouthSmile correction process.
[0149] When creating setting data to be used in the MouthSmile correction process, three input nodes, two setting nodes, and two output nodes are arranged on the setting screen as shown in FIG.
[0150] Next, EmoJ to be used for correction, MSmL to be corrected, and MSmR are input to the three input nodes, respectively. Furthermore, two input parts A and B are set to one of the setting nodes, the input node of EmoJ is connected to input part A of the one setting node, and the input node of MSmL is connected to input part B of the one setting node. Furthermore, two input parts A and B are set to the other setting node, the input node of EmoJ is connected to input part A of the other setting node, and the input node of MSmR is connected to input part B of the other setting node. Furthermore, MSmLN to be output to one output node is set, the output part of one setting node is connected to one output node, MSmRN to be output to the other output node is set, and the output part of the other setting node is connected to the other output node. Furthermore, B+A*0.5 is input to one setting node as a calculation formula for correcting MSmL input from input unit B by EmoJ input from input unit A, and B+A*0.5 is input to the other setting node as a calculation formula for correcting MSmR input from input unit B by EmoJ input from input unit A. By performing a decision operation in this state, setting data for correcting the facial expression parameters MSmL, MSmR using the emotion value EmoJ is created.
[0151] FIG. 23 is a diagram for explaining a procedure for setting the setting data used in the MouthFrown correction process and the BrowDown correction process.
[0152] When creating setting data to be used in the MouthFrown correction process and the BrowDown correction process, five input nodes, four setting nodes, and four output nodes are arranged on the setting screen as shown in FIG.
[0153] Next, EmoA to be used for correction, and MFrL, MFrR, BDoL, and BDoR to be corrected are input to the five input nodes. Also, two input parts A and B are set to the first setting node, the input node of EmoA is connected to the input part A of the first setting node, and the input node of MFrL is connected to the input part B of the first setting node. Also, two input parts A and B are set to the second setting node, the input node of EmoA is connected to the input part A of the second setting node, and the input node of MFrR is connected to the input part B of the second setting node. Also, two input parts A and B are set to the third setting node, the input node of EmoA is connected to the input part A of the first setting node, and the input node of BDoL is connected to the input part B of the third setting node. Also, two input parts A and B are set to the fourth setting node, the input node of EmoA is connected to the input part A of the fourth setting node, and the input node of BDoR is connected to the input part B of the fourth setting node. Also, MFrLN is set to be output to the first output node, the output section of the first setting node is connected to the first output node, MFrRN is set to be output to the second output node, the output section of the second setting node is connected to the second output node, BDoLN is set to be output to the third output node, the output section of the third setting node is connected to the third output node, BDoRN is set to be output to the fourth output node, and the output section of the fourth setting node is connected to the fourth output node. Furthermore, B+A*(-1)*0.5 is input to the first setting node as a calculation formula for correcting MFrL input from input unit B by EmoA input from input unit A, B+A*(-1)*0.5 is input to the second setting node as a calculation formula for correcting MFrR input from input unit B by EmoA input from input unit A, B+A*(-1)*0.5 is input to the third setting node as a calculation formula for correcting BDoL input from input unit B by EmoA input from input unit A, and B+A*(-1)*0.5 is input to the fourth setting node as a calculation formula for correcting BDoR input from input unit B by EmoA input from input unit A. By performing a decision operation in this state, setting data for correcting the facial expression parameters MFrL, MFrR, BDoL, and BDoR using the emotion value EmoA is created.
[0154] In this way, in the correction content setting program, input nodes, setting nodes, and output nodes are arranged on a setting screen, the parameters to be corrected and the types of parameters to be used for the correction are input into the input nodes and linked to the setting nodes, the types of parameters to be output are input into the output nodes and linked to the setting nodes, and the correction contents are set in the setting nodes using calculation formulas and conditional expressions that use the parameters input from the input nodes, thereby making it possible to easily set the setting data to be used in the parameter correction program.
[0155] In addition, the correction content setting program makes it possible to read out setting data that has been created once, and by reading out the setting data, the input nodes, setting nodes, and output nodes based on the setting data and their setting contents are displayed. It is possible to change the connections between nodes, change the types of parameters set in the input nodes and output nodes, and change the calculation formulas and conditional formulas set in the setting nodes, and then overwrite and save or create new setting data, making it possible to easily edit existing setting data and easily create new setting data based on existing setting data.
[0156] [Image generation program] As shown in Fig. 3, the image generation program is a program that executes a process of generating an avatar image based on the facial expression parameters N and extension parameters from the parameter correction program, the motion parameters from the motion tracking program, the voice parameters from the voice parameters, and the input parameters from the input detection program, and outputs the generated avatar image to the delivery program, and includes a face image generation process for generating a face image, a face image correction process for correcting the face image, a costume image generation process for generating a costume image, and an ornament image generation process for generating an ornament image. The image generation program, like the parameter correction program, is executed for each frame at intervals according to the frame rate of the video delivered by the delivery program. For example, when the frame rate is 60 fps, the image generation program is executed 60 times per second.
[0157] FIG. 24 is a flow chart showing the control contents of the image generating program.
[0158] As shown in Fig. 24, the image generation program first imports various parameters (Sc1). In step Sc1, the facial expression parameter N and extended parameters output from the parameter correction program, the motion parameters from the motion tracking program, the voice parameters from the voice parameters, and the input parameters from the input detection program are imported.
[0159] Next, a 3D model of the face is created based on the facial expression parameter N captured in step Sc1 (Sc2). In step Sc2, a plurality of 3D models (face data) preset to correspond to the facial expression parameter N are used and blended (blend shape) according to the value of the corresponding facial expression parameter N to create the model. In addition, some parts are created by setting the positions and angles of the joints between a plurality of bones according to the value of the facial expression parameter N. In addition, the 3D model created by blend shape and the 3D model created by bones may be blended to complement each other.
[0160] Next, the motion parameters captured in step Sc1 are used to determine how many degrees the camera 108 is tilted horizontally and vertically with respect to the front of the face, and the contour, part positions, and shape of the 3D model of the face created in step Sc2 are corrected with the correction amount according to the determined photographing direction. As a result, for example, as shown in Fig. 25(a)(b), the same 3D model is not used when photographing the face from the front and when photographing it from an oblique direction, but a 3D model in which the contour, mouth, nose, and eye positions and shapes are corrected according to the photographing direction is used. The correction amount according to the photographing direction is calculated from the maximum correction amount in the left-right and up-down directions that is predetermined for the contour and the part to be corrected, and the horizontal and vertical angles of the photographing direction, and as shown in Fig. 26(a), the correction amount in the left-right or up-down direction gradually increases as the horizontal or vertical angle of the photographing direction increases. Alternatively, an intermediate correction amount may be set corresponding to an intermediate angle smaller than the maximum angle. In this case, as shown in FIG. 26(b), by gradually increasing the correction amount in the left / right or up / down direction so as to form a curve that passes through the intermediate correction amount at the intermediate angle, the correction amount does not change suddenly due to a change in the shooting angle.
[0161] Next, a 3D model of the body and hand (fingers) is created based on the motion parameters captured in step Sc1 (Sc4). In step Sc4, multiple 3D models (body data) that are preset to correspond to the motion parameters are used to create a 3D model of the body and hand (fingers) in the corresponding pose.
[0162] Next, a 3D model of the costume (including accessories such as hairstyle and accessories) is created (Sc5) based on the motion parameters captured in step Sc1. In step Sc5, multiple 3D models (clothing data) that are preset to correspond to the motion parameters are used to create a 3D model of the costume that matches the body and hands (fingers) in the corresponding pose.
[0163] Next, we determine whether the 3D model of the clothing created in step Sc5 overlaps with the forbidden area (Sc6). The forbidden area includes the overlapping area where the clothing overlaps with the face and the facial expression becomes invisible, as well as the surrounding area, and includes areas other than the overlapping area of the clothing and the face.
[0164] If it is determined in step Sc6 that the 3D model of the costume does not overlap with the prohibited area, the process proceeds to step Sc9. If it is determined that the 3D model of the costume overlaps with the prohibited area, the amount of deformation or movement of the costume is calculated based on the distance from the 3D model of the costume to the overlap area of the costume and the face (Sc7), and the shape or position of the relevant part of the costume is corrected based on the amount of deformation or movement calculated in step Sc7 (Sc8).
[0165] As a result, for example, when the shooting direction is in front of the face as shown in Fig. 27(a), even if the clothes do not overlap with the face, and even if the shooting direction changes and a part of the clothes overlaps with the face and hides the facial expression as shown in Fig. 27(b), the part of the clothes will not deform and hide the facial expression. Also, when the shooting direction is in front of the face as shown in Fig. 28(a), even if the accessories do not overlap with the face, and even if the shooting direction changes and a part of the accessories overlaps with the face and hides the facial expression as shown in Fig. 28(b), the accessories will not move and hide the facial expression.
[0166] In addition, the amount of deformation or movement of the clothes after the 3D clothing model overlaps with the prohibited area is set to gradually increase as it approaches the overlap area of the face and clothing, and furthermore, the closer it is to the overlap area of the face and clothing, the greater the increase in the amount of deformation or movement of the clothes, as shown in Fig. 29. Therefore, even if the 3D clothing model overlaps with the prohibited area, if it is far from the overlap area of the face and clothing, the amount of deformation or movement is kept small, and the closer it is to the overlap area of the face and clothing, the greater the amount of deformation or movement of the clothes is, thereby preventing the face and clothing from overlapping.
[0167] Next, it is determined whether or not an additional condition for adding a decorative image is satisfied based on the extended parameters (emotion values) and voice parameters acquired in step Sc1 (Sc9). If it is determined in step Sc9 that the additional condition is not satisfied, the process proceeds to step Sc11. If it is determined that the additional condition is satisfied, a decorative image is set according to the satisfied additional condition (Sc10). The decorative image set in step Sc10 is held for a certain period of time and is cleared after the certain period has elapsed. Therefore, after the additional condition is satisfied, the decorative image is displayed and maintained for a certain period of time. If a new additional condition is satisfied, it is overwritten.
[0168] The additional conditions and the decorative images corresponding to the additional conditions are set in advance. For example, as shown in FIG. 30, the additional conditions are met when the magnitude of the emotion value, the command input by the input device 107 specifying the decorative image, the volume, and the pitch are equal to or greater than specified values. In this embodiment, if the emotion value Emo is 0.7 or greater or a joy (small) command is input, a decorative image of a cracker is set; if the emotion value Emo is 0.9 or greater or a joy (large) command is input, a decorative image of a heart mark is set in addition to a cracker; if the emotion value Emo is -0.8 or less or an anger command is input, a decorative image of an anger mark is set; and if the volume (Vo) or pitch (Pi) is 0.8 or greater, a decorative image of a speaker mark is set.
[0169] For this reason, for example, when the emotion value (Emo) is 0.9 or more or a joy (large) command is input and a happy situation is identified, decorative images of crackers and heart marks are displayed around the avatar image for a certain period of time as shown in Fig. 31, and when the emotion value (Emo) is less than -0.8 or an anger command is input and an angry situation is identified, decorative images of angry marks are displayed around the avatar image for a certain period of time as shown in Fig. 32. Also, when the volume (Vo) or pitch (Pi) is 0.8 or more and a situation in which a loud or high-pitched voice is spoken is identified, a decorative image of a speaker mark is displayed as shown in Fig. 32.
[0170] Next, the 3D model of the face created in step Sc2 and corrected in step Sc3, the 3D model of the body and hands (fingers) created in step Sc4, the 3D model of the costume created in step Sc5, and the decorative images set in Sc10 are arranged to generate avatar image data and output it to the distribution program (Sc11).
[0171] In this manner, in this embodiment, decorative images other than facial images are generated based on some facial expression parameters detected from facial image data captured by the camera 108, and the facial images detected by face tracking can be output as images with a variety of features.
[0172] In addition, in this embodiment, an emotion value is identified based on some facial expression parameters detected from facial image data captured by the camera 108, and a decorative image other than a facial image is generated based on the identified emotion value, so that the decorative image can be displayed reflecting the emotion identified from the facial expression of the facial image detected by face tracking.
[0173] Furthermore, in this embodiment, in addition to some facial expression parameters detected from face image data captured by the camera 108, the emotion value is determined based on parameters other than the facial expression parameters, so that a decorative image can be displayed that reflects emotions more accurately.
[0174] In this embodiment, the emotion value is determined based on both some of the facial expression parameters and parameters other than the facial expression parameters, but it is also possible to determine the emotion value based on only some of the facial expression parameters, or to determine the emotion value based on only parameters other than the facial expression parameters. Even in such a configuration, it is possible to display a decorative image that reflects the determined emotion.
[0175] Furthermore, in this embodiment, when the emotion value exceeds a certain threshold, a corresponding decorative image is generated. With this configuration, the decorative image can be displayed with an appropriate frequency.
[0176] In this embodiment, a decorative image other than a facial image is generated based on some facial expression parameters detected from facial image data captured by camera 108, but it is also possible to change the design and color of the costume, the composition and color tone of the background image, the shape, size, color of accessories, etc. based on some of the parameters detected from the facial image data and the emotional values identified from these parameters.In such a configuration, it is possible to make changes to things other than the facial image by reflecting parameters identified from the facial expression of the facial image detected by face tracking.
[0177] In this embodiment, if a part of the clothing is positioned in a way that prevents the facial expression in the facial image from being visible depending on the shooting direction of camera 108, the part of the clothing is changed to a shape that allows the facial expression in the facial image to be visible, thereby preventing the facial expression in the facial image from being hidden by part of the clothing due to a change in the shooting direction of the subject by camera 108.
[0178] Furthermore, in this embodiment, even if the user is not in a position where the facial expression of the facial image is not visible, when the user approaches a position where the facial expression of the facial image is not visible (prohibited area), part of the clothing is deformed into a shape that makes the facial expression of the facial image visible, so that the part of the clothing can be deformed into a shape that makes the facial expression of the facial image visible even before the facial expression of the facial image is hidden.
[0179] Furthermore, in this embodiment, depending on the distance to the position where the facial expression of the facial image cannot be seen, a part of the costume is gradually deformed into a shape that allows the facial expression of the facial image to be seen, so that the part of the costume can be naturally deformed into a shape that allows the facial expression of the facial image to be seen.
[0180] In this embodiment, even if the subject is not in a position where the facial expression in the facial image is not visible, when the subject approaches a position where the facial expression in the facial image is not visible (prohibited area), part of the clothing is deformed into a shape that makes the facial expression in the facial image visible. However, it is also possible to deform part of the clothing into a shape that makes the facial expression in the facial image visible when the subject reaches a position where the facial expression in the facial image is not visible. Even with this configuration, it is possible to prevent the facial expression in the facial image from being obscured by part of the clothing due to a change in the shooting direction of the subject by camera 108.
[0181] In addition, in this embodiment, a part of the costume is gradually deformed into a shape that makes the facial expression in the facial image visible depending on the distance to the position where the facial expression in the facial image cannot be seen, but it is also possible to configure the costume to deform into a shape that makes the facial expression in the facial image visible when a position where the facial expression in the facial image cannot be seen (prohibited area) is reached.
[0182] In this embodiment, if the attachment is positioned so that the facial expression in the facial image cannot be seen depending on the shooting direction of the camera 108, the attachment is moved to a position where the facial expression in the facial image can be seen, thereby preventing the facial expression in the facial image from being hidden by the attachment due to a change in the shooting direction of the subject by the camera 108.
[0183] Furthermore, in this embodiment, even if the position is not one in which the facial expression of the facial image is not visible, when the user approaches a position in which the facial expression of the facial image is not visible (prohibited area), the accessory is moved toward a position in which the facial expression of the facial image is visible, so that the accessory can be moved toward a position in which the facial expression of the facial image is visible before the facial expression of the facial image is hidden.
[0184] Furthermore, in this embodiment, the accessory is gradually moved toward a position where the facial expression of the facial image can be viewed depending on the distance to the position where the facial expression of the facial image cannot be viewed, so that the accessory can be naturally moved toward a position where the facial expression of the facial image can be viewed.
[0185] In this embodiment, even if the position is not one in which the facial expression of the facial image is not visible, when the subject approaches a position in which the facial expression of the facial image is not visible (prohibited area), the accessory is moved toward a position in which the facial expression of the facial image is visible. However, it is also possible to move the accessory to a position in which the facial expression of the facial image is visible when the subject reaches a position in which the facial expression of the facial image is not visible. Even with this configuration, it is possible to prevent the facial expression of the facial image from being obscured by the accessory due to a change in the shooting direction of the subject by camera 108.
[0186] In addition, in this embodiment, the attachment is gradually moved toward a position where the facial expression of the facial image can be seen depending on the distance to the position where the facial expression of the facial image cannot be seen, but it may also be configured to move the attachment toward a position where the facial expression of the facial image can be seen when it reaches a position where the facial expression of the facial image cannot be seen (prohibited area).
[0187] In this embodiment, the shape or position of the features that make up the facial image is corrected according to the shooting direction of the camera 108, so that the facial expression of the generated facial image can be prevented from becoming unnatural even if the shooting direction changes.
[0188] Furthermore, in this embodiment, the contours of the face image are corrected according to the shooting direction of the camera 108, so that the contours of the generated face image can be prevented from becoming unnatural even if the shooting direction changes.
[0189] In addition, in this embodiment, the position and shape of the mouth constituting the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, it is possible to prevent the position and shape of the mouth constituting the generated facial image from becoming unnatural.
[0190] In addition, in this embodiment, the position and shape of the nose in the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, it is possible to prevent the position and shape of the nose in the generated facial image from becoming unnatural.
[0191] In addition, in this embodiment, the position and shape of the eyes that make up the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, it is possible to prevent the position and shape of the eyes that make up the generated facial image from becoming unnatural.
[0192] In addition, in this embodiment, the amount of correction for the shape or position of the features that make up the facial image is gradually increased depending on the shooting direction of the camera 108, so that the shape or position of the features that make up the facial image does not suddenly change.
[0193] Although an embodiment of the present invention has been described above with reference to the drawings, the present invention is not limited to this embodiment, and it goes without saying that modifications and additions that do not deviate from the spirit of the present invention are also included in the present invention.
[0194] For example, in the above embodiment, an example was described in which an avatar image generated by an image generation program is used for video distribution, but the use of the avatar image generated by an image generation program is arbitrary, and the avatar image may be recorded as an archive or may be used for producing animations, etc.
[0195] In addition, in the above embodiment, an example was described in which the facial expression parameters output from the face tracking program are corrected by a parameter correction program, and the image generation program generates an avatar image using the corrected facial expression parameters. However, a configuration in which the image generation program directly uses the parameters output from the face tracking program to generate an image without using a parameter correction program is also possible. [Explanation of symbols]
[0196] 1 Computer terminal 101 Processor 102 Memory 103 Storage 104 Communication Interface 105 Display device 106 Speaker 107 Input Devices 108 Camera 109 Mike 110 Vital Signs Measuring Instrument 111 Data Bus
Claims
1. A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A method for generating an image using a computer, comprising: an emotion value identification step of identifying an emotion value that varies based on the parameters of at least a portion of the parts detected in the parameter detection step; a decorative image generating step of generating a decorative image other than a face image based on the emotion value specified in the emotion value specifying step, In the image output step, the decorative image generated in the decorative image generating step is output together with the face image.
13. An image generating method comprising:
2. In the emotion value identification step, a first value indicating happiness and a second value indicating anger are identified based on parameters of the partial part, and the emotion value indicating happiness is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. The image generating method according to claim 1 .
3. In the emotion value determination step, the first value and the second value are determined based on a parameter indicating a degree of movement of the corners of the mouth. The image generating method according to claim 2.
4. In the decorative image generating step, when the emotion value specified in the emotion value specifying step exceeds a threshold, the decorative image is generated. The image generating method according to any one of claims 1 to 3.
5. A face image input step of inputting a photographed face image; a parameter detection step of detecting parameters indicating states of a plurality of features from the face image inputted by the face image input step; A facial image generating step of generating a facial image having a facial expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generating program for causing a computer to execute an emotion value identification step of identifying an emotion value that varies based on the parameters of at least a portion of the parts detected in the parameter detection step; a decorative image generating step of generating a decorative image other than a face image based on the emotion value specified in the emotion value specifying step; In the image output step, the decorative image generated in the decorative image generating step is output together with the face image.
13. An image generating program comprising:
Citation Information
Patent Citations
Content distribution server, content distribution system, content distribution method, and program
JP2019186797A
JPP7145556B
JPP7339420B