Image generation method and image generation program
The method ensures facial expressions in avatars remain visible by transforming or repositioning a specific image around the face image to counteract hiding issues caused by shooting direction and face orientation variations.
Patent Information
- Application Number
- JP2024089021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2044-05-31
AI Technical Summary
Conventional face-tracked technologies for reflecting facial expressions in avatars often result in the facial expression being hidden due to variations in shooting direction and face orientation.
The method involves generating a specific image around the face image that is transformed or moved to ensure the facial expression remains visible, regardless of the shooting direction, by deforming or repositioning the image to maintain visibility of the facial expression.
Prevents the facial expression from being hidden by adjusting the specific image's shape or position to ensure it remains visible, enhancing the visibility of the facial expression in various shooting directions.
Smart Images

Figure 2025181192000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation method and an image generation program for generating a facial image with an expression corresponding to an input facial image. [Background technology]
[0002] Conventionally, a technology has been proposed that, for example, uses face tracking to capture a face photographed with a smartphone, allowing the facial expression of the captured face to be reflected in the expression of an avatar, such as a character created using CG (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-186797 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in conventional face-tracked technology that reflects facial expressions in the expressions of an avatar, when a facial image and its accompanying images are output, there is a problem in that the facial expression in the output facial image may be hidden depending on the shooting direction and the orientation of the face being tracked.
[0005] An object of the present invention is to provide an image generation method and an image generation program that prevent the expression of a facial image output by face tracking from being hidden by another image. [Means for solving the problem]
[0006] The image generation method of means 1 is as follows: a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that makes the facial expression of the facial image visible, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0007] The image generation method of means 2 is the image generation method according to means 1, In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image is not visually recognizable, when the specific image approaches a position where the facial expression of the facial image is not visually recognizable, the specific image is deformed into a shape where the facial expression of the facial image is visually recognizable. It is characterized by the following. According to this feature, the specific image can be transformed into a shape that allows the facial expression of the facial image to be visually recognized before the facial expression of the facial image is hidden.
[0008] The image generation method of means 3 is the image generation method according to means 2, In the specific image generating step, the specific image is gradually deformed to a shape in which the facial expression of the facial image is visible, depending on the distance to a position where the facial expression of the facial image is not visible. It is characterized by the following. According to this feature, the specific image can be naturally transformed into a shape that allows the facial expression of the facial image to be visually recognized.
[0009] The image generation method of means 4 is a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0010] The image generating method of means 5 is the image generating method according to means 4, In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image cannot be visually recognized, when the specific image approaches a position where the facial expression of the facial image cannot be visually recognized, the specific image is moved toward a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, the position of the specific image can be moved toward a position where the expression of the facial image can be seen before the expression of the facial image is hidden.
[0011] The image generating method of means 6 is the image generating method according to means 5, In the specific image generating step, the position of the specific image is gradually moved toward a position where the facial expression of the facial image can be visually recognized, depending on the distance to the position where the facial expression of the facial image cannot be visually recognized. It is characterized by the following. According to this feature, the position of the specific image can be naturally moved toward a position where the expression of the facial image can be visually recognized.
[0012] The image generating program of means 7 comprises: a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that makes the facial expression of the facial image visible, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0013] The image generating program of means 8 comprises: a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0014] Furthermore, the present invention may have only the invention-specific matters set forth in the claims of the present invention, or may have the invention-specific matters set forth in the claims of the present invention as well as configurations other than the invention-specific matters. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 2 is a block diagram showing an example of the configuration of a computer terminal used in an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing the configuration of a program installed in a computer terminal used in an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing the relationship between programs installed in a computer terminal in an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram showing the configuration of parameters used for image generation in an embodiment of the present invention. [Figure 5] 10 is a flowchart showing the control content of a parameter correction program executed by a computer terminal in the embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing the contents of emotion value calculation processing performed by a computer terminal in an embodiment of the present invention. [Figure 7] FIG. 10 is a diagram showing the contents of a MouthSmile correction process performed by a computer terminal in an embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing the contents of a MouthFrown correction process performed by a computer terminal in an embodiment of the present invention. [Figure 9] 10A and 10B are diagrams illustrating the content of a BrowDown correction process performed by a computer terminal in an embodiment of the present invention. [Figure 10] 10A and 10B are diagrams showing the contents of EyeBlink switching processing performed by a computer terminal in an embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing the contents of a configuration correction process performed by a computer terminal in an embodiment of the present invention. [Figure 12] 10A to 10C are diagrams illustrating examples of facial expressions detected by a face tracking program in an embodiment of the present invention. [Figure 13] 10A to 10C are diagrams illustrating examples of changes in facial expressions according to emotion values in an embodiment of the present invention. [Figure 14] 10A to 10C are diagrams showing examples of changes in facial expression in response to switching of facial expression parameters in an embodiment of the present invention. [Figure 15] 10 is a flowchart showing the control content of a correction content setting process performed by a computer terminal in the embodiment of the present invention. [Figure 16] FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 17] FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 18] FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 19] FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 20] FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 21]FIG. 10 is a diagram illustrating the procedure for setting setting data used in emotion value setting processing in an embodiment of the present invention. [Figure 22] 10A to 10C are diagrams for explaining a procedure for setting setting data used in the MouthSmile correction process in an embodiment of the present invention. [Figure 23] 10A and 10B are diagrams for explaining a procedure for setting setting data used in MouthFrown correction processing and BrowDown correction processing in an embodiment of the present invention. [Figure 24] 10 is a flowchart showing the control of an image generation program executed by a computer terminal in the embodiment of the present invention. [Figure 25] 10A to 10C are diagrams illustrating examples of changes in facial expressions resulting from correction according to the shooting direction in an embodiment of the present invention. [Figure 26] 10A and 10B are diagrams illustrating changes in correction amount depending on the shooting direction in an embodiment of the present invention. [Figure 27] 10A to 10C are diagrams illustrating examples of how clothing changes depending on the shooting direction in an embodiment of the present invention. [Figure 28] 10A to 10C are diagrams illustrating examples of changes in appendages depending on the shooting direction in an embodiment of the present invention. [Figure 29] 10A and 10B are diagrams illustrating the amount of change in clothing or the amount of movement of accessories depending on the shooting direction in an embodiment of the present invention. [Figure 30] 10A and 10B are diagrams illustrating the relationship between additional conditions and decorative images corresponding to the additional conditions in an embodiment of the present invention. [Figure 31] 10A and 10B are diagrams illustrating examples of decorative image display in an embodiment of the present invention. [Figure 32] 10A and 10B are diagrams illustrating examples of decorative image display in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] [Form 1] The image generation method of form 1-1 is a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: an information input step of inputting information other than a facial image; a parameter correction step of correcting the parameters detected in the parameter detection step based on information other than the face image input in the information input step; further comprising In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, a facial image with an expression according to parameters corrected based on information other than the photographed facial image is generated, so that the facial image detected by face tracking can be output as a varied image.
[0017] The image generation method of embodiment 1-2 is the image generation method according to embodiment 1-1, further comprising an emotion value determination step of determining an emotion value based on information other than the facial image; In the parameter correction step, the parameters detected in the parameter detection step are corrected based on the emotion value identified in the emotion value identification step. It is characterized by the following. According to this feature, it is possible to generate a facial image with an expression based on an emotional value specified from information other than the facial image.
[0018] The image generation method of embodiment 1-3 is the image generation method according to embodiment 1-2, In the emotion value determination step, the emotion value is determined based on information other than the face image and parameters of some of the features detected in the parameter detection step. It is characterized by the following. This feature allows for more accurate identification of emotion values.
[0019] The image generation method of embodiment 1-4 is the image generation method according to embodiment 1-3, In the parameter correction step, if the emotion value identified in the emotion value identification step exceeds a threshold, the parameter detected in the parameter detection step is corrected. It is characterized by the following. This feature allows for more variation in facial expression.
[0020] The image generation method of embodiment 1-5 is the image generation method according to any one of embodiments 1-1 to 1-4, The information other than the face image includes information on at least one of voice, vital signs, and operation input. It is characterized by the following. According to this feature, a facial image with an expression based on information on at least one of voice, vital signs, and operational input can be generated.
[0021] The image generation method of embodiment 1-6 is the image generation method according to any one of embodiments 1-1 to 1-5, further comprising a setting value receiving step of receiving setting values related to the amount of change of the plurality of parts; The information other than the face image includes the setting value set in the setting value receiving step. It is characterized by the following. According to this feature, it is possible to generate a facial image with an expression according to parameters corrected based on set values relating to the amount of change in a plurality of facial features.
[0022] A correction content setting method of form 1-7 is a correction content setting method using a computer to set correction content of the parameters used in the image generating method according to any one of forms 1-1 to 1-6, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting correction contents; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting content; Contains It is characterized by the following. According to this feature, the correction content of the parameters can be easily set by simply connecting nodes and setting the correction content.
[0023] The image generation program of form 1-8 is a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute an information input step of inputting information other than a facial image; a parameter correction step of correcting the parameters detected in the parameter detection step based on information other than the face image input in the information input step; Then run In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, a facial image with an expression according to parameters corrected based on information other than the photographed facial image is generated, so that the facial image detected by face tracking can be output as a varied image.
[0024] A correction content setting program of form 1-9 is a correction content setting program for setting, by a computer, the correction content of the parameters used in the image generation program of form 1-8, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting correction contents; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting content; Run It is characterized by the following. According to this feature, the correction content of the parameters can be easily set by simply connecting nodes and setting the correction content.
[0025] [Form 2] The image generation method of form 2-1 is as follows: a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: a parameter correction step of correcting parameters of other parts based on the parameters of the part of parts detected in the parameter detection step; In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, parameters of some features are corrected based on the parameters of other features detected from the captured facial image, and a facial image with an expression corresponding to the corrected parameters is generated, so that the facial image detected by face tracking can be output as a varied image.
[0026] The image generation method of embodiment 2-2 is the image generation method according to embodiment 2-1, further comprising an emotion value determination step of determining an emotion value based on the parameters of the part of the part; In the parameter correction step, the parameters of the other parts detected in the parameter detection step are corrected based on the emotion value identified in the emotion value identification step. It is characterized by the following. According to this feature, it is possible to generate a facial image with an expression based on an emotional value that is specified based on the parameters of some facial features.
[0027] The image generation method of aspect 2-3 is the image generation method according to aspect 2-2, In the emotion value identification step, a first value indicating joy and a second value indicating anger are identified based on the parameters of the part of the body parts, and the emotion value indicating joy is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. It is characterized by the following. According to this feature, it is possible to generate a facial image with an expression based on an emotional value that reflects both joy and anger identified from the facial image.
[0028] The image generation method of aspect 2-4 is the image generation method according to aspect 2-3, In the emotion value identification step, the first value and the second value are identified based on a parameter indicating the degree of movement of the corners of the mouth. It is characterized by the following. According to this feature, emotions such as joy and anger can be identified from the degree of movement of the corners of the mouth, which easily reflects emotions.
[0029] The image generation method of embodiment 2-5 is the image generation method according to any one of embodiments 2-2 to 2-4, In the parameter correction step, when the emotion value identified in the emotion value identification step exceeds a threshold, the parameters of the other parts are corrected. It is characterized by the following. This feature allows for more variation in facial expression.
[0030] A correction content setting method of form 2-6 is a correction content setting method using a computer to set correction content of the parameters used in the image generating method according to any one of forms 2-1 to 2-5, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting correction contents; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting content; Contains It is characterized by the following. According to this feature, the correction content of the parameters can be easily set by simply connecting nodes and setting the correction content.
[0031] The image generation program of form 2-7 is a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a parameter correction step of correcting parameters of other parts based on the parameters of the part of the parts detected in the parameter detection step; In the face image generating step, a face image having an expression corresponding to the parameters corrected in the parameter correcting step is generated. It is characterized by the following. According to this feature, parameters of some features are corrected based on the parameters of other features detected from the captured facial image, and a facial image with an expression corresponding to the corrected parameters is generated, so that the facial image detected by face tracking can be output as a varied image.
[0032] A correction content setting program of form 2-8 is a correction content setting program for setting, by a computer, the correction content of the parameters used in the image generation program of form 2-7, a first node setting step of setting a first node that specifies a parameter to be corrected; a second node setting step of setting a second node that specifies parameters to be used for correction; a third node setting step of setting a third node for setting correction contents; a connecting step of connecting the first node and the second node with the third node; a correction content setting step of setting, in the third node, correction content of the parameters of the first node based on the parameters of the second node; creating setting data based on the setting content; It is characterized by the following. According to this feature, the correction content of the parameters can be easily set by simply connecting nodes and setting the correction content.
[0033] [Form 3] The image generation method of form 3-1 is a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: The method further includes a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step, In the image output step, the decorative image generated in the decorative image generation step is output together with the face image. It is characterized by the following. According to this feature, decorative images other than facial images are generated based on parameters detected from the captured facial image, so that the facial images detected by face tracking can be output as varied images.
[0034] The image generation method of embodiment 3-2 is the image generation method according to embodiment 3-1, further comprising an emotion value determining step of determining an emotion value based on parameters of at least some of the parts; In the decorative image generating step, the decorative image is generated based on the emotion value identified in the emotion value identifying step. It is characterized by the following. According to this feature, a decorative image can be generated in accordance with parameters corrected based on emotion values determined based on parameters of at least some of the features.
[0035] The image generation method of embodiment 3-3 is the image generation method according to embodiment 3-2, In the emotion value identification step, a first value indicating joy and a second value indicating anger are identified based on the parameters of the part of the body parts, and the emotion value indicating joy is identified when the first value is greater than the second value, and the emotion value indicating anger is identified when the second value is greater than the first value. It is characterized by the following. According to this feature, it is possible to generate a decorative image based on an emotional value that reflects both joy and anger identified from a facial image.
[0036] The image generation method of embodiment 3-4 is the image generation method according to embodiment 3-3, In the emotion value identification step, the first value and the second value are identified based on a parameter indicating the degree of movement of the corners of the mouth. It is characterized by the following. According to this feature, emotions such as joy and anger can be identified from the degree of movement of the corners of the mouth, which easily reflects emotions.
[0037] The image generation method of embodiment 3-5 is the image generation method according to any one of embodiments 3-2 to 3-4, In the decorative image generating step, the decorative image is generated when the emotion value specified in the emotion value specifying step exceeds a threshold value. It is characterized by the following. According to this feature, decorative images can be generated at an appropriate frequency.
[0038] The image generation program of form 3-6 is a facial image input step of inputting a photographed facial image; a parameter detection step of detecting parameters indicating the states of a plurality of features from the face image inputted by the face image input step; a facial image generating step of generating a facial image having an expression according to a plurality of parameters; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a decorative image generating step of generating a decorative image other than a face image based on the parameters detected in the parameter detecting step; In the image output step, the decorative image generated in the decorative image generation step is output together with the face image. It is characterized by the following. According to this feature, decorative images other than facial images are generated based on parameters detected from the captured facial image, so that the facial images detected by face tracking can be output as varied images.
[0039] [Form 4] The image generation method of form 4-1 is as follows: a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that makes the facial expression of the facial image visible, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0040] The image generation method of embodiment 4-2 is the image generation method according to embodiment 4-1, In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image is not visually recognizable, when the specific image approaches a position where the facial expression of the facial image is not visually recognizable, the specific image is deformed into a shape where the facial expression of the facial image is visually recognizable. It is characterized by the following. According to this feature, the specific image can be transformed into a shape that allows the facial expression of the facial image to be visually recognized before the facial expression of the facial image is hidden.
[0041] The image generation method of embodiment 4-3 is the image generation method according to embodiment 4-2, In the specific image generating step, the specific image is gradually deformed to a shape in which the facial expression of the facial image is visible, depending on the distance to a position where the facial expression of the facial image is not visible. It is characterized by the following. According to this feature, the specific image can be naturally transformed into a shape that allows the facial expression of the facial image to be visually recognized.
[0042] The image generation method of form 4-4 is as follows: a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0043] The image generation method of embodiment 4-5 is the image generation method according to embodiment 4-4, In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image cannot be visually recognized, when the specific image approaches a position where the facial expression of the facial image cannot be visually recognized, the specific image is moved toward a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, the position of the specific image can be moved toward a position where the expression of the facial image can be seen before the expression of the facial image is hidden.
[0044] The image generation method of embodiment 4-6 is the image generation method according to embodiment 4-5, In the specific image generating step, the position of the specific image is gradually moved toward a position where the facial expression of the facial image can be visually recognized, depending on the distance to the position where the facial expression of the facial image cannot be visually recognized. It is characterized by the following. According to this feature, the position of the specific image can be naturally moved toward a position where the expression of the facial image can be visually recognized.
[0045] The image generation program of form 4-7 is a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the specific image is transformed into a shape that makes the facial expression of the facial image visible, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0046] The image generation program of form 4-8 is a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. It is characterized by the following. According to this feature, when the position of the specific image is such that the facial expression of the facial image cannot be seen depending on the shooting direction, the position of the specific image is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the specific image.
[0047] [Form 5] The image generation method of form 5-1 is a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: In the facial image generating step, the shape or position of the features constituting the facial image is corrected according to the photographing direction. It is characterized by the following. According to this feature, by correcting the shape or position of the features that make up the facial image depending on the shooting direction, it is possible to prevent the facial expression of the facial image from becoming unnatural.
[0048] The image generation method of embodiment 5-2 is the image generation method according to embodiment 5-1, In the face image generating step, the contour of the face image is corrected according to the photographing direction. It is characterized by the following. This feature makes it possible to prevent the contours of a face image from becoming unnatural even if the photographing direction changes.
[0049] The image generation method of embodiment 5-3 is the image generation method according to embodiment 5-1 or 5-2, In the facial image generating step, the position or shape of the mouth constituting the facial image is corrected according to the photographing direction. It is characterized by the following. According to this feature, even if the photographing direction changes, it is possible to prevent the position or shape of the mouth constituting the face image from becoming unnatural.
[0050] The image generation method of embodiment 5-4 is the image generation method according to any one of embodiments 5-1 to 5-3, In the facial image generating step, the position or shape of the eyes constituting the facial image is corrected according to the photographing direction. It is characterized by the following. According to this feature, even if the photographing direction changes, it is possible to prevent the position or shape of the eyes constituting the face image from becoming unnatural.
[0051] The image generation method of embodiment 5-5 is the image generation method according to any one of embodiments 5-1 to 5-4, In the facial image generating step, the correction amount of the shape or position of the features constituting the facial image is gradually increased according to the photographing direction. It is characterized by the following. According to this feature, the amount of correction for the shape or position of the features that make up the face image increases gradually in response to changes in the photographing direction, so that the shape or position does not change suddenly.
[0052] The image generation program of form 5-6 is a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute In the facial image generating step, the shape or position of the features constituting the facial image is corrected according to the photographing direction. It is characterized by the following. According to this feature, by correcting the shape or position of the features that make up the facial image depending on the shooting direction, it is possible to prevent the facial expression of the facial image from becoming unnatural.
[0053] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings based on the embodiments. [Example]
[0054] [Computer Terminal] Fig. 1 is a block diagram showing an example of the configuration of a computer terminal 1 used in an embodiment of the present invention. The computer terminal 1 shown in Fig. 1 may be a desktop PC or workstation with separate display devices and input devices such as a keyboard and a mouse, a notebook PC with an integrated display device and input device, or a mobile terminal such as a tablet terminal or smartphone with a display device equipped with a touch panel or the like, but the following description will be given assuming a mobile terminal such as a tablet terminal or smartphone.
[0055] As shown in FIG. 1, the computer terminal 1 is equipped with a processor 101, a memory 102, and a storage 103 such as a hard disk or SSD, and the processor 101, memory, and storage 103 are connected via a data bus 111, and the processor 101 can execute various processes according to programs stored in the storage 103.
[0056] Furthermore, a display device 105 and a speaker 106 are connected to the data bus 111 via an output interface (not shown), and images and sounds can be output based on the processing executed by the processor 101 .
[0057] In addition, input devices such as an input device 107, a camera 108, a microphone 109, and a vital sign measuring instrument 110 are connected to the data bus 111 via an input interface (not shown), and information can be input via these input devices. The input device 107 includes a touch panel integrated with the display device 105, and commands can be input via the touch panel. Note that commands can also be input via an input device such as a keyboard or a mouse. The camera 108 is a 3D camera and can input data including depth data as captured image data. The camera 108 can be configured to include both an outer camera provided on the opposite side of the display device 105 and an inner camera provided on the display device 105 side, or can include only one of them. The microphone 109 can be a stereo microphone or a monaural microphone and can input ambient sounds, including the subject's voice. The vital sign measuring instrument 110 is a wearable device such as a smart band and can input vital sign data (heart rate, blood pressure, body temperature, etc.).
[0058] Furthermore, a communication interface 104 is connected to the data bus 111, and is configured to be able to communicate with other computer terminals and server computers via a local area network, the Internet, or other public network, either wired or wirelessly.
[0059] While this embodiment illustrates a configuration in which the image generation method and correction content setting method of the present invention are implemented using a single computer terminal, a configuration in which the image generation method and correction content setting method of the present invention are implemented using multiple computer terminals is also possible. Furthermore, while the computer terminal 1 of this embodiment is configured to include a display device 105 and a speaker 106 as output devices, it may also be configured without these output devices, with images and audio output from the output device of another computer terminal. Furthermore, the computer terminal 1 of this embodiment is configured to include input devices such as an input device 107 such as a touch panel, a camera 108, and a microphone 109, but it is sufficient that the computer terminal 1 is configured to include at least a touch panel, a keyboard, a mouse, and other input devices, and a camera capable of capturing images.
[0060] 2, the storage 103 of the computer terminal 1 stores various programs executed by the computer terminal 1, in addition to an operating system (OS) not shown. Specifically, a face tracking program, a motion tracking program, a voice detection program, an input detection program, a vital sign detection program, a parameter correction program, an image generation program, a correction content setting program, and a distribution program are stored, and, although not particularly shown, setting data used by the parameter correction program, configuration setting values, material data used by the image generation program, and the like are also stored.
[0061] 3, computer terminal 1 is configured such that an image generation program creates avatar image data with expressions based on the facial parameters and motion parameters based on image data output from camera 108, and the distribution program distributes videos using the avatar image data created by the image generation program. Furthermore, the facial parameters output by the face tracking program are not output directly to the image generation program, but are output to the image generation program via a parameter correction program, which corrects at least some of the facial parameters using some of the facial parameters and parameters other than the facial parameters (voice parameters, input parameters, vital parameters, and configuration parameters), and outputs the corrected facial parameters to the image generation program.
[0062] Next, the programs executed by the computer terminal 1 will be described.
[0063] As shown in FIG. 3, the face tracking program is a program that uses image data including depth data input from camera 108 to detect the state of multiple facial features constituting a facial image of a subject, and outputs facial parameters that quantify the degree of movement of each feature from the detected state. The facial parameters are parameters that can identify the facial expression of the subject. In this embodiment, as shown in FIG. 4, the facial parameters are parameters that quantify, in the range of 0.0 to 1.0, six types of movement degrees related to left eye movement, six types of movement degrees related to right eye movement, 27 types of movement degrees related to mouth and jaw movement, 10 types of movement degrees related to eyebrow, cheek, and nose movement, and one type of movement degree related to tongue movement. Note that the facial parameters are not limited to these exemplified ones, and more detailed types of facial parameters may be used, or fewer types of facial parameters may be used.
[0064] As shown in Fig. 3, the motion tracking program is a program that detects the body movement of a subject using image data including depth data input from the camera 108, and outputs motion parameters that identify the positions and directions of parts that make up the body. The motion parameters are parameters that can identify the body movement of a subject, and in this embodiment, as shown in Fig. 4, include a head pose that indicates the position coordinates and direction of the head, all finger joints that indicate the position coordinates and directions of all finger joints, and a hand pose that indicates the position coordinates and directions of the hand. Note that, although the motion parameters in this embodiment are configured not to include parameters regarding the movement of the entire body, they may also be configured to include parameters regarding the movement of the entire body.
[0065] The voice detection program is a program that detects the voice of the target person using voice data input from microphone 109 and outputs voice parameters that digitize the components of the detected voice, as shown in Fig. 2. The voice parameters are parameters that can identify the voice of the target person, and in this embodiment, as shown in Fig. 4, are parameters that digitize the voice volume and voice pitch in the range of 0.0 to 1.0.
[0066] The input detection program is a program that converts input data from the input device 107 into input parameters and outputs them, as shown in Fig. 3. The input parameters are parameters that can identify the input status from the input device 107, and in this embodiment, as shown in Fig. 4, are parameters that indicate the type of command input by touch operation corresponding to the command displayed on the display device 105. The command inputs include, for example, command inputs that specify emotions such as joy, anger, sadness, and pleasure in stages, command inputs that specify decorative images, and the like.
[0067] The vital sign detection program is a program that converts vital data (heart rate, blood pressure, body temperature, etc.) input from the vital sign measuring device 110 and outputs vital parameters, as shown in Fig. 3. The vital parameters are parameters that can identify the subject's vital signs, and in this embodiment, as shown in Fig. 4, the heart rate, blood pressure, and body temperature are each converted into a numerical value between 0.0 and 1.0.
[0068] [Parameter correction program] As shown in FIG. 3, the parameter correction program is a program that primarily corrects facial expression parameters created by the face tracking program, and includes a parameter correction process that corrects at least some of the facial expression parameters created by the face tracking program using some of the facial expression parameters, voice parameters from the voice detection program, input parameters from the input detection program, vital parameters from the vital detection program, and configuration parameters based on preset configurations (setting values), and a parameter creation process that creates new facial expression parameters based on input parameters. The parameter correction program is executed for each frame at intervals based on the frame rate of the video distributed by the distribution program. For example, if the frame rate is 60 fps, the parameter correction program is executed 60 times per second.
[0069] The config (setting value) is a value that can set the amount of change in the facial expression parameter, and can be set by displaying a configuration setting screen on the display device 105 and using the input device 107. The config (setting value) can be set for each type of facial expression parameter, and in this embodiment, as shown in FIG. 4, a numerical value between 0.0 and 1.0 can be set individually for each of the facial expression parameters related to eye movement, the facial expression parameters related to mouth and jaw movement, the facial expression parameters related to eyebrow movement, the facial expression parameters related to cheek and nose movement, and the facial expression parameters related to tongue movement, and the set numerical value is used as the config parameter. In this embodiment, the larger the numerical value set as the config, the greater the amount of change in the corresponding facial expression parameter and the greater the movement of the corresponding part.
[0070] FIG. 5 is a flowchart showing the control content of the parameter correction program.
[0071] 5, the parameter correction program first imports various parameters (Sa1). In step Sa1, the program imports facial expression parameters output from the face tracking program, voice parameters output from the voice detection program, input parameters output from the input detection program, vital parameters output from the vital detection program, and configuration parameters based on preset configurations (setting values).
[0072] Next, the setting data is referenced and an emotion value is calculated (Sa2) based on the various parameters acquired in step Sa1. The setting data is data set by a correction content setting program, and defines the type of parameters to be input, the correction content (calculation formula, output conditions, etc.) based on the input parameters, and the type of parameters to be output. The emotion value, which will be described in detail later, is a value calculated from the Joy value, which indicates the degree of joy, and the Anger value, which indicates the degree of anger, calculated from the Joy value and the Anger value, as well as voice parameters and vital parameters, and is a parameter that indicates the degree of joy or anger of the subject, as will be described in detail later.
[0073] Next, the setting data is referenced, and the facial expression parameters are corrected using the various parameters acquired in step Sa1 and the emotion value calculated in step Sa2 (Sa3). In step Sa3, it is sufficient that at least some of the facial expression parameters output from the face tracking program are corrected. Furthermore, the facial expression parameters may be corrected using parameters including emotion values, or may be corrected using parameters not including emotion values. Furthermore, the parameters used to correct the facial expression parameters may be only some of the facial expression parameters, or only parameters other than the facial expression parameters, or both some of the facial expression parameters and parameters other than the facial expression parameters. Furthermore, the parameters used to correct the facial expression parameters are not limited to parameters output from a face tracking program, but may also be parameters corrected using multiple parameters and then used to correct the facial expression parameters. The facial expression parameters to be corrected are not limited to those output from the face tracking program, but may also be further corrected using other parameters after correction.
[0074] Next, by referencing the setting data, new facial expression parameters other than those output from the face tracking program are created (Sa4) using the various parameters acquired in step Sa1 and the emotion values calculated in step Sa2. The parameters created in step Sa4 may be output separately from the facial expression parameters output from the face tracking program, or the facial expression parameters output from the face tracking program may be switched to the newly created parameters.
[0075] Next, the facial expression parameters corrected in step Sa3 and the facial expression parameters created in step Sa4 are output to the image generation program (Sa5). Note that in step Sa5, facial expression parameters output from the face tracking program and used without correction are also output. In step Sa5, it is also possible to output not only facial expression parameters but also parameters such as emotion values created in the process of correcting or creating new facial expression parameters as extended parameters to the image generation program. In this embodiment, emotion values are output as extended parameters. The facial expression parameters and extended parameters output to the image generation program in step Sa5 can be specified by a correction content setting process.
[0076] [Emotion value calculation process] Next, we will explain an example of a method for calculating emotion values using a parameter correction program. As shown in Figure 6, the emotion value calculation process uses BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthFunnel (MFu), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR) from the facial expression parameters output from the face tracking program, volume (Vo) and pitch (Pi) are used as voice parameters output from the voice detection program, and pulse (PR) is used as a vital parameter output from the vital detection program.
[0077] As shown in FIG. 12(a), BrowDownLeft (BDoL) is a parameter representing the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter representing the downward movement of the right eyebrow. BrowDownLeft (BDoL) and BrowDownRight (BDoR) are parameters whose numerical values increase with an angry facial expression. As shown in FIG. 12(b), CheekSquintLeft (CSqL) is a parameter representing the upward movement of the cheek around the left eye and below, and CheekSquintRight (CSqR) is a parameter representing the upward movement of the cheek around the right eye and below. CheekSquintLeft (CSqL) and CheekSquintRight (CSqR) are parameters whose numerical values increase with a happy facial expression. As shown in FIG. 12(c), MouthFunnel (MFu) is a parameter representing the contraction of both lips into an open shape. MouthFunnel (MFu) is a parameter whose numerical value increases with an angry facial expression. As shown in FIG. 12(d), MouthStretchLeft (MStL) is a parameter that represents the left movement of the left corner of the mouth, and MouthStretchRight (MStR) is a parameter that represents the right movement of the right corner of the mouth. MouthStretchLeft (MStL) and MouthStretchRight (MStR) are parameters that increase in value when the facial expression is angry. As shown in FIG. 12(e), MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters that increase in value when the facial expression is happy.
[0078] The emotion value calculation process first calculates (BDoL+BDoR) / 2, (CSqL+CSqR) / 2, (MStL+MStR) / 2, and (MSmL+MSmR) / 2, and then calculates BDo, CSq, MSt, and MSm as the average values of these left and right facial expression parameters.
[0079] Next, using BDo, MFu, and MSt, which have higher numerical values for angry facial expressions, ((BDo-CSq) x 2 + MSt + MFu) / 4 x (-1) is performed to calculate the Anger value, which indicates the degree of anger. In this case, a more accurate degree of anger can be calculated by subtracting not only BDo, MFu, and MSt, which have higher numerical values for angry facial expressions, but also CSq, which has higher numerical values for happy facial expressions, from BDo. The Anger value is a number ranging from 0 to -1, and the closer the number is to -1, the greater the degree of anger.
[0080] Next, using CSq and MSm, which have higher values for happy expressions, we calculate the Joy value, which indicates the degree of joy, by calculating (MSm+CSq) / 2. The Joy value is a value ranging from 0 to +1, and the closer the value is to +1, the greater the degree of joy.
[0081] Next, the Anger value and Joy value are used to perform Anger+Joy to calculate the emotion value AG, which is a numerical value in the range of 1 to -1.
[0082] Next, the voice parameters volume (Vo) and pitch (Pi) and the vital parameter pulse (PR) are converted into Vo1, Pi1, and PR1, which indicate values of 1 or 0. Volume (Vo), pitch (Pi), and pulse (PR) are all numerical values between 0 and 1. If the numerical values of volume (Vo) and pitch (Pi) are 0.8 or greater, they are converted into 1, and if they are less than 0.8, they are converted into 0. Furthermore, if the numerical value of pulse (PR) is 0.6 or greater, it is converted into 1, and if it is less than 0.6, it is converted into 0.
[0083] Next, if the emotion value AG exceeds 0, i.e., if the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, AG x 0.8 + (Vo1 + Pi1 + PR1) / 3 x 0.2 is calculated, and the calculated result, EmoJ indicating the degree of joy, is output as the emotion value. As a result, if the volume (Vo), pitch (Pi), and pulse (PR) are above a certain value and emotional intensity can be read, a larger value, i.e., EmoJ indicating a greater degree of joy, will be output than when the volume (Vo), pitch (Pi), and pulse (PR) are below a certain value.
[0084] On the other hand, if the emotion value AG is less than 0, i.e., if the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, then AG x 0.8 - (Vo1 + Pi1 + PR1) / 3 x 0.2 is calculated, and the calculated result, EmoA indicating the degree of anger, is output as the emotion value. As a result, if the volume (Vo), pitch (Pi), and pulse (PR) are above a certain value and emotional intensity can be read, a smaller value, i.e., an EmoA indicating a greater degree of anger, is output than when the volume (Vo), pitch (Pi), and pulse (PR) are below a certain value.
[0085] In this way, in the emotion value calculation process, the Anger value indicating the degree of anger is determined using facial expression parameters that increase in value for angry expressions, and the Joy value indicating the degree of joy is determined using parameters that increase in value for happy expressions; if the Joy value indicating the degree of joy is greater than the Anger value indicating the degree of anger, the emotion value EmoJ indicating the degree of joy is determined, and if the Anger value indicating the degree of anger is greater than the Joy value indicating the degree of joy, the emotion value EmoA indicating the degree of anger is determined.
[0086] Furthermore, when the volume (Vo), pitch (Pi), and pulse rate (PR) are equal to or greater than a certain value, an emotional value indicating a greater degree is identified as an emotional value EmoJ indicating a degree of joy or an emotional value EmoA indicating a degree of anger.
[0087] While this embodiment is configured to specify an emotional value using voice parameters and vital parameters in addition to facial expression parameters, it is also possible to specify an emotional value using only facial expression parameters, or to specify an emotional value using only parameters other than facial expression parameters such as voice parameters and vital parameters.Furthermore, it is also possible to specify an emotional value using parameters other than facial expression parameters, voice parameters, and vital parameters together with these parameters, or to specify an emotional value using only parameters other than facial expression parameters, voice parameters, and vital parameters.For example, when a command input specifying emotions such as joy, anger, sadness, or pleasure in stages is made by operating the input device 107, it is also possible to specify an emotional value using input parameters that specify the corresponding command input.
[0088] [Parameter correction processing] Next, an example of a method for correcting facial expression parameters using a parameter correction program will be described.
[0089] FIG. 7 is a diagram showing the contents of the MouthSmile correction process for correcting the facial expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) based on emotion values.
[0090] In the MouthSmile correction process, MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are corrected using the emotional value EmoJ, which indicates the degree of joy. Specifically, MSmL + EmoJ × 0.5 and MSmR + EmoJ × 0.5 are calculated, and the corrected facial expression parameters MSmLN and MSmRN corresponding to MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are output. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0091] The correction amounts for MSmLN and MSmRN increase as the emotional value EmoJ, which indicates the degree of joy, approaches +1, and decrease as the emotional value EmoJ, which indicates the degree of joy, approaches 0. MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth; the closer the emotional value EmoJ approaches +1, the more the mouth corners turn upward. For this reason, as shown in FIG. 13(a), compared to when the emotional value is 0, when the emotional value EmoJ is +1, for example, the corners of the mouth turn upward, and the closer the emotional value EmoJ is to +1, the more the facial expression of the output facial image can be made to look happier.
[0092] FIG. 8 is a diagram showing the contents of the MouthFrown correction process for correcting the facial expression parameters MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) based on emotion values.
[0093] In the MouthFrown correction process, MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) are corrected using the emotion value EmoA, which indicates the degree of anger. Specifically, MFrL + EmoA × (-1) × 0.5 and MFrR + EmoA × (-1) × 0.5 are performed, respectively, and the corrected facial expression parameters MFrLN and MFrRN corresponding to MouthFrownLeft (MFrL) and MouthFrownRight (MFrR) are output. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0094] The correction amounts for MFrLN and MFrRN become larger as the emotion value EmoA, which indicates the degree of anger, approaches -1, and smaller as the emotion value EmoA, which indicates the degree of anger, approaches 0. MouthFrownLeft (MFrL) is a parameter that represents the downward movement of the left corner of the mouth, and MouthFrownRight (MFrR) is a parameter that represents the downward movement of the right corner of the mouth; the closer the emotion value EmoA approaches -1, the more the mouth corners turn downward. For this reason, as shown in Figure 13(a), compared to when the emotion value is 0, when the emotion value EmoA is -1, for example, the corners of the mouth turn downward, and the closer the emotion value EmoA is to -1, the more the expression of the output facial image can be made to look angrier.
[0095] FIG. 9 is a diagram showing the contents of the BrowDown correction process for correcting the facial expression parameters BrowDownLeft (BDoL) and BrowDownRight (BDoR) based on emotion values.
[0096] In the BrowDown correction process, BrowDownLeft (BDoL) and BrowDownRight (BDoR) are corrected using the emotion value EmoA, which indicates the degree of anger. Specifically, BDoL + EmoA × (-1) × 0.5 and BDoR + EmoA × (-1) × 0.5 are calculated, and the corrected facial expression parameters BDoLN and BDoRN corresponding to BrowDownLeft (BDoL) and BrowDownRight (BDoR) are output. Note that if the calculated result exceeds 1.0, 1.0 is output.
[0097] The correction amounts for BDoLN and BDoRN increase as the emotion value EmoA, which indicates the degree of anger, approaches -1, and decrease as the emotion value EmoA, which indicates the degree of anger, approaches 0. BrowDownLeft (BDoL) is a parameter that represents the downward movement of the left eyebrow, and BrowDownRight (BDoR) is a parameter that represents the downward movement of the right eyebrow; the closer the emotion value EmoA approaches -1, the more downward the eyebrows point. For this reason, as shown in Figure 13(b), compared to when the emotion value is 0, when the emotion value EmoA is -1, for example, the eyebrows are positioned lower, and the closer the emotion value EmoA is to -1, the more the expression of the output facial image becomes angrier.
[0098] FIG. 10 is a diagram showing the contents of the EyeBlink switching process in which the facial expression parameter corresponding to the facial expression parameter EyeBlinkLeft (EBL) is switched to either the facial expression parameter EyeBlinkLeft (EBL) or the facial expression parameter EyeSmileLeft (ESL) based on the facial expression parameters MouthSmileLeft (MSmL) and MouthSmileRight (MSmR).
[0099] EyeBlinkLeft (EBL) is a parameter that represents the degree of closure of the left eyelid, and EyeSmileLeft (ESL) is a parameter that represents the degree of closure of the left eyelid when the outer corner of the eye is drooping, and is an expression parameter that is used in place of EyeBlinkLeft (EBL).
[0100] In the EyeBlink switching process, MSm is calculated as the average value of these left and right parameters based on MouthSmileLeft (MSmL) and MouthSmileRight (MSmR). If MSm is a value of 0.5 or greater, the value of EyeBlinkLeft (EBL) is output as the value of EyeSmileLeft (ESL) rather than EyeBlinkLeft (EBL). On the other hand, if MSm is a value less than 0.5, the value of EyeBlinkLeft (EBL) is output as is as the value of EyeBlinkLeft (EBL).
[0101] As shown in FIG. 12(e), MouthSmileLeft (MSmL) is a parameter that represents the upward movement of the left corner of the mouth, and MouthSmileRight (MSmR) is a parameter that represents the upward movement of the right corner of the mouth. MouthSmileLeft (MSmL) and MouthSmileRight (MSmR) are parameters that increase in value when the facial expression is smiling. Therefore, in the EyeBlink switching process, when MSm is less than a certain value, the facial expression is one in which the eyelids are normally closed using EyeBlinkLeft (EBL) as shown in FIG. 14(a), whereas when MSm is equal to or greater than a certain value, the facial expression is one in which the eyelids are closed with the outer corners of the eyes drooping as shown in FIG. 14(b) using EyeSmileLeft (ESL).
[0102] FIG. 11 is a diagram showing the contents of the configuration correction process for correcting facial expression parameters based on configuration parameters.
[0103] As shown in Figure 4, configurations (setting values) are set for each of the facial expression parameters related to eye movement, mouth and jaw movement, eyebrow movement, cheek and nose movement, and tongue movement, and the configuration correction process uses each configuration parameter to correct the value of the corresponding configuration parameter. Specifically, (facial expression parameter + corresponding configuration parameter) / 2 is calculated and the corrected facial expression parameter N is output.
[0104] A configuration parameter is a numerical value between 0.0 and 1.0, and the larger the numerical value of a configuration parameter, the larger the numerical value of the corresponding facial expression parameter N, and the smaller the numerical value of a configuration parameter, the smaller the numerical value of the corresponding facial expression parameter N. For this reason, the amount of change in the part corresponding to the facial expression parameter can be adjusted according to the numerical value of the preset configuration parameter.
[0105] In this way, in this embodiment, other facial expression parameters are corrected based on some facial expression parameters detected from facial image data captured by camera 108, and a facial image with an expression corresponding to the corrected facial expression parameter N is generated, so that the facial image detected by face tracking can be output as a varied image.
[0106] Furthermore, in this embodiment, in addition to some of the facial expression parameters detected from the facial image data captured by the camera 108, a facial image with an expression corresponding to the facial expression parameter N corrected based on parameters other than the facial expression parameters is generated, so that the facial image detected by face tracking can be output as an even more varied image.
[0107] In this embodiment, the facial parameters are corrected using both some of the facial parameters and parameters other than the facial parameters. However, it is also possible to generate a facial image with an expression corresponding to the facial parameter N corrected based on only some of the facial parameters, or to generate a facial image with an expression corresponding to the facial parameter N corrected based on only parameters other than the facial parameters. Even in such a configuration, the facial image detected by face tracking can be output as a varied image.
[0108] In this embodiment, parameters other than facial expression parameters include voice parameters based on the voice input from the microphone 109, and it is possible to generate a facial image with an expression that corresponds to, for example, the volume, pitch, etc. of the voice input from the microphone 109.
[0109] In addition, in this embodiment, parameters other than facial expression parameters include input parameters based on command input operation via the input device 107, and it is possible to generate a facial image with an expression corresponding to, for example, a command input via the input device 107.
[0110] Furthermore, in this embodiment, parameters other than facial expression parameters include vital parameters based on vital data input by the vital sign measuring device 110, and it is possible to generate facial images with expressions corresponding to, for example, pulse, blood pressure, body temperature, etc. input by the vital sign measuring device 110.
[0111] In addition, in this embodiment, parameters other than facial expression parameters include configuration parameters based on configurations (setting values) set in advance on the configuration setting screen, and for example, a facial image of an expression can be generated with a change amount adjusted by the configuration.
[0112] In this embodiment, parameters other than facial expression parameters include voice parameters, input parameters, vital parameters, and configuration parameters, but the configuration may include only some of these, or may include other parameters, such as a background image of the subject, a background image of an avatar to be distributed at the same time, and parameters identified from the temperature, humidity, and brightness of the shooting location.
[0113] In this embodiment, an emotion value is identified based on some facial expression parameters detected from face image data captured by the camera 108, and the facial expression parameters are corrected based on the identified emotion value. A face image with an expression corresponding to the corrected facial expression parameter N is generated, so that the emotion identified from the expression of the face image detected by face tracking can be reflected in the generated face image.
[0114] Furthermore, in this embodiment, in addition to some facial expression parameters detected from the facial image data captured by the camera 108, an emotion value is determined based on parameters other than the facial expression parameters, so that the emotion to be reflected in the generated facial image can be determined more accurately.
[0115] In this embodiment, the configuration is such that an emotion value is determined based on both some of the facial expression parameters and parameters other than the facial expression parameters, but it is also possible to determine an emotion value based on only some of the facial expression parameters, or to determine an emotion value based on only parameters other than the facial expression parameters, and even in such a configuration, the determined emotion can be reflected in the generated facial image.
[0116] Furthermore, in this embodiment, a Joy value indicating joy and an Anger value indicating anger are identified based on some facial expression parameters, and when the Joy value is greater than the Anger value, an emotional value EmoJ indicating joy is identified, and when the Anger value is greater than the Joy value, an emotional value EmoA indicating anger is identified. This makes it possible to generate a facial image with an expression based on emotional values that reflect both joy and anger identified from some facial expression parameters or parameters other than the facial expression parameters.
[0117] In addition, in this embodiment, the Joy value and Anger value are determined using MouthStretch and MouthSmile, which are facial expression parameters that indicate the degree of movement of the corners of the mouth, and emotions such as joy and anger can be determined from the degree of movement of the corners of the mouth, which tend to reflect emotions.
[0118] In this embodiment, a Joy value indicating joy and an Anger value indicating anger are identified and reflected in the facial expression parameters, but it is also possible to identify values indicating sadness and values indicating enjoyment and reflect the emotional values of joy, anger, sadness, and happiness in the facial expression parameters, or to identify only some of these emotional values and reflect the identified emotional values in the facial expression parameters.
[0119] Furthermore, in this embodiment, the corresponding facial expression parameter is corrected regardless of the magnitude of the emotional value, but it is also possible to have a configuration in which the corresponding facial expression parameter is corrected when the emotional value exceeds a certain threshold value, and by using such a configuration, it is possible to add emphasis to the changes in the facial expression of the generated facial image.
[0120] Furthermore, in this embodiment, an emotional value is specified using some facial expression parameters and parameters other than the facial expression parameters, and the facial expression parameters are corrected using the specified emotional value. However, an alternative configuration is also possible in which the facial expression parameters are directly corrected using some facial expression parameters that change depending on the emotion or parameters other than the facial expression parameters. Even in such a configuration, if a parameter used for correction exceeds a threshold value, the corresponding facial expression parameter is corrected, thereby making it possible to add emphasis to the changes in the facial expression of the generated facial image.
[0121] [Correction content setting program] The correction content setting program is a program that performs processing to set setting data used by the parameter correction program. In addition to the processing to set setting data, the correction content setting program also performs processing to specify facial expression parameters and extension parameters to be output to the image generation program. A detailed description of the processing to specify facial expression parameters and extension parameters to be output to the image generation program will be omitted. In this embodiment, the computer terminal 1 that executes the parameter correction program and the image generation program is configured to be able to execute the correction content setting program. However, a computer terminal other than the computer terminal 1 may be equipped with a function to execute the correction content setting program, and setting data set in the other computer terminal may be applied as setting data for the parameter correction program of the computer terminal 1.
[0122] FIG. 15 is a flowchart showing the control contents of the correction content setting program.
[0123] 15, in the correction content setting program, first, a parameter to be corrected is specified (Sb1). In step Sb1, for example, an input node is placed on the setting screen, and the type of parameter to be corrected is set in the input node.
[0124] Next, parameters to be used for correction are designated (Sb2). In step Sb2, for example, as in step Sb1, input nodes are placed on the setting screen, and the types of parameters to be used for correction are set in the input nodes.
[0125] Next, correction details such as a calculation formula using a plurality of parameters and output conditions are set (Sb3). In step Sb3, for example, setting nodes and output nodes are arranged on the setting screen, and input nodes in which the types of parameters to be corrected are set and input nodes in which the types of parameters used for correction are set are linked to the input sections of the setting nodes, and output nodes in which the types of parameters after correction are set are linked to the output sections of the setting nodes, and further correction details such as a calculation formula using the parameters input to the setting nodes and output conditions are set.
[0126] Next, based on the settings of the input node, setting node, and output node, setting data is created that defines the type of input parameter, the correction content (calculation formula, output conditions, etc.) based on the input parameter, and the type of output parameter, and is output to the parameter correction program (Sb4).
[0127] Next, we will explain the specific procedure for setting the setting data using the correction content setting program. Here, we will explain the procedure for setting the setting data used in the emotion value setting process, and the setting data used in the MouthSmile correction process, MouthFrown correction process, and BrowDown correction process, which use emotion values for correction.
[0128] 16 to 21 are diagrams for explaining the procedure for setting the setting data used in the emotion value setting process.
[0129] The setting data used in the emotion value setting process consists of multiple pieces of setting data, and first, setting data is set to calculate the left and right averages BDo, CSq, MSt, and MSm of BrowDownLeft (BDoL), BrowDownRight (BDoR), CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR).
[0130] For example, when setting setting data for calculating BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR), two input nodes, one setting node, and one output node are placed on the setting screen as shown in Fig. 16. The number and positions of these nodes can be specified arbitrarily.
[0131] Next, BDoL and BDoR are input to the two input nodes, respectively. Two input parts, A and B, are set to the setting node, and the input node for BDoL is connected to input part A of the setting node, and the input node for BDoR is connected to input part B of the setting node. BDo, which is output to the output node, is input, and the output part of the setting node is connected to the output node. Furthermore, (A+B) / 2 is input to the setting node as a calculation formula for taking the average of the parameters input from input parts A and B. By performing a decision operation in this state, setting data is created that calculates BDo, which is the average of BrowDownLeft (BDoL) and BrowDownRight (BDoR).
[0132] The procedure for creating setting data for calculating CSq, MSt, and MSm, which are the left and right averages of CheekSquintLeft (CSqL), CheekSquintRight (CSqR), MouthStretchLeft (MStL), MouthStretchRight (MStR), MouthSmileLeft (MSmL), and MouthSmileRight (MSmR), is also similar.
[0133] Next, setting data is set to calculate the Anger value, which indicates the degree of anger, using BDo, CSq, MFu, and MSt.
[0134] Here, as shown in FIG. 17, four input nodes, one setting node, and one output node are arranged on the setting screen.
[0135] Next, BDo, CSq, MFu, and MSt are input to the four input nodes, respectively. Furthermore, four input sections A to D are set to the setting node, and the input node of BDo is connected to input section A of the setting node, the input node of CSq is connected to input section B of the setting node, the input node of MSt is connected to input section C of the setting node, and the input node of MFu is connected to input section D of the setting node. Furthermore, Anger, which is output to the output node, is input, and the output section of the setting node is connected to the output node. Furthermore, ((AB)*2+C+D) / 4*(-1) is input to the setting node as a formula for calculating the Anger value using the parameters input from input sections A to D. By performing a decision operation in this state, setting data for calculating the Anger value, which indicates the degree of anger, is created.
[0136] Next, setting data is set to calculate a Joy value that indicates the degree of joy using MSm and CSq.
[0137] Here, as shown in FIG. 18, two input nodes, one setting node, and one output node are arranged on the setting screen.
[0138] Next, MSm and CSq are input to the two input nodes, respectively. Two input parts, A and B, are set to the setting node, and the input node of MSm is connected to input part A of the setting node, and the input node of CSq is connected to input part B of the setting node. Joy, which is output to the output node, is input, and the output part of the setting node is connected to the output node. Furthermore, (A+B) / 2 is input to the setting node as a formula for calculating the Joy value using the parameters input from input parts A and B. By performing a decision operation in this state, setting data is created that calculates the Joy value, which indicates the degree of joy.
[0139] Next, the Anger value and the Joy value are used to calculate an emotion value AG, and setting data is set to output AGJ when the calculated AG exceeds 0 and AGA when the calculated AG is less than 0.
[0140] Here, as shown in FIG. 19, two input nodes, two setting nodes, and two output nodes are arranged on the setting screen.
[0141] Next, Anger and Joy are input to two input nodes, respectively. Two input parts, A and B, are set in one setting node, the Anger input node is connected to input part A of one setting node, the Joy input node is connected to input part B of one setting node, and the output part is connected to the input part of the other setting node. Furthermore, A+B is input to the input part of one setting node as a formula for calculating the AG value using the parameters input from A and B. Next, two output parts, A and B, are set in the other setting node, AGJ and AGA are input to the two output nodes, respectively, and the output part A of the other setting node is connected to the output node of AGJ, and the output part B of the other setting node is connected to the output node of AGA. Furthermore, In>0→A In<0→B is input to the other setting node as a conditional formula for allocating parameters according to the AG value input from the input part. By performing a decision operation in this state, configuration data is created that outputs AGJ, where AG is greater than 0, and AGA, where AG is less than 0.
[0142] Next, the setting data for converting the volume (Vo), pitch (Pi), and pulse rate (PR) into Vo1, Pi1, and PR1, which indicate 1 or 0, is set.
[0143] For example, when converting the volume (Vo) to Vo1, one input node, one setting node, and one output node are placed on the setting screen as shown in FIG.
[0144] Next, Vo is input to the input node. The Vo input node is connected to the input section of the setting node. Vo1, which is output to the output node, is input, and the output section of the setting node is connected to the output node. Furthermore, In≧0.8→1 In<0.8→0 is input to the setting node as a conditional expression for converting parameters according to the Vo value input from the input section. By performing a decision operation in this state, setting data for converting volume (Vo) to Vo1 is created.
[0145] The procedure for creating the setting data for converting pitch (Pi) and pulse (PR) is similar. Note that the condition formula for converting pulse (PR) to PR1 is entered as In≧0.6→1 and In<0.6→0.
[0146] Next, setting data for calculating emotion values EmoJ and EmoA is set using AGJ, AGA, Vo1, Pi1, and PR1.
[0147] Here, as shown in FIG. 21, five input nodes, two setting nodes, and two output nodes are arranged on the setting screen.
[0148] Next, AGJ, AGA, Vo1, Pi1, and PR1 are input to the five input nodes, respectively. Furthermore, four input parts A to D are set on one setting node, and the input node of AGJ is connected to input part A of one setting node, the input node of Vo1 is connected to input part B of one setting node, the input node of Pi1 is connected to input part C of one setting node, and the input node of PR1 is connected to input part D of one setting node. Furthermore, four input parts A to D are set on the other setting node, and the input node of AGA is connected to input part A of the other setting node, the input node of Vo1 is connected to input part B of the other setting node, the input node of Pi1 is connected to input part C of the other setting node, and the input node of PR1 is connected to input part D of the other setting node. Furthermore, EmoJ is set to be output to one output node, the output part of one setting node is connected to one output node, and EmoA is set to be output to the other output node, and the output part of the other setting node is connected to the other output node. Furthermore, A*0.8+(B+C+D) / 3*0.2 is input to one setting node as the formula for correcting the AGJ value using the parameters input from input units A to D, and A*0.8-(B+C+D) / 3*0.2 is input to the other setting node as the formula for correcting the AGA value using the parameters input from input units A to D. By performing a decision operation in this state, setting data is created that calculates emotion values EmoJ and EmoA using AGJ, AGA, Vo1, Pi1, and PR1.
[0149] In this way, by setting all of the setting data shown in FIGS. 16 to 21 using the correction content setting program, the parameter correction program can refer to this setting data and calculate the emotion values EmoJ and EmoA.
[0150] FIG. 22 is a diagram for explaining the procedure for setting the setting data used in the MouthSmile correction process.
[0151] When creating setting data to be used in the MouthSmile correction process, three input nodes, two setting nodes, and two output nodes are arranged on the setting screen as shown in FIG.
[0152] Next, EmoJ, which is used for correction, and MSmL and MSmR, which are corrected, are input to the three input nodes, respectively. Furthermore, two input parts A and B are set to one setting node, and the input node of EmoJ is connected to input part A of one setting node, and the input node of MSmL is connected to input part B of one setting node. Furthermore, two input parts A and B are set to the other setting node, and the input node of EmoJ is connected to input part A of the other setting node, and the input node of MSmR is connected to input part B of the other setting node. Furthermore, MSmLN, which is output to one output node, is set, and the output part of one setting node is connected to one output node, and MSmRN, which is output to the other output node, is set, and the output part of the other setting node is connected to the other output node. Furthermore, B+A*0.5 is input to one setting node as a calculation formula for correcting MSmL input from input unit B using EmoJ input from input unit A, and B+A*0.5 is input to the other setting node as a calculation formula for correcting MSmR input from input unit B using EmoJ input from input unit A. By performing a confirm operation in this state, setting data is created that corrects the facial expression parameters MSmL and MSmR using the emotion value EmoJ.
[0153] FIG. 23 is a diagram for explaining a procedure for setting the setting data used in the MouthFrown correction process and the BrowDown correction process.
[0154] When creating setting data to be used for the MouthFrown correction process and the BrowDown correction process, five input nodes, four setting nodes, and four output nodes are arranged on the setting screen as shown in FIG.
[0155] Next, EmoA to be used for correction, and MFrL, MFrR, BDoL, and BDoR to be corrected are input to the five input nodes, respectively. Two input sections, A and B, are set in the first setting node, with the EmoA input node connected to input section A of the first setting node and the MFrL input node connected to input section B of the first setting node. Two input sections, A and B, are set in the second setting node, with the EmoA input node connected to input section A of the second setting node and the MFrR input node connected to input section B of the second setting node. Two input sections, A and B, are set in the third setting node, with the EmoA input node connected to input section A of the first setting node and the BDoL input node connected to input section B of the third setting node. Two input sections, A and B, are set in the fourth setting node, with the EmoA input node connected to input section A of the fourth setting node and the BDoR input node connected to input section B of the fourth setting node. Also, MFrLN is set to be output to the first output node, and the output unit of the first setting node is connected to the first output node, MFrRN is set to be output to the second output node, and the output unit of the second setting node is connected to the second output node, BDoLN is set to be output to the third output node, and the output unit of the third setting node is connected to the third output node, and BDoRN is set to be output to the fourth output node, and the output unit of the fourth setting node is connected to the fourth output node. Furthermore, B+A*(-1)*0.5 is input into the first setting node as a calculation formula for correcting MFrL input from input unit B using EmoA input from input unit A, B+A*(-1)*0.5 is input into the second setting node as a calculation formula for correcting MFrR input from input unit B using EmoA input from input unit A, B+A*(-1)*0.5 is input into the third setting node as a calculation formula for correcting BDoL input from input unit B using EmoA input from input unit A, and B+A*(-1)*0.5 is input into the fourth setting node as a calculation formula for correcting BDoR input from input unit B using EmoA input from input unit A. By performing a decision operation in this state, setting data is created that corrects the facial expression parameters MFrL, MFrR, BDoL, and BDoR using the emotion value EmoA.
[0156] In this way, in the correction content setting program, input nodes, setting nodes, and output nodes are arranged on the setting screen, the parameters to be corrected and the types of parameters to be used for correction are input into the input nodes and linked to the setting nodes, the types of parameters to be output are input into the output nodes and linked to the setting nodes, and the correction content is set in the setting nodes using calculation formulas and conditional expressions that use the parameters input from the input nodes, making it possible to easily set the setting data to be used in the parameter correction program.
[0157] In addition, the correction content setting program makes it possible to read out setting data that has been created once, and by reading out the setting data, the input nodes, setting nodes, and output nodes based on the setting data, as well as their setting contents, can be displayed.It is also possible to change the connections between nodes, change the types of parameters set in input nodes and output nodes, and change the calculation formulas and conditional expressions set in setting nodes, and then overwrite and save or create new setting data, making it possible to easily edit existing setting data and easily create new setting data based on existing setting data.
[0158] [Image generation program] As shown in FIG. 3, the image generation program executes a process for generating an avatar image based on the facial expression parameter N and extension parameters from the parameter correction program, the motion parameters from the motion tracking program, the audio parameters from the audio parameters, and the input parameters from the input detection program, and outputs the generated avatar image to the distribution program. The image generation program includes a facial image generation process for generating a facial image, a facial image correction process for correcting the facial image, a costume image generation process for generating a costume image, and an ornamental image generation process for generating an ornamental image. Like the parameter correction program, the image generation program is executed for each frame at intervals based on the frame rate of the video distributed by the distribution program. For example, if the frame rate is 60 fps, the image generation program is executed 60 times per second.
[0159] FIG. 24 is a flowchart showing the control contents of the image generating program.
[0160] 24, the image generation program first imports various parameters (Sc1). In step Sc1, the facial expression parameter N and extended parameters output from the parameter correction program, the motion parameters from the motion tracking program, the voice parameters from the voice parameters, and the input parameters from the input detection program are imported.
[0161] Next, a 3D model of the face is created based on the facial expression parameter N captured in step Sc1 (Sc2). In step Sc2, a plurality of 3D models (face data) preset to correspond to the facial expression parameter N are used and blended (blend shape) using the value of the corresponding facial expression parameter N. Some parts are also created by setting the positions and angles of the joints between multiple bones using the value of the facial expression parameter N. Alternatively, a 3D model created using blend shape and a 3D model created using bones may be blended to complement each other.
[0162] Next, the motion parameters acquired in step Sc1 are used to determine the horizontal and vertical inclination angles of the camera 108 relative to the front of the face, and the contours, positions, and shapes of the facial features of the 3D model of the face created in step Sc2 are corrected by a correction amount corresponding to the determined imaging direction. As a result, as shown in Figures 25(a) and 25(b), the same 3D model is not used when the face is photographed from the front and when it is photographed from an oblique angle, but a 3D model in which the facial contours, positions, and shapes of the mouth, nose, and eyes are corrected according to the imaging direction is used. The correction amount according to the imaging direction is calculated based on the maximum horizontal and vertical correction amounts predetermined for the contours and the facial features to be corrected, and the horizontal and vertical angles of the imaging direction. As shown in Figure 26(a), the correction amount in the horizontal and vertical directions gradually increases as the horizontal or vertical angle of the imaging direction increases. Alternatively, an intermediate correction amount may be set corresponding to an intermediate angle smaller than the maximum angle. In this case, as shown in FIG. 26(b), by gradually increasing the correction amount in the left-right or up-down direction so as to form a curve that passes through the intermediate correction amount at the intermediate angle, the correction amount does not change suddenly when the shooting angle changes.
[0163] Next, a 3D model of the body and hands (fingers) is created based on the motion parameters acquired in step Sc1 (Sc4). In step Sc4, multiple 3D models (body data) preset to correspond to the motion parameters are used to create a 3D model of the body and hands (fingers) in the corresponding pose.
[0164] Next, a 3D model of the costume (including accessories such as hairstyle and ornaments) is created (Sc5) based on the motion parameters acquired in step Sc1. In step Sc5, multiple 3D models (clothing data) preset to correspond to the motion parameters are used to create a 3D model of the costume that matches the body and hands (fingers) in the corresponding pose.
[0165] Next, we determine whether the 3D clothing model created in step Sc5 overlaps with the prohibited area (Sc6). The prohibited area includes the overlapping area where the clothing and face overlap and the facial expression becomes invisible, as well as the surrounding area, and also includes areas other than the overlapping area between the clothing and face.
[0166] If it is determined in step Sc6 that the 3D model of the clothing does not overlap with the prohibited area, the process proceeds to step Sc9. If it is determined that the 3D model of the clothing overlaps with the prohibited area, the amount of deformation or movement of the clothing is calculated based on the distance from the 3D model of the clothing to the overlap area between the clothing and the face (Sc7), and the shape or position of the relevant part of the clothing is corrected based on the amount of deformation or movement calculated in step Sc7 (Sc8).
[0167] As a result, for example, when the shooting direction is from the front of the face as shown in Fig. 27(a), even if the clothes do not overlap the face, even if a change in the shooting direction causes part of the clothes to overlap the face and hide the facial expression as shown in Fig. 27(b), the part of the clothes will not deform and hide the facial expression. Also, when the shooting direction is from the front of the face as shown in Fig. 28(a), even if the accessories do not overlap the face, even if a change in the shooting direction causes the accessories to overlap the face and hide the facial expression as shown in Fig. 28(b), the accessories will not move and hide the facial expression.
[0168] Furthermore, the amount of deformation or movement of the clothing after the 3D clothing model overlaps the prohibited area is set to gradually increase as it approaches the overlap area between the face and clothing, as shown in Fig. 29. Furthermore, the closer it gets to the overlap area between the face and clothing, the greater the increase in the amount of deformation or movement of the clothing. Therefore, even if the 3D clothing model overlaps the prohibited area, if it is far from the overlap area between the face and clothing, the amount of deformation or movement is kept small, and as it approaches the overlap area between the face and clothing, the amount of deformation or movement of the clothing increases, thereby preventing overlap of the face and clothing.
[0169] Next, it is determined whether or not an additional condition for adding a decorative image is met based on the extended parameters (emotion values) and voice parameters acquired in step Sc1 (Sc9). If it is determined in step Sc9 that the additional condition is not met, the process proceeds to step Sc11. If it is determined that the additional condition is met, a decorative image according to the met additional condition is set (Sc10). The decorative image set in step Sc10 is retained for a certain period of time and is cleared after the certain period has elapsed. Therefore, after the additional condition is met, the decorative image remains displayed for a certain period of time. If a new additional condition is met, the decorative image is overwritten.
[0170] Additional conditions and decorative images corresponding to the additional conditions are set in advance. For example, as shown in FIG. 30, the additional conditions are met when the magnitude of the emotional value, the command input via the input device 107 specifying the decorative image, the volume, and the pitch are equal to or greater than predetermined values. In this embodiment, if the emotional value Emo is 0.7 or greater or a joy (small) command is input, a cracker decorative image is set; if the emotional value Emo is 0.9 or greater or a joy (large) command is input, a heart mark decorative image is set in addition to the cracker decorative image; if the emotional value Emo is −0.8 or less or an anger command is input, an anger mark decorative image is set; and if the volume (Vo) or pitch (Pi) is 0.8 or greater, a speaker mark decorative image is set.
[0171] For example, if the emotion value (Emo) is 0.9 or higher or a joy (large) command is input, specifying a happy situation, decorative images of crackers and heart marks will be displayed around the avatar image for a certain period of time, as shown in Fig. 31. If the emotion value (Emo) is less than -0.8 or an anger command is input, specifying an angry situation, decorative images of angry marks will be displayed around the avatar image for a certain period of time, as shown in Fig. 32. Also, if the volume (Vo) or pitch (Pi) is 0.8 or higher, specifying a situation in which a person is speaking in a loud or high voice, decorative images of speaker marks will be displayed, as shown in Fig. 32.
[0172] Next, the 3D model of the face created in step Sc2 and corrected in step Sc3, the 3D model of the body and hands (fingers) created in step Sc4, the 3D model of the clothing created in step Sc5, and the decorative images set in Sc10 are arranged to generate avatar image data, which is then output to the distribution program (Sc11).
[0173] In this way, in this embodiment, decorative images other than facial images are generated based on some facial expression parameters detected from facial image data captured by camera 108, and facial images detected by face tracking can be output as varied images.
[0174] In addition, in this embodiment, an emotional value is determined based on some facial expression parameters detected from facial image data captured by the camera 108, and a decorative image other than a facial image is generated based on the determined emotional value, so that the decorative image can be displayed reflecting the emotion determined from the facial expression of the facial image detected by face tracking.
[0175] Furthermore, in this embodiment, in addition to some facial expression parameters detected from the face image data captured by the camera 108, the emotional value is determined based on parameters other than the facial expression parameters, so that decorative images can be displayed that more accurately reflect the emotions.
[0176] In this embodiment, the emotional value is determined based on both some of the facial expression parameters and parameters other than the facial expression parameters, but it is also possible to determine the emotional value based on only some of the facial expression parameters, or to determine the emotional value based on only parameters other than the facial expression parameters. Even in such a configuration, it is possible to display a decorative image that reflects the determined emotion.
[0177] Furthermore, in this embodiment, when the emotion value exceeds a certain threshold, a corresponding decorative image is generated, and by using this configuration, the decorative image can be displayed at an appropriate frequency.
[0178] In this embodiment, decorative images other than facial images are generated based on some facial expression parameters detected from facial image data captured by camera 108, but it is also possible to change the design and color of the costume, the composition and color tone of the background image, the shape, size, and color of accessories, etc. based on some of the parameters detected from the facial image data and the emotional values identified from these parameters.In such a configuration, it is possible to make changes to things other than facial images by reflecting parameters identified from the facial expression of the facial image detected by face tracking.
[0179] In this embodiment, if a part of the clothing is positioned in a way that prevents the facial expression from being visible depending on the shooting direction of the camera 108, the part of the clothing is changed to a shape that allows the facial expression to be visible, thereby preventing the facial expression from being hidden by a part of the clothing when the shooting direction of the subject by the camera 108 changes.
[0180] Furthermore, in this embodiment, even if the user is not in a position where the facial expression of the facial image cannot be seen, when the user approaches a position where the facial expression of the facial image cannot be seen (prohibited area), part of the clothing is deformed into a shape that makes the facial expression of the facial image visible, and the part of the clothing can be deformed into a shape that makes the facial expression of the facial image visible even before the facial expression of the facial image is hidden.
[0181] In addition, in this embodiment, depending on the distance to the position where the facial expression of the facial image cannot be seen, a part of the clothing is gradually deformed into a shape that makes the facial expression of the facial image visible, so that a part of the clothing can be naturally deformed into a shape that makes the facial expression of the facial image visible.
[0182] In this embodiment, even if the subject is not in a position where the facial expression of the facial image cannot be seen, when the subject approaches a position where the facial expression of the facial image cannot be seen (prohibited area), part of the clothing is deformed into a shape that makes the facial expression of the facial image visible.However, it is also possible to deform part of the clothing into a shape that makes the facial expression of the facial image visible when the subject reaches a position where the facial expression of the facial image cannot be seen.Even with this configuration, it is possible to prevent the facial expression of the subject from being hidden by part of the clothing when the direction in which the subject is photographed by camera 108 changes.
[0183] In addition, in this embodiment, a part of the clothing is gradually deformed into a shape that makes the facial expression of the facial image visible depending on the distance to the position where the facial expression of the facial image cannot be seen, but it is also possible to configure the clothing to be deformed into a shape that makes the facial expression of the facial image visible when a position where the facial expression of the facial image cannot be seen (prohibited area) is reached.
[0184] In this embodiment, if the attachment is positioned so that the facial expression of the facial image cannot be seen depending on the shooting direction of the camera 108, the attachment is moved to a position where the facial expression of the facial image can be seen, thereby preventing the facial expression of the facial image from being hidden by the attachment when the shooting direction of the subject by the camera 108 changes.
[0185] Furthermore, in this embodiment, even if the position is not one in which the facial expression of the facial image cannot be seen, when the user approaches a position in which the facial expression of the facial image cannot be seen (prohibited area), the accessory is moved toward a position in which the facial expression of the facial image can be seen, and the accessory can be moved toward a position in which the facial expression of the facial image can be seen before the facial expression of the facial image is hidden.
[0186] In addition, in this embodiment, the accessory is gradually moved toward a position where the facial expression of the facial image can be seen depending on the distance to the position where the facial expression of the facial image cannot be seen, so that the accessory can be naturally moved toward a position where the facial expression of the facial image can be seen.
[0187] In this embodiment, even if the position is not one where the facial expression of the facial image cannot be seen, when the subject approaches a position where the facial expression of the facial image cannot be seen (prohibited area), the accessory is moved toward a position where the facial expression of the facial image can be seen.However, it is also possible to move the accessory to a position where the facial expression of the facial image can be seen when the subject reaches a position where the facial expression of the facial image cannot be seen.Even with this configuration, it is possible to prevent the facial expression of the subject from being hidden by the accessory due to a change in the shooting direction of the subject by camera 108.
[0188] In addition, in this embodiment, the accessory is gradually moved toward a position where the facial expression of the facial image can be seen depending on the distance to the position where the facial expression of the facial image cannot be seen, but it may also be configured to move the accessory toward a position where the facial expression of the facial image can be seen when it reaches a position where the facial expression of the facial image cannot be seen (prohibited area).
[0189] In this embodiment, the shape or position of the features that make up the facial image is corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, the facial expression of the generated facial image can be prevented from becoming unnatural.
[0190] Furthermore, in this embodiment, the contours of the face image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, the contours of the generated face image can be prevented from becoming unnatural.
[0191] In addition, in this embodiment, the position and shape of the mouth constituting the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, the position and shape of the mouth constituting the generated facial image can be prevented from becoming unnatural.
[0192] In addition, in this embodiment, the position and shape of the nose that constitutes the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, the position and shape of the nose that constitutes the generated facial image can be prevented from becoming unnatural.
[0193] In addition, in this embodiment, the position and shape of the eyes that make up the facial image are corrected according to the shooting direction of the camera 108, so that even if the shooting direction changes, the position and shape of the eyes that make up the generated facial image can be prevented from becoming unnatural.
[0194] In addition, in this embodiment, the amount of correction for the shape or position of the features that make up the facial image is gradually increased depending on the shooting direction of the camera 108, so that the shape or position of the features that make up the facial image does not suddenly change.
[0195] Although an embodiment of the present invention has been described above with reference to the drawings, the present invention is not limited to this embodiment, and it goes without saying that the present invention also includes modifications and additions that do not deviate from the spirit of the present invention.
[0196] For example, in the above embodiment, an example was described in which an avatar image generated by an image generation program is used for video distribution, but the use of an avatar image generated by an image generation program is arbitrary, and the avatar image may be recorded as an archive, or may be used for producing animation, etc.
[0197] Furthermore, in the above embodiment, an example was described in which facial expression parameters output from a face tracking program are corrected by a parameter correction program, and an image generation program generates an avatar image using the corrected facial expression parameters. However, a configuration may also be adopted in which the image generation program directly uses parameters output from the face tracking program to generate an image, without using a parameter correction program. [Explanation of symbols]
[0198] 1 computer terminal 101 processors 102 memory 103 Storage 104 Communication Interface 105 Display device 106 Speaker 107 Input Device 108 Camera 109 Mike 110 Vital Signs Meter 111 Data Bus
Claims
1. a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. An image generating method comprising:
2. In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image is not visually recognizable, when the specific image approaches a position where the facial expression of the facial image is not visually recognizable, the specific image is deformed into a shape where the facial expression of the facial image is visually recognizable. The image generating method according to claim 1 .
3. In the specific image generating step, the specific image is gradually deformed to a shape in which the facial expression of the facial image is visible, depending on the distance to a position where the facial expression of the facial image is not visible. The image generating method according to claim 2 .
4. a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; A computer-aided image generation method comprising: further comprising a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. An image generating method comprising:
5. In the specific image generating step, even if the specific image is not at a position where the facial expression of the facial image cannot be visually recognized, when the specific image approaches a position where the facial expression of the facial image cannot be visually recognized, the specific image is moved toward a position where the facial expression of the facial image can be visually recognized. The image generating method according to claim 4.
6. In the specific image generating step, the position of the specific image is gradually moved toward a position where the facial expression of the facial image can be visually recognized, depending on the distance to the position where the facial expression of the facial image cannot be visually recognized. The image generating method according to claim 5 .
7. a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the specific image is positioned at a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the specific image is transformed into a shape where the facial expression of the facial image can be visually recognized. An image generation program comprising:
8. a facial image input step of inputting a photographed facial image; a facial image generating step of generating a facial image having an expression corresponding to the facial image inputted in the facial image input step; an image output step of outputting an output image including the face image generated in the face image generation step; In an image generation program that causes a computer to execute a specific image generating step of generating a specific image to be displayed around the face image; In the image output step, the specific image generated in the specific image generation step is output together with the face image; In the specific image generating step, if the position of the specific image is a position where the facial expression of the facial image cannot be visually recognized depending on the photographing direction, the position of the specific image is moved to a position where the facial expression of the facial image can be visually recognized. An image generation program comprising:
Citation Information
Patent Citations
Content distribution server, content distribution system, content distribution method, and program
JP2019186797A