Avatar insertion system, avatar insertion method, and avatar insertion program
The avatar insertion system addresses the challenge of creating avatars that match real people's appearances by modifying avatars to align with target features, improving immersion by reducing incongruity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- POCKETRD CO LTD
- Filing Date
- 2024-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Existing technologies face challenges in creating avatars that reflect real people's characteristics without causing a sense of incongruity, especially when replacing characters with significant appearance differences, leading to unnatural impressions and cumbersome manual adjustments.
An avatar insertion system that extracts replacement targets, determines feature elements such as shape and color, modifies avatars to match these elements, and inserts them into videos, ensuring consistency with the target's appearance.
Effectively suppresses the sense of incongruity by maintaining the characteristics of the avatar while integrating it into images, enhancing viewer immersion.
Smart Images

Figure 2026084296000001_ABST
Abstract
Description
Technical Field
[0004] , , , , , , , , ,<C<000D
[0005] <00000C00D00002><00000C00D00D003>C<000D000004>C<00D000005>C<00D<000006>
Background Art
[0002] In recent years, with the improvement of processing capabilities in electronic computers such as computers, many computer graphics of human figures, so-called avatars, that reflect the characteristics of real people have been utilized. For example, an avatar is used as one's own icon for use on SNS (Social Networking Service), or one's own avatar is used as the character of the protagonist in an online game or the like. In addition, in existing content such as still images and moving images, a service has been proposed in which some of the characters appearing are replaced with one's own avatar for viewing and the like.
[0003] <
[0004] Patent Documents 1 and 2 both disclose technologies for using avatars that mimic the actual appearances of the player himself or co-players in a computer game that uses a head-mounted display to represent a virtual space.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
[0006] However, since the appearances of real people vary greatly, using avatars that reflect the characteristics of real people requires creating each avatar individually, which presents a problem due to the complexity of avatar generation. Furthermore, for example, when replacing a character in existing content with one's own avatar, if the appearance of the character and the user are too different (for example, replacing an infant character with a middle-aged man), using an avatar that reflects the characteristics of a real person as is would create an unnatural impression and hinder the user's immersion. On the other hand, if the appearance is made to be the same as the character in the content, there is no point in using the user's avatar, and ultimately, the appearance of the avatar would have to be adjusted individually by hand, which is extremely cumbersome. These are the problems with using avatars that reflect the characteristics of real people, but neither Patent Documents 1 nor 2 disclose any technology to solve these problems.
[0007] The present invention has been made in view of the above problems, and aims to provide a technology that can effectively suppress the occurrence of a sense of incongruity for viewers of an image after the avatar has been inserted, while maintaining the characteristics of the avatar, when inserting an avatar into an image such as a still image or a video by replacing at least a part of the person to be replaced with the avatar. [Means for solving the problem]
[0008] To achieve the above objective, the avatar insertion system according to claim 1 is an avatar insertion system that inserts an avatar into a video by replacing at least a portion of a replacement target, which is a human figure contained in the video, with an avatar, and is characterized by comprising: replacement target extraction means for extracting the replacement target placed in the video; feature content determination means for determining the content of feature elements in the extracted replacement target, including at least one of shape features and color features; avatar modification means for modifying at least one of the shape and color of the avatar to match the content of the feature elements in the replacement target determined by the feature content determination means, if the content of the feature elements in the avatar does not match the content of the feature elements in the replacement target; avatar insertion means for inserting the avatar into the video by placing the avatar modified by the avatar modification means in the portion of the video where the replacement target was placed; and output means for outputting the video after the avatar insertion by the insertion means to the outside.
[0009] Furthermore, in order to achieve the above objective, the avatar insertion system according to claim 2 is characterized in that, in the above invention, the feature elements subject to determination by the feature content determination means are the appearance age element and the appearance gender element in the replacement target, and the avatar modification means further comprises: an age element modification means for modifying at least one of the shape and color tone of the avatar so that the content of the appearance age element in the avatar is consistent with the content of the appearance age element in the replacement target when the replacement target and the avatar are inconsistent with respect to the content of the appearance age element; and a gender element modification means for modifying at least one of the shape and color tone of the avatar so that the content of the appearance gender element in the avatar is consistent with the content of the appearance gender element in the replacement target when the replacement target and the avatar are inconsistent with respect to the content of the appearance gender element.
[0010] Furthermore, in order to achieve the above objective, the avatar insertion system according to claim 3 is characterized in that, in the above invention, the feature element extracted by the feature content determination means is a partial color element which is a color feature of one or more parts of the skin, hair, or pupils of the face to be replaced, and the avatar modification means further comprises a color element modification means which modifies the color tone of the avatar so that the content of the partial color element in the avatar matches the content of the partial color element in the replacement target when the content of the partial color element does not match that of the replacement target.
[0011] Furthermore, in order to achieve the above objective, the avatar insertion method according to claim 4 is an avatar insertion method that inserts an avatar into a video by replacing at least a part of a replacement target, which is a human figure contained in the video, with an avatar, and is characterized by including: a replacement target extraction step of extracting the replacement target that is placed in the video; a feature content determination step of determining the content of feature elements in the extracted replacement target, which includes at least one of shape features and color features; an avatar modification step of modifying at least one of the shape and color of the avatar to match the content of the feature elements in the replacement target determined in the feature content determination step, if the content of the feature elements in the avatar does not match the content of the feature elements in the replacement target; an avatar insertion step of inserting the avatar into the video by placing the avatar modified in the avatar modification step in the part of the video where the replacement target was placed; and an output step of outputting the video after the avatar insertion in the insertion step to the outside.
[0012] Furthermore, in order to achieve the above objective, the avatar insertion method according to claim 5 further includes, in the above invention, an appearance feature extraction step of extracting appearance feature portions of the avatar based on the result of comparing the appearance of the avatar with the appearance of a standard avatar consisting of a standard shape and color tone; and a protected area information generation step of setting an area on the avatar including the appearance feature portions extracted in the appearance feature extraction step as a protected area, and generating protected area information including position information of the set protected area and information regarding limitations on the modification content in the protected area, wherein in the avatar modification step, even if the content of the feature element in the replacement target determined in the feature content determination step does not match the content of the appearance features in the avatar, the shape and / or color tone of the avatar is modified with respect to the protected area in accordance with the information regarding limitations on the modification content included in the protected area information.
[0013] Furthermore, in order to achieve the above objective, the avatar insertion program according to claim 6 is an avatar insertion program that causes a computer to insert an avatar into a video by replacing at least a part of a replacement target, which is a human figure contained in the video, with an avatar, and is characterized in that it causes the computer to execute: a replacement target extraction function for extracting the replacement target that is placed in the video; a feature content determination function for determining the content of feature elements in the extracted replacement target, which includes at least one of shape features and color features; an avatar modification function for modifying at least one of the shape and color of the avatar to match the content of the feature elements in the replacement target, if the content of the feature elements in the replacement target determined by the feature content determination function does not match the content of the feature elements in the avatar; an avatar insertion function for inserting the avatar into the video by placing the avatar modified by the avatar modification function in the part of the video where the replacement target was placed; and an output function for outputting the video after the avatar insertion by the insertion function to the outside. [Effects of the Invention]
[0014] According to the present invention, when inserting an avatar into a video, such as a still image or video, by replacing at least a portion of the person to be replaced with the avatar, it is possible to effectively suppress the occurrence of a sense of incongruity for the viewer of the video after the avatar has been inserted, while maintaining the characteristics of the avatar. [Brief explanation of the drawing]
[0015] [Figure 1] This is a schematic diagram showing the configuration of the avatar insertion system according to Embodiment 1. [Figure 2] This is a schematic diagram showing the configuration of the avatar insertion system according to Embodiment 2. [Modes for carrying out the invention]
[0016] The embodiments of the present invention will be described in detail below with reference to the drawings. The following embodiments describe the most appropriate examples of the present invention, and naturally, the content of the present invention should not be limited to the specific examples shown in these embodiments. It goes without saying that any configuration other than the specific configurations shown in the embodiments that produces similar functions and effects is also included in the technical scope of the present invention.
[0017] (Embodiment 1) First, the avatar insertion system according to Embodiment 1 will be described. As shown in Figure 1, the avatar insertion system according to Embodiment 1 includes a video input unit 1 that inputs an insertion target video, which is the video into which the avatar will be inserted; a replacement target extraction unit 2 that extracts replacement targets, which are the objects to be replaced by the avatar, from the video input via the video input unit 1; a feature content determination unit 3 that determines the content of feature elements, which include at least one of the shape features and color features of the appearance of the replacement targets extracted by the replacement target extraction unit 2; an avatar generation unit 4 that generates an avatar to be inserted into the insertion target video; and an image analysis unit that analyzes the shape and color of the appearance of the generated avatar. The system comprises an avatar analysis unit 5, an avatar modification unit 6 that compares the content of the feature elements to be replaced with the content of the avatar's feature elements based on the determination results of the feature content determination unit 3 and the image analysis results by the avatar analysis unit 5, and modifies the shape and color tone of the avatar's appearance to match the content of the feature elements to be replaced if they do not match, an avatar insertion unit 7 that inserts the modified avatar into the video to be inserted by placing the modified avatar in the part of the video to be inserted where the replacement target is located, and a video output unit 8 that outputs the video with the inserted avatar to the outside.
[0018] The video input unit 1 is for inputting the video to be inserted, which is the video subject to avatar insertion processing in the avatar insertion system according to this embodiment 1. The video to be inserted may be either a video or a still image, and may be either a 3D image or a 2D image; there are no particular restrictions as long as it contains visual information. Similarly, there are no particular restrictions on the content of the video to be inserted, but it should at least include video related to the replacement target, which is the object to be replaced by the avatar. Furthermore, the video to be inserted may be specially generated for avatar insertion, but it may also be general video content such as existing movies or landscape photographs.
[0019] The replacement target extraction unit 2 is for extracting a replacement target that is to be replaced with an avatar from the insertion target video input via the video input unit 1. Specifically, the replacement target extraction unit 2 has a function of specifying where the specified replacement target is displayed in the insertion target video and specifying the displayed position and the like. As a method for specifying the replacement target, it may be a method specified by a predetermined person such as a user, or after performing image analysis processing on the insertion target video to extract one or more replacement target candidates, a candidate that conforms to a predetermined criterion such as the user's preference may be selected. Further, the replacement target extraction unit 2 preferably has a function of selecting the display position and the like of the replacement target in the plurality of scenes when the replacement target appears in a plurality of scenes such as video content. Information regarding the extraction process of the replacement target by the replacement target extraction unit 2 is transmitted to the avatar insertion unit 7 described later, and the avatar insertion unit 7 performs the replacement process of the avatar based on the information. Although there is no particular limitation on the replacement target extracted by the replacement target extraction unit 2, it is preferably a human or a creative creature imitating a human. In the following description of Embodiment 1, it is assumed that the replacement target is a human.
[0020] The feature content determination unit 3 is for determining the content of predetermined feature elements including the shape features and color tone features in the appearance of the replacement target extracted by the replacement target extraction unit 2 using techniques such as image analysis. In the present invention, the "feature element" includes, in addition to the specific structure itself consisting of the external shape, external color tone, and combination of shape and color tone in the replacement target or avatar, the characteristics and features of the object inferred based on a plurality of specific aspects and combinations of shapes, color tones, etc. In Embodiment 1, the appearance age element and the appearance gender element are used as examples of the feature elements as the characteristics and features of the object, and the partial color tone element (specific examples include the skin color, hair color, and pupil color on the face) is used as an example of the feature element as the external color tone. However, it is of course possible to use other elements as the feature elements in the present invention.
[0021] Specifically, the feature content determination unit 3 includes a target shape analysis unit 9 that analyzes the shape features of the appearance surface to be replaced, a target color tone analysis unit 10 that analyzes the color tone features of the appearance to be replaced, a target age determination unit 11 that determines the content of the appearance age element of the replacement target, which is one of the feature elements, based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10, a target gender determination unit 12 that determines the content of the appearance gender element of the replacement target, which is one of the feature elements, based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10, and a target color tone determination unit 13 that determines the content of the partial color tone element of the replacement target based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10.
[0022] The target shape analysis unit 9 is for analyzing the appearance shape of the replacement target, that is, the two-dimensional and three-dimensional shapes of the appearance surface of the replacement target.Specifically, the target shape analysis unit 9 in the first embodiment has a function of obtaining information on the shape of the replacement target by extracting the positional relationship of feature points predetermined corresponding to the characteristic parts on the appearance surface of each part such as eyes, nose, mouth, ears, etc. in the human body from the appearance of the replacement target displayed in the insertion target video.To extract the feature points, a specific configuration may be to utilize image recognition technology realized by deep learning, machine learning, etc., or a combination of multiple image recognition technologies.Also, the analysis content may be changed according to the format of the insertion target video (for example, if it is a two-dimensional video, analyze the two-dimensional positional relationship of the feature points), or conversely, regardless of the format of the insertion target video, for example, even if the insertion target video is a two-dimensional video, extract the three-dimensional shape (that is, the three-dimensional positional relationship of the feature points).
[0023] The "feature points" extracted by the target shape analysis unit 9 indicate the position of parts of the human body on the body surface, and the parts to be set as feature points can be arbitrarily determined. For example, feature points may be set only for parts of the face that are visually recognizable (eyes, nose, mouth, ears, eyebrows, etc.) or points corresponding to the contour, or feature points may be set not only for the face but also for parts of the torso and limbs. Furthermore, when setting feature points for the eyes, only the center position of the eye may be set as a feature point, or feature points may be set in detail for the inner corner of the eye, outer corner of the eye, pupil, and white of the eye. In this embodiment 1, the target shape analysis unit 9 generates the analysis results in a format that includes information on the position of the extracted feature points (preferably the position in a normalized coordinate system) and the significance of the feature points (e.g., this is the point corresponding to the outer corner of the eye, this is the point corresponding to the tip of the nose). However, when used as data for the determination process in the target age determination unit 11, target gender determination unit 12, and target color tone determination unit 13, which will be described later, it is not necessary to extract all feature points corresponding to each part. Instead, a configuration may be adopted in which only some feature points are extracted, for example, only the feature points corresponding to facial wrinkles are extracted in relation to the target age determination unit 11.
[0024] The target color tone analysis unit 10 is for analyzing the color tones used in the appearance of the replacement target. Specifically, the appearance of the replacement target is composed of an arrangement of points (pixels) of various color tones (saturation, lightness, hue), and the target color tone analysis unit 10 has the function of analyzing such color tones. As a specific embodiment of the color tone analysis processing by the target color tone analysis unit 10, color tone analysis may be performed on all parts of the appearance of the replacement target, but in this embodiment 1, color tone analysis may be performed only to the extent necessary for the determination processing in the target age determination unit 11, target gender determination unit 12, and target color tone determination unit 13. For example, color tone analysis may be performed only on the face instead of the entire appearance of the replacement target, or, more simply, color tone analysis may be performed only on the parts corresponding to predetermined elements that are to be determined by the target color tone determination unit 13.
[0025] The target age determination unit 11 is for determining the content of the external age element of the replacement target, which is one of the feature elements. Specifically, the target age determination unit 11 has the function of determining the content of the external age element of the replacement target, for example, that the person is 15 years old, 30 years old, etc., based on the shape characteristics of the external surface of the replacement target, which are the analysis results of the target shape analysis unit 9, and / or the color characteristics of the external surface of the replacement target, which are the analysis results of the target color analysis unit 10. The determination mechanism in the target age determination unit 11 preferably makes a comprehensive determination from both shape characteristics and color characteristics, but it may also be determined based only on shape characteristics, such as wrinkles at the corners of the eyes or the depth of nasolabial folds, or based on color characteristics, such as the presence or degree of blemishes on the skin of the face. Furthermore, the target age determination unit 11 may determine the content of the external age element based on specific indicators such as the presence or degree of wrinkles on the face as described above, or it may be equipped with a determination function regarding the content of the external age element by deep learning, machine learning, or reinforcement learning.
[0026] Furthermore, the content of the outward age elements determined by the target age determination unit 11 does not necessarily refer only to the actual age of the person to be replaced, but also includes the concept of age estimated from the outward appearance of the person to be replaced. For example, if the video to be inserted is content such as a movie, the actual age of the actor to be replaced may differ from the age of the character they play in the video to be inserted. In such cases, the target age determination unit 11 may be configured to precisely determine both the actual age and the age of the character, but more preferably, the target age determination unit 11 will determine the content of the outward age elements estimated from the outward appearance, regardless of the actual age or the age of the character.
[0027] The target gender determination unit 12 is for determining the content of the external gender elements of the replacement target, which are one of the characteristic elements. Specifically, the target gender determination unit 12 has the function of determining the external gender elements of the replacement target based on the shape characteristics of the replacement target's surface, which are the analysis results of the target shape analysis unit 9, and / or the color characteristics of the replacement target's surface, which are the analysis results of the target color tone analysis unit 10. The determination mechanism in the target gender determination unit 12 preferably determines the content of the external gender elements comprehensively from both shape characteristics and color tone characteristics, for example, that the person has the characteristics of a male in appearance. However, for example, the determination may be made only from differences in shape characteristics based on skeletal differences between men and women, or, in light of the tendency for women to wear makeup, the determination may be made only from color tone characteristics such as the fineness of skin tone or differences in lip tone. Furthermore, the target gender determination unit 12 may determine the content of the external gender elements based on specific indicators such as lip tone, or it may be equipped with a determination function regarding the content of external gender elements by deep learning, machine learning, or reinforcement learning.
[0028] Furthermore, the content of the outward gender elements determined by the target gender determination unit 12 does not necessarily refer only to the actual gender of the object being replaced, but is a concept that also includes the gender estimated from the outward appearance of the object being replaced. For example, gender can be determined from biological perspectives, from gender identity, and in the case of the video to be inserted, it may include cases where a man plays a woman or a woman plays a man. In such cases, the target gender determination unit 12 may make a precise determination of objective facts, such as that the person is biologically female but identifies as male. However, more preferably, the target gender determination unit 12 will determine the content of the outward gender elements based on which gender's appearance the person possesses, without being bound by biological or gender identity considerations.
[0029] The target color determination unit 13 is for determining the content of the partial color element to be replaced, which is one of the characteristic elements. Any element can be used as a partial color element as long as it is an element in which a tonal characteristic is likely to occur depending on the attributes of the replacement target and in which the tonal characteristic is likely to attract visual attention. However, in this embodiment 1, the skin area, hair, and eyes of the face to be replaced are used as partial color elements, and the target color determination unit 13 determines their content. For example, the color of the skin area varies greatly from one replacement target to another depending on differences in race, whether or not one is tanned, etc., and it also leaves a strong visual impression. Similarly, there are various types of hair color such as blonde, black, and brown, and various types of eye color such as black, blue, and brown. Since both hair and eyes leave a strong visual impression, the target color determination unit 13 in this embodiment 1 determines the tonal characteristics of the skin, hair, and eyes of the face to be replaced. The color tone determination process performed by the target color tone determination unit 13 is carried out based on the analysis results of the target color tone analysis unit 10, after the target shape analysis unit 9 has identified the area corresponding to the element to be determined (skin, hair, pupils, etc. of the face). However, it is not mandatory for the target color tone analysis unit 10 to determine the specific values of the color tone (saturation, lightness, hue) of the entire area of the element to be determined as part of the determination process. For example, it may extract only the most frequently occurring combination of the three elements of saturation, lightness, and hue, or it may derive the frequently occurring or average values of saturation, lightness, and hue.
[0030] The avatar generation unit 4 is for generating an avatar to be inserted into the target video by replacing the replacement target. The avatar may consist of skeletal information (bones), surface information (skin), and weight information that defines the relationship between the two, but it is also preferable to consist only of surface information including information on the three-dimensional shape and color tone on the surface, and in an even simpler configuration, it may consist of the image data itself. Furthermore, the avatar generation unit 4 may generate an avatar corresponding to the full body image of a person, but in this embodiment 1, an avatar consisting only of the head from the neck up is generated. The avatar generation unit 4 extracts feature points set according to the position and shape of surface features (eyes, eyebrows, nose, mouth, ears, hairstyle, etc.) and internal features (joints, etc.) based on the face image of the person to be used as a model (which may be two-dimensional or three-dimensional), and by reflecting the positional relationship between the extracted feature points in the avatar, it generates a realistic avatar that reflects the physical characteristics of the model. However, if a simpler configuration is adopted, it may be possible to apply only the minimum necessary modifications, such as adjusting the size and orientation of the model's face image.
[0031] The Avatar Analysis Unit 5 analyzes the shape and color tone of the avatar's appearance generated by the Avatar Generation Unit 4, and determines the content of feature elements that include at least one of the shape features and color tone features. The Avatar Analysis Unit 5 has the function of analyzing the appearance features of the avatar by performing image analysis processing on the appearance of the generated avatar. Specifically, it comprises an Avatar Shape Analysis Unit 15 that analyzes the shape of the avatar's appearance, an Avatar Color Tone Analysis Unit 16 that analyzes the color tone of the avatar's appearance, an Avatar Age Determination Unit 17 that determines the content of the avatar's appearance age elements based on the analysis results of the Avatar Shape Analysis Unit 15 and the Avatar Color Tone Analysis Unit 16, an Avatar Gender Determination Unit 18 that determines the content of the avatar's appearance gender elements based on the analysis results of the Avatar Shape Analysis Unit 15 and the Avatar Color Tone Analysis Unit 16, and an Avatar Color Tone Determination Unit 19 that determines the content of partial color tone elements in the avatar based on the analysis results of the Avatar Shape Analysis Unit 15 and the Avatar Color Tone Analysis Unit 16.
[0032] The avatar shape analysis unit 15 is for analyzing the external shape of the avatar, that is, the two-dimensional or three-dimensional shape of the avatar's external surface. In this embodiment 1, the avatar shape analysis unit 15 performs shape analysis by extracting feature points based on the avatar's appearance using a mechanism similar to that of the target shape analysis unit 9. However, performing the same analysis process as the target shape analysis unit 9 is limited to cases where the avatar does not contain information about feature points, such as when an avatar is simply generated by making minimal modifications to a model's face photograph. For example, if feature points have already been extracted during avatar generation in the avatar generation unit 4, the shape analysis process may be completed by simply reusing the information extracted at that time.
[0033] The avatar color tone analysis unit 16 is for analyzing the color tones (saturation, brightness, and hue) used in the appearance of the avatar. The avatar color tone analysis unit 16 performs color tone analysis based on the appearance of the avatar using the same mechanism as the target color tone analysis unit 10. However, performing the same analysis process as the target color tone analysis unit 10 is limited to cases where the avatar does not contain color tone information as surface information, such as a simple avatar with only minimal modifications made to a model's face photograph. If color tone information on the appearance surface is generated during avatar generation in the avatar generation unit 4, that information may be reused.
[0034] The avatar age determination unit 17 is for determining the content of the avatar's outward age elements. Specifically, the avatar age determination unit 17 has the function of determining the content of the avatar's outward age elements based on the shape characteristics and / or color characteristics of the avatar's outward appearance. In this embodiment 1, the avatar age determination unit 17 may determine the content of the outward age elements (for example, the age estimated from the outward appearance of the person the avatar is targeting is 30 years old) based on specific indicators such as shape characteristics on the avatar's face, such as wrinkles at the corners of the eyes and the depth of nasolabial folds, and color characteristics such as the presence and degree of blemishes on the skin, similar to the target age determination unit 11. Alternatively, it may be equipped with a determination function regarding the content of the outward age elements by deep learning, machine learning, or reinforcement learning. Furthermore, if the avatar age determination unit 17 has obtained information about the age of a specific person that will serve as the model when the avatar is generated in the avatar generation unit 4, it may also determine the content of the outward age elements based on that information. On the other hand, even if information regarding the model's age is available, this information may be ignored, and the content of the apparent age element may be determined solely based on the shape and / or color characteristics of the avatar's appearance.
[0035] The avatar gender determination unit 18 is for determining the content of the avatar's outward age elements. Specifically, the avatar gender determination unit 18 has the function of determining the content of the avatar's outward gender elements based on the shape and / or color characteristics of the avatar's appearance. In this embodiment 1, the avatar gender determination unit 18 may determine the content of the outward age elements (for example, that the avatar has the outward characteristics of a male) based on specific indicators such as shape differences based on skeletal differences between males and females, the fineness of skin tone, and color differences such as the redness of the lips, similar to the target gender determination unit 12, or it may be equipped with a determination function regarding the content of outward gender elements by deep learning, machine learning, or reinforcement learning. Furthermore, if the avatar gender determination unit 18 has obtained information regarding the gender of a specific person to be used as a model when the avatar is generated in the avatar generation unit 4, it may also determine the content of the outward gender elements based on that information. On the other hand, even if information regarding the model's gender is available, this information may be ignored, and the content of the outward gender element may be determined solely based on the shape and / or color characteristics of the avatar's appearance.
[0036] The avatar color tone determination unit 19 is for determining the content of partial color elements of the avatar. In this embodiment 1, specific examples of partial color elements are the color characteristics of the skin, hair, and eyes on the avatar's face, similar to the partial color elements determined by the target color tone determination unit 13. The determination process in the avatar color tone determination unit 19 is performed based on the analysis results of the avatar shape analysis unit 15, which identify the area corresponding to the partial color element to be determined (skin, hair, eyes, etc. on the face), and then based on the analysis results of the avatar color tone analysis unit 16 regarding that area. However, it is not essential for the avatar color tone analysis unit 16 to determine the specific values of the color tone (saturation, lightness, hue) of the entire area of the partial color element to be determined as part of the determination process. For example, it may extract only the most frequently occurring combination of the three elements saturation, lightness, and hue, or it may derive the frequently occurring values or average values of saturation, lightness, and hue.
[0037] The Avatar Modification Unit 6 determines, based on the analysis results of the Feature Content Determination Unit 3 and the Avatar Analysis Unit 5, whether the content of the feature elements in the target of replacement matches the content of the feature elements in the avatar generated by the Avatar Generation Unit 4. If they do not match, it modifies at least one of the shape or color tone of the avatar's appearance so that the content of the avatar's feature elements matches the content of the feature elements in the target of replacement. Specifically, the Avatar Modification Unit 6 has the function of determining whether there is consistency in content for each feature element on the appearance of the target of replacement and the avatar, while referring to the analysis results of the Feature Content Determination Unit 3 and the Avatar Analysis Unit 5, and modifying the appearance of the avatar for items that are determined not to match. To realize this function, the avatar modification unit 6 includes a modification element determination unit 20 that determines whether or not to modify the feature elements of the avatar based on the analysis results of the feature content determination unit 3 and the avatar analysis unit 5; an age element modification unit 21 that modifies the appearance age element of the avatar, which is one of the feature elements, according to the determination result of the modification element determination unit 20; a gender element modification unit 22 that modifies the appearance gender element of the avatar, which is one of the feature elements, according to the determination result of the modification element determination unit 20; and a color tone element modification unit 23 that modifies the color tone of the avatar's color tone element, which is one of the feature elements, according to the determination result of the modification element determination unit 20.
[0038] The modification element determination unit 20 is for determining whether or not modifications should be made to the feature elements of the avatar based on the determination results of the feature content determination unit 3 and the analysis results of the avatar analysis unit 5. Specifically, the modification element determination unit 20 compares the content of each feature element in the replacement target, which is the determination result of the feature content determination unit 3, with the content of each feature element in the avatar, which is the analysis result of the avatar analysis unit 5. For feature elements that do not match, it has the function of determining whether or not to modify at least one of the avatar's shape or color tone so that the content of the feature element in the avatar matches the content of the feature element in the replacement target. First, the modification element determination unit 20 compares the determination result of the target age determination unit 11 and the determination result of the avatar age determination unit 17 for the appearance age element, which is one of the feature elements, and determines whether or not the appearance age element of the avatar matches the content of the appearance age element of the replacement target. If the appearance is consistent, it is determined that there is no need to modify the avatar regarding the appearance age element. If the appearance is inconsistent, it is determined that at least one of the shape and color parts of the avatar's appearance that affect the content of the appearance age element should be modified to be consistent with the content of the appearance age element to be replaced. Note that "consistency in appearance age element" does not only mean that the age indicated by the appearance age element matches, but also includes cases where the age does not match but is similar. The modification element determination unit 20 pre-determines criteria for whether or not they are "similar" (for example, an age difference of 10 years or less), and if these criteria are met, it determines that "the appearance age element is consistent" even if the age does not match.
[0039] Similarly, the modification element determination unit 20 compares the determination result of the target gender determination unit 12 and the determination result of the avatar gender determination unit 18 with respect to the appearance gender element, which is one of the characteristic elements, and determines whether the content of the avatar's appearance gender element is consistent with the content of the appearance gender element to be replaced. If they are consistent, it is determined that there is no need to modify the avatar regarding the appearance gender element. On the other hand, if they are not consistent, it is determined that at least one of the shape and color parts that affect the content of the avatar's appearance gender element should be modified to be consistent with the content of the appearance gender element to be replaced. Note that "consistency in appearance gender elements" is a concept that includes not only cases where the gender indicated by the appearance gender element matches, but also cases where it is similar. Specifically, when gender is to include LGBT gender concepts in addition to male and female, a standard is set in advance regarding gender combinations that are similar to each other, and if the standard is met, it is determined that "the appearance gender elements are consistent" even if the gender does not match.
[0040] Furthermore, the modification element determination unit 20 compares the determination result of the target color tone determination unit 13 with the determination result of the avatar color tone determination unit 19 for a partial color tone element, which is one of the feature elements. If the determination results are consistent, it determines that there is no need to modify the avatar regarding the partial color tone element. On the other hand, if the determination results are inconsistent, it determines that the content of the partial color tone element in the avatar should be modified to be consistent with the content of the partial color tone element in the replacement target. In this embodiment 1, examples of partial color tone elements include the skin tone, hair tone, and eye tone of the face. Note that "consistency in partial color tone elements" is a concept that includes not only cases where the color tone (saturation, brightness, hue) of the partial color tone element matches, but also cases where it is similar. Specifically, for example, regarding skin tone, three categories—black, white, and yellow—can be set according to the range of saturation, lightness, and hue. If the skin tone of both the replacement target and the avatar belongs to the same category, even if the tones do not match, they can be considered similar and judged as having "partially consistent color elements." Similarly, regarding hair tone, four categories—black, blonde, brown, and white—can be set according to the range of saturation, lightness, and hue. If the hair tone of both the replacement target and the avatar belongs to the same category, even if the tones do not match, they can be considered similar and judged as having "partially consistent color elements." The same applies to eye tone; three categories—black, blue, and brown—can be set according to the range of saturation, lightness, and hue. If they belong to the same category, even if the tones do not match, they can be considered similar and judged as having "partially consistent color elements."
[0041] The age element modification unit 21 modifies at least one of the avatar's shape and color tone so that the avatar's appearance age element matches the appearance age element to be replaced, when the modification element determination unit 20 determines that the appearance age element should be modified. Specifically, the age element modification unit 21 has the function of adding or removing shape features that change with age, such as wrinkles around the eyes, nasolabial folds, and drooping eyelids, which affect the content determination of the appearance age element, and adding or removing color features that change with age, such as skin spots and luster on the face, which affect the content determination of the appearance age element. In a more preferred embodiment, the age element modification unit 21 may, for example, acquire information in advance regarding the average positional change of each feature point due to aging, and based on this information, change the position of the feature points on the avatar's face according to the content of the appearance age element to be matched. The specific modification method may be determined based on the results of deep learning, machine learning, or reinforcement learning. Furthermore, the concept of "consistency" in the modification by the age element modification unit 21 is the same concept as "consistency in the appearance age element" in the modification element determination unit 20, and includes not only cases where they are exactly the same, but also cases where they are similar. Therefore, as a result of the modification by the age element modification unit 21, there is no necessity for the appearance age element of the modified avatar to exactly match the appearance age element to be replaced. If the specific ages are close, even if the content of the appearance age element of the modified avatar differs from the content of the appearance age element to be replaced, it will be determined that it has been "modified to be consistent". Preferably, the avatar age determination unit 17 performs a determination process regarding the appearance age element on the modified avatar, and if it is determined that it is consistent with the appearance age element to be replaced, the modification process is considered complete. If it is determined that it is not consistent with the appearance age element to be replaced, the modification process is repeated.
[0042] The gender element modification unit 22 modifies at least one of the avatar's shape and color tone so that the avatar's appearance gender element matches the appearance gender element to be replaced, when the modification element determination unit 20 determines that the appearance gender element should be modified. Specifically, the gender element modification unit 22 has the function of adding or deleting shape features such as specific shapes based on skeletal differences between men and women that affect the content determination of the appearance gender element, and adding or deleting color features such as lip color, skin tone fineness, and the presence or absence of other makeup that affect the content determination of the appearance gender element. In a more preferred embodiment, for example, information on the average position change of feature points on the face corresponding to changes in shape features based on skeletal differences between men and women may be acquired in advance, and the gender element modification unit 22 may change the position of the feature points on the avatar's face according to gender based on this information. The specific modification method may be determined based on the results of deep learning, machine learning, or reinforcement learning. Furthermore, the concept of "consistency" in the modification by the gender element modification unit 22 is the same concept as "consistency in outward appearance gender elements" in the modification element determination unit 20. More preferably, after modification by the gender element modification unit 22, the avatar gender determination unit 18 performs a determination regarding the outward appearance gender elements of the avatar. If it is determined that the modification is consistent with the outward appearance gender elements to be replaced, the modification is considered complete. If it is determined that the modification is not consistent with the outward appearance gender elements to be replaced, the modification is repeated.
[0043] The color tone element correction unit 23 is for correcting the color tone of the avatar so that the partial color tone element of the avatar matches the partial color tone element to be replaced, when the correction element determination unit 20 determines that a partial color tone element should be corrected. Specifically, if the partial color tone element is the pupil, for example, the color tone element correction unit 23 will perform a process to correct the pupil color so that the pupil color of the avatar matches the pupil color of the replacement target. Note that the concept of "matching" in the color tone correction process by the color tone element correction unit 23 is the same concept as "matching in color tone" in the correction element determination unit 20, and includes not only cases where the colors are exactly the same but also cases where they are similar (for example, when the degree of difference is within a certain range). Regarding skin color, three categories—black, white, and yellow—can be set for each certain color range, and elements belonging to the same category can be treated as "matching" as partial color elements. Similarly, for hair, four categories—black hair, blonde hair, brown hair, and white hair—can be set for each certain color range, and for eyes, three categories—black, blue, and brown—can be set, and elements belonging to the same category can be treated as "matching" as partial color elements. Preferably, the avatar color determination unit 19 performs a determination process regarding partial color elements on the modified avatar, and if the determination result is that the color tone of the element to be modified matches the color tone to be replaced, the modification is considered complete; if the determination result is that they do not match, the modification is repeated.
[0044] The avatar insertion unit 7 is for inserting an avatar into the target video after the avatar modification unit 6 has performed the necessary modification processing, specifically by the age element modification unit 21, the gender element modification unit 22, and the color tone element modification unit 23. Specifically, the avatar insertion unit 7 has the function of inserting the modified avatar into the target video by replacing the replacement target with the modified avatar. If the target video is a still image, the avatar insertion unit 7 acquires information regarding the size and direction of the replacement target in the target video, and after appropriately changing the size and direction of the avatar based on that information, inserts it into the target video in a manner that replaces the replacement target. If the target video is a video, the avatar insertion unit 7 acquires information regarding the operation of the replacement target in the target video, for example by detecting changes in the position of feature points, and after appropriately changing the size, direction and operation of the avatar based on that information, inserts it into the target video in a manner that replaces the replacement target.
[0045] The video output unit 8 is for outputting the video into which the avatar has been inserted by the avatar insertion unit 7 to an external source. The output method of the video output unit 8 may be to output the video into which the avatar has been inserted as is, or to extract and output only a part of the video into which the inserted avatar appears, such as the scene in which the inserted avatar appears. Furthermore, the output format may not be limited to outputting the video in video format, but may also be, for example, printing out a scene from the video and outputting it in the form of a postcard or a small-sized photograph.
[0046] Next, the advantages of the avatar insertion system according to this embodiment 1 will be explained. First, the avatar insertion system according to this embodiment 1 has a configuration in which, regarding feature elements which are elements related to external appearance, the content of both the replacement target and the avatar is extracted and compared, and if the content of the feature elements of each does not match, the shape and color tone of the avatar are modified to match the content of the feature elements of the replacement target, and then inserted into the video. With this configuration, the avatar insertion system according to this embodiment 1 has the advantage that, because an avatar that matches the replacement target in terms of feature elements is inserted, the avatar harmonizes naturally with the entire video after insertion, and the occurrence of a sense of incongruity for the viewer of the video can be effectively suppressed.
[0047] Furthermore, the avatar insertion system according to this embodiment 1 uses outward appearance age elements and outward appearance gender elements as specific examples of characteristic elements, and has a configuration in which the replacement target and the avatar are inserted into the video in a state in which they are consistent with these elements. In the case where the avatar consists only of the head and not the whole body, as in this embodiment 1, the torso and limbs of the replacement target remain as they are in the video, and only the head is replaced. The remaining torso and limbs usually have an appearance that strongly reflects the age and gender of the replacement target (for example, behavior in videos, or physique and clothing in videos and still images, etc.), and if there is no consistency between the replacement target and the avatar in the content of these characteristic elements, it will give an extremely unnatural impression in terms of appearance. This embodiment 1 employs a configuration that aligns these characteristic elements, and therefore has the advantage of effectively suppressing the occurrence of a sense of incongruity for viewers of the video after avatar insertion.
[0048] Furthermore, the avatar insertion system according to this embodiment 1 uses partial color elements related to the skin, hair, and eye color of the face as characteristic elements, and has a configuration in which the replacement target and the avatar are inserted into the video in a state where they are consistent with these elements. These elements will have different content reflecting the race, social status, etc. of the replacement target, and often serve as indicators of the role of the replacement target in the video (for example, a person visiting Japan from a foreign country). Therefore, in the avatar insertion system according to this embodiment 1, these elements are also treated as characteristic elements, thereby achieving the advantage of effectively suppressing the occurrence of a sense of incongruity for viewers of the video after the avatar has been inserted.
[0049] (Embodiment 2) Next, an avatar insertion system according to Embodiment 2 will be described. In Embodiment 2, components that have the same name and reference numerals as those in Embodiment 1 will perform the same functions as those in Embodiment 1 unless otherwise specified.
[0050] As shown in Figure 2, the avatar insertion system according to Embodiment 2 newly includes an appearance feature extraction unit 24 that extracts appearance feature portions of an avatar based on the results of a comparison with the appearance of a standard avatar consisting of a standard shape and color tone, a standard avatar database 25 that stores information on the appearance structure of a standard avatar used during the extraction process in the appearance feature extraction unit 24, and a protected area information generation unit 26 that sets the area on the avatar corresponding to the appearance feature portion extracted by the appearance feature extraction unit 24 as a protected area and generates protected area information including information on said area. The age element correction unit 27, gender element correction unit 28, and color tone element correction unit 29 have the function of performing correction processing on the appearance surface of the avatar according to the content of the protected area information generated by the protected area information generation unit 26.
[0051] The appearance feature extraction unit 24 is for extracting appearance feature parts, which are the characteristic features of the appearance of the avatar. Specifically, the appearance feature extraction unit 24 compares numerical information regarding the shape and color tone of the appearance recorded in the standard avatar database 25 (described later) with numerical information regarding the shape and color tone of the avatar's appearance, and has the function of extracting shape and color tone elements with a large discrepancy in the numerical information as appearance feature parts. More specifically, the appearance feature extraction unit 24 extracts shape elements and / or color tone elements corresponding to numerical information regarding the shape and color tone of the avatar that deviates significantly from the numerical information regarding the shape and color tone of the corresponding standard avatar, for example, when the difference value between the two is greater than a certain threshold, as appearance feature parts. As a specific extraction method, the appearance feature extraction unit 24 generates, for example, location information of the region on the surface of the avatar where the appearance feature parts exist, as well as information regarding the content of the appearance feature parts (at least information regarding whether it is a shape feature or a color feature).
[0052] In this embodiment 2, numerical information relating to the shape is used, specifically numerical information relating to the position of feature points (e.g., position coordinates) and numerical information relating to the positional relationship between multiple feature points (e.g., 3D vector information indicating the direction from one feature point to another and the distance between the two feature points). For color tone, numerical information representing saturation, brightness, and hue is used, although numerical information derived by other calculation methods may also be used. Furthermore, in this embodiment 2, the numerical information relating to the shape and color tone of the avatar's appearance is derived based on the analysis results of the avatar shape analysis unit 15 and the avatar color tone analysis unit 16, but numerical information may also be derived by other methods, such as using the information used during avatar generation.
[0053] The standard avatar database 25 stores numerical information regarding the shape and color tone of a standard avatar's appearance, which is used during the extraction process in the appearance feature extraction unit 24. Specifically, the standard avatar database 25 has the function of storing numerical information regarding the shape and color tone of a standard avatar, which is an avatar having a standard shape and color tone. As for the method of calculating the "numerical information regarding the shape and color tone of a standard avatar's appearance," the average value of numerical information regarding the feature points and color tone of a large number of avatars generated in the past may be used, or the numerical information may be calculated based on the appearance of a single avatar that is estimated to have a standard shape. Preferably, the standard avatar database 25 sets multiple avatars with different attributes (age group, gender, race, etc.) as standard avatars and stores numerical information regarding the shape and color tone of the appearance of each. The appearance feature extraction unit 24 extracts numerical information related to the shape and color tone of standard avatars that match the attributes of the avatar targeted for extraction of appearance features from the numerical information stored in the standard avatar database 25, and performs comparison processing, thereby enabling more accurate comparison processing that reflects the actual situation. It is preferable to use the determination results from the avatar age determination unit 17, the avatar gender determination unit 18, etc., as a method for determining the attributes of the avatar.
[0054] The protected area information generation unit 26 sets the area on the avatar corresponding to the location of the appearance feature portion extracted by the appearance feature extraction unit 24 as a protected area, generates protected area information which is information related to the setting, and outputs it to the age element modification unit 27, etc. Specifically, the protected area information generation unit 26 has the function of setting the area including the area where the appearance feature portion is located as a protected area based on the extraction result of the appearance feature extraction unit 24, and generating protected area information which includes information about the location of the protected area and information about restrictions on modification processing for the protected area. The information about the location of the protected area consists of information about the location of a certain area range which includes the entire appearance feature portion, and this area range may include the surrounding area of the appearance feature portion. As for the information about restrictions on modification processing, it can be anything as long as it protects the appearance feature portion, such as not allowing any modification processing on the avatar portion within the protected area, not allowing only shape modification or color tone modification, or not allowing only a part of the shape modification, and there are no particular restrictions on the content.
[0055] The age element modification unit 27, gender element modification unit 28, and color tone element modification unit 29 have functions in addition to those of the age element modification unit 21, gender element modification unit 22, and color tone element modification unit 23 in Embodiment 1, respectively, to perform modification processing on the appearance of the avatar according to the information regarding the limitations of modification processing defined in the protected area information set and generated by the protected area information generation unit 26. By performing modification processing on the appearance of the avatar according to the protected area information which contains content that protects the appearance feature portion, the avatar insertion system according to Embodiment 2 makes it possible to insert an avatar into the target video while retaining the appearance feature portion of the generated avatar, and which is consistent with the content of the appearance age element and the appearance gender element to be replaced.
[0056] Next, the advantages of the avatar insertion system according to this second embodiment will be described. In this second embodiment, the external features of the avatar are extracted, and protective area information is generated, which includes location information of the region containing the part having said features and the content of the modification restrictions on the external features. The avatar modification unit then modifies the appearance of the avatar according to this information. By adopting this configuration, the appearance of the avatar is modified so that the feature elements of the avatar are consistent with the feature elements to be replaced, similar to the first embodiment. At the same time, modifications are restricted to the avatar's unique external features that deviate from the standard structure of the avatar's appearance. This results in the advantage that the appearance of the avatar (and by extension, the real person who served as the model for the avatar) can still be displayed while suppressing the occurrence of incongruity in the video after avatar insertion. [Industrial applicability]
[0057] The present invention can be used as a technique for inserting an avatar into a video, such as a still image or a video, by replacing at least a portion of the person to be replaced with the avatar. [Explanation of Symbols]
[0058] 1. Video Input Section 2. Extraction section for replacement targets 3. Feature content determination unit 4. Avatar Generation Section 5. Avatar Analysis Department 6. Avatar Modification Section 7 Avatar insertion section 8. Video output section 9. Target Shape Analysis Section 10 Target Color Analysis Section 11 Target Age Determination Section 12 Target Gender Determination Unit 13 Target color tone determination unit 15 Avatar Shape Analysis Department 16 Avatar Color Analysis Department 17 Avatar Age Determination Section 18 Avatar Gender Determination Unit 19 Avatar color tone determination unit 20 Modification element determination section 21, 27 Age element correction section 22, 28 Gender element correction section 23, 29 Tonal element correction section 24 External feature extraction unit 25 Standard Avatar Database 26 Protected area information generation unit
Claims
1. An avatar insertion system that inserts an avatar into a video by replacing at least a portion of the human figures contained in the video with an avatar, A means for extracting the replacement target placed in the aforementioned video, A feature content determination means for determining the content of a feature element in the extracted replacement target that includes at least one of a shape feature and a color feature, If the content of the feature element in the replacement target determined by the feature content determination means does not match the content of the feature element in the avatar, the avatar modification means modifies at least one of the shape and color tone of the avatar to match the content of the feature element in the replacement target. An avatar insertion means inserts the avatar into the video by placing the avatar, which has been modified by the avatar modification means, in the portion of the video where the object to be replaced was located. Output means for outputting the video after the avatar has been inserted by the insertion means to the outside, An avatar insertion system characterized by having the following features.
2. The feature elements subject to determination by the feature content determination means are the apparent age element and apparent gender element in the replacement target, The aforementioned avatar modification means is If the content of the aforementioned appearance age element does not match that of the replacement target and the avatar, the age element modification means modifies at least one of the shape and color tone of the avatar so that the content of the aforementioned appearance age element in the avatar matches the content of the aforementioned appearance age element in the replacement target. A gender element modification means for modifying at least one of the shape and color tone of the avatar so that the content of the gender element in the avatar is consistent with the content of the gender element in the replacement target when the content of the gender element in the replacement target is inconsistent with the content of the gender element in the avatar, The avatar insertion system according to claim 1, further comprising the above.
3. The feature elements extracted by the feature content determination means are partial color elements which are color features of one or more parts of the skin, hair, or pupils of the face to be replaced. The avatar insertion system according to claim 1, further comprising a color tone element correction means for correcting the color tone of the avatar so that the content of the partial color tone element in the avatar matches the content of the partial color tone element in the replacement target when the content of the partial color tone element does not match that of the replacement target.
4. An avatar insertion method for inserting an avatar into a video by replacing at least a portion of the human figures included in the video with an avatar, A step of extracting the replacement target that is placed in the aforementioned video, A feature content determination step of determining the content of a feature element in the extracted replacement target that includes at least one of a shape feature and a color feature, If the content of the feature element in the replacement target determined in the feature content determination step does not match the content of the feature element in the avatar, an avatar modification step is performed to modify at least one of the shape and color tone of the avatar so that it matches the content of the feature element in the replacement target. An avatar insertion step involves inserting the avatar into the video by placing the avatar that has been modified in the avatar modification step into the portion of the video where the replacement target was located. An output step for outputting the video after the avatar insertion in the insertion step to the outside, A method for inserting an avatar, characterized by including [a specific element].
5. An appearance feature extraction step in which the appearance features of the avatar are extracted based on the results of comparing the appearance of the aforementioned avatar with the appearance of a standard avatar consisting of a standard shape and color tone, A protected area information generation step includes setting the area on the avatar that includes the visual feature portion extracted in the visual feature extraction step as a protected area, and generating protected area information that includes location information of the set protected area and information regarding the limitations on the modifications made within the protected area. It further includes, The avatar insertion method according to claim 4, characterized in that, in the avatar modification step, even if the content of the feature element in the replacement target determined in the feature content determination step does not match the content of the appearance features in the avatar, the shape and / or color tone of the avatar is modified with respect to the protected area in accordance with the information regarding the limitations on modification content included in the protected area information.
6. An avatar insertion program that causes a computer to insert an avatar into a video by replacing at least a portion of the human figures contained in the video with an avatar, To the aforementioned computer, A replacement target extraction function that extracts the replacement targets placed in the aforementioned video, A feature content determination function that determines the content of feature elements in the extracted replacement target that include at least one of the shape features and the color features, An avatar modification function that modifies at least one of the shape and color tone of the avatar to match the content of the feature element in the replacement target determined by the feature content determination function, if the content of the feature element in the avatar does not match the content of the feature element in the replacement target. An avatar insertion function that inserts the avatar into the video by placing the avatar, which has been modified by the avatar modification function, in the portion of the video where the replacement target was located, An output function that outputs the video after the avatar has been inserted using the insertion function to an external source, An avatar insertion program characterized by causing the following to be executed.