Avatar Insertion System, Avatar Insertion Method, and Avatar Insertion Program

The avatar insertion system addresses the challenge of inserting avatars into videos by correcting the avatar's shape and color to match the replacement target's features, thereby reducing viewer discomfort and maintaining the avatar's characteristics.

JP7699875B1Active Publication Date: 2025-06-30POCKETRD CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024196526
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-06-30
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing technologies face challenges in effectively inserting avatars into videos while maintaining the characteristics of the avatar and avoiding a sense of incongruity in viewers, particularly due to the diversity of human appearances and the need for individual adjustments.

Method used

An avatar insertion system that extracts and compares feature elements such as shape and color features of the replacement target and the avatar, and corrects the avatar's shape and color to match the replacement target's features, while maintaining the avatar's original characteristics through protection area information.

Benefits of technology

The system effectively suppresses the occurrence of a sense of incongruity in viewers by ensuring the avatar is harmonized with the video, while maintaining the avatar's original characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699875000001_ABST
    Figure 0007699875000001_ABST
Patent Text Reader

Abstract

Replace the person in the image with an avatar. 【Solution means】 The avatar insertion system includes a video input unit 1 for inputting a video to be inserted with an avatar, a replacement target extraction unit 2 for extracting a replacement target that is to be replaced with an avatar from the video input through the video input unit 1, a feature content determination unit 3 for determining the content of a feature element including at least one of the shape feature and the color feature in the appearance of the replacement target extracted by the replacement target extraction unit 2, an avatar generation unit 4 for generating an avatar to be inserted into the insertion target video, an avatar analysis unit 5 for performing image analysis on the shape and color of the appearance of the generated avatar, and comparing the content of the feature element of the replacement target with the content of the feature element of the avatar based on the determination result of the feature content determination unit 3 and the result of the image analysis by the avatar analysis unit 5, and an avatar correction unit 6 for correcting the shape and color of the appearance of the avatar to match the content of the feature element of the replacement target when they do not match each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for inserting an avatar into a video by replacing at least a part of a replacement target, which is a person in the video such as a still image or a moving image, with the avatar.

Background Art

[0002] In recent years, with the improvement of processing capabilities in electronic computers such as computers, a large number of computer graphics of human figures, so-called avatars, that reflect the characteristics of real people have been widely used. For example, an avatar is used as one's own icon for use on SNS (Social Networking Service), or one's own avatar is used as the character of the protagonist in an online game or the like. In addition, in existing content such as still images and moving images, a service has been proposed in which some of the characters appearing are replaced with one's own avatar for viewing and the like.

[0003] By using an avatar that reflects one's own characteristics in this way, for example, when a protagonist consisting of an avatar that expresses the characteristics of a user in a game battles against an enemy character, an effect of improving the user's sense of immersion in the game world is produced. By using an avatar that abstractly represents the user himself / herself as an icon indicating the user on SNS, it is expected that an effect such as promoting communication between users in a virtual space in the same sense as the real world will occur.

[0004] Both Patent Documents 1 and 2 disclose a technique of using an avatar that imitates the actual appearance of the player himself / herself or a co-player in a computer game that expresses a virtual space using a head-mounted display.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

[0006] However, since the appearances of real people are highly diverse, in order to use an avatar based on the characteristics of a real person, it is necessary to create the avatar individually, and the complexity of avatar generation becomes a problem. Also, for example, when replacing the characters appearing in existing content with one's own avatar, if the appearance of the character and one's own appearance are too different (for example, when replacing a character of an infant with a middle-aged man), using an avatar based on the characteristics of a real person as it is will rather create an unnatural impression and prevent the user from getting immersed. On the other hand, if the appearance is the same as that of the characters in the content, there is no meaning in using the user's avatar, and ultimately it is necessary to individually adjust the appearance of the avatar manually or the like, but such a response is too complicated. Although there are these problems regarding the use of an avatar based on the characteristics of a real person, neither Patent Document 1 nor 2 discloses any technology for solving such problems.

[0007] The present invention has been made in view of the above problems, and when inserting an avatar into a video such as a still image or a moving image by replacing at least a part of the replacement target, which is a person in the video, with the avatar, an object of the present invention is to provide a technology capable of effectively suppressing the occurrence of a sense of incongruity in the viewer of the video after the avatar insertion while maintaining the characteristics of the avatar. [Means for Solving the Problems]

[0008] To achieve the above object, an avatar insertion system according to claim 1 is an avatar insertion system that inserts an avatar into a video by replacing at least a part of a replacement target, which is a human figure included in the video, with the avatar. The system includes: a replacement target extraction means for extracting the replacement target arranged in the video; a feature content determination means for determining the content of a feature element including at least one of a shape feature and a color feature in the extracted replacement target; an avatar correction means for correcting at least one of the shape and color of the avatar so that the content of the feature element in the replacement target matches the content of the feature element in the avatar when the content of the feature element in the replacement target determined by the feature content determination means does not match the content of the feature element in the avatar; an avatar insertion means for inserting the avatar into the video by arranging the avatar corrected by the avatar correction means at the part where the replacement target was arranged in the video; and an output means for outputting the video after the avatar is inserted by the insertion means to the outside. Appearance feature extraction means for extracting the appearance feature portion of the avatar based on the result of comparing the appearance of the avatar with the appearance of a standard avatar having a standard shape and color tone; Protection area information generation means for setting, as a protection area, an area on the avatar that includes the appearance feature portion extracted by the appearance feature extraction means, and generating protection area information including the position information of the set protection area and information regarding restrictions on modification content in the protection area; comprising , even when the content of the feature element in the replacement target determined by the feature content determination means does not match the content of the appearance feature of the avatar, the avatar modification means modifies the shape and / or color tone of the avatar according to the information regarding restrictions on modification content included in the protection area information with respect to the protection area. characterized in that.

[0009] Further, to achieve the above object, an avatar insertion system according to claim 2 is, in the above invention, wherein the feature element to be determined by the feature content determination means is an appearance age element and an appearance gender element in the replacement target. The avatar correction means further includes an age element correction means for correcting at least one of the shape and color of the avatar so that the content of the appearance age element in the avatar matches the content of the appearance age element in the replacement target when the replacement target and the avatar do not match with respect to the content of the appearance age element, and a gender element correction means for correcting at least one of the shape and color of the avatar so that the content of the appearance gender element in the avatar matches the content of the appearance gender element in the replacement target when the replacement target and the avatar do not match with respect to the content of the appearance gender element.

[0010] In order to achieve the above object, in the avatar insertion system according to claim 3, in the above invention, the feature elements extracted by the feature content determination means are partial color tone elements that are color tone features in one or more of the skin, hair, or pupils of the face to be replaced, and when the avatar modification means determines that the avatar does not match the replacement target regarding the content of the partial color tone element, the avatar modification means further includes color tone element modification means for modifying the color tone of the avatar so that the content of the partial color tone element in the avatar matches the content of the partial color tone element in the replacement target.

[0011] In order to achieve the above object, an avatar insertion method according to claim 4 is an avatar insertion method for inserting an avatar into a video by replacing at least a part of a replacement target, which is a portrait included in the video, with the avatar. The method includes a replacement target extraction step of extracting the replacement target arranged in the video, a feature content determination step of determining the content of feature elements including at least one of the shape feature and the color tone feature in the extracted replacement target, an avatar modification step of modifying at least one of the shape and color tone of the avatar so that the content of the feature elements in the replacement target matches the content of the feature elements in the avatar when the content of the feature elements in the replacement target determined in the feature content determination step does not match the content of the feature elements in the avatar, and an avatar insertion step of inserting the avatar modified in the avatar modification step into the part where the replacement target was arranged in the video, and an output step of outputting the video after the avatar insertion in the insertion step to the outside. An appearance feature extraction step of extracting the appearance feature portion of the avatar based on the result of comparing the appearance of the avatar with the appearance of a standard avatar having a standard shape and color tone; and a protection area information generation step of setting, as a protection area, an area on the avatar that includes the appearance feature portion extracted in the appearance feature extraction step, and generating protection area information including the position information of the set protection area and information regarding restrictions on modification content in the protection area; and Including, in the avatar modification step, even when the content of the feature element in the replacement target determined in the feature content determination step does not match the content of the appearance feature of the avatar, modifying the shape and / or color tone of the avatar according to the information regarding restrictions on modification content included in the protection area information with respect to the protection area. is characterized by.

[0013] Also, to achieve the above object, the avatar insertion program according to claim 6 is an avatar insertion program for causing a computer to insert an avatar into a video by replacing at least a part of a replacement target, which is a human figure included in the video, with the avatar. The program includes: a replacement target extraction function for extracting the replacement target arranged in the video from the computer; a feature content determination function for determining the content of a feature element including at least one of a shape feature and a color tone feature in the extracted replacement target; an avatar correction function for correcting at least one of the shape and color tone of the avatar so that the content of the feature element in the replacement target matches the content of the feature element in the avatar when the content of the feature element in the replacement target determined by the feature content determination function does not match the content of the feature element in the avatar; an avatar insertion function for inserting the avatar into the video by arranging the avatar corrected by the avatar correction function at a part where the replacement target was arranged in the video; and an output function for outputting the video after the avatar is inserted by the insertion function to the outside. An appearance feature extraction function for extracting the appearance feature portion of the avatar based on the result of comparing the appearance of the avatar with the appearance of a standard avatar having a standard shape and color tone; and a protection area information generation function for setting, as a protection area, an area on the avatar that includes the appearance feature portion extracted by the appearance feature extraction function, and generating protection area information including the position information of the set protection area and information regarding restrictions on modification content in the protection area; to execute, In the execution of the avatar correction function, even if the content of the feature element in the replacement target determined by the feature content determination function does not match the content of the appearance feature in the avatar, with respect to the protection area, the shape and / or color tone of the avatar is corrected according to the information regarding the restriction of the correction content included in the protection area information characterized by.

Effect of the Invention

[0014] According to the present invention, when inserting an avatar into a video such as a still image or a moving image by replacing at least a part of a replacement target, which is a human in the video, with the avatar, it is possible to effectively suppress the occurrence of a sense of incongruity in the viewer of the video after the avatar is inserted while maintaining the characteristics of the avatar.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Modes for Carrying Out the Invention

[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following embodiments, examples considered to be the most appropriate as embodiments of the present invention are described. Of course, the content of the present invention should not be construed as being limited to the specific examples shown in these embodiments. Any configuration that exhibits the same operations and effects, even if it is other than the specific configuration shown in the embodiments, is of course included in the technical scope of the present invention.

[0017] (Embodiment 1) First, an avatar insertion system according to Embodiment 1 will be described. As shown in FIG. 1, the avatar insertion system according to Embodiment 1 of the present invention includes a video input unit 1 that inputs an insertion target video, which is a video into which an avatar is to be inserted, a replacement target extraction unit 2 that extracts a replacement target, which is a target to be replaced, from the video input through the video input unit 1, a feature content determination unit 3 that determines the content of a feature element including at least one of the shape feature and the color tone feature in the appearance of the replacement target extracted by the replacement target extraction unit 2, an avatar generation unit 4 that generates an avatar to be inserted into the insertion target video, an avatar analysis unit 5 that performs image analysis on the shape and color tone of the appearance of the generated avatar, a comparison is made between the content of the feature element of the replacement target and the content of the feature element of the avatar based on the determination result of the feature content determination unit 3 and the result of the image analysis by the avatar analysis unit 5, and when they do not match each other, an avatar correction unit 6 that corrects the shape and color tone of the appearance of the avatar to match the content of the feature element of the replacement target, and an avatar insertion unit 7 that inserts the corrected avatar by the avatar correction unit 6 into the insertion target video by arranging the corrected avatar in the portion where the replacement target is arranged in the insertion target video, and a video output unit 8 that outputs the video with the avatar inserted to the outside.

[0018] The video input unit 1 is for inputting an insertion target video, which is a video for which avatar insertion processing is to be performed in the avatar insertion system according to Embodiment 1. The composition of the insertion target video may be either a moving image or a still image, and may be either a three-dimensional image or a two-dimensional image, and there is no particular limitation as long as it is visual information. Similarly, there is no particular limitation on the content of the insertion target video, but it is assumed that the video includes at least a video related to a replacement target that is to be replaced with an avatar. Furthermore, the insertion target video may be one specially generated for avatar insertion, but it may also be possible to use general video contents such as off-the-shelf movies and landscape photos.

[0019] The replacement target extraction unit 2 is for extracting a replacement target, which is a target to be replaced with an avatar, from the insertion target video input via the video input unit 1. Specifically, the replacement target extraction unit 2 has a function of specifying where in the insertion target video the specified replacement target is displayed and specifying the displayed position and the like. As a method for specifying the replacement target, it may be a method specified by a predetermined person such as a user, or after performing image analysis processing on the insertion target video to extract one or more replacement target candidates, a candidate that conforms to a predetermined criterion such as the user's preference may be selected. In addition, the replacement target extraction unit 2 preferably has a function of selecting the display position and the like of the replacement target in a plurality of scenes when the replacement target appears in a plurality of scenes such as video contents. Information regarding the extraction process of the replacement target by the replacement target extraction unit 2 is transmitted to the avatar insertion unit 7 described later, and the avatar insertion unit 7 performs avatar replacement processing based on the information. Although there is no particular need to particularly limit the replacement target extracted by the replacement target extraction unit 2, it is preferably a human or a creative creature imitating a human. In the following description of Embodiment 1, it is assumed that the replacement target is a human.

[0020] The feature content determination unit 3 is for determining the content of predetermined feature elements including the shape features and color tone features in the appearance of the replacement target extracted by the replacement target extraction unit 2 by using technologies such as image analysis. In the present invention, the "feature element" includes, in addition to the specific structure itself composed of the external shape, external color tone, and the combination of shape and color tone in the replacement target or avatar, the characteristics and features of the object inferred based on a plurality of specific aspects and combinations of shapes, color tones, etc. In the first embodiment, the appearance age element and the appearance gender element are used as examples of the feature elements as the characteristics and features of the object. Also, as an example of the feature element as the external color tone, the partial color tone element (specific examples include the skin color, hair color, and pupil color on the face) is used. However, it is of course possible to use other elements as the feature elements in the present invention.

[0021] Specifically, the feature content determination unit 3 includes a target shape analysis unit 9 that analyzes the shape features on the appearance surface of the replacement target, a target color tone analysis unit 10 that analyzes the color tone features in the appearance of the replacement target, a target age determination unit 11 that determines the content of the appearance age element of the replacement target, which is one of the feature elements, based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10, a target gender determination unit 12 that determines the content of the appearance gender element of the replacement target, which is one of the feature elements, based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10, and a target color tone determination unit 13 that determines the content of the partial color tone element of the replacement target based on the analysis results of the target shape analysis unit 9 and the target color tone analysis unit 10.

[0022] The target shape analysis unit 9 is for analyzing the external shape of the replacement target, that is, the two-dimensional and three-dimensional shapes on the external surface of the replacement target. Specifically, the target shape analysis unit 9 in the first embodiment acquires information regarding the shape of the replacement target by extracting the positional relationship of feature points predetermined corresponding to the characteristic parts on the external surfaces of each part such as eyes, nose, mouth, ears, etc. in the human body from the external appearance of the replacement target displayed in the video to be inserted. As a specific configuration for extracting the feature points, image recognition technology realized by deep learning, machine learning, etc. may be used, or a combination of multiple image recognition technologies may be used. Also, it is possible to change the analysis content according to the format of the video to be inserted (for example, analyze the two-dimensional positional relationship of the feature points if it is a two-dimensional video), or conversely, regardless of the format of the video to be inserted, for example, extract the three-dimensional shape (that is, the three-dimensional positional relationship of the feature points) even if the video to be inserted is a two-dimensional video.

[0023] The "feature points" extracted by the target shape analysis unit 9 indicate the positions on the body surface of the human body parts. However, it can be arbitrarily determined which parts should be set as feature points. For example, it is also possible to handle setting feature points only for visually recognizable parts (eyes, nose, mouth, ears, eyebrows, etc.) on the face or points corresponding to the contour, or to set feature points not only for the face but also for parts on the torso and limbs. Also, when setting feature points for the eyes, it is possible to use only the center position of the eyes as the feature points, or to set feature points in detail for each of the inner corner of the eye, outer corner of the eye, pupil, and sclera. The target shape analysis unit 9 in the first embodiment of the present invention generates the analysis result in a format including information on the positions of the extracted feature points (preferably the positions in the normalized coordinate system) and the meanings of the feature points (points corresponding to the outer corners of the eyes, points corresponding to the top of the nose, etc.). However, when used as materials for the determination processes in the target age determination unit 11, target gender determination unit 12, and target color tone determination unit 13 described later, it is not necessary to extract all the feature points corresponding to each part. For example, in relation to the target age determination unit 11, it is also possible to adopt a configuration in which only some feature points, such as only the feature points of the parts corresponding to the wrinkles on the face, are extracted.

[0024] The target color tone analysis unit 10 is for analyzing the color tone used in the appearance of the replacement target. Specifically, since the appearance of the replacement target is composed of an arrangement of points (pixels) of various color tones (chroma, lightness, hue), the target color tone analysis unit 10 has a function of analyzing such color tones. As a specific aspect of the color tone analysis process by the target color tone analysis unit 10, it is also possible to perform color tone analysis on all parts of the appearance of the replacement target. However, in the first embodiment of the present invention, it is also possible to perform color tone analysis only within the range necessary for the determination processes in the target age determination unit 11, target gender determination unit 12, and target color tone determination unit 13. For example, it is also possible to perform color tone analysis only on the face instead of the entire appearance of the replacement target, or more simply, to perform color tone analysis only on the part corresponding to a predetermined element to be determined by the target color tone determination unit 13.

[0025] The target age determination unit 11 is for determining the content of the appearance age element to be replaced, which is one of the characteristic elements. Specifically, the target age determination unit 11 determines the content of the appearance age element of the replacement target, such as being 15 years old, 30 years old, etc., based on the shape characteristics on the appearance surface of the replacement target, which is the analysis result of the target shape analysis unit 9, and / or the color characteristics on the appearance surface of the replacement target, which is the analysis result of the target color tone analysis unit 10. The determination mechanism in the target age determination unit 11 preferably makes a comprehensive determination from both the shape characteristics and the color characteristics. However, for example, it may be determined only by the shape characteristics such as the wrinkles at the outer corners of the eyes and the depth of nasolabial folds, or it may be determined by the color characteristics such as the presence and degree of spots on the facial skin. Further, the target age determination unit 11 may determine the content of the appearance age element based on specific indicators such as the presence and degree of facial wrinkles described above, or may be provided with a determination function regarding the content of the appearance age element by deep learning, machine learning, or reinforcement learning.

[0026] Note that the content of the appearance age element determined by the target age determination unit 11 does not necessarily only refer to the actual age of the replacement target, but is a concept that also includes the age estimated from the appearance surface of the replacement target. For example, when the video to be inserted is content such as a movie, for the actor who is the replacement target, the actual age may be different from the age of the role played in the video to be inserted. In such a case, the target age determination unit 11 may be configured to precisely determine both the actual age and the age of the role. However, more preferably, the target age determination unit 11 determines the content of the appearance age element that is externally estimated regardless of the actual age or the age of the role.

[0027] The target gender determination unit 12 is for determining the content of the appearance gender element to be replaced, which is one of the characteristic elements. Specifically, the target gender determination unit 12 has a function of determining the appearance gender element of the replacement target based on the shape characteristics on the appearance surface of the replacement target, which is the analysis result of the target shape analysis unit 9, and / or the color tone characteristics on the appearance surface of the replacement target, which is the analysis result of the target color tone analysis unit 10. The determination mechanism in the target gender determination unit 12 preferably comprehensively determines the content of the appearance gender element, such as having characteristics as a man in appearance, etc., from both the shape characteristics and the color tone characteristics. However, for example, it may also be determined only from the difference in shape characteristics based on the skeletal difference between men and women, or only from the color tone characteristics such as the fineness of the skin color tone and the difference in the lip color tone in view of the tendency of women to wear makeup. Further, the target gender determination unit 12 may determine the content of the appearance gender element based on specific indicators such as the lip color tone, or may be provided with a determination function regarding the content of the appearance gender element by deep learning, machine learning, or reinforcement learning.

[0028] Note that the content of the appearance gender element determined by the target gender determination unit 12 does not necessarily only refer to the actual gender of the replacement target, but is a concept that also includes the gender estimated from the appearance surface of the replacement target. For example, regarding gender, biological ones, those based on gender self-identification, and further, in the case where the insertion target video is content such as a movie, cases where a man plays a woman, a woman plays a man, etc. are assumed. In such cases, the target gender determination unit 12 may, for example, accurately determine objective facts such as being biologically female but having a male gender self-identification. More preferably, the target gender determination unit 12 does not get stuck in biological gender self-identification, etc., and determines the content of the appearance gender element based on the criterion of which gender appearance is possessed.

[0029] The target color tone determination unit 13 is for determining the content of the partial color tone element of the replacement target, which is one of the characteristic elements. As the partial color tone element, any element can be used as long as it is likely to produce a color tone characteristic according to the attributes of the replacement target and the like, and the color tone characteristic is likely to attract visual attention. In the first embodiment, the skin area, hair, and pupils of the face of the replacement target are used as the partial color tone elements, and the target color tone determination unit 13 determines the content thereof. For example, the color of the skin area varies greatly for each replacement target depending on differences in race, presence or absence of sunburn, etc., and has a strong visual impression. Similarly, there are various forms of hair color such as blond, black, and brown, and the color of the pupils is also various such as black, blue, and brown. Since both the hair and pupils have a strong visual impression, the target color tone determination unit 13 in the first embodiment determines the color tone characteristics of the skin, hair, and pupils on the face of the replacement target. Note that the content of the determination process of the color tone characteristics by the target color tone determination unit 13 is performed based on the analysis result of the target color tone analysis unit 10 for the region corresponding to the element to be determined (such as the skin, hair, and pupils of the face) after grasping the region according to the analysis result of the target shape analysis unit 9. However, it is not essential to determine the specific values of the color tone (chroma, lightness, hue) of the entire region of the element to be determined as the determination process of the target color tone analysis unit 10. For example, only the most frequently occurring combination among the combinations of the three elements of chroma, lightness, and hue may be extracted, or the frequently occurring values or average values of chroma, lightness, and hue may be derived respectively.

[0030] The avatar generation unit 4 is for generating an avatar to be inserted into the target video to be inserted by replacing the target to be replaced. As the configuration of the avatar, it may be composed of skeleton information (bones), surface information (skin), and weight information that defines the relationship between the two. However, it is also preferable to adopt a configuration consisting only of surface information including information on three-dimensional shapes and color tones on the surface. As an even simpler configuration, it may be composed of the image data itself. Also, although the avatar generation unit 4 may generate an avatar corresponding to a full-body image of a person, in the first embodiment, an avatar consisting only of the upper head from the neck is generated. The avatar generation unit 4 extracts feature points set according to the positions and shapes of features on the surface (such as eyes, eyebrows, nose, mouth, ears, hairstyle, etc.) and internal features (such as joints) based on the face image of the model person (which may be two-dimensional or three-dimensional), and reflects the positional relationship between the extracted feature points in the avatar, thereby generating a realistic avatar that reflects the physical characteristics of the model. However, when adopting a simpler configuration, for example, it may be configured to perform only the minimum necessary corrections such as adjusting the size and orientation of the face image of the model.

[0031] The avatar analysis unit 5 is for analyzing the shape and color tone of the appearance of the avatar generated by the avatar generation unit 4 and determining the content of the feature elements including at least one of the shape feature and the color tone feature. The avatar analysis unit 5 has a function of analyzing the appearance features of the generated avatar by performing image analysis processing or the like on the appearance of the avatar. Specifically, it includes an avatar shape analysis unit 15 for analyzing the shape of the appearance of the avatar, an avatar color tone analysis unit 16 for analyzing the color tone of the appearance of the avatar, an avatar age determination unit 17 for determining the content of the appearance age element of the avatar based on the analysis results of the avatar shape analysis unit 15 and the avatar color tone analysis unit 16, an avatar gender determination unit 18 for determining the content of the appearance gender element of the avatar based on the analysis results of the avatar shape analysis unit 15 and the avatar color tone analysis unit 16, and an avatar color tone determination unit 19 for determining the content of the partial color tone element in the avatar based on the analysis results of the avatar shape analysis unit 15 and the avatar color tone analysis unit 16.

[0032] The avatar shape analysis unit 15 is for analyzing the external shape of the avatar, that is, the two-dimensional or three-dimensional shape on the external surface of the avatar. The avatar shape analysis unit 15 in the first embodiment performs shape analysis by extracting feature points based on the appearance of the avatar using the same mechanism as the target shape analysis unit 9. However, performing the same analysis process as the target shape analysis unit 9 is limited to cases where the avatar does not contain information regarding feature points, such as when an avatar is simply generated by making only minimal corrections to a model's face photo. For example, when feature points have already been extracted during avatar generation in the avatar generation unit 4, the shape analysis process may be completed by directly using the information at the time of extraction.

[0033] The avatar color tone analysis unit 16 is for analyzing the color tone (chroma, lightness, hue) used in the appearance of the avatar. The avatar color tone analysis unit 16 performs color tone analysis based on the appearance of the avatar using the same mechanism as the target color tone analysis unit 10. However, performing the same analysis process as the target color tone analysis unit 10 is limited to cases where the avatar does not contain information regarding the color tone as surface information, such as in the case of a simple avatar with only minimal corrections made to a model's face photo. When information regarding the color tone on the external surface is generated during avatar generation in the avatar generation unit 4, such information may be directly used.

[0034] The avatar age determination unit 17 is for determining the content of the appearance age elements of the avatar. Specifically, the avatar age determination unit 17 has a function of determining the content of the appearance age elements of the avatar based on the shape features and / or color tone features on the appearance surface of the avatar. The avatar age determination unit 17 in the first embodiment of the present invention may determine the content of the appearance age elements (for example, the age estimated from the appearance surface of the person indicated by the avatar is 30 years old, etc.) based on specific indicators such as the crow's feet at the outer corners of the eyes and the depth of nasolabial folds on the face of the avatar and color tone features such as the presence or absence and degree of freckles on the skin, similar to the target age determination unit 11. It may also be configured to have a determination function regarding the content of the appearance age elements by deep learning, machine learning, or reinforcement learning. Further, when the avatar age determination unit 17 has acquired information regarding the age of a specific person who serves as a model at the time of avatar generation in the avatar generation unit 4, it may determine the content of the appearance age elements based on the information. On the other hand, even when information regarding the age of the model has been acquired, it may be determined to ignore the information and determine the content of the appearance age elements based solely on the shape features and / or color tone features on the appearance surface of the avatar.

[0035] The avatar gender determination unit 18 is for determining the content of the appearance age element of the avatar. Specifically, the avatar gender determination unit 18 has a function of determining the content of the appearance gender element of the avatar based on the shape characteristics and / or color tone characteristics on the appearance surface of the avatar. The avatar gender determination unit 18 in the first embodiment of the present invention may determine the content of the appearance age element (for example, the avatar has appearance characteristics as a male, etc.) based on specific indicators such as the shape differences based on the skeletal differences between men and women, the fineness of the skin color tone, and the color tone differences such as the redness of the lips, similar to the target gender determination unit 12. It may also be provided with a determination function regarding the content of the appearance gender element by deep learning, machine learning, or reinforcement learning. Furthermore, when the avatar gender determination unit 18 acquires information regarding the gender of a specific person who serves as a model during avatar generation in the avatar generation unit 4, it may determine the content of the appearance gender element based on the information. On the other hand, even when information regarding the gender of the model is acquired, it may be possible to ignore the information and determine the content of the appearance gender element based solely on the shape characteristics and / or color tone characteristics on the appearance surface of the avatar.

[0036] The avatar color tone determination unit 19 is for determining the content of the partial color tone element of the avatar. As a specific example of the partial color tone element, in the first embodiment of the present invention, it is the color tone characteristics of the skin, hair, and pupils on the face of the avatar, similar to the partial color tone element determined by the target color tone determination unit 13. The determination process in the avatar color tone determination unit 19 is performed based on the analysis result of the avatar color tone analysis unit 16 regarding the region corresponding to the partial color tone element (such as the skin, hair, and pupils of the face) to be determined, after grasping the region corresponding to the partial color tone element to be determined by the analysis result of the avatar shape analysis unit 15. However, it is not essential to determine the specific values of the color tone (saturation, lightness, hue) of the entire region of the partial color tone element to be determined as the determination process of the avatar color tone analysis unit 16. For example, it may be possible to extract only the most frequently occurring combination among the combinations of the three elements of saturation, lightness, and hue, or to derive the frequently occurring values or average values of saturation, lightness, and hue respectively.

[0037] Based on the analysis results of the feature content determination unit 3 and the avatar analysis unit 5, the avatar correction unit 6 determines whether the content of the feature elements in the replacement target matches the content of the feature elements in the avatar generated by the avatar generation unit 4. When they do not match, it is for correcting at least one of the shape or color tone on the appearance surface of the avatar so that the content of the feature elements of the avatar matches the content of the feature elements in the replacement target. Specifically, while referring to the analysis results of the feature content determination unit 3 and the avatar analysis unit 5, the avatar correction unit 6 determines the presence or absence of consistency in terms of content for each feature element on the appearance surfaces of the replacement target and the avatar, and has a function of correcting the appearance of the avatar for the items determined to be inconsistent. To realize such a function, the avatar correction unit 6 includes a correction element determination unit 20 that determines whether to perform correction on the feature elements in the avatar based on the analysis results of the feature content determination unit 3 and the avatar analysis unit 5, an age element correction unit 21 that corrects the appearance age element of the avatar, which is one of the feature elements, according to the determination result of the correction element determination unit 20, a gender element correction unit 22 that corrects the appearance gender element of the avatar, which is one of the feature elements, according to the determination result of the correction element determination unit 20, and a color tone element correction unit 23 that corrects the color tone in the color tone element of the avatar, which is one of the feature elements, according to the determination result of the correction element determination unit 20.

[0038] The modification element determination unit 20 is for determining whether or not to modify the characteristic elements in the avatar based on the determination result of the characteristic content determination unit 3 and the analysis result of the avatar analysis unit 5. Specifically, the modification element determination unit 20 compares the content of each characteristic element in the replacement target, which is the determination result by the characteristic content determination unit 3, with the content of each characteristic element in the avatar, which is the analysis result by the avatar analysis unit 5, and for the characteristic elements that do not match, it has the function of determining to modify at least one of the shape and color tone of the avatar so that the content of the characteristic element in the avatar matches the content of the characteristic element in the replacement target. First, the modification element determination unit 20 compares the determination result of the target age determination unit 11 and the determination result of the avatar age determination unit 17 for the appearance age element, which is one of the characteristic elements, and determines whether or not the appearance age element of the avatar matches the content of the appearance age element of the replacement target. If they match, it is determined that there is no need to modify the avatar regarding the appearance age element, and if they do not match, it is determined that at least one of the shape and color tone parts that affect the content of the appearance age element in the appearance of the avatar should be modified to match the content of the appearance age element of the replacement target. Note that "matching in the appearance age element" does not mean only the case where the ages indicated by the appearance age element are the same, but is a concept that includes cases where the ages are not the same but are similar. The modification element determination unit 20 determines in advance the criteria (for example, the age difference is 10 years or less, etc.) for whether or not they are "similar", and if the criteria are met, it determines that "the appearance age elements match" even if the ages do not match.

[0039] Similarly, the modification element determination unit 20 compares the determination result of the target gender determination unit 12 and the determination result of the avatar gender determination unit 18 for the appearance gender element, which is one of the characteristic elements, and determines whether the content of the appearance gender element of the avatar matches the content of the appearance gender element to be replaced. If they match, it is determined that there is no need to modify the avatar with respect to the appearance gender element. On the other hand, if they do not match, it is determined that at least one of the shape and color tone portions that affect the content of the appearance gender element of the avatar should be modified to match the content of the appearance gender element to be replaced. Note that "matching in the appearance gender element" is a concept that includes not only the case where the genders indicated by the appearance gender elements are the same but also the case where they are similar. Specifically, when including LGBT-like gender concepts in addition to male and female for gender, criteria for combinations of genders that are mutually similar in advance are defined, and when the criteria are met, it is determined that "the appearance gender elements match" even if the genders do not match.

[0040] Further, the correction element determination unit 20 compares the determination result of the target color tone determination unit 13 and the determination result of the avatar color tone determination unit 19 for a partial color tone element which is one of the characteristic elements. When the determination results match, it is determined that there is no need to correct the avatar regarding the partial color tone element. On the other hand, when the determination results do not match, it is determined that the content of the partial color tone element in the avatar should be corrected to match the content of the partial color tone element in the replacement target. In the first embodiment, examples of the partial color tone element include the skin color tone, hair color tone, and eye color tone on the face. Note that "matching in the partial color tone element" is a concept that includes not only the case where the color tones (saturation, lightness, hue) in the partial color tone element are the same but also the case where they are similar. Specifically, for example, regarding the skin color tone, three categories of black, white, and yellow are set according to the ranges of saturation, lightness, and hue. When the skin color tones of both the replacement target and the avatar belong to the same category, even if the color tones of the two do not match exactly, they are considered similar, and it may be determined that "the partial color tone elements match". Similarly, for the hair color tone, four categories of black hair, blond hair, brown hair, and white hair are set according to the ranges of saturation, lightness, and hue. When the hair color tones of both the replacement target and the avatar belong to the same category, even if the color tones of the two do not match exactly, they are considered similar to each other, and it may be determined that "the partial color tone elements match". The same applies to the eye color tone. Three categories of black, blue, and brown are set according to the ranges of saturation, lightness, and hue. When they belong to the same category, even if the color tones of the two do not match exactly, they are considered similar, and it may be determined that "the partial color tone elements match".

[0041] When the modification element determination unit 20 determines that the external age element should be modified, the age element modification unit 21 modifies at least one of the shape and color tone of the avatar so that the external age element of the avatar matches the external age element to be replaced. Specifically, the age element modification unit 21 adds or deletes shape features that change according to age, such as crow's feet, nasolabial folds, and eyelid drooping, which affect the determination of the content of the external age element, or adds or deletes color tone features that change according to age, such as facial skin spots and luster, which affect the determination of the content of the external age element. In a more preferred embodiment, the age element modification unit 21 may, for example, acquire in advance information on the average positional change of each feature point due to aging, and change the position of the feature points on the avatar's face according to the content of the external age element to be matched based on the information. Regarding the specific modification mode, it may be determined based on the results of deep learning, machine learning, or reinforcement learning. Note that the concept of "matching" in the modification by the age element modification unit 21 is the same concept as "matching in the external age element" in the modification element determination unit 20, and includes not only the case of complete coincidence but also the case of similarity. Therefore, as a result of the modification by the age element modification unit 21, there is no necessity for the external age element of the modified avatar to completely match the external age element to be replaced. Even if the content of the external age element of the modified avatar is different from the content of the external age element to be replaced as long as the specific ages are close, it is determined that the modification has been made "to match". Preferably, the avatar age determination unit 17 performs a determination process on the external age element of the modified avatar. If it is determined that the modification matches the external age element to be replaced, the modification process is completed, and if it is determined that the modification does not match the external age element to be replaced, the re-modification is repeated.

[0042] When the modification element determination unit 20 determines that the external gender element should be modified, the gender element modification unit 22 modifies at least one of the shape and color tone of the avatar so that the external gender element of the avatar matches the external gender element to be replaced. Specifically, the gender element modification unit 22 adds or deletes shape features such as specific shapes based on the skeletal differences between men and women that affect the determination of the content of the external gender element, and adds or deletes color tone features such as the color of the lips, the fineness of the skin color, and the presence or absence of other makeup that affect the determination of the content of the external gender element. As a more preferable aspect, for example, information regarding the average positional change of the feature points on the face corresponding to the change in the shape feature based on the skeletal difference between men and women is acquired in advance, and the gender element modification unit 22 changes the positions of the feature points on the face of the avatar according to the gender based on the information. Regarding the specific modification aspect, it may be determined based on the results of deep learning, machine learning, or reinforcement learning. Note that the concept of "matching" in the modification by the gender element modification unit 22 is the same concept as "matching in the external gender element" in the modification element determination unit 20. More preferably, after the modification by the gender element modification unit 22, the avatar gender determination unit 18 determines the external gender element of the avatar. If it is determined that it matches the external gender element to be replaced, the modification is completed, and if it is determined that it does not match the external gender element to be replaced, the re-modification may be repeated.

[0043] When the correction element determination unit 20 determines that a partial color tone element should be corrected, the color tone element correction unit 23 corrects the color tone of the avatar so that the partial color tone element of the avatar matches the partial color tone element to be replaced. Specifically, for example, when the partial color tone element is the pupil, the color tone correction unit 23 performs a process of correcting the color tone of the pupil so that the color tone of the pupil in the avatar matches the color tone of the pupil in the replacement target. Note that the concept of "matching" in the color tone correction process by the color tone element correction unit 23 is the same concept as "matching in color tone" in the correction element determination unit 20, and includes not only the case where the color tones completely match but also the case where they are similar (for example, when the degree of difference is within a certain range). For the skin color, three categories of black, white, and yellow can be set for each certain color tone range, and when belonging to the same category, it may be treated as "matching" as a partial color tone element. Similarly, for the hair, four categories of black hair, blond hair, brown hair, and white hair can be set for each certain color tone range, and for the pupils, three categories of black, blue, and brown can be set, and when belonging to the same category, it may be treated as "matching" as a partial color tone element. Preferably, the avatar color tone determination unit 19 performs a determination process on the partial color tone element for the corrected avatar. If the determination result is that the color tone of the element to be corrected matches the color tone of the replacement target, the correction is completed, and if the determination result is that they do not match, the re-correction may be repeated.

[0044] The avatar insertion unit 7 is for inserting, into the video to be inserted, the avatar that has been subjected to the necessary correction processing by the avatar correction unit 6, specifically, the age element correction unit 21, the gender element correction unit 22, and the color tone element correction unit 23. Specifically, the avatar insertion unit 7 has a function of inserting into the video to be inserted through the process of replacing the avatar that has been subjected to the necessary correction processing with the replacement target. When the video to be inserted is a still image, the avatar insertion unit 7 acquires information regarding the size and direction of the replacement target in the video to be inserted, appropriately changes the size and direction of the avatar based on the information, and then inserts it into the video to be inserted in a manner of replacing the replacement target. When the video to be inserted is a moving image, the avatar insertion unit 7 acquires information regarding the motion content of the replacement target in the video to be inserted by detecting, for example, the positional variation of feature points, appropriately changes the size, direction, and motion of the avatar based on the information, and then inserts it into the video to be inserted in a manner of replacing the replacement target.

[0045] The video output unit 8 is for externally outputting the video to be inserted into which the avatar has been inserted by the avatar insertion unit 7. As the output mode by the video output unit 8, it may be to output the video to be inserted as it is, or it may be to extract and output only a part of the video to be inserted such as the scene where the inserted avatar appears. Also, as the output form, not only when outputting the video in the form of a video, but for example, it may be to print out one scene of the video and output it in the form of a postcard or a small-sized photo.

[0046] Next, the advantages of the avatar insertion system according to Embodiment 1 will be described. First, the avatar insertion system according to Embodiment 1 extracts and compares the contents of both the replacement target and the avatar with respect to the feature elements, which are elements related to the external characteristics, and corrects the shape and color tone of the avatar when the contents of the feature elements of each other do not match so as to match the contents of the feature elements of the replacement target, and then inserts the corrected avatar into the video. By having such a configuration, the avatar insertion system according to Embodiment 1 has the advantage that the inserted avatar is harmonized with the entire video naturally in the video after insertion because the avatar that is matched with the replacement target in terms of the feature elements, and the generation of a sense of incongruity in the viewer who watches the video can be effectively suppressed.

[0047] In addition, the avatar insertion system according to Embodiment 1 uses the external age element and the external gender element as specific examples of the feature elements, and has a configuration in which the replacement target and the avatar are inserted into the video in a state where they are matched in terms of these elements. When the avatar is composed of only the head instead of a full-body image as in Embodiment 1, the body and limbs of the replacement target remain as they are in the video and only the head is replaced. Usually, the remaining body and limbs have an appearance that strongly reflects the age and gender of the replacement target (for example, behavior in a video, physique, clothing, etc. in a video or a still image). If the consistency between the replacement target and the avatar is not achieved in the contents of these feature elements, it will give a very unnatural impression externally. Since Embodiment 1 adopts a configuration for matching these feature elements, from this aspect as well, it has the advantage that the generation of a sense of incongruity in the viewer who watches the video after the avatar is inserted can be effectively suppressed.

[0048] Furthermore, the avatar insertion system according to Embodiment 1 uses partial color tone elements related to the skin, hair, and eye color of the face as characteristic elements, and has a configuration in which the replacement target and the avatar are inserted into the video in a state where they are matched by these elements. These elements have different contents reflecting the race, social status, etc. of the replacement target, and often serve as indicators showing the role of the replacement target in the video (for example, a person who visits Japan from abroad). Therefore, in the avatar insertion system according to Embodiment 1, these elements are also treated as characteristic elements, thereby realizing the advantage of effectively suppressing the occurrence of discomfort in viewers who watch the video after avatar insertion.

[0049] (Embodiment 2) Next, the avatar insertion system according to Embodiment 2 will be described. In Embodiment 2, with respect to the components having the same name and the same reference numerals as those in Embodiment 1, unless otherwise specified, they exhibit the same functions as the components in Embodiment 1.

[0050] As shown in FIG. 2, the avatar insertion system according to Embodiment 2 newly includes an appearance feature extraction unit 24 that extracts the appearance feature part of the avatar based on the result of comparing with the appearance of a standard avatar having a standard shape and color tone, a standard avatar database 25 that stores information regarding the appearance structure of the standard avatar used during the extraction process in the appearance feature extraction unit 24, and a protection area information generation unit 26 that sets the area on the avatar corresponding to the appearance feature part extracted by the appearance feature extraction unit 24 as a protection area and generates protection area information including information regarding the area. The age element correction unit 27, the gender element correction unit 28, and the color tone element correction unit 29 have a function of performing correction processing on the appearance surface of the avatar according to the content of the protection area information generated by the protection area information generation unit 26.

[0051] The appearance feature extraction unit 24 is for extracting the appearance feature part, which is the appearance feature part of the avatar. Specifically, the appearance feature extraction unit 24 compares the numerical information regarding the shape and color tone in the appearance recorded in the standard avatar database 25 described later with the numerical information regarding the shape and color tone in the appearance of the avatar, and has a function of extracting the shape and color tone elements with a large deviation in numerical information as the appearance feature part. More specifically, the appearance feature extraction unit 24 extracts, as the appearance feature part, the shape element or / and the color tone element corresponding to the numerical information with a large deviation from the numerical information regarding the shape and color tone of the corresponding standard avatar among the numerical information regarding the shape and color tone of the avatar, for example, the difference value between the two companies is equal to or greater than a certain threshold. As a specific extraction mode, the appearance feature extraction unit 24 generates, for example, information regarding the content of the appearance feature part (information regarding at least whether it is a shape feature or a color tone feature) in addition to the position information of the area where the appearance feature part exists on the surface of the avatar.

[0052] In addition, in the second embodiment, as the numerical information regarding the shape, the numerical information regarding the position of the feature points (for example, position coordinates) and the numerical information regarding the positional relationship between a plurality of feature points (for example, the three-dimensional vector information indicating the direction from one feature point to another feature point and the distance between the two feature points) are used, and regarding the color tone, the numerical information obtained by quantifying the saturation, lightness, and hue is used. However, it is also possible to use the numerical information derived by other calculation methods. Also, in the second embodiment, the derivation of the numerical information regarding the shape and color tone in the appearance of the avatar is based on the analysis results of the avatar shape analysis unit 15 and the avatar color tone analysis unit 16. However, it is also possible to derive the numerical information by other methods, such as deriving the numerical information using the information used at the time of avatar generation.

[0053] The standard avatar database 25 is for storing numerical information regarding the shape and color tone in the appearance of the standard avatar, which is used during the extraction process in the appearance feature extraction unit 24. Specifically, the standard avatar database 25 has a function of storing numerical information regarding the shape and color tone on the appearance surface for the standard avatar, which is an avatar having a standard shape and color tone. As a method for calculating the "numerical information regarding the shape and color tone in the appearance of the standard avatar", it may be to use the average value of the numerical information regarding the feature points and color tone of a large number of avatars generated in the past, or it may be to calculate the numerical information based on the appearance of a single avatar estimated to have a standard shape. Preferably, the standard avatar database 25 sets a plurality of avatars with different attributes (age group, gender, race, etc.) as the standard avatar, and stores the numerical information regarding the shape and color tone in each appearance. The appearance feature extraction unit 24 extracts and compares the numerical information regarding the shape and color tone of the standard avatar with an attribute that matches the attribute of the avatar to be extracted for the appearance feature part among the numerical information stored in the standard avatar database 25, thereby making it possible to perform a more accurate comparison process that conforms to the actual situation. As a method for grasping the attributes of the avatar, it is preferable to use the determination results in the avatar age determination unit 17, avatar gender determination unit 18, etc.

[0054] The protection area information generation unit 26 sets, as a protection area, an area on the avatar corresponding to the position of the appearance feature part extracted by the appearance feature extraction unit 24, generates protection area information which is information regarding the setting, and outputs it to the age element correction unit 27 and the like. Specifically, based on the extraction result of the appearance feature extraction unit 24, the protection area information generation unit 26 sets, as a protection area, an area including the area where the appearance feature part is located, and has a function of generating protection area information including information regarding the position of the protection area and information regarding the restriction of the correction process for the protection area. The information regarding the position of the protection area shall consist of information regarding the position of a certain area range including the whole of the appearance feature part, and the area range may include the peripheral area of the appearance feature part. As the information regarding the restriction of the correction process, anything can be adopted as long as it can protect the appearance feature part, such as not allowing any correction process for the avatar part within the protection area, allowing only one of shape correction or color tone correction, or allowing only some modes of shape correction, and there is no particular restriction on the content aspect.

[0055] In addition to the functions respectively provided by the age element correction unit 21, the gender element correction unit 22, and the color tone correction unit 23 in the first embodiment, the age element correction unit 27, the gender element correction unit 28, and the color tone correction unit 29 have a function of performing a correction process on the appearance surface of the avatar according to the information regarding the restriction of the correction process defined in the protection area information set and generated by the protection area information generation unit 26. By performing a correction process on the appearance of the avatar according to the protection area information for protecting the appearance feature part, the avatar insertion system according to the second embodiment can insert, into the video to be inserted, an avatar that is consistent with the content of the appearance age element and the content of the appearance gender element of the replacement target while maintaining the appearance feature part of the generated avatar.

[0056] Next, the advantages of the avatar insertion system according to the second embodiment will be described. In the second embodiment, the appearance features of the avatar are extracted, and protection area information including the position information of the area including the part having the features and the content of the correction restriction for the appearance features is generated. The avatar correction unit corrects the appearance of the avatar according to the information. By adopting such a configuration, while correcting the appearance of the avatar so that the feature elements to be replaced match the content of the feature elements of the avatar as in the first embodiment, the correction of the appearance feature part unique to the avatar that deviates from the standard structure among the appearance of the avatar is restricted. In the video after the avatar is inserted, while suppressing the occurrence of discomfort, the features of the avatar (and thus the real person who is the model of the avatar) can still be displayed.

Industrial Applicability

[0057] The present invention can be used as a technology for inserting an avatar into a video by replacing at least a part of a replacement target, which is a person in the video such as a still image or a moving image, with the avatar.

Explanation of Signs

[0058] 1 Video input unit 2 Replacement target extraction unit 3 Feature content determination unit 4 Avatar generation unit 5 Avatar analysis unit 6 Avatar correction unit 7 Avatar insertion unit 8 Video output unit 9 Target shape analysis unit 10 Target color tone analysis unit 11 Target age determination unit 12 Target gender determination unit 13 Target color tone determination unit 15 Avatar shape analysis unit 16 Avatar color tone analysis unit 17 Avatar age determination unit 18 Avatar gender determination unit 19 Avatar color tone determination unit 20 Correction Element Determination Unit 21, 27 Age Element Correction Unit 22, 28 Gender Element Correction Unit 23, 29 Hue Element Correction Unit 24 Appearance Feature Extraction Unit 25 Standard Avatar Database 26 Protection Region Information Generation Unit

Claims

1. An avatar insertion system that inserts an avatar into a video by replacing at least a part of a replacement target, which is a human figure included in the video, with the avatar, comprising: a replacement target extraction means for extracting the replacement target arranged in the image; a feature content determining means for determining the content of a feature element including at least one of a shape feature and a color feature in the extracted replacement target; an avatar modification means for modifying at least one of a shape and a color of the avatar so as to match the content of the feature element of the replacement target when the content of the feature element of the replacement target determined by the feature content determination means does not match the content of the feature element of the avatar; an avatar inserting means for inserting the avatar into the video by arranging the avatar modified by the avatar modifying means in a portion of the video where a replacement target was located; an output means for outputting the video after the avatar is inserted by the inserting means to an outside; an appearance characteristic extraction means for extracting an appearance characteristic part of the avatar based on a result of comparing an appearance of the avatar with an appearance of a standard avatar having a standard shape and color tone; a protection area information generating means for setting an area on the avatar including the external feature portion extracted by the external feature extracting means as a protection area, and generating protection area information including position information of the set protection area and information regarding restrictions on modifications to the protection area; Equipped with The avatar modification means of the avatar insertion system is characterized in that, even if the content of the feature element of the replacement target determined by the feature content determination means does not match the content of the external features of the avatar, it modifies the shape and / or color of the avatar in accordance with information regarding restrictions on modification content contained in the protection area information for the protection area.

2. The feature elements to be judged by the feature content judging means are an appearance age element and an appearance gender element of the replacement object, The avatar modification means includes: an age element correcting means for correcting at least one of a shape and a color of the avatar when the content of the appearance age element of the avatar does not match that of the replacement target, so that the content of the appearance age element of the avatar matches that of the replacement target; a gender element correction means for correcting at least one of a shape and a color of the avatar when the content of the gender element of the appearance of the avatar does not match that of the replacement target, so that the content of the gender element of the appearance of the avatar matches that of the replacement target; 2. The avatar insertion system of claim 1, further comprising:

3. The avatar insertion system described in claim 1, characterized in that the feature elements extracted by the feature content determination means are partial color tone elements which are color tone features in one or more parts of the skin, hair or pupils of the face of the replacement target, and the avatar modification means further includes color tone element modification means for modifying the color tone of the avatar so that the content of the partial color tone element in the avatar matches the content of the partial color tone element in the replacement target when the replacement target and the avatar do not match in terms of the content of the partial color tone element.

4. 1. An avatar insertion method for inserting an avatar into a video by replacing at least a part of a replacement target, which is a human figure included in the video, with an avatar, comprising: a replacement target extraction step of extracting the replacement target arranged in the image; a feature content determination step of determining the content of a feature element including at least one of a shape feature and a color feature in the extracted replacement target; an avatar modification step of modifying at least one of a shape and a color of the avatar so as to match the content of the feature element of the replacement target when the content of the feature element of the replacement target determined in the feature content determination step does not match the content of the feature element of the avatar; an avatar insertion step of inserting the avatar modified in the avatar modification step into the video by arranging the avatar in a portion of the video where the replacement target was located; an output step of outputting the video after the avatar is inserted in the inserting step to an outside; an appearance feature extraction step of extracting an appearance feature portion of the avatar based on a result of comparing an appearance of the avatar with an appearance of a standard avatar having a standard shape and color tone; a protection area information generating step of setting an area on the avatar including the external feature portion extracted in the external feature extracting step as a protection area, and generating protection area information including position information of the set protection area and information regarding restrictions on modification content in the protection area; Including, An avatar insertion method characterized in that, in the avatar modification step, even if the content of the feature element of the replacement target determined in the feature content determination step does not match the content of the external features of the avatar, the shape and / or color of the avatar is modified in accordance with information regarding restrictions on the modification content contained in the protection area information for the protection area.

5. An avatar insertion program for causing a computer to insert an avatar into a video by replacing at least a part of a replacement target, which is a human figure included in the video, with the avatar, The computer is a replacement target extraction function for extracting the replacement target arranged in the video; a feature content determination function for determining the content of a feature element including at least one of a shape feature and a color feature in the extracted replacement target; an avatar correction function that corrects at least one of a shape and a color of the avatar so as to match the content of the feature element of the replacement target when the content of the feature element of the replacement target determined by the feature content determination function does not match the content of the feature element of the avatar; an avatar insertion function for inserting the avatar into the video by placing the avatar modified by the avatar modification function in a portion of the video where a replacement target was located; an output function for outputting the video after the avatar is inserted by the insertion function to an outside; an appearance feature extraction function for extracting an appearance feature portion of the avatar based on a result of comparing an appearance of the avatar with an appearance of a standard avatar having a standard shape and color tone; a protection area information generating function that sets an area on the avatar that includes the external feature portion extracted by the external feature extracting function as a protection area, and generates protection area information that includes position information of the set protection area and information regarding restrictions on modifications to the protection area; Run the command, An avatar insertion program characterized in that, when executing the avatar modification function, even if the content of the feature element of the replacement target determined by the feature content determination function does not match the content of the external features of the avatar, the shape and / or color of the avatar is modified in accordance with information regarding restrictions on modification content contained in the protection area information for the protection area.

Citation Information

Patent Citations

  • Image processing device

    JP2019009752A

  • Program for providing virtual space with head-mounted display, method, and information processing apparatus for executing program

    JP2019012509A

  • Information processing apparatus, information processing method, and computer program

    JP2019139673A

  • The avatar mask up-loading service based on facial recognition

    KR102198844B1