Composite image generation system, composite image generation method, and composite image generation program
The composite video generation system addresses the challenge of selecting appropriate insertion areas for avatars by using a system that sets multiple insertion regions and selects based on similarity, resulting in high-quality and artistically valuable composite videos.
Patent Information
- Application Number
- JP2025011829
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-28
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2045-01-28
AI Technical Summary
Existing technologies face challenges in appropriately selecting insertion areas for avatars in composite video generation, leading to issues with completion degree and artistic value, especially when multiple avatars are involved.
A composite video generation system that sets multiple insertion regions in an insertion target video, acquires region and avatar information, and selects insertion areas based on similarity between region and avatar attributes and appearances, ensuring appropriate placement of avatars.
The system effectively selects insertion areas for avatars, resulting in a natural and high-quality composite video that maintains artistic value and completion degree, even with multiple avatars.
Smart Images

Figure 0007699880000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for generating a composite video by inserting an avatar generated based on a person video into an insertion target video composed of one or more video materials.
Background Art
[0002] In recent years, with the improvement of processing capabilities in electronic computers such as computers, a large number of computer graphics of human figures, so-called avatars, reflecting the characteristics of real people have been widely used. For example, an avatar is used as one's own icon when using SNS (Social Networking Service), or one's own avatar is used as the character of the protagonist in an online game or the like. In addition, in contents such as still images and moving images, a service has been proposed in which some of the characters appearing are replaced with one's own avatar for viewing and the like.
[0003] By using an avatar that reflects one's own characteristics in this way, for example, when a protagonist consisting of an avatar expressing the characteristics of a user in a game battles against an enemy character, an effect of improving the user's sense of immersion in the game world occurs. By using an avatar that abstractly represents the user himself / herself as an icon indicating the user in SNS, it is expected that an effect such as promoting communication between users in a virtual space in the same sense as the real world will occur.
[0004] Both Patent Documents 1 and 2 disclose a technique of using an avatar that imitates the actual appearance of the player himself or a co-player in a computer game that expresses a virtual space using a head-mounted display.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, when generating a composite video in which some of the characters appearing in contents such as still images and moving images are replaced with avatars, the problem arises as to which character appearing in the content should be replaced with one's own avatar. While the mode in which the video generator selects the replacement target each time of creation is complicated, in the mode of randomly selecting the target, problems will occur in the completion degree, artistic value, etc. of the generated composite video. This problem becomes even more serious when a plurality of avatars are replaced with a plurality of characters appearing in the content, but neither Patent Document 1 nor 2 discloses any technology for solving such a problem.
[0007] The present invention has been made in view of the above problems, and an object thereof is to provide a technology capable of appropriately selecting an insertion area for an avatar when selecting an insertion area for inserting one or more avatars from among a plurality of insertion areas set in an insertion target video in composite video generation.
Means for Solving the Problems
[0008] To achieve the above object, the composite video generation system according to claim 1 is a composite video generation system that generates a composite video by inserting an avatar generated based on a predetermined person video into an insertion target image in which a plurality of insertion regions are set, including: an insertion target video generation means for generating an insertion target video based on one or more video materials; an insertion region setting means for setting a plurality of insertion regions for the insertion target video generated by the insertion target video generation means; an insertion region information acquisition means for acquiring insertion region information that is information regarding a display target in the insertion region set by the insertion region setting means and includes at least one of region attribute information that is information regarding an attribute of the display target and region appearance information that is information regarding an appearance of the display target; an avatar generation means for generating an avatar based on the person video; an avatar information acquisition means for acquiring avatar information that is information regarding the avatar and includes at least one of avatar attribute information that is information regarding an attribute of a representation target that the avatar represents and avatar appearance information that is information regarding an appearance of the avatar; an insertion region selection means for selecting an insertion region for inserting the avatar from among the plurality of insertion regions based on at least one of a similarity between the region attribute information regarding the plurality of insertion regions and the avatar attribute information and a similarity between the region appearance information regarding the plurality of insertion regions and the avatar appearance information; and a composite video generation means for generating the composite video by inserting the avatar into the insertion region selected by the insertion region selection means , the insertion area selection means selects a plurality of insertion areas for inserting the avatar, and the composite video generation means generates a plurality of composite videos inserted into each of the plurality of insertion areas selected by the insertion area selection means, and further includes video selection means for selecting a predetermined number of the composite videos from the plurality of composite videos based on browsing information which is information regarding the browsing mode for the plurality of composite videos characterized by comprising the above
[0010] Also, to achieve the above object, in the invention described above, the composite video generation system according to claim 2 further comprises a viewing count information acquisition means for acquiring viewing count information that is information regarding the number of times the composite video has been viewed as one of the viewing information, and a viewing reaction information acquisition means for acquiring viewing reaction information that is information regarding the reaction mode of a viewer during viewing of the composite video as one of the viewing information, and the video selection means is characterized by selecting one composite video based on the viewing count information and the viewing reaction information
[0011] Also, to achieve the above object, the synthetic video generation method according to claim 3 is a synthetic video generation method for generating a synthetic video in which an avatar generated based on a predetermined person video is inserted into an insertion target image in which a plurality of insertion regions are set. The method includes an insertion target video generation step of generating an insertion target video based on one or more video materials, an insertion region setting step of setting a plurality of insertion regions for the insertion target video generated in the insertion target video generation step, and an insertion region information acquisition step of acquiring insertion region information including at least one of region attribute information which is information regarding a display target in the insertion region set in the insertion region setting step and which is information regarding an attribute of the display target, and region appearance information which is information regarding an appearance of the display target. Based on the person video, a plurality of an avatar generation step of generating an avatar, an avatar information acquisition step of acquiring avatar information including at least one of avatar attribute information which is information regarding an attribute of a representation target which is a target represented by the avatar and avatar appearance information which is information regarding an appearance of the avatar, an insertion region selection step of selecting an insertion region for inserting the avatar from among the plurality of insertion regions based on at least one of a similarity between the region attribute information regarding the plurality of insertion regions and the avatar attribute information and a similarity between the region appearance information regarding the plurality of insertion regions and the avatar appearance information, and a synthetic video generation step of generating the synthetic video in which the avatar is inserted into the insertion region selected in the insertion region selection step , in the insertion area selection step, a matching order calculation step of assigning, for each avatar, a matching order of the insertion area according to the magnitude of a fitness which is numerical information calculated based on the similarity between the avatar information and the insertion area information for each of the plurality of avatars to a plurality of the insertion areas, and an insertion area determination step of determining a combination of the plurality of avatars and the plurality of insertion areas such that the total value of the matching order in the avatar for the insertion area combined with each avatar is minimized under the condition that the same insertion area is not selected for different avatars characterized by the above.
[0013] Also, to achieve the above object, the synthetic video generation method according to claim 4The composite video generation program according to the present invention causes a computer to generate a composite video in which an avatar generated based on a predetermined person video is inserted into an insertion target image in which a plurality of insertion areas are set. The program includes: an insertion target video generation function for generating an insertion target video based on one or more video materials for the computer; an insertion area setting function for setting a plurality of insertion areas for the insertion target video generated by the insertion target video generation function; an insertion area information acquisition function for acquiring insertion area information including at least one of area attribute information which is information about a display target in the insertion area set by the insertion area setting function and which is information about an attribute of the display target, and area appearance information which is information about an appearance of the display target; an avatar generation function for generating an avatar based on the person video; an avatar information acquisition function for acquiring avatar information including at least one of avatar attribute information which is information about an attribute of a representation target which is a target represented by the avatar and avatar appearance information which is information about an appearance of the avatar; an insertion area selection function for selecting an insertion area for inserting the avatar from among the plurality of insertion areas based on at least one of a similarity between the area attribute information for the plurality of insertion areas and the avatar attribute information and a similarity between the area appearance information for the plurality of insertion areas and the avatar appearance information; and a composite video generation function for generating a composite video in which the avatar is inserted into the insertion area selected by the insertion area selection function. a plurality of An avatar generation function for generating an avatar; an avatar information acquisition function for acquiring avatar information including at least one of avatar attribute information which is information about an attribute of a representation target which is a target represented by the avatar and avatar appearance information which is information about an appearance of the avatar; an insertion area selection function for selecting an insertion area for inserting the avatar from among the plurality of insertion areas based on at least one of a similarity between the area attribute information for the plurality of insertion areas and the avatar attribute information and a similarity between the area appearance information for the plurality of insertion areas and the avatar appearance information; and a composite video generation function for generating a composite video in which the avatar is inserted into the insertion area selected by the insertion area selection function. , in the insertion area selection function, a matching order calculation function of assigning, for each avatar, a matching order of the insertion area according to the magnitude of a fitness which is numerical information calculated based on the similarity between the avatar information and the insertion area information for each of the plurality of avatars to a plurality of the insertion areas, and an insertion area determination function of determining a combination of the plurality of avatars and the plurality of insertion areas such that the total value of the matching order in the avatar for the insertion area combined with each avatar is minimized under the condition that the same insertion area is not selected for different avatars characterized by causing the above to be executed.
Advantages of the Invention
[0014] According to the present invention, when selecting an insertion area for inserting one or more avatars from among a plurality of insertion areas set in an insertion target video in composite video generation, there is an effect that an appropriate insertion area for the avatar can be selected.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Embodiments for Carrying Out the Invention
[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following embodiments, examples considered to be the most appropriate as embodiments of the present invention are described. Of course, the content of the present invention should not be construed as being limited to the specific examples shown in these embodiments. Any configuration that exhibits the same operations and effects, even if it is different from the specific configuration shown in the embodiments, is of course included in the technical scope of the present invention.
[0017] (Embodiment 1) First, the composite video generation system according to Embodiment 1 will be described. As shown in FIG. 1, the composite video generation system according to Embodiment 1 of the present invention includes a video material input unit 1 for inputting one or more video materials used as materials for the composite video, an insertion target video generation unit 2 for generating an insertion target video that is the target for inserting an avatar based on the input video materials, an insertion area setting unit 3 for setting a plurality of insertion areas that are areas where the avatar can be inserted in the insertion target video, an insertion area information acquisition unit 4 for acquiring insertion area information that is information regarding the display content of the insertion target video in the set insertion areas, a person video input unit 5 for inputting one or more person videos that are the source of the avatar to be inserted into the insertion target video, an avatar generation unit 6 for generating one or more avatars from the one or more person videos, an avatar information acquisition unit 7 for acquiring avatar information that is information regarding the generated avatars, an insertion area selection unit 8 for selecting into which of the set insertion areas the generated one or more avatars are to be inserted, a composite video generation unit 9 for inserting the avatar corresponding to the area selected by the insertion area selection unit 8 to generate a plurality of composite videos, a composite video output unit 10 for outputting the generated plurality of composite videos, a browsing information acquisition unit 11 for acquiring browsing information that is information regarding the browsing mode for each of the output plurality of composite videos, and a video selection unit 12 for selecting, based on the acquired browsing information, a composite video with a high degree of viewer interest from among the generated plurality of composite videos.
[0018] The video material input unit 1 is for inputting video materials which are image data used for generating a composite video. The composite video is composed of an avatar inserted into all or part of the video to be inserted. The video materials input through the video material input unit 1 are directly used for generating the video to be inserted. The video to be inserted is composed of one or more video materials. As a simple configuration, a single video material can be directly used as the video to be inserted, or a single video material can be used as the video to be inserted after performing image processing such as enlargement, reduction, rotation, and partial modification. When the video to be inserted is composed of multiple video materials, image processing is performed on each video material as necessary, and then each video material is arranged at a predetermined position, or multiple video materials are superimposed, etc., to form the video to be inserted. The form of the video material can be either a 2D video or a 3D video, and can also be either a moving image or a still image. Also, the content of the video material can be realistic (for example, a video of a real scene taken using an imaging device such as a camera), or can be pictorial (for example, an illustration created by an illustrator).
[0019] The video to be inserted generation unit 2 is for generating a video to be inserted based on one or more video materials. The video to be inserted is the original video of the composite video. More specifically, it has the relationship that the composite video is the result of performing avatar insertion processing on the video to be inserted. The video to be inserted generation unit 2 performs necessary processing on one or more video materials, and then uses a single video material as the video to be inserted as it is, or performs processing such as arranging each of multiple video materials at a predetermined position to generate the video to be inserted. As a specific configuration of the video to be inserted generation unit 2, for example, it can have a configuration with an image processing function using the prior art, and the configuration of the generated video to be inserted can also be a moving image, a still image, a color video, a black and white video, a 2D video, a 3D video, etc., according to the form of the completed video.
[0020] The insertion area setting unit 3 is for setting a plurality of insertion areas, which are areas in the insertion target video generated by the insertion target video generation unit 2 where an avatar can be inserted. Specifically, the insertion area setting unit 3 has a function of extracting a human part in the insertion target video by image recognition processing and setting it as an insertion area. For example, it may be configured to pre-set areas that can become insertion areas when video materials are input by the video material input unit 1, or the user, operator, etc. may visually set the insertion area for either the video material or the insertion target video. Note that the number of insertion areas set by the insertion area setting unit 3 may be any number as long as it is plural, but it is preferably the same as or more than the number of avatars generated by the avatar generation unit 7. Also, the shape of the area set by the insertion area setting unit 3 is determined corresponding to the shape of the avatar inserted into the area. For example, when the avatar has a shape reflecting the full body of a person, the insertion area should also have a shape corresponding to the full body of the person, and when the avatar is composed only of the head of a person, the insertion area should also have a shape corresponding to the head of the person. However, the insertion area is not limited to the human video in the insertion target video. For example, when the avatar is composed only of the facial part of the head, it is also possible to configure the front of the locomotive in the insertion target video as the insertion area.
[0021] The insertion area information acquisition unit 4 is for acquiring insertion area information, which is information about display targets such as people in the insertion areas set by the insertion area setting unit 3. Specifically, the insertion area information acquisition unit 4 includes an attribute information acquisition unit 13 that acquires area attribute information, which is information about the attributes of the display target in the insertion area, and an appearance information acquisition unit 14 that acquires area appearance information, which is information about the appearance of the display target in the insertion area, and has a function of acquiring at least one of the area attribute information and the area appearance information as the insertion area information. The insertion area information acquisition unit 4 performs the acquisition process of the insertion area information for the plurality of insertion areas set by the insertion area setting unit 3.
[0022] The attribute information acquisition unit 13 is for acquiring area attribute information, which is information regarding the attributes of the display target in the insertion area set by the insertion area setting unit 3. Specifically, the attribute information acquisition unit 13 has a function of acquiring information regarding attributes such as the gender, nationality, race, age, educational background, work experience, etc. of the display target as the area attribute information. As a mode of acquiring the area attribute information, it may be possible to input information regarding the display target when inputting via the video material input unit 1 and extract the area attribute information from among the said information, or to acquire the area attribute information based on the attributes inferred from the display content (appearance, etc.) in the insertion area. Furthermore, for example, when the display target is a specific person (not limited to an actual person, but may be a character appearing in a creative work, etc.), it may be possible to extract the attribute information of the said specific person from information publicly available on the Internet or the like. Note that, depending on the items of the area attribute information, there may be cases where it also corresponds to the area external appearance information described later (for example, information regarding the race can be used to infer information regarding the appearance such as the eyes and skin color), but it may be possible to exclude such items from the area attribute information (= all information related to the appearance is treated as appearance information), or it may be configured to include both in the area attribute information and the area external appearance information.
[0023] The appearance information acquisition unit 14 is for acquiring out-of-region appearance information, which is information regarding the appearance of the display target in the insertion region set by the insertion region setting unit 3. Specifically, the appearance information acquisition unit 14 has a function of extracting shape features and color tone features on the surface of a display target such as a person, and acquiring the extracted information as shape information and color tone information respectively. More specifically, for example, as shape features, the appearance information acquisition unit 14 has a function of extracting feature points of the display target and acquiring information regarding the positional relationship between the feature points. The feature points refer to points predetermined corresponding to characteristic parts in the display target, such as each part of eyes, nose, mouth, ears, etc. on the face of a human body, and connection points of skeletons in the limbs. As a mechanism for extracting feature points, for example, image recognition technology realized by deep learning, machine learning, etc. may be used, or a combination of multiple image recognition technologies may be used. Also, the feature points in the first embodiment of the present invention indicate the positions of body surface or internal parts in the display target, but in the present invention, it may be arbitrarily determined which parts should be set as feature points. For example, it may be handled that feature points are set only for visually recognizable parts (eyes, nose, mouth, ears, eyebrows, etc.) on the face surface and points corresponding to the contour, or feature points may be determined corresponding to facial expression muscles inside the face and positions of joints in the whole body including the face. Also, for example, when setting feature points for eyes, only the center position of the eyes may be set as feature points, or feature points may be set in detail for each of the inner corner of the eye, outer corner of the eye, pupil, and white part of the eye. The appearance information acquisition unit 14 has a function of acquiring, as part of the out-of-region appearance information, feature point information including information regarding the positions of the extracted feature points (which may be the position coordinates themselves in the insertion region, but preferably the relative positional relationship in a normalized coordinate system) and the significance of the feature points (a point corresponding to the outer corner of the eye, a point corresponding to the knee joint, etc.). However, it is also possible to adopt a simpler configuration as shape information. For example, the shape information may be constituted by character information describing shapes such as the contour of the face of a person as the display target (round face, long face, etc.) and the appearance features of eyes (narrow eyes, droopy eyes, etc.).
[0024] In addition, the appearance information acquisition unit 14 has a function of acquiring color tone information, which is information regarding color tone characteristics on the surface of the display target as out-of-region appearance information. Specifically, the appearance information acquisition unit 14 acquires, as color tone information regarding the color tone surface, the extraction results of color tone characteristics (such as fair skin, black hair, and brown eyes) of each region (skin, hair, eyes, etc.) in the display target by means of image recognition technology or the like. As an extraction mode of the color tone characteristics, among the color tones of the entire target region (for example, "skin"), the one that occupies the largest area (= the one that appears most frequently) may be extracted as the color tone characteristic, or the average of the color tones of the entire target region may be extracted as the color tone characteristic. The appearance information acquisition unit 14 acquires out-of-region appearance information, which is information including at least one of shape information, which is information regarding the shape surface of the display target, and color tone information, which is information regarding the color tone surface, and outputs it to the insertion region selection unit 8.
[0025] The person video input unit 5 is for inputting a person video, which is a material used for generating an avatar that constitutes a composite video. Specifically, the person video input unit 5 has a function of inputting a person video, which is a video regarding the whole or a part of the person who is the original of the avatar to be inserted into the insertion target video. As a specific configuration, it may be configured to input a video from the outside, or may be configured to include an imaging mechanism for acquiring a video. As a specific aspect of the person video, it may be a full-body video of the target person or a partial video such as a face image. Also, it may be either a still image or a moving image, and may be either a 2D video or a 3D video. In the first embodiment, an explanation is given using a 2D and still image video regarding the face of the target person as the person video, but it goes without saying that person videos other than such a configuration can also be used in the composite video generation system according to the first embodiment.
[0026] The avatar generation unit 6 is for generating an avatar to be inserted into and replace all or part of the video materials constituting the video to be inserted, based on the person video input through the person video input unit 5. Specifically, as the configuration of the avatar, the avatar generation unit 6 may be configured from skeleton information (bones), surface information (skin), and weight information that defines the relationship between the two. However, it is also preferable to adopt a configuration consisting only of surface information including two-dimensional or three-dimensional shape and information on color tone on the surface. As an even simpler configuration, it may be configured by the image data itself. Also, as the avatar generated by the avatar generation unit 6, it is possible to generate an avatar corresponding to a full body image of a person, but it is also possible to generate an avatar consisting of only a part of a person, for example, only a part of the face. In the first embodiment, an avatar consisting only of the head above the neck is generated. The avatar generation unit 6 extracts feature points set according to the features (eyes, eyebrows, nose, mouth, ears, hairstyle, etc.) on the surface and internal features (joints, etc.) such as the position and shape of the person video input through the person video input unit 5, and reflects the positional relationship between the extracted feature points in the avatar, thereby generating a realistic avatar that reflects the physical characteristics of the model. However, when adopting a simpler configuration, for example, it is also possible to generate as an avatar a configuration that uses the person video as it is or that only makes the minimum necessary corrections such as adjusting the size and orientation of the person video.
[0027] The avatar information acquisition unit 7 is for acquiring avatar information which is information about the avatar generated by the avatar generation unit 6. Specifically, the avatar information acquisition unit 7 includes an attribute information acquisition unit 15 that acquires avatar attribute information which is information about the attributes of the representation target (such as a person) represented by the avatar, and an appearance information acquisition unit 16 that acquires avatar appearance information which is information about the appearance of the avatar, and has a function of acquiring at least one of the avatar attribute information and the avatar appearance information as the avatar information. When a plurality of avatars are generated, the avatar information acquisition unit 7 performs acquisition processing of the avatar attribute information and the avatar appearance information for each generated avatar.
[0028] The attribute information acquisition unit 15 has a function of acquiring avatar attribute information, which is information regarding the attributes of the represented object (such as a person) represented by the avatar, that is, information regarding attributes such as the gender, nationality, race, age, educational background, work history, etc. of the person or the like. Here, the "object represented by the avatar" is in principle the person in the person video from which the avatar is generated. However, when a new character is given to the generated avatar (for example, the person in the person video is female, but the character of the generated avatar is male, etc.), the new character corresponds to the "object represented by the avatar". As a mode of acquiring the avatar attribute information, when inputting a person video via the person video input unit 5, it may be configured to input the attribute information of the person who is the object of the person video at the same time, or it may be configured to input the setting information regarding the attributes of the "new character" to be given to the generated avatar. Further, it may be configured to estimate information regarding attributes based on the appearance of the person video or the avatar, or it may be configured to identify the target person from the appearance of the person video and acquire the avatar attribute information by collecting and extracting public information regarding the person. Note that the items, formats, etc. of the avatar attribute information acquired by the attribute information acquisition unit 15 may be arbitrary. However, from the viewpoint of convenience in comparison processing and the like in the insertion area selection unit 8, it is preferable that the items, formats, etc. be the same as those of the area attribute information regarding the display target in the insertion area acquired by the attribute information acquisition unit 13.
[0029] The appearance information acquisition unit 16 is for acquiring avatar appearance information, which is information regarding the appearance of the avatar generated by the avatar generation unit 6. Specifically, the appearance information acquisition unit 16 extracts at least one of the shape features and color tone features of the avatar, and has a function of acquiring avatar appearance information including at least one of shape information and color tone information (regarding the definition of these information, it shall be the same as the shape information and color tone information in the out-of-domain appearance information). The appearance information acquisition unit 16 has a function of, for example, extracting feature points of the avatar as shape information, and acquiring feature point information, which is information regarding the positional relationship of each extracted feature point and the significance of each feature point (a feature point corresponding to the outer corner of the eye, a feature point corresponding to the knee joint, etc.), and acquiring shape information based on this. In addition, when a configuration is adopted in which feature points are set when the avatar generation unit 6 generates an avatar, it may be an aspect of directly acquiring the information regarding the feature points set at the time of avatar generation as the feature point information. Further, as a simpler configuration, the appearance information acquisition unit 16 may acquire shape information consisting of character information (the contour of the face is a round face, etc.) depicting the shape of each part of the avatar.
[0030] In addition, the appearance information acquisition unit 16 has a function of acquiring color tone information, which is information regarding color tone characteristics on the surface of the avatar as avatar appearance information. Specifically, the appearance information acquisition unit 16 acquires, as color tone information regarding the color tone surface, the extraction results of color tone characteristics (such as fair skin color, black hair, and brown eyes) of each region (skin, hair, eyes, etc.) on the avatar by means of image recognition technology or the like. When such information is set during avatar generation, it may be used as the set information. As an extraction mode of color tone characteristics, among the color tones of the entire target region (for example, "skin"), the one that occupies the largest area (= the one that appears most frequently) may be extracted as the color tone characteristic, or the average of the color tones of the entire target region may be extracted as the color tone characteristic. The items, formats, etc. of the avatar appearance information acquired by the appearance information acquisition unit 16 may be arbitrary, but from the viewpoint of convenience in the comparison process and the like in the insertion region selection unit 8, it is preferable that they be the same as the items, formats, etc. of the region appearance information regarding the display target in the insertion region acquired by the appearance information acquisition unit 14.
[0031] The insertion region selection unit 8 is for selecting a region in which to insert the avatar generated by the avatar generation unit 6 with respect to the video to be inserted. Specifically, the insertion region selection unit 8 includes an attribute information comparison unit 17 that compares the avatar attribute information regarding the avatar representation target with the region attribute information regarding the display target in the insertion region to determine the similarity, an appearance information comparison unit 18 that compares the avatar appearance information regarding the avatar with the region appearance information of the display target in the insertion region to determine the similarity, and an insertion region determination unit 19 that determines the insertion region for inserting the avatar based on the comparison determination results of the attribute information comparison unit 17 and the appearance information comparison unit 18. Note that the insertion region selection unit 8 may select a single insertion region for each individual avatar, but in the context of generating a plurality of composite videos in the first embodiment, it is assumed that a plurality of insertion regions corresponding to the number of composite videos are selected for each individual avatar.
[0032] The attribute information comparison unit 17 compares the avatar attribute information regarding the object represented by the avatar with the region attribute information regarding the display object in each of the plurality of insertion regions in the insertion target video, and is for determining the similarity between the two. Specifically, the attribute information comparison unit 17 has a function of comparing the information of each item (gender, nationality, race, age, etc.) included in the avatar attribute information and the region attribute information, and determining the mutual similarity. Here, in the comparison process of the attribute information comparison unit 17, it may be possible to equally determine the similarity for all items of the avatar attribute information and the region attribute information, or it may be possible to compare only some items, or it may be possible to determine the similarity after attaching different weightings to each item. Regarding the determination of similarity, it may be possible to determine only whether they match or not. However, for example, regarding the nationality which is one item of the attribute information, even if they are different countries, if they belong to the same group (Europe, East Asia, etc.) that is geographically close (not the same), it is preferable to set a determination criterion regarding similarity in advance and then perform the determination of similarity. Also, when deriving a determination result other than whether they match or not as the similarity (fairly similar, partially similar, etc.), it is desirable to set a criterion in advance for quantifying the similarity according to the degree of similarity, and configure the attribute information comparison unit 17 to output the similarity quantified according to the criterion as the determination result. When a plurality of avatars are generated, the attribute information comparison unit 17 performs a comparison process on the avatar attribute information regarding each avatar with the region attribute information regarding the plurality of insertion regions set by the insertion region setting unit 3.
[0033] The appearance information comparison unit 18 compares the avatar appearance information regarding the appearance of the avatar with the outer area appearance information regarding the display targets of a plurality of insertion areas in the video to be inserted, and is for determining the similarity between the two. Specifically, the appearance information comparison unit 18 has a function of comparing the information of each item included in the shape information and color tone information that constitute the avatar appearance information and the outer area appearance information, for the avatar and the insertion area respectively, and determining the mutual similarity. As the content of the comparison process of the appearance information comparison unit 18, it may be to equally determine the similarity for all items of the shape information and color tone information that constitute the avatar appearance information and the outer area appearance information, but it may also be to compare only some items, or to perform different weightings for each item (for example, increasing the evaluation value of the information regarding the shape of the eyes among the shape information, etc.) and then determine the similarity. Regarding the determination of the similarity, for example, it may be to determine only whether the positional relationships of the feature points all match, but for example, regarding the positional relationships of the feature points, it is preferable to set a determination criterion regarding the similarity in advance, such that the smaller the difference value of the distance between specific feature points, the higher the similarity, and the larger the difference value, the lower the similarity, and then perform the determination of the similarity. Also, when deriving a determination result other than whether they match as the similarity, it is preferable to set a criterion in advance for quantifying the similarity according to the degree of similarity, and configure the appearance information comparison unit 18 to output the similarity quantified according to the criterion as the determination result. When a plurality of avatars are generated, the appearance information comparison unit 18 performs a comparison process on the avatar appearance information regarding each avatar with the outer area appearance information regarding the plurality of insertion areas set by the insertion area setting unit 3.
[0034] The insertion area determination unit 19 is for determining in which insertion area of the target video for insertion a specific avatar should be inserted based on the results of the comparison processes in the attribute information comparison unit 17 and the appearance information comparison unit 18. Specifically, the insertion area determination unit 19 has a function of determining in which insertion area the avatar should be inserted based on at least one of the similarity of the attribute information and the similarity of the appearance information between the avatar and the insertion area. Note that the insertion area determined by the insertion area determination unit 19 may be a single one for each avatar, but in the first embodiment, a configuration is adopted in which the composite video generation unit 9 generates a plurality of patterns of composite videos. Correspondingly, the insertion area determination unit 19 in the first embodiment has a function of determining a plurality of patterns of insertion areas for a single avatar according to the similarity of each piece of information.
[0035] As the simplest configuration of the determination algorithm by the insertion area determination unit 19, in terms of the relationship with the avatar, an insertion area with a high similarity in each item of attribute information and appearance information (shape information and color tone information) (for example, the average value of the similarities calculated for each item, or the average value after multiplying the weighting coefficient for each item (for example, a large coefficient for important items and a small coefficient for items with low importance), or even the average value of the similarities calculated only for some items is high) is determined as the insertion area of the avatar. On the other hand, as the determination algorithm of the insertion area by the insertion area determination unit 19, instead of the insertion area with the highest similarity in each item (or a part of each item), it is also possible to select the insertion area with the lowest similarity in each item (or a part of each item), or to select the insertion area with the second highest similarity. Also, when determining multiple patterns of insertion areas in relation to generating multiple composite images, the insertion area determination unit 19, for example, selects the insertion area with the highest similarity only for the attribute information as the first pattern, selects the insertion area with the lowest similarity only for some items of the shape information as the second pattern, and selects the insertion area with the second highest similarity only for the color tone information as the third pattern. For each insertion area determination process, it is also possible to vary the information and items to be compared. When multiple avatars are generated, the insertion area determination unit 19 determines the insertion area for the first avatar according to a predetermined order, and then, using the insertion areas excluding the determined insertion area as candidates, determines the insertion area for the second avatar. By repeating the same process thereafter, the insertion area determination unit 19 determines the insertion area for all avatars. The insertion area determination unit 19 outputs information regarding the combination of the avatar and the insertion area determined by itself to the composite image generation unit 9, and the composite image generation unit 9 generates a composite image based on the information output from the insertion area determination unit 19.
[0036] The composite video generation unit 9 is for generating a composite video with an avatar inserted into the insertion area selected by the insertion area selection unit 8. Specifically, the composite video generation unit 9 has a function of inserting an avatar into the video to be inserted by replacing the avatar generated by the avatar generation unit 6 with the insertion area selected by the insertion area selection unit 8 among a plurality of insertion areas in the video to be inserted. In the first embodiment, a plurality of patterns of insertion areas are selected by the insertion area selection unit 8, and accordingly, the composite video generation unit 9 has a function of generating a plurality of composite videos with an avatar inserted into each of the plurality of insertion areas selected by the insertion area selection unit 8. When the video to be inserted is a still image, the composite video generation unit 9 acquires information on the size and direction of the object to be replaced in the video to be inserted, appropriately changes the size and direction of the avatar based on the information, and then inserts the avatar into the video to be inserted in a manner of replacing the object to be replaced. When the video to be inserted is a moving image, the composite video generation unit 9 acquires information on the motion content of the object to be replaced in the video to be inserted by detecting, for example, the positional variation of feature points, appropriately changes the size, direction, and motion of the avatar based on the information, and then inserts the avatar into the video to be inserted in a manner of replacing the object to be replaced. As the object to be replaced by the avatar, it is usually desirable to use the video related to the person in the video to be inserted, but it is not necessary to be limited thereto. For example, when the content video of a locomotive is arranged, an avatar consisting of a part of a face image may be inserted in front of the locomotive.
[0037] The composite video output unit 10 is for outputting externally the composite video generated by the composite video generation unit 9. The composite video output by the composite video output unit 10 is provided for viewing by a predetermined viewer, and viewing information, which is information regarding the viewing mode of the viewer, is acquired by the viewing information acquisition unit 11, and a process is performed to determine a single composite video from among a plurality of composite videos based on the acquired viewing information. The video output mode by the composite video output unit 10 may be a general one using existing technologies, but from the perspective of avoiding the output mode from affecting the viewing mode of the viewer, when sequentially outputting a plurality of composite videos, it is preferable to randomly set the output order, and when simultaneously outputting to a single web page or the like, it is preferable to randomly set the arrangement order, etc. Also, although the output destination of the composite video by the composite video output unit 10 can be arbitrarily set, in the first embodiment, from the perspective of viewing information acquisition by the viewing information acquisition unit 11, it is assumed that the output is directed to a video display terminal equipped with a mechanism for calculating the number of times the composite video is displayed and a mechanism for detecting the reaction of the viewer at the time of displaying the composite video (such as a camera, an acceleration sensor, etc.).
[0038] The viewing information acquisition unit 11 is for acquiring, for each of the plurality of composite videos output by the composite video output unit 10, viewing information, which is information regarding the viewing mode of a predetermined viewer. Specifically, the viewing information acquisition unit 11 includes a viewing count information acquisition unit 20 that acquires viewing count information, which is information regarding the number of times the composite video is viewed, and a viewing reaction information acquisition unit 21 that acquires viewing reaction information, which is information regarding the reaction mode of the viewer at the time of viewing the composite video.
[0039] The view count information acquisition unit 20 is for acquiring view count information, which is information regarding the view counts of each of the plurality of synthesized videos output by the synthesized video output unit 10. Specifically, the view count information acquisition unit 20 calculates the view count of each synthesized video based on the calculation result by the view count calculation mechanism provided in the display terminal of the output synthesized video. Note that as the criterion for determining whether or not it has been "viewed", it may be determined that it has been viewed only when all the contents of the synthesized video are played from the beginning to the end on the display terminal, and the view count is calculated for the target. However, when only a part of the synthesized video is played, the view count may be calculated according to the ratio of the time of partial playback to the total playback time (for example, when played for half of the total playback time, it is calculated as 0.5 times). Also, as the content of the view count information, it may be configured to include only information regarding the view count. However, in the first embodiment, the configuration includes information regarding the view count, information regarding the viewing date and time and the video display terminal used for each viewing that constitutes the view count, and information linking the two.
[0040] The browsing reaction information acquisition unit 21 is for acquiring browsing reaction information, which is information regarding the reaction mode of a viewer when browsing each of a plurality of synthesized videos output by the synthesized video output unit 10. Specifically, the browsing reaction information acquisition unit 21 generates browsing reaction information based on measurement values of the physical reactions of the viewer during browsing, such as the vibration of the viewer's body, the variation of the viewer's expression, and the sound emitted by the viewer's actions (the voice of the viewer and the sound generated along with the actions, etc.). For the vibration of the viewer's body, it is desirable to measure it with a 3D acceleration sensor or the like that detects the actions of the viewer. For the variation of the expression, it is preferable to perform expression analysis on the face image of the viewer captured by a camera and then derive the variation mode. Also, for the sound emitted by the actions, it is desirable to measure it with a voice input mechanism such as a microphone. It is also possible to calculate the intensity of the browsing reaction only based on the magnitude of these measurement values to generate browsing reaction information. However, more specifically, after determining the semantic content of the body vibration, the content of the expression change, and the semantic content of the voice, classify them into four modes of joy, anger, sorrow, and pleasure, calculate the intensity of the reaction for each mode, perform weighting according to the mode on the calculation result, and then calculate the total value of the reaction intensity, etc., to calculate the intensity of the browsing reaction. Also, as the content of the browsing reaction information, it may be configured to include only information regarding the intensity of the above physical reactions of the viewer during browsing. However, in the first embodiment, it is configured to include information regarding the intensity of the physical reaction, the browsing date and time of the synthesized video for which the physical reaction was made, information regarding the video display terminal used for browsing, and information linking the two.
[0041] The video selection unit 12 is for selecting a predetermined number of composite videos with high viewer interest from among the plurality of composite videos generated by the composite video generation unit 9 based on the viewing information acquired by the viewing information acquisition unit 11. Specifically, the video selection unit 12 acquires interest degree information, which is information regarding the viewer interest degree for each composite video, based on at least one of the viewing count information acquired by the viewing count information acquisition unit 20 as one of the viewing information and the viewing reaction information acquired by the viewing reaction information acquisition unit 21 also as one of the viewing information, and has a function of selecting a composite video with a high viewer interest degree based on the interest degree information. The content of the interest degree information may be either only one of the viewing count information and the viewing reaction information or information obtained by simply adding the viewing count information and the viewing reaction information. However, in the first embodiment, the value corresponding to the intensity of the viewing reaction is multiplied for each individual viewing constituting the viewing count information, and the content is such that the multiplied values are added. For example, for a certain composite video, if there are 4 viewings, namely viewing A, viewing B, viewing C, and viewing D, and the intensities of the viewing reactions in each are α, β, γ, and δ, respectively, the interest degree information when composed only of the viewing count information is 4, whereas when the interest degree information is such that the value corresponding to the intensity of the viewing reaction is multiplied and the multiplied values are added, the value of the interest degree information is α + β + γ + δ. Note that the number of composite videos selected by the video selection unit 12 may be arbitrary, but in the first embodiment, one composite video is selected.
[0042] Next, the advantages of the composite video generation system according to Embodiment 1 will be described. The composite video system according to Embodiment 1 compares region information regarding the appearance and attributes of a display target that was displayed in an insertion region in the video to be inserted, with avatar information regarding the appearance of the avatar and the attributes of the representation target of the avatar, and adopts a configuration for determining a combination of the avatar and the insertion region according to the degree of similarity. By adopting such a configuration, for example, when a combination with a high degree of similarity is selected, an avatar with high affinity to the display target in the insertion region will be inserted, and an advantage is produced in that a natural composite video without a sense of incongruity can be generated as compared with the original video material and the video to be inserted. Conversely, when a combination with a low degree of similarity is selected, an advantage is produced in that a composite video with an unexpectedness that gives an impression completely different from the original video material and the video to be inserted can be generated.
[0043] Also, the composite video generation system according to Embodiment 1 adopts a configuration in which the insertion region selection unit 8 performs multiple selections of insertion regions, generates a plurality of composite videos according to the number of selections, and then selects a predetermined number of composite videos from among the plurality based on the viewing information regarding the plurality of composite videos. By adopting such a configuration, not only is a combination of the avatar and the insertion region selected based on the degree of similarity, but also a composite video is selected based on viewing information according to the preferences of the viewer, thereby producing an advantage in that a more appropriate composite video can be generated according to the preferences of the viewer. In particular, in Embodiment 1, by using the number of viewing times information and the viewing reaction information as the viewing information, an advantage is produced in that appropriate composite video generation is possible based on information regarding the quantity and quality of viewing.
[0044] (Embodiment 2) Next, the composite video generation system according to Embodiment 2 will be described. In Embodiment 2, with respect to components having the same name and the same reference numerals as those in Embodiment 1, unless otherwise specified, they shall exhibit the same functions as the components in Embodiment 1.
[0045] As shown in FIG. 2, in addition to the configuration shown in the first embodiment, the composite video generation system according to the second embodiment further includes a fitness rank calculation unit 23 that calculates the fitness rank of a plurality of insertion regions for each avatar, and an insertion region determination unit 24 that determines a combination of an avatar and an insertion region in which the total value of the fitness ranks calculated by the fitness rank calculation unit 23 is the smallest.
[0046] The fitness rank calculation unit 23 is for assigning the fitness rank of the insertion region according to the magnitude of the fitness, which is numerical information calculated based on the similarity between the avatar information and the insertion region information for each of the plurality of avatars generated by the avatar generation unit 6, to a plurality of insertion regions for each avatar. Specifically, the fitness rank calculation unit 23 calculates the fitness of each insertion region for the avatar based on at least one of the similarity between the avatar attribute information and the region attribute information calculated by the attribute information comparison unit 17 and the similarity between the avatar appearance information and the region appearance information calculated by the appearance information comparison unit 18, and then has a function of assigning the fitness ranks in order from the insertion region with the highest fitness, such as the first rank, the second rank, the third rank, and so on. As a specific example, for instance, when avatars A, B, C, etc. are generated while insertion regions a, b, c, d, etc. are specified in the video to be inserted, the fitness rank calculation unit 23 calculates each fitness value and then assigns the first fitness rank to insertion region a, the second rank to insertion region c, the third rank to insertion region e, etc. for avatar A, assigns the first fitness rank to insertion region a, the second rank to insertion region b, the third rank to insertion region d, etc. for avatar B, and assigns the first fitness rank to insertion region c, the second rank to insertion region e, the third rank to insertion region f, etc. for avatar C, and so on for the fitness ranks.
[0047] Note that the "degree of fitness", which is an index used for calculating the fitness rank by the fitness rank calculation unit 23, is a value based on the similarity of the attribute information and the appearance information, similar to the determination algorithm in the insertion area determination unit 19. However, it does not necessarily have to be a value that coincides with the similarity. For example, when preferentially selecting combinations of avatars and insertion areas where the attribute information and the appearance information are similar to each other, it is possible to directly use the similarity as the degree of fitness. However, when preferentially selecting combinations with low similarity, it is preferable to use the reciprocal of the similarity or a numerical value obtained by multiplying the similarity by a negative coefficient as the degree of fitness. Also, in relation to generating a plurality of composite videos, when generating composite video A, the calculation process by the fitness rank calculation unit 23 may be performed with the degree of fitness = similarity, and when generating composite video B, the degree of fitness may be calculated as the reciprocal of the similarity. It is also possible to vary the definition of the degree of fitness, such as this.
[0048] The insertion area determination unit 24 is for determining a combination of a plurality of avatars and a plurality of insertion areas such that the total value of the fitness ranks of the insertion areas combined with each avatar in the avatar is minimized under the condition that the same insertion area is not selected for different avatars based on the fitness ranks calculated by the fitness rank calculation unit 23. Specifically, the insertion area determination unit 24 has a function of selecting, from among a large number of achievable combination candidates between a plurality of avatars and a plurality of insertion areas, the combination with the smallest total value of fitness ranks. For example, when the insertion areas ranked first in terms of fitness for each of the plurality of generated avatars A, B, C... do not overlap, the combination of each avatar and the insertion area ranked first in terms of fitness for each is achievable, and since the total value of the fitness ranks is the smallest, the combination of inserting each of avatars A, B, C... into the insertion area ranked first in terms of fitness is selected. On the other hand, as in the example shown in the
[0046] paragraph, when the insertion areas ranked first in terms of fitness are the same for a plurality of avatars, it is impossible to insert a plurality of avatars into the same insertion area, so while excluding the combination of inserting each avatar into the insertion area ranked first in terms of fitness, a combination that gives the smallest total value of fitness ranks is selected. In the above example, where the first-ranked fitness for both avatars A and B is insertion area a, the combination that gives the smallest total value of fitness ranks is the combination where avatar A is inserted into insertion area a, avatar B is inserted into insertion area b, and avatar C is inserted into insertion area c. Therefore, the insertion area determination unit 24 selects such a combination for the plurality of avatars and the plurality of insertion areas.
[0049] Next, the advantages of the composite video generation system according to the second embodiment will be described. When a plurality of avatars are inserted into the video to be inserted in the composite video generation system according to the second embodiment, the system avoids different avatars being inserted into the same insertion area, and determines the combination of avatars and insertion areas so that the total value of the ranking of the insertion areas for each avatar according to the similarity of each avatar information and insertion area information becomes small. By adopting such a configuration, the composite video generation system according to the second embodiment has the advantage that it can select a combination of an avatar and an insertion area with a high degree of fitness calculated based on similarity while avoiding a plurality of avatars being combined in the same insertion area.
Industrial Applicability
[0050] The present invention can be used as a technique for appropriately selecting an insertion area for an avatar when selecting one or more insertion areas for inserting an avatar from a plurality of insertion areas set in the video to be inserted in composite video generation.
Explanation of Signs
[0051] 1 Video material input unit 2 Insertion target video generation unit 3 Insertion area setting unit 4 Insertion area information acquisition unit 5 Character video input unit 6 Avatar generation unit 7 Avatar information acquisition unit 8, 22 Insertion area selection unit 9 Composite video generation unit 10 Composite video output unit 11 Browsing information acquisition unit 12 Video selection unit 13, 15 Attribute information acquisition unit 14, 16 Appearance information acquisition unit 17 Attribute information comparison unit 18 Appearance information comparison unit 19, 24 Insertion area determination unit 20 Browsing count information acquisition unit 21 Browsing reaction information acquisition unit 23 Fitness rank calculation unit
Claims
1. A synthetic video generation system for generating a synthetic video in which an avatar generated based on a predetermined person video is inserted into an insertion target image having a plurality of insertion areas, the synthetic video generation system comprising: an insertion target image generating means for generating an insertion target image based on one or more image materials; an insertion area setting means for setting a plurality of insertion areas in the insertion target image generated by the insertion target image generating means; an insertion area information acquiring means for acquiring information about a display object in the insertion area set by the insertion area setting means, the information including at least one of area attribute information, which is information about attributes of the display object, and area appearance information, which is information about an appearance of the display object; an avatar generating means for generating an avatar based on the person image; an avatar information acquiring means for acquiring avatar information, the avatar information including at least one of avatar attribute information, which is information on attributes of a representation target that is a target represented by the avatar, and avatar appearance information, which is information on an appearance of the avatar; an insertion area selection means for selecting an insertion area into which the avatar is to be inserted from among the plurality of insertion areas based on at least one of a similarity between the region attribute information and the avatar attribute information for the plurality of insertion areas and a similarity between the region appearance information and the avatar appearance information for the plurality of insertion areas; a composite image generating means for generating the composite image by inserting the avatar into the insertion area selected by the insertion area selecting means; Equipped with the insertion area selection means selects a plurality of the insertion areas into which the avatar is to be inserted; the composite video generation means generates a plurality of the composite videos inserted into the respective insertion areas selected by the insertion area selection means; A composite image generation system further comprising an image selection means for selecting a predetermined number of the composite images from among the plurality of composite images based on viewing information which is information relating to viewing patterns for the plurality of composite images.
2. a view count information acquiring means for acquiring view count information, which is information regarding the number of times the composite image has been viewed, as one of the view information; a viewing reaction information acquiring means for acquiring viewing reaction information, which is information regarding a reaction state of a viewer when viewing the composite image, as one of the viewing information; Further equipped with 2. The synthetic image generating system according to claim 1, wherein said image selecting means selects one synthetic image based on said number of views information and said viewing response information.
3. A synthetic video generation method for generating a synthetic video in which an avatar generated based on a predetermined person video is inserted into an insertion target image having a plurality of insertion areas, the method comprising: an insertion target video generating step of generating an insertion target video based on one or more video materials; an insertion area setting step of setting a plurality of insertion areas in the insertion target video generated in the insertion target video generating step; an insertion area information acquisition step of acquiring information about a display object in the insertion area set in the insertion area setting step, the information including at least one of area attribute information, which is information about an attribute of the display object, and area appearance information, which is information about an appearance of the display object; an avatar generating step of generating a plurality of avatars based on the person image; an avatar information acquisition step of acquiring avatar information, the avatar information including at least one of avatar attribute information, which is information about attributes of a representation object that is an object represented by the avatar, and avatar appearance information, which is information about an appearance of the avatar; an insertion area selection step of selecting an insertion area into which the avatar is to be inserted from among the plurality of insertion areas based on at least one of a similarity between the region attribute information and the avatar attribute information for the plurality of insertion areas and a similarity between the region appearance information and the avatar appearance information for the plurality of insertion areas; a composite image generating step of generating the composite image by inserting the avatar into the insertion area selected in the insertion area selecting step; Including, In the insertion region selection step, a matching rank calculation step of allocating a matching rank of the insertion area to each of the plurality of avatars according to a degree of matching, the degree being numerical information calculated based on a similarity between the avatar information and the insertion area information; and an insertion region determination step of determining a combination of a plurality of avatars and a plurality of the insertion regions such that the total value of the matching rank of the insertion regions combined with each of the avatars is minimized under a condition that the same insertion region is not selected for different avatars; The synthetic image generating method further comprises:
4. A composite image generating program that causes a computer to generate a composite image by inserting an avatar generated based on a predetermined person image into an insertion target image having a plurality of insertion areas, The computer is an insertion target image generating function that generates an insertion target image based on one or more image materials; an insertion area setting function for setting a plurality of insertion areas in the insertion target image generated by the insertion target image generation function; an insertion area information acquisition function that acquires information about a display object in the insertion area set by the insertion area setting function, the information including at least one of area attribute information, which is information about attributes of the display object, and area appearance information, which is information about an appearance of the display object; an avatar generation function for generating a plurality of avatars based on the person image; an avatar information acquisition function for acquiring avatar information, the avatar information including at least one of avatar attribute information, which is information about attributes of a representation target that is a target represented by the avatar, and avatar appearance information, which is information about an appearance of the avatar; an insertion area selection function that selects an insertion area into which the avatar is to be inserted from among the plurality of insertion areas based on at least one of a similarity between the region attribute information and the avatar attribute information for the plurality of insertion areas and a similarity between the region appearance information and the avatar appearance information for the plurality of insertion areas; a composite image generating function for generating the composite image by inserting the avatar into the insertion area selected by the insertion area selecting function; In the insertion area selection function, a matching rank calculation function for allocating a matching rank of the insertion area to each of the plurality of avatars according to a degree of matching, the degree being numerical information calculated based on a similarity between the avatar information and the insertion area information for each of the plurality of avatars; An insertion area determination function that determines a combination of a plurality of avatars and a plurality of the insertion areas so that the total value of the matching rank of the insertion areas combined with each of the avatars is minimized under the condition that the same insertion area is not selected for different avatars; and A synthetic image generating program for causing a user to execute the above steps.
Citation Information
Patent Citations
Image processing device
JP2019009752A
Program for providing virtual space with head-mounted display, method, and information processing apparatus for executing program
JP2019012509A
Information processing apparatus, information processing method, and computer program
JP2019139673A
Method and system for providing avatar service
JP2021157800A
Monitoring system
JP2023093912A