Computer program, system, server device and control method of server device
The system generates and manages derivative images of characters at different angles and poses, addressing limitations in existing technologies by enabling more varied character movements and expressions, enhancing the realism and engagement of distributed content.
Patent Information
- Application Number
- JP2024020547
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-02-14
AI Technical Summary
Existing technologies for generating character images are limited in their ability to create avatars that perform actions beyond facial expressions and require pre-drawn body parts in a database, restricting the variety of character movements.
A system that generates and registers derivative images of characters at different angles or poses using an image generation model, allowing for more varied character movements by extracting and managing facial and body features, and associating them with emotion tags and artistic styles.
Enables characters to move and express a wider variety of facial expressions and gestures, enhancing the realism and engagement of distributed content.
Smart Images

Figure 2025124468000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for generating character images. [Background technology]
[0002] Patent Document 1 discloses that derived images of an input character image are generated with the eyes and mouth open and closed, and that these derived images are used to generate moving images with changing facial expressions.
[0003] Patent Document 2 discloses that the skeletal features of a character in an input image are extracted, similar images having features similar to the extracted features are extracted from a database, and an image is generated when the skeleton is applied to the character in the input image based on values related to the angle of the character's skeleton drawn in the similar image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-111102 [Patent Document 2] Patent Publication No. 2021-071843 Summary of the Invention [Problem to be solved by the invention]
[0005] The technology disclosed in Patent Document 1 is applicable only to faces, and is therefore of limited applicability to avatars that perform actions such as turning around, for example.
[0006] The technology disclosed in Patent Document 2 requires that the input image character has an image in which the body parts corresponding to the desired posture are drawn, and that such images be stored in a database.
[0007] In view of the above-described conventional techniques, the present invention provides a technique for generating images for distribution that allow characters to move in a more varied manner. [Means for solving the problem]
[0008] One aspect of the present invention is characterized in that a computer is made to execute a generation step of generating, based on an image including an input character, a plurality of derivative images in which the character is drawn at different angles or poses, and a registration step of registering the derivative images as images to be distributed for moving the character. [Effects of the Invention]
[0009] According to the configuration of the present invention, it is possible to provide a technique for generating images for distribution that allow characters to move in a more varied manner. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram showing an example of a system configuration. [Figure 2] 10 is a flowchart of a process performed by the system for allowing viewers to view content distributed by a distributor. [Figure 3] 10 is a flowchart of a process performed by a system to generate avatar data corresponding to a character from an image including the character (character image). [Figure 4] 10A to 10C are diagrams showing examples of derived images of each character's facial expression. [Figure 5] FIG. 10 is a diagram showing an example of a character's facial features. [Figure 6] FIG. 2 is a diagram showing an example of a character image (input image) of a character. [Figure 7] 10A and 10B are diagrams showing examples of a plurality of derived images in which characters are drawn at different angles or poses. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant explanations will be omitted.
[0012] First, an example of the configuration of a system according to one embodiment of the present invention will be described using the block diagram of Fig. 1. As shown in Fig. 1, the system includes a distributor terminal 100, a server device 120, and a viewer terminal 140, each of which is connected to a network such as the Internet. For simplicity of explanation, Fig. 1 shows only one distributor terminal 100, one server device 120, and one viewer terminal 140, but the number of each may be two or more.
[0013] First, we will explain the broadcaster terminal 100. The broadcaster terminal 100 is a terminal device operated by a broadcaster such as a Vtuber, and is a computer device such as a PC, a tablet terminal device, or a smartphone.
[0014] The CPU 101 executes various processes using computer programs and data stored in the RAM 102. As a result, the CPU 101 controls the overall operation of the distributor terminal 100 and executes various processes that will be described as processes performed by the distributor terminal 100.
[0015] RAM 102 provides various areas as appropriate, such as an area for storing computer programs and data loaded from ROM 103 or storage device 107, an area for storing computer programs and data received from the outside via I / F 108, and a work area used by CPU 101 when executing various processes.
[0016] The ROM 103 stores setting data for the distributor terminal 100, computer programs and data related to the startup of the distributor terminal 100, computer programs and data related to the basic operation of the distributor terminal 100, and the like.
[0017] The operation unit 104 is a user interface such as a keyboard, a mouse, a touch panel screen, etc., and can be operated by the distributor to input various instructions and information to the distributor terminal 100.
[0018] The imaging unit 105 captures a moving image and outputs an image of each frame of the moving image. Note that, although the imaging unit 105 is built into the distributor terminal 100 in FIG. 1, it may be external to the distributor terminal 100.
[0019] The display unit 106 has a liquid crystal screen or a touch panel screen, and displays the results of processing by the CPU 101 as images, text, and the like.
[0020] The storage device 107 is a large-capacity information storage device such as a hard disk drive, etc. The storage device 107 stores an OS, computer programs and data for causing the CPU 101 to execute various processes described as processes performed by the distributor terminal 100.
[0021] The I / F 108 is a communication interface for performing data communication with an external device via a network such as the Internet.
[0022] The CPU 101 , RAM 102 , ROM 103 , operation unit 104 , imaging unit 105 , display unit 106 , storage device 107 , and I / F 108 are all connected to a system bus 109 .
[0023] Next, a description will be given of the server device 120. The server device 120 is a computer device such as a PC, a tablet terminal device, or a smartphone.
[0024] The CPU 121 executes various processes using computer programs and data stored in the RAM 122. As a result, the CPU 121 controls the overall operation of the server device 120 and executes various processes that will be described as processes performed by the server device 120.
[0025] RAM 122 provides various areas as appropriate, such as an area for storing computer programs and data loaded from ROM 123 or storage device 125, an area for storing computer programs and data received from the outside via I / F 126, and a work area used by CPU 121 when executing various processes.
[0026] The ROM 123 stores setting data for the server device 120, computer programs and data related to the startup of the server device 120, computer programs and data related to the basic operation of the server device 120, and the like.
[0027] The operation unit 124 is a user interface such as a keyboard, a mouse, a touch panel screen, etc., and can be operated by the user of the server device 120 to input various instructions and information to the server device 120.
[0028] The storage device 125 is a large-capacity information storage device such as a hard disk drive, etc. The storage device 125 stores an OS, computer programs and data for causing the CPU 121 to execute various processes described as processes performed by the server device 120.
[0029] The I / F 126 is a communication interface for performing data communication with an external device via a network such as the Internet.
[0030] The CPU 121 , RAM 122 , ROM 123 , operation unit 124 , storage device 125 , and I / F 126 are all connected to a system bus 127 .
[0031] Next, we will explain the viewer terminal 140. The viewer terminal 140 is a terminal device operated by a viewer who views the content distributed by a distributor, and is a computer device such as a PC, a tablet terminal device, or a smartphone.
[0032] The CPU 141 executes various processes using computer programs and data stored in the RAM 142. As a result, the CPU 141 controls the overall operation of the viewer terminal 140 and executes various processes that will be described as processes performed by the viewer terminal 140.
[0033] RAM 142 provides various areas as appropriate, such as an area for storing computer programs and data loaded from ROM 143 or storage device 146, an area for storing computer programs and data received from the outside via I / F 147, and a work area used by CPU 141 when executing various processes.
[0034] The ROM 143 stores setting data for the viewer terminal 140, computer programs and data related to the startup of the viewer terminal 140, computer programs and data related to the basic operation of the viewer terminal 140, and the like.
[0035] The operation unit 144 is a user interface such as a keyboard, a mouse, or a touch panel screen, and the viewer can operate it to input various instructions and information to the viewer terminal 140.
[0036] Display unit 145 has a liquid crystal screen or a touch panel screen, and displays the results of processing by CPU 141 using images, text, and the like.
[0037] The storage device 146 is a large-capacity information storage device such as a hard disk drive, etc. The storage device 146 stores an OS, computer programs and data for causing the CPU 141 to execute various processes described as processes performed by the viewer terminal 140.
[0038] The I / F 147 is a communication interface for performing data communication with an external device via a network such as the Internet.
[0039] The CPU 141 , RAM 142 , ROM 143 , operation unit 144 , display unit 145 , storage device 146 , and I / F 147 are all connected to a system bus 148 .
[0040] The hardware configuration of each device shown in FIG. 1 is an example, and the configuration is not limited to that shown in FIG.
[0041] Next, the process performed by the system to enable viewers to view content distributed by a distributor will be described with reference to the flowchart of FIG.
[0042] In step S201, the broadcaster terminal 100 detects the broadcaster's movements and state. An example of the processing in step S201 will be described below. The imaging unit 105 captures a video of the broadcaster, and the images of each frame of the video are stored in RAM 102. The CPU 101 detects the broadcaster's movements and state based on the images of each frame of the broadcaster stored in RAM 102. A technique for detecting the broadcaster's movements from the broadcaster's image may be, for example, a well-known image recognition technique that uses a machine learning model trained to output the position, posture, open / closed state, and shape of a person's face, body parts, eyes, mouth, and other parts in the input image. The detected movements and states may be, for example, skeletal features of the broadcaster in the image and time-series displacement information of the features, or the position, posture, open / closed state, etc. of each part of the broadcaster (head, arms, legs, upper body, lower body, etc.) and each part of the face (eyes, nose, mouth, ears, outline, hair, beard, eyelashes, eyebrows, accessories, etc.), and time-series displacement information thereof. Such detection is performed using conventional face recognition techniques.
[0043] In step S202, the distributor terminal 100 transmits operation information indicating the operation or state detected in step S201 to the server device 120 via the I / F .
[0044] In step S221, the server device 120 receives the operation information transmitted from the distributor terminal 100 via the I / F 126.
[0045] In step S222, the server device 120 transmits (distributes) the movement information received in step S221 and avatar data generated by the process described below to the viewer terminal 140 via the I / F 126. Specifically, the movement information is numerical data indicating the positions, movements (amounts of change), and orientations of each feature point and bone of the broadcaster's limbs, torso, head, and even the eyes and mouth that make up the face, and is movement information that can identify the open / closed state of the eyes, mouth, and other components of the face. The avatar data is a set of multiple images used to render a two-dimensional character (avatar) on the viewer terminal 140 that displays and moves in accordance with the broadcaster's movements and state. The avatar data is managed so that multiple images are linked to one character to create a single character. When sending the avatar data to the viewer terminal 140, all of the avatar data may be transmitted to the viewer terminal 140 around the time the viewer terminal 140 starts viewing the broadcast provided by the broadcaster, or the avatar data may be transmitted each time the broadcaster transmits the movement information.
[0046] In step S241, the viewer terminal 140 receives the action information and avatar data transmitted from the server device 120 via the I / F 147.
[0047] In step S242, based on the motion information and avatar data received in step S241, the viewer terminal 140 generates, as a display image, an image including a character with motions and states similar to those of the streamer represented by the motion information. Specifically, based on the streamer's motion information, the viewer terminal 140 determines the orientation and angle of each body part, such as the face, as well as the state of each body part (e.g., open / closed state), and selects an appropriate derivative image included in the avatar data based on the determination results to select the display image. For example, if it is determined based on the streamer's motion information that the face is tilted 15 degrees to the left, the eyes are open, and the mouth is closed, a derivative image in which the face is tilted 15 degrees to the left, the eyes are open, and the mouth is closed is selected, and a display image is generated based on the derivative image. Similarly, when body parts or parts other than the face are included, the state of each body part or part is determined from the streamer's motion information, and a derivative image corresponding to the determined state of each body part or part is extracted, thereby generating a display image of a character appropriate for the streamer's motion.
[0048] Furthermore, a specific emotion may be estimated based on the degree of eye opening, the degree of mouth opening, changes in facial contours, etc., and a derived image may be acquired to which a corresponding emotion tag is assigned based on the estimated emotion. For example, this emotion estimation may use a trained model that is prepared in advance in the server device 120 and that is trained to output the type of emotion of a person using as input human facial expressions or changes in human facial expressions, such as the degree of eye opening, the degree of mouth opening, and changes in facial contours. If each feature of the broadcaster's face meets predetermined conditions, the server device 120 may estimate that the broadcaster is showing a specific emotion or an expression intended to express a specific emotion, and may acquire multiple derived part images to which tags for the estimated emotions are assigned.
[0049] In step S243, the viewer terminal 140 causes the display unit 145 to display the display image generated in step S242.
[0050] Next, the process performed by the system to generate avatar data corresponding to a character from an image including the character (character image) will be described with reference to the flowchart in Figure 3. Note that the avatar data is generated in advance in response to operations by the distributor before the distributor starts distribution, and the generated avatar data is stored on server device 120 or the distributor terminal, thereby realizing the distribution process described above.
[0051] In step S301, the distributor terminal 100 transmits a character image to the server device 120 via the I / F .
[0052] The character image may be an image including the character from the top of the head to the toes, an image including part or all of the character's upper body, or an image including the character's upper body but not the lower body. The character image may be a character image selected by the distributor operating the operation unit 104 of the distributor terminal 100, a character image generated by the distributor terminal 100, or a preset character image.
[0053] Furthermore, the broadcaster terminal 100 may transmit parameters that are conditions for generating a derivative image to the server device 120. The parameters, for example, specify conditions regarding avatar data to be generated based on a character image input by the broadcaster operating the operation unit 104 of the broadcaster terminal 100.
[0054] In step S321, the server device 120 receives the character image and parameters transmitted from the distributor terminal 100 via the I / F 126. Note that the character image may be designated on the server device 120 side.
[0055] In step S322, server device 120 generates, based on the parameters received in step S321, a plurality of derived images in which the character is depicted at different angles or poses from the character image received in step S321.
[0056] In this embodiment, the server device 120 inputs the character image received in step S321 as an input image into an image generation model (for example, a model such as the well-known Stable Diffusion) that has been trained to output at least a derived image, and the image generation model processes the input image to generate a plurality of derived images in which the character is drawn at different angles or poses. Specifically, the server device 120 generates a prompt that instructs the image generation model to generate a derived image based on the image received in step S321 and the parameters received from the distributor or parameters related to derived image generation that are set in advance in the server, and inputs the generated prompt into the image generation model to generate the derived image.
[0057] In step S323, server device 120 extracts character features from each of the derived images generated in step S322. For example, server device 120 uses well-known techniques such as Anime Face Detector or Segment Anything to recognize and extract the facial features of the character from the derived image. Note that server device 120 may also delete unintended pixels that occur during segmentation (the process of separating features from the derived image).
[0058] The processes in steps S322 and S323 will be described in more detail using a specific example.
[0059] In step S322, if multiple types of facial expressions are set in the parameters, such as "fawning, sadness, anger, disgust, condescending / taunting, looking up, and a stunned face," the server device 120 generates derived images of each of the character's facial expressions from the character image (input image), as shown in Fig. 4. Note that the types of facial expressions are not limited to those listed here.
[0060] Furthermore, for example, the parameters may include instructions regarding the angle at which the character is drawn in the derived image to be generated. As one example, the parameters may specify the tilt up / down or left / right with respect to the front view of the character's face (the direction in which the face is facing straight ahead) as the reference. Specifically, a parameter specifying an angle in 10-degree increments within an angle range from when the character's face is facing 45 degrees left to when it is facing 45 degrees right with respect to the front view of the face is set, and the server device 120 generates a derived image of the face from the character image (input image) at an angle changed in 10-degree increments within that angle range based on the parameter, for example, with the vertical direction of the character's head as the central axis of rotation (an axis passing through the head from the apex of the head downward). Note that the numerical values given here are merely examples and are not limited to these numerical values.
[0061] Similarly, if the parameter is set to "5 degree increments within an angle range from when the character's face is facing up 20 degrees to when it is facing down 20 degrees," server device 120 will use the character image (input image) as the center of rotation, for example, at the base of the character's neck or the base of the head, and generate derived images of the character's face as the angle of the character's face changes in 5 degree increments in the forward and backward directions around the center of rotation within that angle range. Note that the numerical values given here are merely examples and are not limited to these numerical values.
[0062] In addition, the parameters may specify not only up, down, left, and right, but also diagonal angles such that the character's face faces diagonally upward and diagonally to the left, such as 15 degrees up, 20 degrees left, or 15 degrees down, 20 degrees right.
[0063] In addition, the server device 120 may generate derivative images when the character's head moves forward, backward, left, right, or tilted (derived images including the face in positions such as looking left, right, up, down, bowing, fanning, or tilting the head left and right).
[0064] In this way, the server device 120 generates derivative images of parts (in the case of the face, the face facing left, right, up, down, bowing, tilting the head left and right, etc.) when the head moves in the θ (rotation) direction, forward, backward, left and right tilting directions, with the central axis being the vertical axis of the character's head.
[0065] For example, when the server device 120 receives a character image (input image) of the character shown in Fig. 6 and a parameter specifying an angle range of 45 degrees to the left and 45 degrees to the right and a range of 5 degrees upward and 20 degrees downward, based on the state in which the character's face is facing forward, as shown in Fig. 7(a) to (d), the server device 120 can generate multiple derived images in which the character is depicted at different angles or poses depending on the parameter specification, such as an image in which the character is (a) facing left, (b) facing diagonally downward to the left, (c) facing diagonally downward to the right, or (d) facing right. The image depicted in Fig. 7 is merely an example and is not limited to this.
[0066] Note that server device 120 may set the number of derived images in the range where the character's face is visible to be greater than the number of derived images in the range where the character's face is not visible. Images in the range where the face is not visible refer to, for example, images that do not include facial features such as the eyes and mouth that make up facial expressions, or images that do not include the facial contours. Furthermore, the posture of the character in a derived image differs from the posture of the character in the character image, and the angle of the character in a derived image differs from the angle of the character in the character image.
[0067] The server device 120 then recognizes the character's facial features from each of the derived images thus generated and extracts the features. For example, as shown in Fig. 5, the server device 120 extracts the character's "Body (head, torso, lower body)," "Face," "EyebrowsL (left eyebrow)," "EyebrowsR (right eyebrow)," "EyesL (outline of the left eye, white of the left eye, left eyelashes)," "PupilL (left pupil)," "EyesR (outline of the right eye, white of the right eye, right eyelashes)," "PupilR (right pupil)," "Nose," and "Mouth" from the character image. The server device 120 manages the extracted features hierarchically by associating each feature extracted from the original derived image with the original image as follows:
[0068] Root (a) Body (b) Face (c) EyebrowsL (d) EyebrowsR (e) EyesL (f) PupilL (g) EyesR (h) PupilR (i) Nose (j) Mouth In such a hierarchical structure, the parent part (parent node) of (b) is (a), and the parent part (parent node) of (c) to (j) is (b).
[0069] In step S324, server device 120 generates multiple derived part images of the parts extracted in step S323 using a known image generation model such as stable diffusion. In this case, as with the derived images, server device 120 inputs a prompt containing an image of each extracted part into the image generation model, and obtains the derived part images generated by the image generation model. For example, server device 120 may generate derived part images of eyebrows in multiple states between the open and closed states, or may generate derived part images of eyes at 5-degree intervals within an angle range of 45 degrees left to 45 degrees right, or may generate derived part images of eyes at 5-degree intervals within an angle range of 20 degrees up to 20 degrees down, or may generate derived part images of eyes at 5-degree intervals within an angle range of 15 degrees up, 20 degrees left to 15 degrees down, and 20 degrees right.
[0070] Furthermore, when the generated derived part images are managed in a hierarchical structure as described above, they are managed in association with parameter information such as angle, posture, and type of facial expression that was input at the time of generation. By managing each derived part image in this way, for example, when selecting an image that matches the broadcaster's movements on a viewer terminal, an appropriate derived part image can be selected by checking the broadcaster's angle information and the angle information of the generated derived part image.
[0071] In addition, the server device 120 may generate a derived hair part image while maintaining the characteristic information of the character in the derived image, such as the lines and contours of the character's hair, and / or the contours and shape of the face.
[0072] Similarly to the derived images, server device 120 may generate fewer derived part images that are not related to the character's face than the number of derived part images that are related to the face. This allows for maintaining and improving the satisfaction of streamers and viewers by generating more images related to faces, which are more likely to affect the quality of the avatar and stream, even if the upper limit on the number of images that can be generated is set in advance.
[0073] Furthermore, the server device 120 generates derived part images for the eyes and mouth of a character when the character expresses joy, anger, sadness, happiness, or a unique character trait (personality: yandere, tsundere, little sister type, etc.). The server device 120 may acquire / estimate such character traits by searching the web, and may also generate derived part images for derived emotions, which are emotions that further subdivide joy, anger, sadness, and happiness, based on the character traits (basic emotion: sadness → derived emotions: dirty eyes and watery eyes; further emotions such as smiling, laughing, and impatience). The derived part images generated based on the character traits are defined by a rule base for each character trait, and the server device 120 may generate the derived part images based on such a rule base.
[0074] The server device 120 may generate derivative images in which the characters express emotions or unique character traits (personality: yandere, tsundere, little sister type, etc.).
[0075] In addition, the server device 120 always generates derived part images relating to joy, anger, sadness, and happiness, and may generate derived part images relating to specific character traits for character traits specified by parameters, or may generate them for all pre-set character traits.
[0076] The derived images and / or derived part images expressing these character characteristics are managed with tags attached to each image indicating which character characteristic the image relates to, and the broadcaster may set a specific character characteristic before broadcasting for the derived images and / or derived part images expressing these character characteristics, and server device 120 may set the derived images and / or derived part images corresponding to the selected character characteristic as avatar data. Furthermore, when the broadcaster's action information for the derived images and derived part images meets predetermined conditions (mouth opening of 70% or more, relative position of face and hands within a predetermined distance, etc.), the derived images and / or derived part images with the predetermined character characteristic may be selected and displayed on the viewer's screen.
[0077] Furthermore, the server device 120 may generate a derived part image so as to maintain the artistic style of the derived image. The artistic style refers to, for example, the characteristics of character images commonly seen in animations from a particular animation company or the characteristics commonly seen in character images drawn by a particular illustrator.
[0078] Server device 120 may also increase or decrease the number of derived part images for a predetermined facial expression or a predetermined angle compared to other facial expressions or angles, or may generate derived part images by making changes to a predetermined part (for example, changing teeth to jagged teeth or changing hair to a different color). Such processing content is defined by parameters.
[0079] It should be noted that some of the various methods for generating derived part images described above may be available for a fee.
[0080] Then, for each extracted part, the server device 120 generates and manages a set of {multiple derived part images of the part, information identifying the parent part of the part, the relative position of the part with respect to the parent part, and the playback speed of the multiple derived part images} as metadata.
[0081] In step S325, the server device 120 generates a Json file including the derived image generated in step S322 and the derived part image and metadata generated in step S324. The Json file may include the above-mentioned feature information.
[0082] In step S326, server device 120 stores (registers) the Json file generated in step S325 as avatar data in storage device 125. This allows server device 120 to register the derived image as an image for distribution to cause the character to move.
[0083] Thereafter, when server device 120 receives a request to transmit avatar data of a predetermined character from viewer terminal 140 via I / F 126, server device 120 reads the avatar data of the predetermined character from storage device 125 and transmits the read avatar data to viewer terminal 140 via I / F 126. The avatar data may be transmitted as a binary file.
[0084] In step S242, the viewer terminal 140 uses a well-known technology such as TalkingHead to generate a display image of a character that corresponds to the streamer's movements based on the avatar data and the movement information. Because the avatar data includes derived images in which the character is depicted at different angles and poses, the server device 120 generates a display image using derived images and derived part images that correspond to the movements and states indicated by the movement information. The server device 120 may generate a character by placing a derived part image on the derived image, referring to the metadata. For example, when placing a derived part image of the left eye on the derived image, the server device 120 refers to the metadata and places the derived part image of the left eye at a position relative to the parent part of the left eye, the "face." The server device 120 then refers to the metadata and plays back the derived part image of the left eye at the playback speed of the derived part image of the left eye.
[0085] This allows, for example, a display image of a front-facing character blinking and lip-syncing to be displayed on the display unit 145, making it possible to express the character's facial expressions with a wider variety of facial expressions and gestures than the broadcaster's.
[0086] As a display image displayed on the display unit 145, the original character image may be displayed in the default state.
[0087] In addition, when the server device 120 receives an instruction to modify or regenerate a derived image or a derived part image from an external device such as the viewer terminal 140 or the distributor terminal 100, the server device 120 may modify or regenerate the derived image or the derived part image.
[0088] For example, when server device 120 receives an instruction from an external device to adjust the parameters (dimensions, angle, position, etc.) of a derived part image, it regenerates the derived part image based on the parameters and updates the above metadata accordingly, for example, adjusting the position of the face relative to the character's body, and adjusting the angles and dimensions of the facial features to match the face.
[0089] Furthermore, the image for distribution may be an image used by the distributor for distribution, or an image displayed when a viewer makes a comment.
[0090] The numerical values, processing timing, processing order, processing subject, data (information) acquisition method / destination / source / storage location, etc. used in the above embodiment are given as examples to provide a concrete explanation, and are not intended to be limited to these examples.
[0091] In addition, some or all of the above-described embodiments may be used in appropriate combination, and some or all of the above-described embodiments may be selectively used.
[0092] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention.
Claims
1. On the computer, a generating step of generating a plurality of derivative images in which the character is drawn at different angles or poses based on an image including the input character; a registration step of registering the derivative image as a distribution image for moving the character; A computer program for executing
2. 2. The computer program according to claim 1, further causing the computer to execute a step of specifying the different angles or attitudes.
3. 2. The computer program according to claim 1, wherein the generating step generates the derived image from the image using an image generation model that has been trained to output at least a derived image.
4. the computer program causes the computer to execute a receiving step of receiving input conditions input by a user as generation conditions for the derivative image; 4. The computer program according to claim 3, wherein the generating step generates the derived image from the image using the image generation model in accordance with the input conditions.
5. 2. The computer program according to claim 1, wherein the generating step generates a derivative image for each facial expression of the character.
6. 2. The computer program according to claim 1, wherein the number of derivative images in the area where the character's face is visible is greater than the number of derivative images in the area where the character's face is not visible.
7. 2. The computer program according to claim 1, wherein the generating step extracts parts of the character from the derived image, and generates a plurality of derived part images of the parts based on the extracted parts.
8. 8. The computer program according to claim 7, wherein said generating step generates at least one of said derived image and said derived part image based on the personality of said character.
9. 9. The computer program according to claim 8, wherein the number of derived part images generated from a derived image in which the character faces forward is greater than the number of derived part images generated from a derived image in which the character faces directly backward.
10. 8. The computer program according to claim 7, wherein the registration step registers the derived image, the plurality of derived part images, and information representing positions of the plurality of derived part images in the derived image.
11. 9. The computer program according to claim 8, wherein the features include the eyes, mouth, and eyebrows of the character.
12. 2. The computer program of claim 1, wherein the angle is the angle of the character's face.
13. 2. The computer program according to claim 1, wherein the angle is within a range of angles when the face of the character is facing left, right, up, or down.
14. the pose of the character in the derived image is different from the pose of the character in the image; 2. The computer program of claim 1, wherein an angle of the character in the derived image is different from an angle of the character in the image.
15. 2. The computer program according to claim 1, wherein said registration step registers characteristic information of said character in said image.
16. 16. The computer program product of claim 15, wherein the feature information includes the character's hairline and / or facial contour.
17. 2. The computer program according to claim 1, wherein the image includes a part or all of the upper body of the character.
18. 2. The computer program according to claim 1, wherein the computer program causes the computer to execute a distribution step of distributing the derivative image and action information representing an action of a distributor.
19. A system having a distributor terminal, a server device, and a viewer terminal, The distributor terminal An acquisition means for acquiring motion information representing a distributor's motion; a first transmitting means for transmitting the operation information to the server device; Equipped with The server device a first generating means for generating, based on an image including an input character, a plurality of derivative images in which the character is depicted at different angles or poses; a second transmitting means for transmitting the action information received from the distributor terminal and the derived image generated by the first generating means to the viewer terminal; Equipped with The viewer terminal includes: a second generating means for generating an image of a character corresponding to the movement of the distributor based on the derived image and the action information received from the server device; a display control means for displaying the image generated by the second generation means on a display unit; A system comprising:
20. a generating means for generating, based on an input image including a character, a plurality of derivative images in which the character is drawn at different angles or postures; a registration means for registering the derivative image as a distribution image for causing the character to move; A server device for executing the above.
21. A method for controlling a server device, comprising: a generation step in which the generation means of the server device generates, based on an image including an input character, a plurality of derivative images in which the character is depicted at angles or poses different from each other; a registration step in which the registration means of the server device registers the derivative image as a distribution image for moving the character; A control method comprising:
Citation Information
Patent Citations
Moving image generation device and live communication system
JP2021111102A
Information processing system, information processing method and information processing program
JP2023004585A
Personalized shopping avatar
US20110022965A1
Character image generation device and learning model generation device
WO2022244120A1
Information processing device, information processing method, and information processing program
WO2023176210A1