Programs and Systems
Patent Information
- Application Number
- JP2025065230
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-04-10
AI Technical Summary
【0035】 本発明に係るプログラム及びシステムによれば、ユーザの好みが反映されたアバターを容易に生成することができる。
Smart Images

Figure 0007917209000001_ABST
Abstract
Description
[[Technical Field]]
[0001] The present invention relates to a program and a system. [[Background Art]]
[0002] Conventionally, programs and systems for generating an avatar to be displayed in a video distributed by a distributing user on a video site or the like have been known.
[0003] For example, Patent Document 1 discloses an avatar establishing apparatus (avatar generation system) capable of generating an avatar. This avatar establishing apparatus executes the steps of: receiving a picture including a human face; acquiring an initial parameter corresponding to the picture; receiving a plurality of adjustments input by a user; acquiring adjusted avatar parameters according to the initial parameters and the adjustments; and generating an adjusted avatar according to the avatar parameters. [[Prior Art Document]] [[Patent Document]]
[0004] [[Patent Document 1]] Japanese Unexamined Patent Publication No. 2020-52992 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]
[0005] However, with the technique described in Patent Document 1, since avatar parameters are acquired based on a picture including a human face, the generated avatar is limited to a humanoid avatar. Further, since the technique does not include a step of setting the motion of the avatar, it is difficult to set the motion of the generated avatar according to the user's preference. For this reason, it has not been possible to easily generate an avatar that appropriately reflects the user's preference.
[0006] This invention has been made in view of these problems, and its purpose is to provide a program and system that can easily generate avatars that reflect the user's preferences. [Means for solving the problem]
[0007] To solve the above problems, the program according to the present invention causes a computer to function as an acquisition means for acquiring a first image of an avatar, and a generation means for generating an avatar model of the avatar using the first image. The acquisition means acquires adjustment information for adjusting the avatar model, and the generation means generates motion of the avatar model based on the adjustment information, and generates a second image of the avatar from the avatar model and the motion.
[0008] With this configuration, a first image tailored to the user's preferences is provided to the acquisition means, enabling the generation of various avatar models, not limited to humanoid forms. Furthermore, adjustment information tailored to the user's preferences is provided to the acquisition means, allowing the generation of a second avatar image reflecting the user's preferences without requiring any effort from the user. Therefore, avatars that reflect the user's preferences can be easily generated.
[0009] In another embodiment of the present invention, the generation means adjusts the parameters relating to the avatar model based on the adjustment information.
[0010] This configuration allows the avatar model to more effectively reflect user preferences through parameters based on adjustment information.
[0011] In another embodiment of the present invention, the generation means sets bones, which are the skeleton of the avatar, for the first image.
[0012] This configuration allows users to set bones for the first image according to their preferences.
[0013] In another embodiment of the present invention, the generation means sets a mesh for forming the avatar with respect to the first image on which the bones are set.
[0014] With this configuration, the mesh can be set for the first image according to the user's preference.
[0015] In another embodiment of the present invention, the generation means assigns parameters relating to weight information that set the ease of movement of each part constituting the avatar to the first image on which the mesh is set.
[0016] This configuration allows users to assign parameters related to weight information to the first image according to their preferences.
[0017] In another embodiment of the present invention, the generation means performs at least one of the following based on the adjustment information: setting the bones, setting the mesh, and assigning parameters related to the weight information.
[0018] This configuration allows for effective reflection of user preferences through adjustment information when assigning parameters related to bone settings, mesh settings, and weight information.
[0019] In another embodiment of the present invention, the adjustment information includes at least information regarding the parameter of the intensity of the motion.
[0020] This configuration allows for more detailed settings regarding the avatar model's motion, according to the user's preferences.
[0021] In another embodiment of the present invention, the generation means generates a derived image of the avatar based on the first image, and generates the avatar model using the first image and the derived image.
[0022] According to this configuration, by using derived images to generate an avatar model, it is possible to generate an avatar model that effectively reflects the user's preferences.
[0023] A program according to another aspect of the present invention causes a computer to function as acquisition means for acquiring a first image of an avatar and generation means for generating an avatar model of the avatar, wherein the acquisition means acquires adjustment information for adjusting the avatar model, the adjustment information including attribute information related to attributes of the avatar, and the generation means generates the avatar model of the avatar using the first image and the attribute information.
[0024] According to this configuration, by providing the first image corresponding to the user's preferences to the acquisition means, various avatar models can be generated without being limited to humanoid forms. Furthermore, by providing attribute information corresponding to the user's preferences to the acquisition means, it is possible to generate a second image of an avatar that is set in detail according to the user's preferences. Therefore, an avatar that reflects the user's preferences can be easily generated.
[0025] In a program according to another aspect of the present invention, the generation means generates motion for the avatar model based on the attribute information.
[0026] According to this configuration, based on the attribute information, it is possible to generate motion for an avatar model that reflects the user's preferences in detail.
[0027] In a program according to another aspect of the present invention, the generation means generates a second image of the avatar from the avatar model and the motion.
[0028] According to this configuration, based on the attribute information, it is possible to generate a second image of an avatar that reflects the user's preferences in detail.
[0029] In another embodiment of the present invention, the acquisition means displays the generated avatar model on the screen of the user terminal and acquires the attribute information based on operations performed on the screen on which the avatar model is displayed.
[0030] With this configuration, users can provide attribute information that appropriately reflects their preferences to the acquisition means by interacting with the screen while viewing the generated avatar model on the screen.
[0031] In another embodiment of the present invention, the acquisition means analyzes the first image to acquire the attribute information.
[0032] With this configuration, even if the user does not provide attribute information to the acquisition means, the acquisition means can acquire attribute information from the first image according to the user's preferences.
[0033] The system according to the present invention comprises acquisition means for acquiring a first image of an avatar, and generation means for generating an avatar model of the avatar using the first image, wherein the acquisition means acquires adjustment information for adjusting the avatar model, and the generation means generates motion of the avatar model based on the adjustment information, and generates a second image of the avatar from the avatar model and the motion.
[0034] Another embodiment of the present invention provides a system comprising: acquisition means for acquiring a first image of an avatar; and generation means for generating an avatar model of the avatar, wherein the acquisition means acquires adjustment information for adjusting the avatar model, which includes attribute information relating to the attributes of the avatar; and the generation means generates an avatar model of the avatar using the first image and the attribute information. [Effects of the Invention]
[0035] According to the program and system of the present invention, avatars that reflect the user's preferences can be easily generated. [Brief explanation of the drawing]
[0036] [Figure 1] This is a block diagram showing the overall configuration of the avatar generation system according to the first embodiment. [Figure 2] This block diagram shows the electrical configuration of the server device shown in Figure 1. [Figure 3] This figure shows a portion of multiple basic images uploaded to the server device in the first embodiment. [Figure 4] This figure shows an avatar image in the first embodiment. [Figure 5] This figure shows an example of a screen for setting parameter information for an avatar image in the first embodiment. [Figure 6] This is a magnified view of the upper body of the avatar model in the first embodiment. [Figure 7] This figure shows the avatar's distribution image in the first embodiment. [Figure 8] This flowchart shows the processing flow performed by each functional unit shown in Figure 2 in the avatar generation system according to the first embodiment. [Figure 9] This is a flowchart showing the flow of the avatar model generation process in the first embodiment. [Figure 10] This flowchart shows the flow of the distribution process in the first embodiment. [Figure 11] This flowchart shows the process flow for acquiring a basic image in the avatar generation system according to the second embodiment. [Modes for carrying out the invention]
[0037] <First Embodiment> The avatar generation system (system) 1 according to the first embodiment of the present invention will be described with reference to Figures 1 to 10. The avatar generation system 1 is a system that can easily generate avatars that reflect the user's preferences. Hereinafter, a user who generates an avatar using the avatar generation system 1 will be simply referred to as a "user," and a user who watches a video stream of the generated avatar will be referred to as a "viewing user." As shown in Figure 1, the avatar generation system 1 according to the first embodiment comprises a server device 10 and a plurality of terminal devices 12a to 12c. These devices are configured to communicate with each other via a communication network NT such as the Internet or a telephone network.
[0038] In Avatar Generation System 1, an avatar image is created based on a base image, and an avatar model AM (see Figure 6) modeled for video distribution is generated from the created avatar image. Then, by adjusting the generated avatar model AM based on various information, the final avatar distribution image (second image) SI (see Figure 7) for video distribution is generated. In the following, the generation of the avatar distribution image SI will also be simply referred to as "generating an avatar." Furthermore, the "avatar model AM" mentioned above refers to an avatar image modeled for video distribution based on a base image, and is the version before adjustments for video distribution based on various information.
[0039] The server device 10 is an information processing device (computer) for providing a service that generates avatars and allows users to distribute videos to one or more viewers using the generated avatars. The generated avatars express facial expressions and movements according to the user's preferences. Examples of videos using avatars include live streaming and on-demand streaming in which the user distributing the video has the avatar displayed on the screen perform for one or more viewers. Examples of performances include game commentary, singing, playing musical instruments, and dancing.
[0040] Terminal devices 12a to 12c are information processing devices used by users who distribute videos and users who view the distributed videos. Of these, terminal device 12a will be described as the user's terminal device (hereinafter also referred to as "user terminal 12a"), and terminal devices 12b and 12c will be described as terminal devices used by viewing users (hereinafter also referred to as "viewing terminals 12b and 12c"). Examples of these terminal devices 12a to 12c include mobile phones, smartphones, tablets, personal computers, etc., equipped with input devices and display devices. Examples of input devices include keyboards, touch panels, cameras, microphones, etc. Examples of display devices include monitors, touch panels, etc.
[0041] Next, the functional configuration of the server device 10 will be described. As shown in Figure 2, the server device 10 comprises a control means 20, a storage means 30, a communication means 32, and a display means 34. The control means 20 includes an acquisition means 22 and a generation means 24, and is mainly composed of a CPU (Central Processing Unit) and memory. Furthermore, in addition to the CPU, the control means 20 is equipped with an NPU (Neural Processing Unit) specialized for AI (Artificial Intelligence) processing. The control means 20 functions as various functional units by having the CPU and NPU execute predetermined programs stored in memory or the storage means 30, etc.
[0042] The storage means 30 is composed of a hard disk or the like. The storage means 30 stores various programs and information necessary for executing the processing in the control means 20, processing result information, and various images.
[0043] The communication means 32 consists of a communication interface for communicating with an external device. The communication means 32, for example, sends and receives various types of information with the user terminal 12a.
[0044] The display means 34 displays various images necessary for generating an avatar on the screen of the user terminal 12a. The display means 34 also displays the screen of the streaming video using the user's avatar on the screens of the viewing terminals 12b and 12c.
[0045] The server device 10 can be implemented using an information processing device such as a dedicated or general-purpose server computer. Furthermore, the server device 10 may consist of a single information processing device or multiple information processing devices distributed across the communication network NT.
[0046] The acquisition means 22 performs an acquisition process to acquire various types of information. For example, the acquisition means 22 acquires a 2D basic image (first image) BI (see Figure 3) of an avatar uploaded by the user via the communication means 32. The basic image BI referred to here is the basic image information of the avatar that the user generates using the avatar generation system 1, that is, the avatar that the user wishes to use for distribution. The basic image BI uploaded to the acquisition means 22 may be one image or multiple images. Furthermore, the basic image BI is not limited to human-shaped images, but may also be images of animals or fictional creatures. Examples of basic image BI that are not limited to human-shaped images include images of angels, demons, or characters with cat ears and tails. In addition, the basic image BI may be an image of a character wearing equipment or accessories, such as a staff, sword, hat, or helmet.
[0047] The orientation of the basic image BI uploaded to acquisition means 22 is not limited. It may be a front-facing image, a side-facing image, or a back-facing image. Furthermore, the pose of the basic image BI uploaded to acquisition means 22 is not limited. For example, if a human-shaped basic image BI is uploaded, it may be in a standing pose, a sitting pose, or the like.
[0048] As shown in Figure 3, when a user uploads multiple basic image BIs of a humanoid avatar of their choice (for example, a girl's avatar in the first embodiment), they can upload images BI1 to BI6 for each part (hair, face, torso, both hands, both legs, and accessories) that make up the basic image BI, one image at a time. Although Figure 3 shows six images BI1 to BI6 as an example, the basic image BI that makes up the girl's avatar includes images of multiple parts other than those shown (for example, images of eyebrows, eyes, nose, and mouth that make up the face).
[0049] In the first embodiment, layers are pre-set for each part that constitutes the basic image BI of a humanoid avatar. The user can specify and upload the images of the required layers on the screen displayed on the user terminal 12a.
[0050] The acquisition means 22 overlays the uploaded images of each part according to the set layers and acquires them as a single basic image BI. For example, when images I1 to I6 of each part shown in Figure 3 are uploaded, the acquisition means 22 overlays these images I1 to I6 according to the layers to create a single image, which is acquired as the avatar image shown in Figure 4 (hereinafter referred to as "avatar image CI"). The avatar image CI acquired by the acquisition means 22 is stored in the storage means 30.
[0051] The storage means 30 has several basic avatar image BIs pre-stored in it. The acquisition means 22 may acquire basic image BIs from the user, read these basic image BIs stored in the storage means 30, or acquire basic image BIs from the internet via the communication means 32.
[0052] Furthermore, for example, acquisition means 22 acquires adjustment information 40 for adjusting the avatar model AM generated by generation means 24, which will be described later. The adjustment information 40 includes, for example, parameter information 40A related to avatar adjustment. Parameter information 40A includes information regarding the avatar's position parameters when distributing a video, information regarding the avatar's movement parameters, information regarding parameters for adjusting the position and shape of each part constituting the avatar, and information regarding the motion strength parameters of the avatar model AM. Information regarding the avatar's movement parameters includes, for example, information regarding parameters for controlling the avatar's facial expressions (joy, anger, sadness, emotions, blending rate, intensity and duration of expressions), information regarding physical simulation parameters (physical movement of clothing and accessories (swaying, etc.)), information regarding lip-sync parameters, and information regarding the range of motion parameters for each part constituting the avatar. Information regarding parameters for adjusting the position and shape of each part constituting the avatar includes, for example, information regarding parameters for adjusting the texture (color, texture, pattern, etc.) of the avatar's skin, hair, clothing, etc., and information regarding camera viewpoint parameters (camera position, angle, field of view, tracking ability, etc.).
[0053] The "strength of motion" mentioned above refers to the "weight" of the motion. Specifically, the weight of the motion refers to the reaction speed, ease of movement, and magnitude of movement of parts that follow the movement of another part, such as how much the hair moves when the head of an avatar with long hair is moved. The weight of the motion can be set for each part that makes up the avatar. For example, when setting the weight of the motion for the hair of an avatar with ponytail hair, the hair can be divided into a front bangs section and a back ponytail section, and the movement of the ponytail section can be set to be greater than that of the bangs section.
[0054] Parameter information 40A includes various types of information in addition to the information mentioned above, such as information regarding the parameters of the avatar model AM's hairstyle, the parameters of the avatar model AM's face, the parameters of the avatar model AM's body, the parameters of the avatar model AM's hands, and the parameters of the avatar model AM's legs.
[0055] Information regarding the avatar's position parameters when streaming a video specifically refers to parameters for adjusting the avatar's position in the X and Y axes, as well as the avatar's zoom ratio, on the video streaming screen. Information regarding the avatar's movement parameters specifically refers to parameters for adjusting how much the avatar's movement is amplified in relation to the user's movement, and the strength of the avatar's movement, when the avatar is linked to the user's movement (the streamer). Adjusting how much the avatar's movement is amplified in relation to the user's movement is also called "sensitivity adjustment" or "acceleration adjustment." Adjusting the strength of the avatar's movement in relation to the user's movement means adjusting the weight of the avatar's motion in relation to the user's movement. The strength of the avatar's movement can be adjusted according to differences in body characteristics, such as the difference between a male character image and a female character image.
[0056] Referring to Figure 5, an example of setting the motion strength parameter of the avatar model AM from the parameter information 40A is shown. The avatar model AM generated by the avatar model generation process described later has the motion of each part that makes up the avatar model AM set. In the example shown in Figure 5, the avatar model AM is displayed on the screen of the user terminal 12, the playback bar R1 and strength meter M1 are displayed at the bottom of the screen, the Y-direction adjustment meter M2 and zoom meter M3 are displayed on the left and right sides of the screen, and the play button R2 is displayed in the center of the screen. In this example, the strength of the motion can be set by sliding the strength meter M1. The vertical position of the displayed avatar model AM can be adjusted by sliding the Y-direction adjustment meter M2. The displayed avatar model AM can be zoomed in and out by sliding the zoom meter M3. Then, by tapping the play button R2, the avatar model AM will move with the set motion strength for the length of the playback bar R1. This allows the user to check the degree of the set motion strength.
[0057] The acquisition means 22 can, for example, acquire parameter information 40A from the user. When the acquisition means 22 acquires parameter information 40A from the user, the display means 34 displays the avatar model AM of the avatar, which will be described later, on the screen of the user terminal 12a. The user can input parameter information 40A by operating on the screen with the avatar model AM displayed on the screen of the user terminal 12a as the adjustment target. The input parameter information 40A is acquired by the acquisition means 22 and stored in the storage means 30.
[0058] The storage means 30 has the basic parameter information 40A of the avatar pre-stored in it. The acquisition means 22 may acquire adjustment information 40 from the user, read and acquire the parameter information 40A stored in the storage means 30, or acquire the parameter information 40A from the internet via the communication means 32.
[0059] Furthermore, the adjustment information 40 includes attribute information 40B. For example, the acquisition means 22 acquires the attribute information 40B contained in the adjustment information 40. The attribute information 40B is information about the avatar's attributes and includes basic information and pose information.
[0060] Basic information is fundamental information that constitutes the characteristics of an avatar, and includes information such as gender, personality and characteristics, age, and nationality. Basic information may be automatically determined and obtained by the acquisition means 22 from the avatar image CI, or it may be obtained by the acquisition means 22 from the user. When the acquisition means 22 automatically determines and obtains attribute information 40B from the avatar image CI, the acquisition means 22 obtains attribute information 40B by analyzing the avatar image CI.
[0061] On the other hand, when the acquisition means 22 acquires attribute information 40B from the user, the display means 34, for example, displays selectable buttons (for example, buttons to select male or female) on the user terminal 12a to select gender, age, etc. The user can set the basic information (attribute information 40B) of their avatar by touching or selecting these buttons.
[0062] Pose information is information about the avatar's pose desired by the user (e.g., standing, sitting, arms raised). The acquisition means 22 acquires pose information from the avatar image CI. When the acquisition means 22 acquires the avatar's pose, for example, the control means 20 performs a pose determination process that determines the avatar's pose from the avatar image CI based on known image processing techniques. The acquisition means 22 acquires the avatar's pose determined by the control means 20. The avatar's pose acquired by the acquisition means 22 is stored in the storage means 30.
[0063] The generation means 24 performs generation processing to generate various images. For example, the generation means 24 performs avatar model generation processing to generate a 3D avatar model AM of the avatar using the basic image BI acquired by the acquisition means 22. The avatar model AM referred to here is image information that has been modified from the avatar image CI, including the identification of components described later, the setting of bones which are the skeleton of the avatar, the setting of a mesh to form the avatar, the addition of parameters related to weight information that set the ease of movement of each part that makes up the avatar, and the integration of these settings.
[0064] Component identification refers to identifying each part (face, both hands, both legs, etc.) that makes up an avatar image CI.
[0065] Bone setting refers to setting the avatar's bones as a combination of lines and points for the avatar image CI. The storage means 30 has pre-stored (default) bone reference information 30A, which serves as the basis for setting the bones (see Figure 2). The information for setting the bones includes information on the position and orientation of each bone, information on the hierarchical structure (e.g., the arm bone is set in a lower hierarchy than the shoulder bone), and information on the weight of the bones. The generation means 24 automatically sets the bones based on the avatar image CI, using the bone reference information 30A as a reference. Then, it makes changes and modifications based on attribute information 44, etc., and sets the bones for the avatar image CI. In addition to the bone reference information 30A, the storage means 30 also stores special bone information for when the avatar image CI corresponds to a specific attribute, and the generation means 24 automatically determines and sets whether to apply the bone reference information 30A or the special bone information to the avatar image CI based on adjustment information 40 and attribute information 44. When dedicated bone information is applied, for example, dedicated bone information for women is prepared for a female avatar image CI.
[0066] Mesh setting refers to setting a mesh pattern on the surface of each part that makes up the avatar, based on the bones, for an avatar image CI with the bone settings described above. The textures mentioned above are images (skin, clothing, etc.) that are applied to this mesh. The storage means 30 has pre-stored (default) mesh reference information 30B, which serves as the basis for setting the mesh (see Figure 2). The information for setting the mesh includes information about the shape of the mesh (triangle, quadrilateral, etc.), information about the fineness of the mesh (which may be set as the number of vertices, faces, polygons, etc.), and information about the material of the mesh (texture of the mesh surface (e.g., transparency, reflectivity, surface roughness), etc.). The generation means 24 automatically sets the mesh based on the avatar image CI, using the mesh reference information 30B as a reference. Then, it makes changes and modifications based on attribute information 40B, etc., and sets the mesh for the avatar image CI. In addition, the storage means 30 stores dedicated mesh information for cases where the avatar image CI corresponds to a specific attribute, separate from the mesh reference information 30B. The generation means 24 automatically determines and sets whether to apply the mesh reference information 30B or the dedicated mesh information to the avatar image CI based on the adjustment information 40 and attribute information 44. When dedicated mesh information is applied, for example, dedicated mesh information for women is prepared for a female avatar image CI.
[0067] Assigning parameters related to weight information means setting a parameter indicating weight for each vertex of the mesh (for each section separated by the mesh) of the avatar image CI with the above mesh settings. The storage means 30 has (default) weight setting information 30C for setting parameters related to weight information stored in advance (see Figure 2). The generation means 24 assigns parameters related to weight information based on the avatar image CI according to this weight setting information 30C. In addition to the weight setting information 30C, the storage means 30 also stores special weight information for when the avatar image CI corresponds to a specific attribute, and the generation means 24 determines and automatically sets whether to apply the special weight information to the avatar image CI based on the adjustment information 40 and attribute information 44.
[0068] The generation means 24 may change and modify the parameters related to weight information based on the attribute information 44. For example, since male avatars and female avatars have different physical characteristics, it is better to change the parameters related to weight information according to gender. By assigning parameters related to weight information according to gender, for example, if the avatar is a woman with long hair, when the avatar moves in a video, it is possible to express femininity in the way the hair sways and the weight of the hair.
[0069] The integration of settings refers to integrating information regarding the identification of the above components, bone settings, mesh settings, and the assignment of parameters related to weight information. The generation means 24 integrates this information to generate the avatar model AM.
[0070] Furthermore, when the generation means 24 sets bones, sets meshes, and assigns parameters related to weight information, it may prioritize setting and modifying based on attribute information 40B over setting and assigning based on avatar image CI. When setting and assigning based on avatar image CI, there is a risk of misidentification of the image, that is, a discrepancy may occur between the settings based on avatar image CI and attribute information 40B. However, by prioritizing attribute information 40B that reflects the user's preferences and instructions, the user can make final adjustments in the generation of the avatar model AM.
[0071] In this way, the avatar generation system 1 of the first embodiment automatically adjusts the bone settings, mesh settings, and the assignment of parameters related to weight information according to various variables. Therefore, it is possible to generate an avatar model AM that reflects the user's preferences without the user having to take the time to make various adjustments.
[0072] Furthermore, the generation means 24 may perform at least one of the following actions based on the adjustment information 40 acquired by the acquisition means 22: setting bones, setting meshes, and assigning parameters related to weight information.
[0073] Furthermore, for example, the generation means 24 performs motion generation processing (so-called rigging) to set motion for the avatar model AM. In the avatar generation system 1, motion generation is performed as one of the adjustments to the avatar model AM. In the motion generation processing, the generation means 24 uses the avatar model AM to generate avatar motion data relating to predetermined basic movements. The motion referred to here is not limited to motion that accurately reflects the physical movements of the avatar, but also includes motion that has been adjusted to reflect the characteristics of the avatar, etc.
[0074] Types of motion include motion for individual parts, motion when the user is stationary, and motion when the user performs a specific action. Motion for individual parts includes hair motion, eyebrow motion, eye motion, mouth motion, breathing motion, head motion, and body motion. Hair motion includes, for example, the swaying of bangs, side hair, and back hair. Eyebrow motion includes, for example, up-and-down movement and changes in eyebrow shape. Eye motion includes, for example, blinking, eyeball movement, and pupil movement. Mouth motion includes, for example, opening and closing of the mouth. Breathing motion includes, for example, automatic breathing. Head motion and body motion include, for example, rotation around a specific axis.
[0075] The motion settings for the avatar model AM may be automatically set by the generation means 24 based on attribute information 40B, or the generation means 24 may set the motion based on user input. When the generation means 24 sets the motion based on attribute information 40B, it automatically increases the movement of characteristic parts of the avatar model AM, for example. In this case, if the avatar model AM has cat ears and a tail, the generation means 24 will set the cat ears and tail to move more easily (setting the motion for each part). On the other hand, when the motion is set based on user input, the display means 34 displays a setting screen for setting the motion of the avatar model AM on the screen of the user terminal 12a. The user can set the motion of the avatar model AM by entering setting information related to the motion settings on the setting screen displayed on the screen of the user terminal 12a. The storage means 30 stores dedicated motion information for when the avatar model AM corresponds to a specific attribute, and the generation means 24 determines and automatically sets whether to set the motion for the avatar image CI based on attribute information 44 or apply the dedicated motion information.
[0076] Referring to Figure 6, we will now explain how a user can set the motion of the avatar model AM's head AM1. As shown in Figure 6, the generation means 24 sets the avatar model AM's head AM1 to rotate around an axis perpendicular to the plane of Figure 6. Specifically, the user can set the maximum angle of rotation for both the clockwise direction (the + side shown in Figure 6) and the counterclockwise direction (the - side shown in Figure 6) when viewing Figure 6 from the front. Similarly, the user can also set the rotation angle around an axis horizontal in the vertical direction of Figure 6, and the rotation angle around an axis horizontal in the left-right direction of Figure 6.
[0077] Furthermore, for example, the generation means 24 adjusts parameters related to the avatar model AM based on the adjustment information 40 acquired by the acquisition means 22. Specifically, the generation means 24 adjusts the avatar model AM using the parameter information 40A of the adjustment information 40, which includes information on the parameters of the avatar's position when distributing the video, information on the parameters of the avatar's movement, information on parameters for adjusting the position and shape of each part that makes up the avatar, and information on the parameters of the motion strength of the avatar model AM.
[0078] Furthermore, for example, the generation means 24 generates a distribution image SI from the avatar model AM and the motion set for the avatar model AM (see Figure 7). The distribution image SI is the final avatar image for video distribution. When the generation means 24 generates the distribution image SI, the display means 34 displays the avatar model AM on the screen of the user terminal 12a. The user inputs parameter information 40A as the target for adjustment of the avatar model AM displayed on the screen of the user terminal 12a and adjusts the avatar model AM. As shown in Figure 7, the adjusted avatar model AM is displayed on the screen of the user terminal 12a as a distribution image SI and stored in the storage means 30.
[0079] Each time a distribution image SI is generated in the manner described above, it is stored in the storage means 30. When streaming a video, the user can select and use an avatar from among the multiple distribution image SI avatars stored in the storage means 30.
[0080] Next, referring to Figure 8, we will explain the processing flow performed by each functional unit shown in Figure 2 during the process in which a user generates an avatar for video distribution using the avatar generation system 1.
[0081] (Step S12) Acquisition means 22 acquires basic image BI from the user. Specifically, the user uploads one or more basic image BIs. If there is one basic image BI uploaded, acquisition means 22 acquires the basic image BI as an avatar image CI. If there are multiple basic image BIs uploaded, acquisition means 22 combines them into a single image and acquires it as an avatar image CI. Then, the process proceeds to step S14.
[0082] (Step S14) The acquisition means 22 acquires attribute information 40B. Specifically, the acquisition means 22 acquires the basic information contained in the attribute information 40B by automatically determining it from the avatar image CI, or by acquiring it from the user. Next, the control means 20 performs a pose determination process on the pose information contained in the attribute information 40B, and the acquisition means 22 acquires the determined avatar pose as pose information. Then, the process moves on to step S16.
[0083] (Step S16) The generation means 24 executes an avatar model generation process to generate an avatar model AM. The sequence of steps in the avatar model generation process will now be explained with reference to Figure 9.
[0084] (Step S16A) The generation means 24 performs a component identification process to identify components in the avatar image CI. The process then proceeds to step S16B.
[0085] (Step S16B) The generation means 24 performs bone processing to set bones for the avatar image CI whose components have been identified. Then the processing proceeds to step S16C.
[0086] (Step S16C) The generation means 24 performs a meshing process to set a mesh on the avatar image CI with bones set. Then the process proceeds to step S16D.
[0087] (Step S16D) The generation means 24 performs a weight assignment process to assign parameters related to weight information to the avatar image CI with the mesh set. The process then proceeds to step S16E.
[0088] (Step S16E) The generation means 24 performs an integration process that combines information regarding component identification, bone settings, mesh settings, and the assignment of parameters related to weight information. Through this process, the avatar model AM is generated. The process then returns to Figure 8 and proceeds to step S18.
[0089] (Step S18) The generation means 24 executes a motion setting process to set motion for the avatar model AM generated in the avatar model generation process. Then the process proceeds to step S20.
[0090] (Step S20) The control means 20 displays an image of the avatar model AM, for which motion settings have been applied, on the screen of the user terminal 12a. The avatar model AM displayed on the screen of the user terminal 12a performs the performance according to the set motions, etc. This allows the user to confirm whether or not the desired motion settings have been applied to the avatar model AM. Then, the process proceeds to step S22.
[0091] (Step S22) The acquisition means 22 acquires the parameter information 40A contained in the adjustment information 40. Specifically, the acquisition means 22 acquires the parameter information 40A from the user. Then, the process proceeds to step S24.
[0092] (Step S24) The generation means 24 adjusts the avatar using the parameter information 40A acquired by the acquisition means 22. Specifically, the generation means 24 adjusts the parameters for the avatar's position when distributing the video, the parameters for the avatar's movement, parameters for adjusting the position and shape of each part that makes up the avatar, and parameters for the intensity of the motion of the avatar model AM. Then, the process moves on to step S26.
[0093] (Step S26) The control means 20 displays an image of the adjusted avatar model AM on the screen of the user terminal 12a. The avatar model AM displayed on the screen of the user terminal 12a performs the performance according to the adjustments made in step S24. This allows the user to confirm whether the desired adjustments have been made to the avatar. Then, the process proceeds to step S28.
[0094] (Step S28) The control means 20 determines whether parameter information 40A has been input again for the image of the avatar model AM displayed on the screen of the user terminal 12a, that is, whether an adjustment instruction has been given. If the control means 20 determines that an adjustment instruction has been given by the user, the process returns to step S22. If the control means 20 determines that no adjustment instruction has been given by the user, the process proceeds to step S30.
[0095] (Step S30) The generation means 24 generates a distribution image SI from the adjusted avatar model AM and the motion set for the avatar model AM. The generated distribution image SI is stored in the storage means 30. Then, the control means 20 completes the series of processes shown in Figure 8.
[0096] Through the above sequence of steps performed by Avatar Generation System 1, an avatar, which is a streaming image SI for video distribution, is generated.
[0097] Next, referring to Figure 10, we will explain the distribution process when a user performs video distribution using an avatar of the distribution image SI generated by the avatar generation system 1.
[0098] (Step S42) In the distribution process, first, the display means 34 displays avatars of multiple distribution images SI stored in the storage means 30 on the screen of the user terminal 12a. The user can select an avatar to be used for video distribution from among the multiple distribution image SI avatars. If an avatar is selected by the user, the process moves to step S44. If no avatar is selected by the user, the control means 20 repeatedly executes the process of step S42.
[0099] (Step S44) The control means 20 displays the avatar of the selected distribution image SI on the screen of the user terminal 12a for video distribution and starts video distribution. In the case of live distribution, the avatar of the distribution image SI is simultaneously displayed on the screen of the viewer's terminal device. The avatar displayed on the screen performs a performance according to the motion and other settings set in the series of steps shown in Figure 8. Then, the process moves to step S46.
[0100] (Step S46) The display means 34 displays a screen for inputting parameter information 40A on the user terminal 12a's screen while the avatar is performing a performance on the user terminal 12a's screen. The control means 20 then determines whether the user has input the parameter information 40A again, that is, whether an adjustment instruction has been given. If the control means 20 determines that an adjustment instruction has been given by the user, the process proceeds to step S48. If the control means 20 determines that there was no adjustment instruction from the user, the process proceeds to step S50.
[0101] (Step S48) The control means 20 modifies (adjusts) the distributed image SI according to the parameter information 40A input by the user. The modified distributed image SI is stored (stored) as a modified image in the storage means 30. Then the process proceeds to step S50.
[0102] (Step S50) The control means 20 stops displaying the avatar in the distributed image SI and terminates the video distribution. Then, the control means 20 terminates the series of processes shown in Figure 10.
[0103] As described above, the program provided by the avatar generation system 1 of the first embodiment causes the computer to function as an acquisition means 22 for acquiring a basic image BI of an avatar, and a generation means 24 for generating an avatar model AM of an avatar using an avatar image CI based on the basic image BI. The acquisition means 22 acquires adjustment information 40 for adjusting the avatar model AM, and the generation means 24 generates motion for the avatar model AM based on the adjustment information 40, and also generates a distribution image SI of the avatar from the avatar model AM and motion.
[0104] With this configuration, a basic image BI tailored to the user's preferences is provided to the acquisition means 22, enabling the generation of various avatar models AM, not limited to humanoid forms. Furthermore, adjustment information 40 tailored to the user's preferences is provided to the acquisition means 22, enabling the generation of avatar distribution images SI that reflect the user's preferences without requiring any effort from the user. As a result, avatars that reflect the user's preferences can be easily generated.
[0105] Furthermore, in the program provided by the avatar generation system 1, the generation means 24 adjusts parameters related to the avatar model AM based on the adjustment information 40.
[0106] With this configuration, user preferences can be more effectively reflected in the avatar model AM through parameters (parameter information 40A) based on the adjustment information 40.
[0107] Furthermore, in the program provided by the avatar generation system 1, the generation means 24 sets bones, which are the skeleton of the avatar, for the avatar image CI based on the basic image BI, sets a mesh to form the avatar for the avatar image CI with bones set, and assigns parameters related to weight information that set the ease of movement of each part constituting the avatar for the avatar image CI with mesh set. With this configuration, bones and meshes can be set and parameters related to weight information can be assigned to the avatar image CI based on the basic image BI according to the user's preferences.
[0108] Furthermore, in the program provided by the avatar generation system 1, the generation means 24 performs at least one of the following based on the adjustment information 40: setting bones, setting meshes, and assigning parameters related to weight information.
[0109] This configuration allows for effective reflection of user preferences through adjustment information 40 in the setting of bones, mesh settings, and the assignment of parameters related to weight information.
[0110] Furthermore, in the program provided by the avatar generation system 1, the adjustment information 40 includes at least information regarding the motion intensity parameter.
[0111] This configuration allows for more detailed settings regarding the avatar model's (AM) motion, according to the user's preferences.
[0112] Furthermore, the program provided by the avatar generation system 1 causes the computer to function as an acquisition means 22 for acquiring a basic image BI of the avatar, and a generation means 24 for generating an avatar model AM of the avatar. The acquisition means 22 acquires adjustment information 40 for adjusting the avatar model AM, which includes attribute information 40B relating to the attributes of the avatar, and the generation means 24 generates an avatar model AM of the avatar using the basic image BI and the attribute information 40B.
[0113] With this configuration, a basic image BI tailored to the user's preferences is provided to the acquisition means 22, enabling the generation of various avatar models AM, not limited to humanoid forms. Furthermore, attribute information 40B tailored to the user's preferences is provided to the acquisition means 22, enabling the generation of a detailed avatar distribution image SI tailored to the user's preferences. As a result, avatars that reflect the user's preferences can be easily generated.
[0114] Furthermore, in the program provided by the avatar generation system 1, the generation means 24 generates motion for the avatar model AM based on attribute information 40B.
[0115] With this configuration, it is possible to generate motions for the avatar model AM that reflect the user's preferences in detail, based on attribute information 40B.
[0116] Furthermore, in the program provided by the avatar generation system 1, the acquisition means 22 displays the generated avatar model AM on the screen of the user terminal 12a and acquires attribute information 40B based on the operation on the screen where the avatar model AM is displayed.
[0117] With this configuration, the user can provide the acquisition means 22 with attribute information 40B that appropriately reflects the user's preferences by operating on the screen while viewing the generated avatar model AM on the screen.
[0118] Furthermore, in the program provided by the avatar generation system 1, the acquisition means 22 analyzes the basic image BI to obtain attribute information 40B.
[0119] With this configuration, even if the user does not provide attribute information 40B to the acquisition means 22, the acquisition means 22 can acquire attribute information 40B from the avatar image CI based on the basic image BI according to the user's preferences.
[0120] <Second Embodiment> Next, a second embodiment of the present invention will be described with reference to Figure 11.
[0121] In the avatar generation system according to the second embodiment, the function of the generation means differs in part from that of the generation means 24 in the first embodiment. Other configurations are the same as in the first embodiment, and therefore their explanation will be omitted or simplified.
[0122] The generation means of the second embodiment, in addition to the functions of the generation means 24 of the first embodiment, generates derived images of the avatar based on the basic image BI. Furthermore, the generation means generates an avatar model AM using the basic image BI and the generated derived images.
[0123] The derived images referred to here are images generated by being derived from the base image BI. Specifically, for example, if the base image BI uploaded by the user is only one image of an avatar viewed from the front, the generation means will generate derived images of that avatar with different angles and poses. In this case, for example, the user may input prompts to the NPU of the control means 20 to change the base image BI and the angle and pose, and the generation means may then generate the derived images.
[0124] For example, if a single basic image BI uploaded by a user is an image of a humanoid avatar that does not include hair, face, or legs, the generation means generates derived images of the hair, face, and legs corresponding to the basic image BI based on that basic image BI. In this case, for example, the user may input prompts to the NPU of the control means 20 to specify the basic image BI and the parts to be generated, and the generation means may then generate the derived images.
[0125] Next, in the processing flow shown in Figure 8, the processing flow executed by the acquisition means 22 and the generation means of the second embodiment in step S12 will be explained with reference to Figure 11.
[0126] (Step S12A) The acquisition means 22 acquires a basic image BI from the user and acquires an avatar image CI based on the basic image BI. Then, the process proceeds to step S12B.
[0127] (Step S12B) The generation means generates derived images based on the basic image BI or avatar image CI acquired by the acquisition means 22. The generation means may generate one or more derived images. For example, the user may input the desired number of derived images, and the generation means may generate derived images corresponding to that number. Then, the process proceeds to step S14 shown in Figure 8, and the subsequent processing is executed.
[0128] As described above, in the avatar generation system of the second embodiment, the generation means generates derived images of the avatar based on a basic image BI or avatar image CI, and generates an avatar model AM using the basic image BI and the derived images.
[0129] With this configuration, by using derived images to generate avatar models (AM), it is possible to generate avatar models (AM) that effectively reflect the user's preferences.
[0130] <Variation> The embodiments described above are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and essence of the invention, as well as in the claims of the invention and its equivalents.
[0131] In the embodiments described above, examples were shown in which a 2D base image BI is uploaded by the user. However, the base image BI uploaded by the user may also be a 3D image. In this case, it is possible to generate an avatar image CI that better reflects the user's preferences.
[0132] Furthermore, while the above embodiments illustrate a configuration in which the acquisition means 22 acquires a basic image BI uploaded by a user of the avatar generation system 1, the acquisition means 22 may also acquire a basic image BI uploaded by another user, such as a company. With this configuration, an avatar image CI can be generated based on high-quality basic image BI provided by a company or the like. Alternatively, for example, the control means 20 may determine the user's preference tendencies from among the avatars of multiple distributed image SIs stored in the storage means 30, and the control means 20 may provide the user with a basic image BI of an avatar that reflects the user's preferences. With this configuration, the acquisition means 22 can acquire a basic image BI that reflects the user's preferences without the user having to provide a basic image BI.
[0133] Furthermore, while the above embodiments illustrate a configuration in which the user uploads basic image BIs for each related part, the control means 20 may also be configured to determine whether or not there is a relationship between the uploaded basic image BIs. If the control means 20 determines that there is no relationship between the basic image BIs, the control means 20 may not accept uploads from the user for the unrelated basic image BIs. With this configuration, by having the control means 20 accept only related basic image BIs, it is possible to generate a high-quality avatar image CI based on multiple basic image BIs.
[0134] Furthermore, while the above embodiments illustrate a configuration in which the user uploads basic image BIs for each related part, the user may upload multiple basic image BIs that are not related to each other. For example, if the user uploads two images that are unrelated and have completely different properties, such as an image of a human and an image of a dog, the control means 20 may generate an image of a fictional creature by combining the human and the dog based on the two images, and the acquisition means 22 may acquire the generated image of the fictional creature as an avatar image CI. With this configuration, the user can enjoy the fun of generating an image of a fictional creature by providing two images with different properties.
[0135] Furthermore, for example, if a user uploads two images, the control means 20 may generate an image unrelated to the two images, and the acquisition means 22 may acquire the generated image as an avatar image CI. With this configuration, the user can enjoy the fun of generating an unexpected image that cannot be imagined from the two images.
[0136] Furthermore, while the above embodiments illustrate a configuration in which the generation means 24 sets bones and meshes for the avatar image CI and assigns parameters related to weight information, it is also possible to configure the system in which, for example, the user sets bones and meshes for the avatar image CI and assigns parameters related to weight information. Alternatively, after the generation means sets bones and meshes for the avatar image CI and assigns parameters related to weight information, the user may be able to modify the bone and mesh settings and weight information parameters for the avatar image CI. With this configuration, for example, it is possible to set feminine bones for a male avatar image CI, and the user's preferences can be better reflected in the bone and mesh settings and the assignment of weight information parameters.
[0137] Furthermore, in each of the above embodiments, the generation means 24 is shown to set bones for the avatar image CI, then set meshes, and then assign parameters related to weight information, but the order of these steps is not limited. For example, the bones may be set after the mesh is set for the avatar image CI.
[0138] Furthermore, in each of the above embodiments, the generation means 24 is shown as generating a distribution image SI of an avatar from an avatar model AM and the motion of the avatar model AM, but the generated image is not limited to the distribution image SI. The generation means 24 may also generate images that can be used for purposes other than video distribution.
[0139] Furthermore, while the above embodiments illustrate a configuration in which the distributed image SI is modified by adjustment instructions from the user during the distribution process, the system is not limited to this. For example, the storage means 30 may store past distributed videos, and the control means 20 may identify prominent areas or areas with significant movement in the past distributed image SI during distribution, and the acquisition means 22 may acquire information on these areas. The control means 20 may then modify the distributed image SI for these areas. With this configuration, past distributed image SI can be utilized, eliminating the need for the user to take extra effort to modify the distributed image SI, and allowing for easy generation of modified images.
[0140] Furthermore, when utilizing past distributed image SIs as described above, the acquisition means 22 may acquire information on areas that exhibit characteristic movements in the past distributed image SIs, and the generation means 24 may generate a modified image that makes the parts related to the movement of those areas appear more beautiful. With this configuration, the visual effect of the distributed image SIs can be enhanced by utilizing past distributed image SIs.
[0141] Alternatively, the generation means 24 may not generate motion for the avatar model AM, but instead generate the avatar's distribution image SI based on the motion of the avatar model AM and past distribution image SI. In this case, the acquisition means 22 does not need to acquire adjustment information 40 to adjust the avatar model AM. Alternatively, the generation means 24 may generate motion for the avatar model AM from the motion of multiple past distribution image SI. [Explanation of symbols]
[0142] 1: Avatar generation system, 20: Control means, 22: Acquisition means, 24: Generation means, 40: Adjustment information, AM: Avatar model, : Basic image (first image), SI: Distribution image (second image)
Claims
1. Computers, A means for obtaining the first image of an avatar. A generation means for generating an avatar model of the avatar using the first image, A storage means for storing bone reference information that serves as a basis for setting the bones, which are the skeleton of the aforementioned avatar. To make it function as, The acquisition means acquires adjustment information for adjusting the avatar model, which includes at least one of the following: information regarding the parameters of the avatar's movement, information regarding the parameters of the motion strength of the avatar model, and information regarding the parameters of the avatar's position when distributing a video. The generation means applies the bone reference information to the first image based on the adjustment information to automatically set the bones, generates motion for the avatar model based on the adjustment information, and generates a second image of the avatar from the avatar model and the motion. program.
2. The generation means sets the motion of the avatar model based on at least one of the information relating to the strength parameter of the motion of the avatar model included in the adjustment information and the information relating to the position parameter of the avatar when distributing the video. The program according to claim 1.
3. The storage means stores the bone reference information and dedicated bone information when the first image corresponds to a specific attribute. The generation means automatically sets the bone by applying the bone reference information and the dedicated bone information. The program according to claim 1.
4. The information for setting the bones includes, in addition to the bone reference information or the dedicated bone information, at least one of the following: information on the position and orientation of each bone, information on the hierarchical structure, or information on the weight of the bones. The program according to claim 3.
5. The generation means adjusts the parameters relating to the avatar model based on the adjustment information. The program according to claim 1.
6. The generation means sets a mesh for forming the avatar on the first image on which the bones are set. The program according to claim 1.
7. The generation means assigns parameters relating to weight information to the first image on which the mesh is set, which set the ease of movement of each part constituting the avatar. The program according to claim 6.
8. The generation means performs at least one of the following based on the adjustment information: setting the mesh and assigning parameters related to the weight information. The program according to claim 7.
9. The generation means generates a derived image of the avatar based on the first image, and generates the avatar model using the first image and the derived image. The program according to claim 1.
10. The acquisition means displays the generated avatar model on the screen of the user terminal and acquires the adjustment information based on the operation on the screen on which the avatar model is displayed. The program according to any one of claims 1 to 9.
11. Computers, A means for obtaining the first image of an avatar. A generation means for generating the avatar model of the aforementioned avatar, A storage means for storing bone reference information that serves as a basis for setting the bones, which are the skeleton of the aforementioned avatar. To make it function as, The acquisition means acquires adjustment information for adjusting the avatar model, which includes attribute information relating to the attributes of the avatar. The generation means applies the bone reference information to the first image based on the attribute information to automatically set the bones, sets parameters related to weight information that set the ease of movement of each part constituting the avatar, modifies these parameters based on the attribute information, and generates an avatar model of the avatar using the first image and the attribute information. program.
12. The storage means stores the bone reference information and dedicated bone information when the first image corresponds to a specific attribute. The generation means automatically sets the bone by applying the bone reference information and the dedicated bone information. The program according to claim 11.
13. The information for setting the bones includes, in addition to the bone reference information or the dedicated bone information, at least one of the following: information on the position and orientation of each bone, information on the hierarchical structure, or information on the weight of the bones. The program according to claim 12.
14. The program according to claim 11, wherein the generation means generates motion of the avatar model based on the attribute information.
15. The program according to claim 14, wherein the generation means generates a second image of the avatar from the avatar model and the motion.
16. The acquisition means displays the generated avatar model on the screen of the user terminal and acquires the attribute information based on the operation on the screen on which the avatar model is displayed. The program according to claim 15.
17. The acquisition means analyzes the first image to acquire the attribute information. The program according to claim 15.
18. A means for obtaining the first image of an avatar, A generation means for generating an avatar model of the avatar using the first image, The system includes a storage means for storing bone reference information that serves as a basis for setting the bones, which are the skeleton of the avatar, The acquisition means acquires adjustment information for adjusting the avatar model, which includes at least one of the following: information regarding the parameters of the avatar's movement, information regarding the parameters of the motion strength of the avatar model, and information regarding the parameters of the avatar's position when distributing a video. The generation means applies the bone reference information to the first image based on the adjustment information to automatically set the bones, generates motion for the avatar model based on the adjustment information, and generates a second image of the avatar from the avatar model and the motion. system.
19. A means for obtaining the first image of an avatar, A generation means for generating the avatar model of the aforementioned avatar, The system includes a storage means for storing bone reference information that serves as a basis for setting the bones, which are the skeleton of the avatar, The acquisition means acquires adjustment information for adjusting the avatar model, which includes attribute information relating to the attributes of the avatar. The generation means applies the bone reference information to the first image based on the attribute information to automatically set the bones, sets parameters related to weight information that set the ease of movement of each part constituting the avatar, modifies these parameters based on the attribute information, and generates an avatar model of the avatar using the first image and the attribute information. system.
Citation Information
Patent Citations
Animation character creation method and device, equipment, storage medium and program product
CN118799457A
Method and device for generating avatars
JP2020052992A
Computer Method and Apparatus for Rotating 2D Cartoons Using 2.5D Cartoon Models
US20120075284A1
Video generation device
WO2021039857A1