Avatar generation device, learning device, and program
The avatar generation device addresses the inefficiency of manual avatar setup by learning transformation relationships between virtual spaces, enabling automatic avatar generation across different virtual environments, thus improving user experience and reducing setup effort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KDDI CORP
- Filing Date
- 2023-03-24
- Publication Date
- 2026-06-01
AI Technical Summary
Users are required to manually set up avatars for each virtual space they enter, which is cumbersome and inefficient as the settings differ for each space, leading to a significant obstacle in using multiple virtual spaces.
An avatar generation device that utilizes a pre-learned conversion relationship to automatically generate a second avatar for a user in a second virtual space based on the first avatar set up in a first virtual space, reducing the need for manual setup by learning transformation relationships between avatars across different virtual environments.
Reduces the effort required to set up avatars in new virtual spaces, enhancing user experience and reducing user drop-off by automating the avatar setup process across multiple virtual environments.
Smart Images

Figure 0007868003000001 
Figure 0007868003000002 
Figure 0007868003000003
Abstract
Description
[Technical Field]
[0001] This invention relates to an avatar generation device, a learning device, and a program. [Background technology]
[0002] Virtual spaces and services within them are also called metaverses, where each user uses an avatar, which is their virtual representation, to communicate with other users through these avatars. Examples of metaverses include various social networking services (SNS) that operate in virtual spaces, virtual offices, and online games. Users participating in the metaverse need to prepare their own avatars, and prior art related to preparing such avatars can be found in patent documents 1-3, etc.
[0003] Patent Document 1 relates to a technology for generating a caricature image of a user as an avatar. Patent Document 1 improves upon existing methods, which involve preparing a database of images of various facial parts and their feature quantities as template images, taking a face image taken from the user as input, and generating a caricature image by combining similar template images for each facial part of the face image using feature quantity matching. Specifically, regarding the facial part images of the resulting caricature image, if they are judged to be dissimilar, the technology accepts the user's selection of modifications from within the template images, and then updates the database to reflect the correspondence between the feature quantities extracted from the input face image and the facial parts of template images that were judged to be similar as a result of generation and necessary modifications. By improving the accuracy of feature quantity matching in the database from which the caricature image is generated, the similarity of the generated caricature image is improved.
[0004] Patent Document 2 relates to a technology for recommending a list of appropriate face avatars for a user's social networking service (SNS) or other platform, which are linked in real time as videos to the user's facial orientation and changes in facial expression using face tracking. Specifically, by analyzing user attributes, text messages such as the user's inventions, and SNS attributes, if it is estimated that the platform is discussing topics such as the Soccer World Cup, then an avatar of a soccer player is prepared.
[0005] Patent document 3 relates to a technology that "analyzes an image containing a human face and automatically generates an animal-shaped avatar corresponding to the human face." [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2014-006873 [Patent Document 2] Special Publication No. 2018-505462 [Patent Document 3] Japanese Patent Publication No. 2022-060420 [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] However, with the conventional technologies described above, there was a problem in that if a user was using multiple virtual spaces, they would have to go through the trouble of setting up an avatar appropriate for each virtual space each time.
[0008] In other words, as mentioned above, when users enjoy services through a virtual space (metaverse), especially in a virtual space where multiple people exist simultaneously, it is necessary to set up an avatar to distinguish oneself from others as a digital representation of oneself.
[0009] On the other hand, virtual spaces require their settings to be adjusted according to the services they offer, that is, the value users receive. As the number of services offered through virtual spaces increases, the number of virtual spaces will also increase, and it is anticipated that users will have to go through the extra step of setting up an avatar that is appropriate for each virtual space (specifically, an avatar that matches the worldview of that virtual space) each time. Since the settings differ for each virtual space, having users have to make such settings each time will be a major obstacle to using the service.
[0010] Patent Document 1 describes a technology for automatically selecting parts to be used by an avatar from an image, and Patent Document 2 describes a technology for assisting avatar selection based on the avatar's image, profile information, and communication content with others. However, both require the use and configuration of the user's information in the real world each time. Patent Document 3 describes a technology for automatically creating animal-shaped avatars from human face photographs. While this technology can be converted into avatars other than human ones, it still requires the use of the user's information in the real world.
[0011] In other words, none of the conventional technologies had taken measures to automatically generate a second avatar for a user that was suitable for a second virtual space, separate from the first virtual space, by utilizing information such as the first avatar already set up for that user in the first virtual space. As a result, users had to manually set up a suitable avatar for each user in both the first and second virtual spaces each time.
[0012] In view of the problems of the prior art described above, the first objective of the present invention is to provide an avatar generation device that can reduce the effort required for avatar setup. The second objective is to provide a learning device that learns the conversion relationships used by this avatar generation device. [Means for solving the problem]
[0013] To achieve the first objective described above, the present invention is an avatar generation device that takes information of a first avatar used by a user in a first virtual space as input, and outputs information of a second avatar used by the user in the second virtual space by applying a pre-learned conversion relationship using the information of the avatar used in the first virtual space and the information of the avatar used in the second virtual space to the input.
[0014] To achieve the second objective described above, the present invention provides a learning device for learning a transformation relationship used by an avatar generation device that takes information of a first avatar used by a user in a first virtual space as input, and outputs information of a second avatar used by the user in a second virtual space by applying a pre-learned transformation relationship using information of an avatar used in the first virtual space and information of an avatar used in a second virtual space to the input, wherein the transformation relationship is learned as a relationship that takes setting values of individual items for generating each avatar used in the first virtual space as input and outputs setting values of individual parts for generating each avatar used in the second virtual space, and the data for learning is characterized in that it uses first setting values of individual items for generating each avatar set for a first set consisting of a plurality of avatars already used in the first virtual space, and second setting values of individual items for generating each avatar set for a second set consisting of a plurality of avatars already used in the second virtual space. [Effects of the Invention]
[0015] According to the first feature, by applying a transformation relationship to the information of the first avatar already set up by the user in the first virtual space and outputting information of the second avatar to be used in the second virtual space, the effort required to set up the second avatar again is reduced, and the first objective can be achieved. According to the second feature, the second objective can be achieved by learning the transformation relationship using training data. [Brief explanation of the drawing]
[0016] [Figure 1] This is a functional block diagram of an avatar generation device according to one embodiment. [Figure 2] This diagram schematically illustrates the overall processing performed by the avatar generation device. [Figure 3] This is a diagram illustrating the configuration of an avatar generation system according to one embodiment. [Figure 4] This is a flowchart showing the operation of an avatar generation device according to one embodiment. [Figure 5] This figure shows an example of the first embodiment regarding the learning of the transformation relation f. [Figure 6] This diagram shows the configuration for learning the inverse function R-1 used in the feature setting item conversion unit in the first embodiment (and the second embodiment). [Figure 7] This figure shows an example illustrating the second embodiment regarding the learning of the transformation relation f. [Figure 8] This is a diagram showing the configuration of the avatar generation device according to the second embodiment. [Figure 9] This diagram shows the hardware configuration of a typical computer. [Modes for carrying out the invention]
[0017] Figure 1 is a functional block diagram of an avatar generation device 10 according to one embodiment. As shown in the figure, the avatar generation device 10 comprises a first avatar information input unit 11, a first avatar feature extraction unit 12, a second avatar setting item input unit 21, a feature setting item conversion learning unit 22, a virtual space information input unit 31, a learning unit 30 comprising a virtual space feature extraction unit 32 and a conversion relationship calculation unit 33, a second avatar feature calculation unit 41, a feature setting item conversion unit 42, and a second avatar presentation unit 43.
[0018] Figure 2 is a schematic diagram illustrating the overall processing performed by the avatar generation device 10. If a user U has already set up a first avatar AV1 for use in the first virtual space VS1, the avatar generation device 10 can automatically generate a second avatar AV2 for use in the second virtual space VS2. This generation process uses information about the first avatar AV1, information about the first virtual space VS1, and information about the second virtual space VS2 as inputs to automatically generate the second avatar AV2. For example, if the first virtual space VS1 is composed of a realistic atmosphere and worldview, the first avatar AV1 is prepared to match that atmosphere and worldview. On the other hand, if the second virtual space VS2 is composed of a stylized atmosphere and worldview, it is desirable to generate the second avatar AV2 to match that atmosphere and worldview, and the avatar generation device 10 can perform such generation automatically. Furthermore, the first avatar AV1 is pre-prepared by user U with its appearance reflecting user U's desired settings, and the avatar generation device 10 can automatically generate the second avatar AV2 with an appearance that inherits at least some of these settings.
[0019] Figure 3 is a diagram of the configuration of an avatar generation system 100 according to one embodiment, and the avatar generation device 10 can be realized as a partial configuration of such an avatar generation system 100. The avatar generation system 100 comprises a terminal device 200 used by user U, a first server 201 that operates the first virtual space VS1 and provides its services to participating users, a second server 202 that operates the second virtual space VS2 and provides its services to participating users, and a third server 203 that handles any part or all of the functions of the avatar generation device 10, and each of these configurations 200 to 203 is configured to communicate with each other via a network NW such as the Internet. Each of the configurations 200 to 203 may consist of one or more computer devices that can communicate with each other via a network NW, and for example, the terminal device 200 may be configured as all or part of a smartphone, personal computer, head-mounted display, etc., so that user U can participate in the first virtual space VS1 and the second virtual space VS2 by using the terminal device 200.
[0020] Each functional unit of the avatar generation device 10 can be implemented in any configuration, provided that it is included in at least one of the configurations 200 to 203. For example, among the functional units of the avatar generation device 10, the first avatar information input unit 11 and the second avatar presentation unit 43 may be provided in the terminal device 200, the configuration of the virtual space information input unit 31 that receives information about the first virtual space VS1 may be provided in the first server 201, the configuration of the virtual space information input unit 31 that receives information about the second virtual space VS2 and the second avatar setting item input unit 21 may be provided in the second server 202, and the other functional units 11, 12, 22, 31, 32, 33, 41, and 42 may be provided in the third server 203, but any other configuration is also possible. In addition, although the first server 201, the second server 202, and the third server 203 are depicted as separate entities, it is also possible for any two or more of these three to be implemented as the same server device. For example, the functions of the third server 203 can be implemented within the first server 201 and / or the second server 202, thus enabling a configuration in which the third server 203 does not exist. Alternatively, for example, the first server 201 and the second server 202 may be implemented as the same server device.
[0021] Figure 4 is a flowchart of the operation of an avatar generation device 10 according to one embodiment, and the overall outline of the processing content of the avatar generation device 10 will be explained by referring to each step S1 to S4 in Figure 4. As a premise before the start of the flow in Figure 4, the first virtual space VS1 is operated by the first server 201, and the second virtual space VS2 is operated by the second server 202, and both virtual spaces are pre-configured to accommodate a large number of users. In step S1, the information of the first virtual space VS1 and the information of the second virtual space VS2 is read to learn the correspondence relationship for automatically generating the second avatar AV2 from the first avatar AV1. In step S2, the information of the first avatar is received from user U. In step S3, the input of the second avatar setting items is received. In step S4, the second avatar AV2 for user U is automatically generated using the information obtained in steps S1, S2, and S3, and this second avatar AV2 is presented to user U, and the flow in Figure 4 ends.
[0022] Each step in Figure 4 can be specifically implemented as follows: step S1 is implemented by the processing of the learning unit 30, step S2 by the processing of the first avatar information input unit 11, step S3 by the processing of the second avatar setting item input unit 21 and the feature quantity setting item conversion learning unit 22, and step S4 by the processing of the other functional units 12, 41, 42, and 43. Note that steps S1, S2, and S3 are not limited to this order and may be performed in any order or in parallel. Details of the processing of each of these functional units will be explained below.
[0023] <Step S1...Functional parts 31, 32, 33 of the learning unit 30> The virtual space information input unit 31 acquires information about the virtual spaces by receiving information input from the operators of the first virtual space VS1 and the second virtual space VS2, and outputs it to the virtual space feature extraction unit 32. The virtual space information input unit 31 can acquire information about the virtual spaces, such as background video and avatar activity video within the virtual space, avatar settings and setting values, and all or part of the activity log.
[0024] Regarding these videos and activity logs, when many users are participating in each virtual space, the virtual space screen provided to each user (a first-person view where each user's avatar is not visible as themselves, or a third-person view where each user's avatar is visible, etc.) and its activity log (an activity log recording the movement and other behavior of each user's avatar within the virtual space) can be used, which have been saved on the first server 201 and the second server 202, etc. (The virtual space screen actually provided to each user may be saved as a video, or a video reconstructed from the activity log may be used.) Furthermore, the avatar settings and settings (the types of possible settings for each setting and the range of possible settings, etc.) can be those that have been predefined and saved on the first server 201 and the second server 202, etc. for the provision of services by the first virtual space VS1 and the second virtual space VS2.
[0025] Furthermore, if the operators of the first virtual space VS1 and the second virtual space VS2 are able to provide such information, information about the virtual space may be acquired in addition to / instead of the background video of the virtual space, such as 3D model information of the virtual space (polygon shape and surface texture information) or light source information for the 3D model.
[0026] Furthermore, the virtual space information input unit 31 may accept all or part of the following types of information in addition to the above-mentioned video and other images as input. ● Maximum number of polygons / Minimum number of polygons to represent in that virtual space ● Types and performance of usable rendering engines ●Types / number of usable colors (based on the assumption that each virtual space will exhibit unique characteristics depending on the colors used) ● Configuration of the avatar's movable parts (types, arrangement, and range of motion of movable parts such as the head and arms) ●Character proportions, facial features, etc. (may differ from those of a human) ●The movement itself (what kind of movement is intended, what kind of movement is pre-set) ●Audio (character voices, BGM (background music), SE (sound effects)) ● The types of items available for use with avatars according to the official settings (settings made by the operators) (what kinds of items are available, such as clothing and accessories, and which are most common), and the extent to which those items can be used to change the avatar's settings (degree of freedom). ●Movement (e.g., walking speed, jumping height) ● Is the avatar's movement based on real-world synchronization (linked to each user's real-world location and behavior), or is it a control-based system?
[0027] In the above context, "character" includes not only avatars that users can use as their alter egos, but also characters provided as content within each virtual space.
[0028] The virtual space feature extraction unit 32 analyzes the information of the first virtual space VS1 and the second virtual space VS2 obtained by the virtual space information input unit 31 for each virtual space, extracting features such as avatars, images within the virtual space, and settings, and outputs them to the conversion relationship calculation unit 33. For the purpose of explanation, the feature extracted by the virtual space feature extraction unit 32 from the information of the first virtual space VS1 will be called feature A, and the feature extracted from the information of the second virtual space VS2 will be called feature B.
[0029] The transformation relationship calculation unit 33 takes the features A and B obtained by the virtual space feature extraction unit 32 as input, derives a transformation relationship as a calculation formula to transform from the feature space of the first virtual space VS1, which is the source of the settings, to the feature space of the second virtual space VS2, which is the destination of the settings, and outputs it to the second avatar feature extraction unit 41. This transformation relationship is denoted as transformation relationship f(f: A → B). Here, the feature space consists of the range and frequency of items that can be set in each virtual space (for example, hue, brightness, saturation used), as well as combinations with other setting items.
[0030] The conversion relationship f, etc., will be described later with reference to explanatory examples in Figures 5 and 7, etc., regarding the first and second embodiments.
[0031] <Step S2...First Avatar Information Input Section 11> The first avatar information input unit 11 receives input from user U regarding the information of the first avatar that user U is currently using in the first virtual space VS1, and outputs this information of the first avatar to the first avatar feature extraction unit 12.
[0032] Information about the first avatar can be obtained as video footage of the first avatar while it is active in the first virtual space VS1 (for example, video footage of the first avatar being filmed from a third-person perspective), avatar settings (settings for each part that makes up the avatar), and all or part of the activity log. As mentioned above, the information input to the first virtual space VS1 in the virtual space information input unit 31 can include video footage and activity logs of all users participating in the first virtual space VS1 as avatars. Therefore, by selecting and obtaining the information of the first avatar used by user U from this information of all users, it is possible to avoid having user U directly input the information of the first avatar.
[0033] <Step S3... Second avatar setting item input unit 21 and feature quantity setting item conversion learning unit 22> The second avatar setting item input unit 21 acquires training data containing information on items that can be set for avatars used in the second virtual space VS2 by receiving input from the operator of the second virtual space VS2, or by referring to setting information stored in the second server 202, and outputs this setting item information to the feature setting item conversion training unit 22.
[0034] The feature setting item conversion learning unit 22 converts each item that can be set for the avatar in the second virtual space VS2 and its value into a feature, and the inverse function R of the relationship (let's call it function R) -1 This is calculated by learning, and its inverse function R -1 The output is sent to the feature setting item conversion unit 42.
[0035] For example, if the eye part XX is one of the configurable items for identifying an individual avatar, the function R converts it into a feature quantity such as (color, size, shape) = (x, y, z). For example, if there are three possible values for the eye part XX, XX=XX1, XX2, and XX3, which can be selected from a menu in the second virtual space VS, then the function R gives the feature quantities (color, size, shape) = (x1, y1, z1), (x2, y2, z2), and (x3, y3, z3) for each of the three values of the XX=XX1, XX2, and XX3 configurable items. The learning unit 22 for feature quantity configurable item conversion uses the inverse function R of this function R. -1 This is learned using the method described later and output to the feature setting item conversion unit 42.
[0036] <Step S4...Other functional parts 12, 41, 42, 43> The first avatar feature extraction unit 12 analyzes the information of the first avatar obtained from the first avatar information input unit 11 and outputs its features (let's call them feature a) to the second avatar feature calculation unit 41.
[0037] The second avatar feature calculation unit 41 takes the feature quantity a obtained by the first avatar feature extraction unit 12 as input to the transformation relationship f obtained by the transformation relationship calculation unit 33, obtains the feature quantity f(a) of the second avatar as its output f(a), and outputs this feature quantity f(a) of the second avatar to the feature quantity setting item transformation unit 42.
[0038] The feature setting item conversion unit 42 takes the feature f(a) of the second avatar obtained by the second avatar feature calculation unit 41 and the inverse function R obtained by the feature setting item conversion learning unit 22. -1 By applying this, the setting value R of the second avatar -1 { f(a)} is obtained and output to the second avatar presentation unit 43.
[0039] The second avatar display unit 43 displays the setting value R of the second avatar obtained by the feature setting item conversion unit 42. -1 The result of rendering the second avatar using { f(a)} is presented to user U by displaying it on the screen.
[0040] Regarding the learning unit 22 for feature setting item conversion, as shown in the example above, for example, for the eye part XX, if XX=XX1 out of 3 possible settings, the feature (x1, y1, z1) is not defined as a setting in the second virtual space VS2, but is merely an arbitrary value, so it is not possible to draw the eye part XX using the feature (x1, y1, z1) in its original form. In contrast to this, the inverse function R -1 By applying R -1 By setting (x1, y1, z1) = XX1, we obtain the information that "eye part XX = XX1". Using this information, which is defined as a setting in the second virtual space VS2, it becomes possible to render the second avatar. By rendering each part other than the eye part in the same way, the second avatar can be rendered as a whole.
[0041] The rendering of the second avatar to be presented to user U may be done as a single still image viewed from a specific camera position, such as the front, or, in the case of a 3D avatar, it may be rendered as a video showing the avatar rotating while standing in a specific position, in a format that allows the entire 3D avatar to be viewed from all sides in a 360° range. Similarly, as a format that allows viewing from all sides in a 360° range, a so-called 360° rotating image may be used, in which user U can interactively rotate the image by dragging the mouse cursor on the image using a browser or the like to see the rotated state. In addition, any existing rendering method that can be used to represent the avatar may be used.
[0042] The details of the processing content of each functional unit in Figure 1 have been explained above. Figure 5 is a diagram illustrating an example of the first embodiment concerning the learning of the conversion relationship f, and shows schematic examples of the processing content of each unit 31, 32, and 33 of the learning unit 30. As shown in Figure 5, the first embodiment can be realized as follows.
[0043] In the virtual space information input unit 31, as information on the first virtual space VS1, information of a large number of registered user avatars A i (i = 1, 2, …, N) (including at least information on the appearance of each avatar in the form of video, etc.) is read. Similarly, as information on the second virtual space VS2, information of a large number of registered user avatars B j (j = 1, 2, …, M) (including at least information on the appearance of each avatar in the form of video, etc.) is read. If available, the information of each avatar A i , B j may be acquired as information on a three-dimensional model (for example, a three-dimensional model determined as a plurality of polygons and textures attached to each polygon) that directly represents its appearance.
[0044] In the virtual space feature amount extraction unit 32, as shown in the explanation column CL11, from the information (image-like information of the appearance) of each avatar A i , B j , for the predefined main setting items (setting items for avatar parts), each part is cut out by applying object recognition of existing methods. Further, as shown in the explanation column CL12, the parts as the cut-out setting items are quantified.
[0045] The setting items for each cut-out part may be hierarchically defined as, for example, large items = human, object, …, medium items (for large item = human) = face, hair, …, and small items (for medium item = face) = eyes, nose, mouth. When cutting out these parts by object recognition, hierarchical recognition may also be applied. In this way, from the image-like appearance information of each avatar A i , B j , a set of K partial region images {r i}, {r j} can cut out each part as follows. A i → {r i} = {r 1i , r 2i , …, r Ki}, Bj→{r j}={r 1j , r 2j , …,r Kj}
[0046] Furthermore, if the avatar is a 3D avatar, each part will have a 3D shape. However, by applying object recognition to an image of each part viewed from the front (or an image of each polygon element viewed from the front if the 3D model is given as polygons), the set of sub-region images described above can be used as a 2D image region {r i},{r j You can cut out the}.
[0047] Furthermore, as described above, the numerical values of the setting items (each of the K parts) extracted as K sub-region images can be quantified by applying dimensionality reduction or the like using any existing method to each of the K sub-region images. As shown in the figure, for example, with respect to one part, the eye, its size, color, eyebrows, eyelashes, and outer corner of the eye are all quantified, and the quantified results can be obtained in the form of a vector or the like. That is, the sub-image set {r i},{r j From}, information {V} is obtained by listing the features of K vectors, etc. i},{V j} can be obtained as follows. {r i}→{V i}={V 1i , V 2i , …,V Ki}, {r j}→{V j}={V 1j , V 2j , …,V Kj}
[0048] For example, regarding the partial image set {r1} of avatar A1 with i=1, the first partial image r 11When the feature is an eye, its various characteristics such as size, color, eyebrows, eyelashes, and outer corner of the eye are ultimately quantified as a feature in vector form, V. 11 We can obtain (0.2, 0.1, 1.0, 0.5, 0.2).
[0049] In the conversion relation calculation unit 33, the avatar set A = {A} is calculated as described above. i |i=1,2,…,N} and B={A j The feature vectors {V} obtained for each element of |j=1,2,…,M} i} and {V j Using}, the transformation relation f: A → B is calculated. Specifically, for example, as shown in the explanation section CL13, each avatar A i Feature {V i}={V 1i , V 2i , …,V Ki} and each avatar B j Feature {V j}={V 1j , V 2j , …,V Kj The expression} is an enumerated list of K vectors. Further dimensionality reduction is applied to these vectors to obtain a one-dimensional value as follows: (For example, principal component analysis is applied to avatar sets A and B to find their first principal components, thereby obtaining a one-dimensional value.) {V 1i , V 2i , …,V Ki}→{v 1i , v 2i , …,v Ki}, {V 1j , V 2j , …,V Kj}→{v 1j , v 2j , …,v Kj}
[0050] Thus, Avatar A i ,B j The features are expressed in a K-dimensional vector {v 1i , v 2i , …,v Ki,{v 1j , v 2j , …,v Kj} is obtained, and then it can be histogrammed by assigning predetermined bins to each element of the K-dimensional vectors within avatar set A and within avatar set B, respectively. Here, let the histograms created for each element of the K-dimensional feature vectors of avatar sets A and B be Hist_A and Hist_B as follows. Hist_A = {hist_A1, hist_A2, …, hist_A K}, Hist_B = {hist_B1, hist_B2, …, hist_B K}
[0051] Avatar A i , B j 's K-dimensional vectors of each element {v 1i , v 2i , …, v Ki}, {v 1j , v 2j , …, v Kj}, for the information on the frequency ranking of which bin each belongs to in descending order of frequency, it is possible to obtain it by collating with the histograms Hist_A and Hist_B. Therefore, the information on the frequency ranking obtained by this collation is defined as rank_A i , rank_B j as follows. rank_A i = {rank 1i , rank 2i , …, rank Ki}, rank_B j = {rank 1j , rank 2j , …, rank Kj}
[0052] As shown in the description column CL13, in the conversion relationship calculation unit 33, as the conversion relationship f: A → B, the avatar B i corresponding to avatar A j can be calculated such that the above frequency rankings are the same. f(A i )=B j rank_A i =rank_B j
[0053] Regarding the conversion relationship f: A → B, a certain avatar A i Avatar B such that the frequency ranking is the same for all of the following j If there are two or more, one avatar Bj may be randomly selected from among them, or multiple avatars B that have the same frequency ranking may be selected. j The average (its feature {V j Alternatively, you could calculate the average of}.
[0054] In the second avatar feature calculation unit 41 according to the first embodiment, by using the transformation relationship f:A→B as explained in Figure 5 above, the first avatar A in the feature space i Feature {V i From}, the corresponding feature of the second avatar Bj {V j The feature vector f(a) of the second avatar in Figure 1 can be calculated.
[0055] Furthermore, in the first embodiment, the learning configuration shown in Figure 6 is used by the feature setting item conversion unit 42, and the inverse function R -1 This can be learned in the feature setting item conversion learning unit 22 (referred to as learning configuration 22 in relation to Figure 6). The learning configuration 22 includes an encoder 302 and a feature setting item conversion unit 42 (under learning), and the encoder 302 reads information from a large number of second avatars for learning, and each feature f(a) (the feature {V} of the second avatar Bj explained in Figure 5) j The output is}), and this feature f(a) is further read by the feature setting item conversion unit 42 (during training), and the inverse function R identified by the parameters during training is obtained. -1This is applied to output the setting value of each setting item in the second virtual space (for example, for the eye part XX mentioned above, which setting value is XX=XX1, XX2, or XX3). By comparing this output with the correct values given as training data, the parameters of the feature setting item conversion unit 42, which is composed of a deep learning network, are learned, and the inverse function R corresponding to the parameters at the time of training completion is used. -1 This can be obtained as a learning result. For example, if the correct answer for a certain avatar is eye part XX=XX1, but the system outputs XX=XX2, then by assigning this incorrect answer as a cost, the parameters of the feature setting item conversion unit 42 can be learned to output the correct answer using any existing method such as backpropagation.
[0056] Figure 7 is a diagram illustrating an example of a second embodiment concerning the learning of the transformation relationship f. In the first embodiment described above, the correspondence of feature quantities was obtained using frequency ranking information, so each avatar A in the first virtual space VS1 i And each avatar B in the second virtual space VS2 j In the first embodiment, information on which user is actually using the system was not required, whereas in the second embodiment, it is assumed that the user ID information is known and this information is used. That is, the user is identified by index i and is referred to as "user i". In the second embodiment, user i is avatar A in the first virtual space VS1 i This user i is using Avatar B in the second virtual space VS2. i Use the information that they are using it.
[0057] Furthermore, as shown in Figure 8, the second embodiment can be configured such that, in place of the first avatar feature extraction unit 12 and the second avatar feature calculation unit 41 of the avatar generation device 10 of the first embodiment shown in Figure 1, a second avatar feature conversion unit 124 is provided which performs these processes (indirectly) as internal processing in a single unit.
[0058] As shown in columns CL21 to CL23 of Figure 7, the virtual space information input unit 31 receives information from the same user i in the first virtual space VS1, and avatar A i It utilizes and, in the second virtual space VS2, avatar B i Based on the information used, this avatar pair {A i ,B i After generating input information pairs from each setting item of}, the virtual space feature extraction unit 32 uses the first encoder 301 to extract avatar A i Feature vector {VA i} is generated, and Avatar B is generated by the second encoder 302 i Feature vector {VB i Generates}.
[0059] In the first embodiment described with reference to Figure 4, the input information pair is "feature quantity {V i} and {V j Alternatively, the settings may be converted into vectors for each setting item using the same method as described above and input to the first encoder 301 and the second encoder 302, or they may be vectorized using any other existing preprocessing method.
[0060] Then, as shown in column CL25, the conversion relationship calculation unit 33 calculates the avatar A output by the first encoder 301 for the same user i. i Feature vector {VA i} and the feature vector of avatar Bi output by the second encoder 302 {VB i} and are identical ({VA i}={VB iThe parameters of the first encoder 301 and the second encoder 302 are learned so that {VAi} = {VBi}, and the first encoder 301 composed of the learned parameters can be configured as the learned second avatar feature conversion unit 124. That is, by configuring the second avatar feature conversion unit 124 as the first encoder 301 learned to be identical ({VAi} = {VBi}), the first encoder 301 can read the information of the first avatar as input and output information corresponding to the feature quantity f(a) of the second avatar, which is both the feature vector of the first avatar and the feature quantity of the second avatar, and is identical ({VAi} = {VBi}).
[0061] Furthermore, both the first encoder 301 and the second encoder 302 are prepared as deep learning networks of a predetermined structure, and the same ({VA i}={VB i The parameters of the network (such as the weights for the convolution process) should be learned using existing methods so that the result is}).
[0062] The inverse function R in the feature setting item conversion unit 42 in the second embodiment -1 Learning can be performed in the same manner as the configuration shown in Figure 6 in the first embodiment, as can be understood from the fact that the (second) encoder 302 is given a common reference number.
[0063] As described above, according to each embodiment of the present invention, users can easily move between multiple virtual spaces and enjoy services. That is, if an avatar has already been set up in at least one first virtual space VS1, that avatar can be automatically set up in any one or more second virtual spaces VS2 that are different from the first one. On the other hand, service providers operating virtual spaces can reduce user drop-off due to users disliking the hassle of setting up avatars in the initial setup step.
[0064] The following sections will explain various supplementary examples, alternative examples, and additional examples.
[0065] (1) According to embodiments of the present invention, by reducing the effort required to set up avatars, it becomes possible to encourage participation in services that provide immersive remote communication using avatars. This makes it possible to conduct remote meetings etc. without necessarily requiring actual travel to remote locations, and by saving energy resources required for user travel, carbon dioxide emissions can be reduced, thus contributing to Goal 13 of the United Nations Sustainable Development Goals (SDGs), "Take urgent action to combat climate change and its impacts."
[0066] (2) A configuration in which only the learning unit 30 of the avatar generation device 10 is extracted may be provided as a learning device for learning the conversion relationship f: A → B.
[0067] (3) With respect to the second avatar, the first avatar information input unit 11 may reflect the user U's requests through all or part of the UI (user interface) listed below, in order to accurately acquire and generate the user U's requests.
[0068] ● UI that allows you to select the parts / elements you want to keep when converting to an avatar. ● (Related to the above) A UI that allows you to select parts / elements that do not need to be retained during avatar conversion (parts that do not need to be converted). To keep / not to keep these (to carry over from the first avatar to the second avatar / not to carry over) The settings for the first and second avatars should be configured so that the settings are the same (however, the range of possible values for common settings may differ), and the user can select a setting from a predetermined menu. For settings that are to be retained, the features extracted from the first avatar are input into the transformation relationship, and the setting value can be obtained from the features of the second avatar. On the other hand, for settings that are not to be retained, the features extracted from the first avatar will not be used as input to the transformation relationship, so a random number or the like can be used as a substitute value.
[0069] ● A UI that notifies the user in advance whether the amount of user input data (images, videos, settings, activity logs, conversation logs, etc.) is sufficient to display an avatar. ● Especially when the user input data is an image or video, the UI detects whether there are any missing data acquisitions for the entire avatar and 360 degrees, and notifies the user in advance. Whether the data is sufficient or whether any data has been missed can be automatically detected by applying rule-based processing to the input data. For example, activity logs and conversation logs can be detected based on whether the data volume or data storage period exceeds a certain amount, and similarly for video and other data, detection can be done based on whether the data has been stored for a certain period of time or longer.
[0070] ● In addition to video information that the user consciously captures, the system allows input of settings, activity logs, and conversation logs, and if the user has not selected these, it prompts the user to input them (permission to use each of these data). ● A UI that remembers information (selected items) previously entered by the user and prompts the user to choose whether to enter the same information again or to add or delete information before entering it. In other words, the settings used when user U generates a second avatar in the second virtual space using the first avatar in the first virtual space VS as input may be stored, and these settings may be made available when generating a third avatar to be used in another third virtual space.
[0071] (4) In the second embodiment of Figure 6, a common ID is accessible to each user in the first virtual space VS1 and each user in the second virtual space VS2, and input information pairs are obtained using this ID. However, if such an ID is not accessible, input information pairs may be defined as users who are presumed to be the same or have common attributes, and the following methods may be used.
[0072] ● Users are categorized (e.g., using the BIG5 model) and paired based on their attributes, input commands, and conversation content (for example, the number and duration of input commands used to move the avatar, and the frequency of certain words). ● Pair the avatar settings of users who have moved directly from the first virtual space VS1 to the second virtual space VS2 (or vice versa), or who have recently been active in both virtual spaces (if each avatar can be identified as belonging to the same user by some key). In other words, if the activity period of each avatar in the first virtual space VS1 and the activity period of each avatar in the second virtual space VS2 are in a relationship where one succeeds the other through direct movement (for example, if an avatar in the second virtual space VS2 logs in immediately after an avatar logs out of the first virtual space VS1 (determined by the fact that no more than a threshold time has elapsed)), or if there is a login history of an avatar in both virtual spaces VS1 and VS2 within a certain period, it may be assumed that the avatars belong to the same user and be paired.
[0073] (5) In the above description, the avatar generation device 10 takes one first avatar as input for use by a user U in the first virtual space VS1, and outputs only one second avatar corresponding to this for use by the same user U in the second virtual space VS2. However, measures such as those listed below may be taken to encourage user U to make choices that are highly satisfactory to them.
[0074] ● A UI that presents multiple options is provided in the second avatar display unit 43. To achieve this, the internal processing involves applying a certain random number when creating candidates and generating variations with a constant difference (for example, outputting multiple models with the same part shape but different colors). In other words, for the transformation relationship f: A → B used in the second avatar feature calculation unit 41, instead of using only one f obtained from the transformation relationship calculation unit 33, additional f1, f2, ... etc., which are variations within a certain range (such as variations in the transformation parameters that determine the transformation relationship f), can be used to generate a second avatar with multiple variations.
[0075] ● When presenting multiple options (suggestions), the second avatar presentation unit 43 is provided with a UI to assist the user in making a selection. For example, the following UI may be used. ◆ UI that checks if similar avatars include paid items and adopts them if OK (for example, if the avatar's clothing includes both paid and free items) ◆ UI that displays paid and free options side by side ◆ A UI that allows you to rearrange some of the suggestions from multiple options. ◆ Combinations of the face from Proposal 1 and the body from Proposal 2, etc. ◆ A UI that shuffles the settings of multiple proposals and resubmits them. ◆ A UI where the user selects the closest option from the suggestions, and then the variations are recalculated and suggested again.
[0076] (6) Figure 9 is a diagram showing an example of the hardware configuration of a typical computer device 70. The avatar generation device 10 can be realized as one or more computer devices 70 having such a configuration. When the avatar generation device 10 is realized with two or more computer devices 70, information necessary for processing may be sent and received via a network. The computer device 70 includes a CPU (Central Processing Unit) 71 that executes predetermined instructions, a GPU (Graphics Processing Unit) 72 as a dedicated processor that executes some or all of the execution instructions of the CPU 71 on behalf of or in cooperation with the CPU 71, RAM 73 as main memory that provides a work area to the CPU 71 (and GPU 72), ROM 74 as auxiliary memory, a communication interface 75, a display 76, an input interface 77 that accepts user input via a mouse, keyboard, touch panel, etc., and a bus BS for sending and receiving data between these.
[0077] Each functional unit of the avatar generation device 10 can be implemented by a CPU 71 and / or GPU 72 that read and execute a predetermined program corresponding to the function of each unit from ROM 74. Both the CPU 71 and GPU 72 are types of arithmetic units (processors). When display-related processing is performed, the display 76 also operates in conjunction, and when communication-related processing related to data transmission and reception is performed, the communication interface 75 also operates in conjunction. [Explanation of Symbols]
[0078] 10...Avatar generation device, 11...First avatar information input unit, 12...First avatar feature extraction unit, 21...Second avatar setting item input unit, 22...Setting item feature conversion unit, 30...Learning unit, 31...Virtual space information input unit, 32...Virtual space feature extraction unit, 33...Conversion relationship calculation unit, 41...Second avatar feature extraction unit, 42...Feature setting item conversion unit, 43...Second avatar presentation unit
Claims
1. An avatar generation device characterized by taking information of a first avatar used by a user in a first virtual space as input, and applying a pre-learned conversion relationship using the information of the avatar used in the first virtual space and the information of the avatar used in the second virtual space to the input, thereby outputting information of a second avatar used by the user in the second virtual space, The aforementioned conversion relationship is pre-trained as a relationship in which the feature quantities of each avatar extracted based on at least the setting values of the individual items for generating each avatar used in the first virtual space are taken as input, and the feature quantities of each avatar extracted based on at least the setting values of the individual items for generating each avatar used in the second virtual space are output. Between the first feature vector of each avatar extracted from each avatar drawn based at least on the first setting value of the individual item for generating each avatar set for the first set of multiple avatars already used in the first virtual space, and the second feature vector of each avatar extracted from each avatar drawn based at least on the second setting value of the individual item for generating each avatar set for the second set of multiple avatars already used in the second virtual space, The second feature corresponding to the first feature is determined such that the frequency rank in the first histogram obtained from the distribution of the first feature within the first set is the same as the frequency rank in the second histogram obtained from the distribution of the second feature within the second set. An avatar generation device characterized in that the aforementioned conversion relationship is pre-learned.
2. An avatar generation device characterized by taking information of a first avatar used by a user in a first virtual space as input, and applying a pre-learned conversion relationship using the information of the avatar used in the first virtual space and the information of the avatar used in the second virtual space to the input, thereby outputting information of a second avatar used by the user in the second virtual space, Between a first feature quantity extracted based at least on a first setting value for generating each avatar set for a first set of multiple avatars already used in the first virtual space, and a second feature quantity extracted based at least on a second setting value for generating each avatar set for a second set of multiple avatars already used in the second virtual space, A correspondence between each avatar in the first set and each avatar in the second set is predetermined, where the user using the avatar or the avatar having the same avatar attributes is given in advance. With respect to the same avatar, the first encoder and the second encoder are pre-trained so that the first output of the first encoder, which encodes the first setting value and uses it as the first feature, and the second output of the second encoder, which encodes the second information and uses it as the second feature, are the same. The avatar generation device is characterized in that the conversion relationship is configured to include processing by the first encoder.
3. The correspondence between each avatar in the first set and each avatar in the second set, where the user using the avatar is the same, is given in advance. The avatar generation device according to claim 2, characterized in that, with respect to the activity period of each avatar in the first virtual space and the activity period of each avatar in the second virtual space, a correspondence between avatars is given such that it is presumed that the user is the same when one succeeds the other, or when there is an activity period in both the first virtual space and the second virtual space within a certain period.
4. An avatar generation device characterized by taking information of a first avatar used by a user in a first virtual space as input, and applying a pre-learned conversion relationship using the information of the avatar used in the first virtual space and the information of the avatar used in the second virtual space to the input, thereby outputting information of a second avatar used by the user in the second virtual space, The aforementioned conversion relationship is pre-trained as a relationship in which the feature quantities of each avatar extracted based on at least the setting values of the individual items for generating each avatar used in the first virtual space are taken as input, and the feature quantities of each avatar extracted based on at least the setting values of the individual items for generating each avatar used in the second virtual space are output. An avatar generation device characterized by applying the aforementioned conversion relationship to each of a plurality of conversion relationships with variations within a certain range, thereby outputting information on a second avatar used by the user in the second virtual space across multiple options, and enabling the user to select from among the plurality of candidates.
5. An avatar generation device characterized by taking information of a first avatar used by a user in a first virtual space as input, and applying a pre-learned conversion relationship using the information of the avatar used in the first virtual space and the information of the avatar used in the second virtual space to the input, thereby outputting information of a second avatar used by the user in the second virtual space, The aforementioned conversion relationship is pre-trained as a relationship in which the feature quantities of each avatar extracted based at least on the setting values of individual items for generating each avatar used in the first virtual space are input, and the feature quantities corresponding to the setting values of individual parts for generating each avatar used in the second virtual space are output. The avatar generation device according to claim 1, characterized in that, in order to allow the user to input only a portion, rather than all, of the information of the first avatar used in the first virtual space, the device accepts from the user that the user will select only a portion of the individual items for generating each avatar used in the first virtual space.
6. The individual items for generating each avatar used in the first virtual space and the individual items for generating each avatar used in the second virtual space are the same. The avatar generation device according to claim 5, characterized in that, by receiving a designation from the user of which individual items of the first avatar to be carried over to the second avatar, the conversion relationship is applied to feature quantities corresponding to the values set in the first avatar for the setting values of the individual items to be carried over, and to feature quantities corresponding to random numbers for the setting values of the individual items that are not carried over, thereby outputting information of the second avatar that the user will use in the second virtual space.
7. A learning device for learning a transformation relationship, used by an avatar generation device that takes information of a first avatar used by a user in a first virtual space as input, and outputs information of a second avatar used by the user in a second virtual space by applying a pre-learned transformation relationship using the information of the avatar used in the first virtual space and the information of the avatar used in a second virtual space to the input, The aforementioned conversion relationship is learned as a relationship in which the feature quantities of each avatar extracted based at least on the setting values of individual items for generating each avatar used in the first virtual space are input, and the feature quantities corresponding to the setting values of individual parts for generating each avatar used in the second virtual space are output. As data for the learning process, A learning device characterized by using: a first feature quantity for each avatar extracted based at least on a first setting value of an individual item for generating each avatar set for a first set of multiple avatars already used in the first virtual space; and a second feature quantity for each avatar extracted based at least on a second setting value of an individual item for generating each avatar set for a second set of multiple avatars already used in the second virtual space.
8. A program characterized by causing a computer to function as an avatar generation device according to any one of claims 1 to 6 or a learning device according to claim 7.