Information processing system, information processing method, and program

The system enables intuitive motion data conversion between avatars with different bone structures using high-level metadata in colloquial terms, addressing the challenge of transferring motion data across platforms and formats.

WO2025204168A1PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003926
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-02-06
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

The challenge of transferring motion data between avatars with different bone structures across various platforms and formats is hindered by differences in bone settings, requiring specialized knowledge and labor-intensive manual conversion processes.

Method used

An information processing system and method that utilizes high-level metadata to convert motion data between avatars with different bone structures by defining control content in colloquial terms, allowing intuitive motion generation without specialized knowledge.

Benefits of technology

Facilitates seamless motion data transfer between avatars with varying bone settings, enhancing convenience and reducing the need for manual conversion expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003926_02102025_PF_FP_ABST
    Figure JP2025003926_02102025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To improve the convenience of generation of a motion to be applied to another avatar. [Solution] An information processing system including a control unit that performs: processing for converting first motion data into second motion data on the basis of first metadata in which control content corresponding to the motion of a first avatar is defined and inputted first motion data of the first avatar; and processing for converting the second motion data into third motion data for a second avatar by using second metadata in which control content corresponding to the motion of the second avatar is defined.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and program

[0001] The present disclosure relates to an information processing system, an information processing method, and a program.

[0002] In recent years, a technology has become known that provides a virtual space in which various objects, such as 3D models or 2D images, are arranged. A user can experience the virtual space using various information processing terminals, such as a smartphone, a head-mounted display (HMD), or a PC.

[0003] Objects placed in a virtual space include avatars that can move within the virtual space as the user's avatar. The user can experience the virtual space by operating the avatar and viewing the virtual space from the avatar's perspective. Objects placed in a virtual space may also include avatars that do not require an operator and are known as non-player characters (NPCs).

[0004] Here, each movement of the avatar is realized by controlling (moving) the bones set for the avatar according to a script (motion data), for example. Such control involves specifying the rotation angle of the bones. When specifying the rotation angle of the bones, rotations that are impossible for a normal human can occur. However, Non-Patent Document 1 below discloses that it is possible to set limits on the rotation angle of the bones.

[0005] "Unity User Manual (2019.4 LTS)", Avatar Muscle & Settings tab, [online], April 2019, Unity Technologies, [Retrieved March 22, 2024], Internet <URL: https: / / docs.unity3d.com / ja / 2019.4 / Manual / MuscleDefinitions.html>

[0006] However, the bone structure of an avatar (the number of bones, their positions, how they are connected, etc.) and the local coordinate axis settings of the bones differ depending on the platform and format, making it difficult to transfer motions generated for one avatar to other avatars on different platforms or in different formats.

[0007] Therefore, the present disclosure proposes an information processing system, an information processing method, and a program that can improve the convenience of generating motions to be applied to other avatars.

[0008] According to the present disclosure, an information processing system is provided that includes a control unit that performs a process of converting first motion data into second motion data based on first metadata that defines control content corresponding to the motion of a first avatar and input first motion data of the first avatar, and a process of converting the second motion data into third motion data for the second avatar using second metadata that defines control content corresponding to the motion of a second avatar.

[0009] Furthermore, according to the present disclosure, an information processing method is provided, which includes a processor converting, based on first metadata defining control content corresponding to the motion of a first avatar and input first motion data of the first avatar, the first motion data into second motion data, and converting, using second metadata defining control content corresponding to the motion of a second avatar, the second motion data into third motion data for the second avatar.

[0010] Furthermore, according to the present disclosure, a program is provided that causes a computer to function as a control unit that performs a process of converting first motion data into second motion data based on first metadata that defines control content corresponding to the motion of a first avatar and input first motion data of the first avatar, and a process of converting the second motion data into third motion data for the second avatar using second metadata that defines control content corresponding to the motion of a second avatar.

[0011] FIG. 1 is a diagram for explaining differences in bone structure settings of two avatars. FIG. 2 is a diagram for explaining knee bending motion control. FIG. 3 is a block diagram showing an example of the configuration of an information processing device 10 according to an embodiment of the present disclosure. FIG. 4 is a diagram showing examples of cases where the pose state indicated in colloquial expression according to this embodiment is 0 and 1. FIG. 5 is a diagram showing an example of a generation interface for high-level metadata according to this embodiment. FIG. 6 is a diagram showing an example of source high-level metadata according to this embodiment. FIG. 7 is a diagram showing an example of destination high-level metadata according to this embodiment. FIG. 8 is a diagram showing an example of high-level motion data according to this embodiment. FIG. 9 is a diagram showing an example of a generation interface for generating destination motion data according to this embodiment. FIG. 10 is a diagram explaining a motion of tilting the head diagonally forward and to the right. A flowchart showing the flow of motion conversion processing by the information processing device 10 of this embodiment. FIG. 11 is a diagram showing an example of the configuration of an information processing system 1 according to an embodiment of the present disclosure.

[0012] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0013] The explanation will be given in the following order: 1. Overview 2. Structure 3. Details of motion conversion 3-1. Generation of motion data for source avatar 3-2. Generation of high-level metadata 3-3. Conversion to high-level motion 3-4. Generation of motion data for destination avatar 4. Movement processing 5. Other 6. Supplementary information

[0014] <1. Overview> As one embodiment of the present disclosure, a mechanism for converting the motion of one avatar into the motion of an avatar for a different platform or format will be described.

[0015] An avatar has a bone structure consisting of many connected bones that correspond to body parts and joint positions, and the avatar's movement (posture) is controlled by moving the bones. The bone structure settings may differ depending on the platform (virtual space execution environment) on which the avatar operates and the avatar format. More specifically, the bone structure settings include the number of bones, the positions of the bones, how the bones are connected, and the settings of the local coordinate axes of the bones.

[0016] Fig. 1 is a diagram illustrating the difference in bone structure settings between the two avatars. As shown in Fig. 1, even for the same humanoid avatar, the local coordinate axes set at the knee joint of avatar 210 shown on the left in Fig. 1 are different from the local coordinate axes set at the knee joint of avatar 220 shown on the right in Fig. 1. Specifically, the +X direction at the knee joint of avatar 210 is opposite to the +X direction at the knee joint of avatar 220, and the +Z direction is also opposite.

[0017] The movement (posture) of the avatar is controlled by specifying the rotation angles of the bones (X-axis rotation angle, Y-axis rotation angle, Z-axis rotation angle). The connection between two connected bones corresponds to a joint, and the angle of the joint changes when at least one of the bones rotates around the connection point.

[0018] For example, if one wishes to bend the knees of the avatar 210, the angle of the knee joint is changed. Specifically, the upper end connector J-t of the Lower Leg, which corresponds to the part of the avatar 210 below the knee, is rotated from 0 degrees to 90 degrees in a positive direction around the Z axis. FIG. 2 is a diagram for explaining knee bending motion control. As shown on the left side of FIG. 2, in the avatar 210, by specifying a Z-axis rotation angle of 90 degrees, the Lower Leg rotates 90 degrees in a positive direction around the Z axis, with the upper end connector J-t as the base point, thereby bending the knees. In this specification, information for controlling the posture and motion of an avatar is referred to as motion data. The motion data may describe posture changes in a time series, or may describe a single pose not in a time series. The motion data may include, for example, joint position information (three-dimensional coordinate position), joint angles (X-axis rotation angle, Y-axis rotation angle, Z-axis rotation angle), bone position information, or bone angles.

[0019] If the motion data for controlling the knee bending movement of avatar 210 (in the example above, this describes changing the Z-axis rotation angle from 0 degrees to 90 degrees) is applied to avatar 220, which has a different bone structure setting, the knee will bend in a physically strange direction, as shown on the right in Figure 2, and the correct movement will not be reproduced. To move the angle of the knee joint of avatar 220 in the same way as avatar 210, it is necessary to change the description so that the Z-axis rotation angle is specified as -90 degrees (changing from 0 degrees to 90 degrees in the negative rotation direction around the Z axis).

[0020] In this way, differences in the direction of rotation around the Z axis (positive rotation direction) make it difficult to share motion data between avatar 210 and avatar 220. Here, different rotation directions have been given as an example of differences in bone structure settings, but this is not limiting, and the number and length of bones, connection method, etc. may also be different.

[0021] When sharing motions between two avatars (characters) with different bone structure settings, manually converting the motion data requires an understanding of both the source and destination avatars (characters), motion editing skills, knowledge of each game engine, and a sense of motion design. Furthermore, the conversion process is labor-intensive. However, avatar creators and those who want to apply motions to other avatars often lack the necessary knowledge, creating a need for a system that allows intuitive conversion without requiring specialized knowledge.

[0022] Therefore, in the present disclosure, by using metadata that defines the control content corresponding to the motion of an avatar, it is possible to improve the convenience of generating (converting) motion to be applied to other avatars.

[0023] More specifically, by using high-level metadata that expresses avatar motions in colloquial terms and redefines motions as higher-level meanings using colloquial expressions rather than as three-dimensional geometric posture changes, it becomes possible to improve the convenience of generating (converting) motions to be applied to avatars with different settings. For example, in the above example, the avatar's movement is defined in colloquial terms such as "change the knee bend from 0 to 1" rather than describing it as "change the knee joint angle (or the "upper end connection of the lower leg") from 0 degrees to 90 degrees." Defining motions in such human-understandable terms allows users without sufficient experience or knowledge in motion creation to intuitively understand the motions when converting motions.

[0024] 2. Configuration Next, an information processing device 10 (an example of an information processing system) that realizes motion conversion according to this embodiment will be described with reference to FIG.

[0025] 3 is a block diagram showing an example of the configuration of the information processing device 10 according to an embodiment of the present disclosure. As shown in FIG. 3, the information processing device 10 includes a communication unit 110, a control unit 120, an operation input unit 130, a display unit 140, and a storage unit 150.

[0026] (Communication Unit 110) The communication unit 110 has a transmission unit that transmits data to an external device and a reception unit that receives data from an external device. The communication unit 110 according to this embodiment may be communicatively connected to an external device or the Internet using, for example, a wired or wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), a mobile communication network (LTE (Long Term Evolution), 4G (fourth generation mobile communication system), 5G (fifth generation mobile communication system)), or the like.

[0027] (Control Unit 120) The control unit 120 functions as an arithmetic processing unit and a control device, and controls the overall operation of the information processing device 10 in accordance with various programs. The control unit 120 is realized by an electronic circuit such as a CPU (Central Processing Unit) or a microprocessor. The control unit 120 may also include a ROM (Read Only Memory) that stores programs to be used, arithmetic parameters, etc., and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.

[0028] The control unit 120 can also function as a motion generation unit 121, a high-order metadata generation unit 122, a high-order motion conversion unit 123, and a motion conversion unit 124.

[0029] The motion generation unit 121 has a function of generating motion data of an avatar. The motion generation unit 121 generates motion of the avatar in accordance with user operations and stores the motion data in the motion data DB 151. Note that motion generation may be performed outside the information processing device 10. The motion generation unit 121 may obtain motion data generated by, for example, an external 3D model creation tool and store the motion data in the motion data DB 151.

[0030] The high-level metadata generation unit 122 has a function of generating high-level metadata (an example of metadata) that colloquially expresses and defines avatar motions at a high level. In this embodiment, the high-level metadata generation unit 122 performs a process of converting motion data generated for a first avatar into motion data that can be used for a second avatar with a different bone structure setting (also referred to as a different platform or format). Here, the first avatar is referred to as the source avatar, and the second avatar is referred to as the destination avatar. In the motion data conversion according to this embodiment, high-level metadata that defines motion at a high level is used. The high-level metadata generation unit 122 generates high-level metadata used in such motion data conversion.

[0031] Higher-order metadata is prepared for each avatar (character) with a different bone structure setting. The high-order metadata generation unit 122 generates first higher-order metadata (an example of first metadata; hereinafter referred to as "source high-order metadata") that defines, at a high level, motion data for moving a source avatar (first avatar), and stores the first higher-order metadata in the source high-order metadata DB 152. The high-order metadata generation unit 122 also generates second higher-order metadata (an example of second metadata; hereinafter referred to as "destination high-order metadata") that defines, at a high level, motion data for moving a destination avatar (second avatar), and stores the second higher-order metadata in the destination high-order metadata DB 153. Details of the high-order metadata will be described later.

[0032] The high-order motion conversion unit 123 converts the motion data of the source avatar (first motion data) into high-order motion data (second motion data) using the source high-order metadata. The high-order motion conversion unit 123 stores the high-order motion data in the motion data DB 151. Details of the high-order motion data will be described later.

[0033] The motion conversion unit 124 converts the high-order motion data into motion data (third motion data) of the conversion destination avatar using the conversion destination high-order metadata. Details of this conversion process will be described later.

[0034] (Operation Input Unit 130 and Display Unit 140) The operation input unit 130 accepts operation input by the user and outputs the input information to the control unit 120. The display unit 140 displays various screens. The display unit 140 may be a display panel such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The operation input unit 130 and the display unit 140 may be integrated. For example, the operation input unit 130 may be a touch sensor stacked on the display unit 140 (e.g., a panel display).

[0035] (Storage unit 150) The storage unit 150 is realized by a ROM that stores programs and calculation parameters used in the processing of the control unit 120, and a RAM that temporarily stores parameters that change as appropriate. The storage unit 150 stores a motion data DB 151, a source high-level metadata DB 152, and a destination high-level metadata DB 153.

[0036] The configuration of the information processing device 10 has been specifically described above. Note that the configuration of the information processing device 10 according to the present disclosure is not limited to the example shown in Fig. 3. For example, the information processing device 10 does not necessarily have all of the components shown in Fig. 3. Furthermore, the information processing device 10 may be realized by multiple devices.

[0037] The functions of information processing device 10 may be realized by separating them into a device (or application) that creates motions for the source avatar and a device (or application) that creates motions for the destination avatar. For example, the motion creation device for the source avatar has at least the functions of high-order motion conversion unit 123, and the motion creation device for the destination avatar has at least the functions of motion conversion unit 124.

[0038] <<3. Details of Motion Conversion>> Next, details of the motion conversion according to this embodiment will be described.

[0039] <3-1. Generation of Motion Data of Source Avatar> The motion generation unit 121 generates motion data of the source avatar in response to user operations. The generation method is not particularly limited, and may be, for example, a generation method using a 3D model creation tool, or a generation method using posture estimation technology that tracks human movement and outputs joint position information as posture estimation.

[0040] <3-2. Generation of High-Order Metadata> The high-order metadata generation unit 122 generates high-order metadata, which is a definition file that colloquially expresses the motions of avatars. High-order metadata is generated for each avatar with a different bone structure setting. That is, in this embodiment, source high-order metadata corresponding to the source avatar and destination high-order metadata corresponding to the destination avatar are generated.

[0041] The high-level metadata includes identification information (colloquial expression ID) of colloquial expressions that express motions (poses) in colloquial language, and corresponding motion control content (information for realizing the pose, more specifically, angle information of bones, etc.). In this embodiment, for example, 45 colloquial expressions are prepared, and avatar motion data is associated with each of them. An example of colloquial expressions according to this embodiment is shown in Table 1 below.

[0042]

[0043] The above 45 colloquial expressions are predefined, but the user can add new colloquial expressions as needed, starting from 46. This allows the use of colloquial expressions to express motions for avatars with wings or tails, or for avatars that are not humanoid. Note that, while the number of colloquial expressions is set to 45 as an example, it may be less than 45.

[0044] When defining the motion data (conversion target) of the source avatar using the colloquial expressions (i.e., generating high-level motion data), a numerical value (e.g., 0 to 1) is taken to represent the degree of the state of the motion (more specifically, for example, pose) indicated by the colloquial expression. Here, 0 indicates that the motion indicated by the colloquial expression is not being performed at all, and 1 indicates that the motion indicated by the colloquial expression is being performed completely. The pose in the 0 state can also be defined in the high-level metadata. Note that the 0 state is common to all colloquial expressions, and in this embodiment, it is also referred to as the standard state (or standard pose).

[0045] FIG. 4 is a diagram showing examples of pose states indicated by colloquial expressions according to this embodiment when they are 0 and 1. In FIG. 4, the case of colloquial expression ID: 0, "head turned to the right," will be described as an example. Pose image 310 on the left side of FIG. 4 is an example of an avatar pose when state: 0 is used. In this case, the avatar is not facing right at all, i.e., facing forward (standard state). In contrast, pose image 320 on the right side of FIG. 4 is an example of a pose when state: 1 is used, in which the avatar is facing directly to the right. Note that, in this example, state 1 is used where the avatar is facing directly to the right (head rotation angle 90 degrees), but the extent (degree) to which the avatar faces right when state 1 is used can be set as appropriate in the definition of high-level metadata.

[0046] Although FIG. 4 shows the cases of state 0 and state 1, for example, state 0.5 is an intermediate state between state 0 and state 1.

[0047] Furthermore, the values ​​that the state can take are not limited to the above 0 to 1, and it is also possible to take -1 to represent a state in the completely opposite direction, or to set a value greater than 1 to represent an exaggerated movement. For example, in the case of colloquial expression ID: 0, it is possible to represent a state in which the head is facing directly left with state -1, and a state in which the head is facing directly backward with state 2.

[0048] An example of a generation interface for high-level metadata will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of a generation interface for high-level metadata according to this embodiment. A user can use the generation interface to generate source high-level metadata and destination high-level metadata.

[0049] In the example shown in Fig. 5, source high-level metadata for a source avatar (e.g., type "Gray Skeleton") is being generated. A setting screen 351 for setting a standard state (standard pose) and a setting screen 352 for setting a pose (motion) of an arbitrary colloquial expression ID are displayed on the generation interface 350. The number of setting screens displayed on the generation interface 350 is not limited to two, but by presenting two or more screens, for example, the user can set a pose while comparing them.

[0050] On each setting screen, the user selects the colloquial expression they want to set using a pull-down menu, and then rotates the displayed bones of the source avatar as desired (X-axis rotation, Y-axis rotation, or Z-axis rotation) to determine the pose they want to associate with the selected colloquial expression. More specifically, for example, the user selects a bone from the bone list on the left side of the screen, or from the bones superimposed on the avatar in the setting screen, and then rotates the selected bone by manipulating the rotation control panel 3511 displayed around it. The rotation control panel 3511 includes an X-axis rotation control panel, a Y-axis rotation control panel, and a Z-axis rotation control panel. The user clicks on the control panel portion they want to rotate. For example, clicking on the Y-axis rotation control panel changes the panel to a Y-axis rotation control panel 3512, as shown in the setting screen 352. The user can then drag on the Y-axis rotation control panel 3512 to rotate the selected bone (the head bone in the example shown in FIG. 5 ) around the Y-axis. When the user has decided on the pose of the source avatar that he / she wants to associate with the colloquial expression, he / she clicks the Confirm button. In this way, the user can intuitively define the pose of the source avatar with the colloquial expression.

[0051] The high-level metadata generation unit 122 obtains the bone angles in the determined avatar pose, associates them with the colloquial expression ID, and generates high-level metadata.

[0052] The user sets the standard pause and the pause for each colloquial expression using the generation interface 350, and in response, the high-level metadata generation unit 122 generates high-level metadata for the standard pause and the pause for each colloquial expression.

[0053] More specifically, the high-level metadata consists of a header section that describes the number of bones of the target avatar, the bone index, and the angle of each bone in the standard pose (X-axis rotation angle, Y-axis rotation angle, Z-axis rotation angle), and a main section that describes the number of colloquial expressions and the angle of each bone (as changed from the standard pose) in the pose of each colloquial expression.

[0054] 6 is a diagram showing an example of source high-level metadata according to this embodiment. As shown in Fig. 6, the header section 411 of the source high-level metadata 410 begins with the number of bones of the source avatar, and below that, on separate lines, are written information about the bones of the source avatar in their standard pose, including the bone index (number), bone name, X-axis rotation angle, Y-axis rotation angle, and Z-axis rotation angle. Each value is separated by a half-width space.

[0055] In the main body 412 of the source high-level metadata 410, the number of colloquial expressions is written at the beginning, and below that, on each line, information on the pause of each colloquial expression is written in this order: the colloquial expression ID, the bone index to be changed, and the X-axis rotation angle, Y-axis rotation angle, and Z-axis rotation angle after the change. The pose of each colloquial expression written in the main body 412 indicates state 1, that is, the case where the pose indicated by that colloquial expression is in the maximum state.

[0056] The second line of the main body 412 indicates, as information for the pose with colloquial expression ID: 0 "head turned to the right," that the changes in the rotation angles of the head bone (bone index 0) are X-axis rotation angle: 90, Y-axis rotation angle: 0, and Z-axis rotation angle: 0. The number of bones that change is not limited to one, and multiple bones may change. For example, the third line of the main body 412 indicates, as information for the pose with colloquial expression ID: 1 "head tilted forward," that the changes in the rotation angles of the head bone (bone index 0) are X-axis rotation angle: 0, Y-axis rotation angle: 90, and Z-axis rotation angle: 0, and that the changes in the rotation angles of the neck_01 bone (bone index 1) are X-axis rotation angle: 0, Y-axis rotation angle: 90, and Z-axis rotation angle: 90. Additionally, the fourth line of the main body 412 indicates that, as information for the pose of colloquial expression ID: 2 "head tilted to the right," the changes in the rotation angles of the head bone (bone index 0) are X-axis rotation angle: 0, Y-axis rotation angle: 0, and Z-axis rotation angle: 90.

[0057] The above provides a specific description of how source high-level metadata is generated using the generation interface 350. Similarly, the user also generates destination high-level metadata for a destination avatar (for example, type "Red Player").

[0058] 7 is a diagram showing an example of destination high-level metadata according to this embodiment. As shown in Fig. 7, destination high-level metadata 420 also consists of a header portion 421 and a body portion 422. The file description format is the same as that of source high-level metadata 410 shown in Fig. 6.

[0059] The angles of each bone in the standard pose shown in FIGS. 6 and 7, and the angles of change of each bone in the poses corresponding to each colloquial expression, are merely examples, and the present embodiment is not limited to these.

[0060] 3-3. Conversion to Higher-Order Motion The higher-order motion conversion unit 123 uses the source avatar motion data generated by the motion generation unit 121 and the source high-order metadata generated by the higher-order metadata generation unit 122 to generate higher-order motion data, which is a file that represents the motion data at a higher level. The higher-order motion data describes changes in higher-order information over time, i.e., changes in the state of a pose expressed in colloquial language. For example, the higher-order motion data may describe information such as, "In this motion, the state of 'tilting the head to the right' changes from 0.2 at 0 seconds, to 0.7 at 1 second, and to 1 at 2 seconds."

[0061] In the motion data of the source avatar, posture changes are expressed by coordinate rotation, so the high-order motion conversion unit 123 compares the coordinate rotation in the motion data with the high-order metadata and performs vector calculations to generate high-order motion data.

[0062] The high-order motion conversion unit 123 can convert all of the motion data into high-order motion data by converting the motion (pose) of each frame into high-order information for all frames in the motion data of the source avatar.

[0063] For example, let us consider a case where the motion (pose) of one frame in the motion data of the source data is "head joint rotation angle (X, Y, Z) = (45, 45, 0), neck joint rotation angle (X, Y, Z) = (0, 45, 45)."

[0064] In contrast, in the source high-level metadata, the pose of colloquial expression ID: 2 "tilting head to the right" is defined as a head bone rotation angle of "X: 0, Y: 0, Z: 90," and the pose of colloquial expression ID: 1 "tilting head forward" is defined as a head bone rotation angle of "X: 0, Y: 90, Z: 0" and a neck_01 bone rotation angle of "X: 0, Y: 90, Z: 90." Here, the head joint corresponds to the head bone, and the neck joint corresponds to the neck_01 bone.

[0065] In this case, if the colloquial expression ID: 2 "head tilted to the right" state is set to 0.5 and the colloquial expression ID: 1 "head tilted forward" state is set to 0.5, the angle of the pose in the above one frame of the motion data can be indicated as shown in the following equation 1.

[0066] (Formula 1) Head bone (head joint) angle (X,Y,Z) = 0.5 x (90,0,0) + 0.5 x (0,90,0) = (45,45,0) Neck_01 bone (neck joint) angle (X,Y,Z) = 0.5 x (0,90,90) = (0,45,45)

[0067] The high-order motion conversion unit 123 may determine the state value ("0.5") of each colloquial expression by, for example, a numerical analysis method such as Newton's method or by vector synthesis calculations. In this case, since it is possible that the value obtained may not be exactly equal to the angle of the pose to be converted depending on how the value is calculated, a threshold value is set, and if the vector distance is equal to or less than the threshold value, the pose is considered to be expressed. Furthermore, since the calculation cost for complex colloquial expressions expressed using multiple bones (or joints) can be enormous, the number of bones (or joints) moved for each colloquial expression may be limited to one.

[0068] For each frame, the high-order motion transformation unit 123 determines and describes the state value of each colloquial expression ID that indicates the pose in that frame, and generates high-order motion data.

[0069] 8 is a diagram showing an example of high-order motion data according to this embodiment. As shown in Fig. 8, high-order motion data 510 consists of a header section 511 describing fps (frames per second) and the number of frames, and a body section 512 describing the pose in each frame.

[0070] 8 , in main body 512, each line corresponds to information for one frame, and describes a colloquial expression ID and its state. For example, the first line of main body 512 indicates that, as information for the first frame (the avatar's pose in the first frame), colloquial expression ID: 0 is in state 0.2, colloquial expression ID: 1 is in state 0, colloquial expression ID: 2 is in state 0, ..., colloquial expression ID: 44 is in state 1. Next, the second line of main body 512 indicates, as information for the second frame (the avatar's pose in the second frame), colloquial expression ID: 0 is in state 0.1, colloquial expression ID: 1 is in state 0, colloquial expression ID: 2 is in state 0, ..., colloquial expression ID: 44 is in state 0.9.

[0071] The format of the high-order motion data described above is an example, and the present embodiment is not limited to this. The high-order motion data may also be output as an intermediate file and used for other purposes.

[0072] <3-4. Generation of motion data for the destination avatar> The motion conversion unit 124 converts the high-order motion data generated by the high-order motion conversion unit 123 into motion data that can be used in the destination avatar, using the high-order metadata of the destination avatar (destination high-order metadata) generated by the high-order metadata generation unit 122.

[0073] The motion conversion unit 124 restores the motion (pose) by multiplying the value of the state of the colloquial expression in each frame described in the high-order motion data by the pose (pose of the destination avatar) defined in the corresponding colloquial expression in the destination high-order metadata and synthesizing the vectors. By restoring the poses for all frames described in the high-order motion data, the motion conversion unit 124 can generate motion data that can be used for the destination avatar.

[0074] For example, let us consider a case where the motion (pose) of one frame described in the high-level motion data is colloquial expression ID: 2 "tilting head to the right" state 0.5 and colloquial expression ID: 1 "tilting head forward" state 0.5.

[0075] In contrast, in the destination high-level metadata, the pose of colloquial expression ID: 2 "head tilted to the right" is defined as having head bone rotation angles of "X: 0, Y: 0, Z: -90," and the pose of colloquial expression ID: 1 "head tilted forward" is defined as having head bone rotation angles of "X: 90, Y: 0, Z: 0." Note that in this embodiment, it is assumed that the bone structure from the neck to the head of the source avatar is made up of the neck_01 bone, neck_02 bone, and head bone, while the bone structure from the neck to the head of the destination avatar is made up of the neck_01 bone and head bone, and that head movement in the destination avatar is realized with only one bone (the head bone).

[0076] In this case, the pose angle in the above one frame of high-level motion data can be reproduced as shown in the following formula 2. The angle of the neck_01 bone (neck joint) is zero because it is not included in either colloquial expression ID: 2 or colloquial expression ID: 1 in the converted high-level metadata.

[0077] (Formula 2) Head bone (head joint) angle (X, Y, Z) = 0.5 × (0, 0, -90) + 0.5 × (90, 0, 0) = (45, 0, -45) Neck_01 bone (neck joint) angle (X, Y, Z) = 0.5 × (0, 0, 0) + 0.5 × (0, 0, 0) = (0, 0, 0)

[0078] An example of a generation interface for generating motion data as a destination of conversion will be described with reference to Fig. 9. Fig. 9 is a diagram showing an example of a generation interface for generating motion data as a destination of conversion according to this embodiment.

[0079] A generation interface 610 shown in FIG. 9 displays a playback screen 611 for playing back the motion of the source avatar, and a playback screen 612 for playing back the motion of the destination avatar.

[0080] When the user selects source high-level metadata, source motion data, and destination high-level metadata in the generation interface 610 and presses the conversion button, the motion data conversion process is performed, and the motion of the destination avatar is played back on the playback screen 612. The motion of the source avatar may also be played back on the playback screen 611. By displaying the two screens side by side on the generation interface 610, the user can compare and confirm the motions of both.

[0081] The avatar motion shown in FIG. 9 is assumed to involve tilting the head diagonally forward and to the right from a standard state (head facing straight ahead). The avatar shown in FIG. 9 is in a pose at the end of the motion, with the head joint rotation angle (X, Y, Z) = (45, 45, 0) and the neck joint rotation angle (X, Y, Z) = (0, 45, 45). FIG. 10 is a diagram illustrating the motion of tilting the head diagonally forward and to the right. Pose image 650 shown in the upper left of FIG. 10 is the standard state, pose image 651 shown in the upper right of FIG. 10 is a state in which the head is tilted 45 degrees to the right, pose image 652 shown in the lower left of FIG. 10 is a state in which the head is tilted 45 degrees forward, and pose image 653 shown in the lower right of FIG. 10 is a state in which the head is tilted diagonally forward and to the right. In FIG. 9, the motion of pose image 653 is reproduced by the converted avatar.

[0082] <<4. Motion Processing>> FIG. 11 is a flowchart showing the flow of motion conversion processing by the information processing device 10 of this embodiment.

[0083] As shown in FIG. 11, first, the information processing device 10 acquires motion data of the source avatar (step S103).

[0084] Next, the information processing device 10 converts the motion data of the source avatar into high-order motion data using the high-order metadata of the source avatar (source high-order metadata) (step S106).

[0085] Next, the information processing device 10 uses the high-level metadata of the conversion destination avatar (conversion destination high-level metadata) to convert the high-level motion data into motion data that can be used by the conversion destination avatar (step S109).

[0086] The information processing device 10 then reproduces the motion of the destination avatar based on the converted motion data (step S112). This allows the motion generated for the source avatar to be properly reproduced on the destination avatar, which has a different bone structure setting.

[0087] <<5. Others>> (5-1) In the above-described embodiment, a configuration in which motion conversion processing is performed by the information processing device 10 has been described, but the present disclosure is not limited to this. The present disclosure may be configured by an information processing system 1 including a server and a terminal. The system configuration will be described below with reference to FIG. 12 .

[0088] 12 is a diagram illustrating a configuration example of an information processing system 1 according to an embodiment of the present disclosure. As illustrated in Fig. 12, the information processing system 1 includes a server 20 and user terminals 30 (30a, 30b, 30c, ...). The server 20 and the user terminals 30 are connected for communication via a network 40.

[0089] 1 , and can display various generated interfaces on the user terminal 30. In addition, a part of the functional configuration of the information processing device 10 may be provided in the user terminal 30.

[0090] The system configuration example shown in FIG. 12 is an example, and the present disclosure is not limited to this.

[0091] (5-2) In the above-described embodiment, it is assumed that pre-generated motion data of a source avatar is used to generate motion data that can be used for a destination avatar, and the generated motion data is saved or output. However, this is not limited to this, and the information processing device 10 may convert motion data in real time. For example, the present disclosure may be applied to real-time motion conversion when a person's movement is reflected from one avatar to another avatar using virtual reality (VR), augmented reality (AR), capture technology, or the like.

[0092] (5-3) In the above embodiment, the motion is represented by the angle of the bone, but this is not limiting and the motion may be represented by the angle of a body part such as a joint.

[0093] <<6. Supplementary Information>> Although preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the present technology is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.

[0094] For example, one or more computer programs can be created for hardware such as a CPU, ROM, and RAM built into the information processing device 10 to perform the functions of the information processing device 10. Also provided is a computer-readable storage medium storing the one or more computer programs.

[0095] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0096] The present technology can also be configured as follows. (1) An information processing system including a control unit that performs the following processes: converting input first motion data of a first avatar into second motion data based on first metadata defining control content corresponding to a motion of the first avatar, and converting the second motion data into third motion data for the second avatar using second metadata defining control content corresponding to a motion of a second avatar. (2) The information processing system according to (1), wherein the first metadata and the second metadata are expressed in colloquial language. (3) The information processing system according to (2), wherein the control content is information regarding angles of bones of the avatar in a pose indicated by the corresponding colloquial language. (4) The information processing system according to (2) or (3), wherein the first metadata is data that defines, for each colloquial expression of a motion, a control content of the motion corresponding to the colloquial expression of the first avatar, corresponding to the setting of the first avatar, and the second metadata is data that defines, for each colloquial expression of the motion, a control content of the motion corresponding to the colloquial expression of the second avatar, corresponding to the setting of the second avatar. (5) The information processing system according to any one of (4) to (3), wherein the setting of the first avatar and the setting of the second avatar are different, and the setting is a setting of the bone structure of the avatar. (6) The information processing system according to any one of (1) to (5), wherein the first and second metadata consist of a header portion and a body portion. (7) The information processing system according to any one of (1) to (6), wherein the second motion data consists of a header portion and a body portion. (8) The information processing system according to any one of (1) to (7), wherein the control unit generates the second motion data by replacing, for each frame in the first motion data, the motion of the first avatar with a value of a pose state indicated by a colloquial expression defined in the first metadata.(9) The information processing system according to (8), wherein the control unit generates the third motion data by multiplying a value of a pose state indicated by a colloquial expression described for each frame in the second motion data by a bone angle of the second avatar in the pose indicated by the corresponding colloquial expression in the second metadata. (10) The information processing system according to any one of (1) to (9), wherein the first metadata includes information regarding bone angles of the first avatar in a standard pose, and the second metadata includes information regarding bone angles of the second avatar in the standard pose. (11) The information processing system according to any one of (1) to (10), wherein the control unit controls displaying a generation screen that accepts user operations related to generation of the first or second metadata. (12) The information processing system according to (11), wherein the generation screen allows specification of a colloquial expression and a pose of the first or second avatar. (13) The information processing system according to (12), wherein the pose specification is specification of X-axis rotation angles, Y-axis rotation angles, and Z-axis rotation angles of bones constituting an avatar. (14) The information processing system according to any one of (1) to (13), wherein the control unit controls displaying a generation screen that accepts user operations related to processing of converting first motion data for the first avatar to generate third motion data for the second avatar. (15) An information processing method, including: a processor converting the first motion data into second motion data based on first metadata defining control content corresponding to a motion of a first avatar and input first motion data of the first avatar; and converting the second motion data into third motion data for the second avatar using second metadata defining control content corresponding to the motion of a second avatar.(16) A program that causes a computer to function as a control unit that performs the following processes: converting first motion data into second motion data based on first metadata that defines control content corresponding to the motion of a first avatar and input first motion data of the first avatar; and converting the second motion data into third motion data for the second avatar using second metadata that defines control content corresponding to the motion of a second avatar.

[0097] REFERENCE SIGNS LIST 1 Information processing system 10 Information processing device 110 Communication unit 120 Control unit 121 Motion generation unit 122 High-level metadata generation unit 123 High-level motion conversion unit 124 Motion conversion unit 130 Operation input unit 140 Display unit 150 Storage unit 151 Motion data DB 152 Source high-level metadata DB 153 Destination high-level metadata DB 20 Server 30 User terminal

Claims

1. An information processing system comprising a control unit that performs the following processes: converting first motion data into second motion data based on first metadata that defines control content corresponding to the motion of a first avatar and input first motion data of the first avatar; and converting the second motion data into third motion data for the second avatar using second metadata that defines control content corresponding to the motion of a second avatar.

2. The information processing system according to claim 1, wherein the first metadata and the second metadata are expressed colloquially.

3. The information processing system according to claim 2, wherein the control content is information regarding angles of the avatar's bones in a pose indicated by a corresponding colloquial expression.

4. The information processing system described in claim 2, wherein the first metadata is data that defines the control content of the motion corresponding to the colloquial expression of the first avatar for each colloquial expression of the motion, corresponding to the setting of the first avatar, and the second metadata is data that defines the control content of the motion corresponding to the colloquial expression of the second avatar for each colloquial expression of the motion, corresponding to the setting of the second avatar.

5. The information processing system according to claim 4, wherein the settings of the first avatar and the settings of the second avatar are different, and the settings are settings of the bone structure of the avatar.

6. The information processing system according to claim 1, wherein the first and second metadata consist of a header portion and a body portion.

7. The information processing system according to claim 1, wherein the second motion data comprises a header portion and a body portion.

8. The information processing system of claim 1, wherein the control unit generates the second motion data by replacing, for each frame in the first motion data, the motion of the first avatar with a pose state value indicated by a colloquial expression defined in the first metadata.

9. The information processing system of claim 8, wherein the control unit generates the third motion data by multiplying the value of the pose state expressed in colloquial expressions described for each frame in the second motion data by the bone angle of the second avatar in the pose expressed in the corresponding colloquial expression in the second metadata.

10. An information processing system as described in claim 1, wherein the first metadata includes information regarding the angles of the bones of a first avatar in a standard pose, and the second metadata includes information regarding the angles of the bones of a second avatar in a standard pose.

11. The information processing system according to claim 1, wherein the control unit controls displaying a generation screen that accepts user operations related to the generation of the first or second metadata.

12. The information processing system according to claim 11, wherein the generation screen allows specification of colloquial expressions and a pose for the first or second avatar.

13. The information processing system according to claim 12, wherein the pose specification is a specification of X-axis rotation angles, Y-axis rotation angles, and Z-axis rotation angles of bones that make up the avatar.

14. The information processing system of claim 1, wherein the control unit controls displaying a generation screen that accepts user operations related to the process of converting first motion data for the first avatar to generate third motion data for the second avatar.

15. An information processing method comprising: a processor converting first motion data into second motion data based on first metadata defining control content corresponding to the motion of a first avatar and input first motion data of the first avatar; and converting the second motion data into third motion data for the second avatar using second metadata defining control content corresponding to the motion of a second avatar.

16. A program that causes a computer to function as a control unit that performs the following processes: converting first motion data into second motion data based on first metadata that defines control content corresponding to the motion of a first avatar and input first motion data of the first avatar; and converting the second motion data into third motion data for the second avatar using second metadata that defines control content corresponding to the motion of a second avatar.

Citation Information

Patent Citations

  • Movement converter for three-dimensional skeleton structure

    JP1997330424A

  • Motion cover method, apparatus, and system for linking animation data between different characters in a compatible manner

    KR102589194B1