Feature amount generation device, information processor, server, information processing method, and program
The feature generation device addresses the limitations of fixed joint count in 3D motion models by rearranging and standardizing skeletal joints, enabling versatile feature representation and cross-modal analysis.
Patent Information
- Application Number
- JP2024035322
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-19
AI Technical Summary
Existing 3D motion models struggle to handle varying numbers of joint points and are limited in expressing local movements and applying to non-human shapes, such as animals, due to fixed joint numbers and structures.
A feature generation device that classifies skeletal joints by part, rearranges them based on distance from a set joint, interpolates to a predetermined number, standardizes position information, and converts it into color information to generate features independent of joint count.
Enables the generation of features that uniformly represent various skeletal structures, allowing for cross-modal analysis between motion sequences and natural language, and supports applications beyond human models.
Smart Images

Figure 2025136621000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a feature generation device, an information processing device, a server, an information processing method, a program, and the like. [Background technology]
[0002] With the rapid development and spread of video technologies such as VFX and 3D animation that can express the motion of the human body, etc., there is a demand for the establishment of a representation format (3D motion model) of the three-dimensional motion of the human body, etc. that can be easily processed by information processing devices. In a 3D motion model, a set of points indicating joints is used for skeletal representation (skeleton model). For example, Non-Patent Document 1 discloses a method for generating a human body model based on principal component features that represent the shape of the human body and 23 joint parameters. However, in the method disclosed in Non-Patent Document 1, the number of joint points of the human body model is fixed in advance, making it difficult to, for example, handle data with different numbers of joint points in an integrated manner and have an AI model learn it. Furthermore, because the shape and number of joints of the human body model are fixed, there are problems such as difficulty in expressing local movements of the human body (e.g., hand movements) and application to animals other than humans whose shapes and joint structures are significantly different. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Matthew Loper, et al. “SMPL: a skinned multi-person linear model”, 2015. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention has been made based on the above-mentioned technical background, and aims to provide an information processing method for generating features that can effectively represent a 3D motion model, regardless of the shape of the object or the number of joints. [Means for solving the problem]
[0005] According to a first aspect of the present invention, a feature generation device includes a control unit that acquires a three-dimensional skeletal partial time series from a three-dimensional skeletal movement sequence, classifies skeletal joints in the three-dimensional skeletal partial time series by skeletal part, rearranges the skeletal joints for each part based on their distance from a set joint, interpolates the rearranged skeletal joints to a predetermined number of skeletal joints, standardizes position information of the skeletal joints in the interpolated three-dimensional skeletal partial time series, generates part patches by converting the standardized position information into color information, and generates a feature that integrates the part patches generated for each part. According to a second aspect of the present invention, an information processing device capable of mutually converting between three-dimensional skeletal movement sequences and sentence sequences comprises a feature extraction device, an image encoding unit, and a sentence encoding unit, wherein the image encoding unit calculates image features based on the features generated from the three-dimensional skeletal movement sequences by the feature extraction device, the sentence encoding unit calculates sentence features based on the sentence sequences, and a control unit of the information processing device optimizes the image encoding unit and the sentence encoding unit so as to be able to calculate appropriate combinations of three-dimensional skeletal movement sequences and sentence sequences based on the image features and the sentence features. According to a third aspect of the present invention, a server that communicates with at least a first terminal includes an information processing device, and a control unit of the server transmits first information regarding authorization to use the information processing device to the first terminal via a communication unit of the server, receives second information regarding a request to use the information processing device from the first terminal via the communication unit, and drives the information processing device if the second information includes the first information. According to a fourth aspect of the present invention, an information processing method in a feature generation device includes obtaining a three-dimensional skeletal partial time series from a three-dimensional skeletal movement sequence, classifying skeletal joints in the three-dimensional skeletal partial time series by skeletal part, rearranging the skeletal joints for each part based on their distance from a set joint, interpolating the rearranged skeletal joints to a predetermined number of skeletal joints, standardizing position information of the skeletal joints in the interpolated three-dimensional skeletal partial time series, generating part patches by converting the standardized position information into color information, and generating features by integrating the part patches generated for each part. According to a fifth aspect of the present invention, a program executed by the feature generation device performs the following operations: acquiring a three-dimensional skeletal partial time series from a three-dimensional skeletal movement sequence; classifying skeletal joints in the three-dimensional skeletal partial time series by skeletal part; rearranging the skeletal joints for each part based on their distance from a set joint; interpolating the rearranged skeletal joints to a predetermined number of skeletal joints; standardizing position information of the skeletal joints in the interpolated three-dimensional skeletal partial time series; generating part patches by converting the standardized position information into color information; and generating features by integrating the part patches generated for each part. [Brief explanation of the drawings]
[0006] [Figure 1-1] FIG. 1 is a diagram showing an example of the configuration of a feature generating device according to a first embodiment. [Figure 1-2] 4 is a flowchart showing an example of the flow of processing executed by the feature generating device according to the first embodiment. [Figure 1-3] FIG. 3 is a conceptual diagram of a body part division process according to the first embodiment. [Figure 1-4] 5A and 5B are conceptual diagrams of joint point rearrangement processing and joint point position color information conversion processing according to the first embodiment. [Figure 1-5] FIG. 2 is a diagram showing an example of motion patch data generated by the feature generating device according to the first embodiment. [Figure 1-6] FIG. 2 is a diagram showing an example of motion patch data generated by the feature generating device according to the first embodiment. [Figure 2-1]FIG. 10 is a diagram showing an example of the configuration of an information processing device according to a second embodiment. [Figure 2-2] 10 is a flowchart showing an example of the flow of a learning process executed by an information processing device according to a second embodiment. [Figure 2-3] FIG. 10 is a diagram showing an example of a text-based 3D motion search result according to the second embodiment. [Figure 2-4] FIG. 10 is a diagram showing another example of a text-based 3D motion search result according to the second embodiment. [Figure 2-5] 10 is a table showing the results of a performance comparison with other methods according to the second embodiment. [Figure 2-6] 10 is a separate table showing the results of a performance comparison with other methods according to the second embodiment. [Figure 3-1] FIG. 10 is a diagram showing an example of the system configuration of a communication system according to a third embodiment. [Figure 3-2] FIG. 11 is a diagram showing an example of functions realized by a control unit of a server according to a third embodiment. [Figure 3-3] FIG. 11 is a diagram showing an example of information stored in a storage unit of a server according to a third embodiment. [Figure 3-4] FIG. 11 is a diagram showing an example of token registration data according to the third embodiment. [Figure 3-5] 10 is a flowchart showing an example of the flow of processing executed by each device according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0007] <Compliance with legal matters> It should be noted that the disclosures set forth herein are subject to compliance with the laws of the country of implementation, such as communications privacy, as required for the implementation of the disclosures.
[0008] <Embodiment> In this specification, for ease of understanding, there are places where the term "for example" is used, but please note that not only those places but also the entire embodiment described below are not limited to the content of the description.
[0009] An embodiment for implementing a program etc. according to the present disclosure will be described with reference to the drawings.
[0010] The production of a terminal of the claimed invention (terminal of the claimed invention) may include, for example, the concept that a state is created in which the functions of the claimed invention can be realized (a state in which the claimed invention can be executed) on a terminal owned (possessed) by a user by receiving (or receiving and storing) a program (for example, an application program) described in this specification.
[0011] Furthermore, the production of the system of the invention claimed in the present application (the system of the invention of the present application) may include, for example, the concept that a state in which the functions of the invention of the system claimed in the present application can be realized (a state in which the invention claimed in the present application can be executed) is created by receiving a program described in this specification (for example, an application program) transmitted from a server included in the system of the present application at a terminal included in the system of the present application (or by storing the received program in the terminal).
[0012] In this specification, a system may be, for example, configured to include a plurality of devices. The plurality of devices may be a combination of devices of the same type, a combination of devices of different types, or a combination of devices of the same type and devices of different types. A system can also be thought of as, for example, a plurality of devices working together to perform some kind of processing.
[0013] Furthermore, a system relating to a client (client device) and a server can be considered to be, for example, at least one of the following: (1) Terminals and servers (2) Server (3) Terminal
[0014] (1) is, for example, a system including at least one terminal and at least one server. One example of this is a client-server system.
[0015] The server is configured by the following devices, for example, and may be a single device or a combination of multiple devices.
[0016] Specifically, a server is configured to have, for example, at least one processor (for example, CPU: Central Processing Unit, GPU: Graphics Processing Unit, APU: Accelerated Processing Unit, DSP: Digital Signal Processor (for example, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array), etc.), computer device (processor + memory), control device, arithmetic device, processing device, etc., and may be configured to have multiple of the same type of device (for example, CPU + CPU, homogeneous multi-core processor, etc.), or multiple of different types of device (for example, CPU + DSP, heterogeneous multi-core processor, etc.), or may be a combination of multiple devices (for example, processor + computer device, processor + arithmetic device, multiple devices made heterogeneous, etc.). The processor may be a virtual processor.
[0017] Furthermore, when a server performs some processing, if the server is configured with a single device, the processing described in the embodiments is performed by the single device. Furthermore, if the server is configured with multiple devices, one device may perform some of the processing, and the other device may perform other processing. For example, if the server is configured with a processor and an arithmetic device, the processor may perform a first processing, and the arithmetic device may perform a second processing. Furthermore, when a plurality of devices are used, the devices may be located at positions physically separated from one another.
[0018] Furthermore, the server functions may be provided in the form of PaaS, IaaS, or SaaS in cloud computing, for example. Some or all of the processes described in this specification may be implemented as a program included in an application installed on a terminal from a server. Furthermore, the manufacturer and manager of the terminal, the manufacturer and manager of the application, and the manufacturer and manager of the server may be different entities (businesses), or some or all may be the same entity (business).
[0019] The control unit of the system can be at least one of the control unit of the terminal and the control unit of the server. That is, for example, the control unit of the system can be any of (1A) only the control unit of the terminal, (1B) only the control unit of the server, or (1C) both the control unit of the terminal and the control unit of the server.
[0020] Furthermore, the control and processing (hereinafter collectively referred to as "control, etc.") performed by the control unit of the system may be performed by (1A) only the control unit of the terminal, (1B) only the control unit of the server, or (1C) both the control unit of the terminal and the control unit of the server. In addition, in (1C), for example, some of the controls performed by the control unit of the system may be performed by the control unit of the terminal, and the remaining controls may be performed by the control unit of the server. In this case, the allocation (allocation) of the controls may be equal or may be allocated in different proportions.
[0021] Furthermore, when referring to the communication unit of a server, if the server is configured with a single device, it may refer to the communication unit itself that the single device has, or if the server is configured with multiple devices, it may be configured to include each communication unit that each device has. For example, if a server comprises a first device and a second device, and the first device has a first communication unit and the second device has a second communication unit, the communication unit of the server may be considered to include the first communication unit and the second communication unit.
[0022] (2) can be, for example, a system consisting of multiple servers (hereinafter referred to as a "server system"). In this case, the configuration of each server can be similarly applied to the configuration described above.
[0023] The control etc. performed by the server system may be performed by only one of the multiple servers (2A), by only the other servers (2B), or by both the one server and the other servers (2C). In addition, in (2C), for example, one server may perform some of the control, etc., performed by the server system, and another server may perform the remaining control, etc. In this case, the allocation (allocation) of the control, etc. may be equal or may be allocated in different proportions.
[0024] (3) can be, for example, a system consisting of multiple terminals. This system can be, for example, the following system. A system that gives server functions to terminals (distributed system). This can be realized, for example, using blockchain technology. A system in which terminals communicate wirelessly with each other. This can be achieved, for example, by using short-range wireless communication technology such as Bluetooth (registered trademark) to communicate in a P2P (peer-to-peer) format.
[0025] The above is not limited to the control unit, but also applies to each functional unit such as an input / output unit, a communication unit, a storage unit, and a clock unit that may be components of the system.
[0026] In the following embodiment, a system including a terminal and a server (a client-server system, for example) will be described as an example. It is also possible to apply the server system described in (2) above as the server.
[0027] Furthermore, instead of a system including a terminal and a server, it is also possible to apply a system that does not include a server, such as the system (3) above. In this case, the embodiment can be configured based on the above-mentioned blockchain technology. Specifically, for example, data stored and managed in a server described in the following embodiment is stored on the blockchain. Then, a terminal generates a transaction to the blockchain, and when the transaction is approved on the blockchain, the data stored on the blockchain is updated.
[0028] It should be noted that even when the term "terminal" is used, this is not limited to the meaning of a terminal as a client device in a client server. That is, a terminal may include the concept of a device that is not in a client-server context.
[0029] Furthermore, in this specification, the expression "through a communication I / F" is used as appropriate. This may, for example, indicate that a device transmits and receives various information and data via a communication I / F (via a communication unit) based on the control of a control unit (such as a processor).
[0030] Furthermore, in this specification, when the terms "related to" or "related to" are used, for example, "B related to A" or "B related to A" may mean "B" that has some kind of relationship with "A." Specific examples of this will be described later.
[0031] Furthermore, in this specification, when a device performs processing on two or more objects, such as "sending A and B" or "receiving A and B," this may include performing "A" and "B" at the same time (hereinafter referred to as "simultaneous"), and performing "A" and "B" at different times (hereinafter referred to as "non-simultaneous"). For example, when referring to transmitting first information and second information, this may include both the concepts of transmitting the first information and the second information at the same time, and transmitting the first information and the second information at different times. In addition, taking into account the lag (time lag), "simultaneous" may include "almost simultaneously."
[0032] Note that even though "A" and "B" are performed at different times, this only needs to be done with "A" and "B" as the processing targets, and the purposes do not necessarily have to be the same. For example, when the first information and the second information are transmitted as described above, it is sufficient to transmit the first information and the second information, and this may include cases where the first information and the second information are transmitted for the same purpose, as well as cases where the first information and the second information are transmitted for different purposes.
[0033] Hereinafter, an example of an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same elements are denoted by the same reference numerals, and duplicated descriptions may be omitted. Furthermore, the components described in this embodiment are merely examples and are not intended to limit the scope of the present invention.
[0034] <First Example> The first embodiment is an embodiment in which a feature generation device (which may also be called an information conversion processing device, a feature conversion device, a feature extraction device, or an image generation device) generates motion patch data, which is a feature independent of the number of joint points, based on motion sequence data, which is time-series data of a human skeleton model including various numbers of joint points. The contents described in the first embodiment can be similarly applied to any of the other embodiments and other modified examples.
[0035] FIG. 1-1 is a block diagram showing an example of a functional configuration of a feature generating device 1 according to an aspect of this embodiment. The feature generation device 1 includes, for example, a partial patch generation unit 2 and a motion patch generation unit 3. Furthermore, for example, the partial patch generation unit 2 includes a motion sequence acquisition unit 4, a body part division unit 5, a joint point placement unit 6, and a joint point coordinate-to-color information conversion unit 7. These are, for example, functional units (functional blocks) included in a control unit (control device) (not shown) of the feature generating device 1. The control unit may also be called a processing unit (processing device).
[0036] The partial patch generating unit 2 has a function of generating a motion patch (referred to as a "partial patch") for each predetermined part of the human body based on the motion sequence data 8 input to the feature generating device 1, for example. Here, a motion patch is a feature that can uniformly express various skeletal structures in a motion sequence. For example, spatial information of each joint point in the motion sequence is converted into color information on a motion patch, and temporal information of each joint point in the motion sequence is converted into positional information on a motion patch. More specifically, spatial information of each joint point in the motion sequence is represented as horizontal color information on a motion patch, and the time progression of each joint point is represented as a vertical change in color information on a motion patch. Details of motion patches will be described later.
[0037] The motion patch generation unit 3 has a function of, for example, integrating the part patches generated by the partial patch generation unit 2 and generating motion patch data 9 corresponding to the motion sequence data 8.
[0038] The motion sequence data 8 is, for example, time-series data of a skeleton model composed of frames with a sequence length of "T" ("T" is any natural number). If the number of joint points of the skeleton model in each frame is "J" ("J" is any natural number) and the position information of each joint point is expressed in a Cartesian coordinate system (x, y, z), the motion sequence data will be "T x J x 3" dimensional data. Hereinafter, the skeleton model in one frame of the motion sequence data 8 will be referred to as "skeleton data." Each piece of skeleton data will be, for example, "J x 3" dimensional data. The position information of each joint point may be expressed in any coordinate system (for example, a polar coordinate system).
[0039] The motion sequence data 8 may be data provided in, for example, a BVH (Biovision Hierarchy) format. Alternatively, the motion sequence data 8 may be data provided based on a skeleton model in an SMPL (A Skinned Multi-Person Linear Model) format.
[0040] The motion patch data 9 is, for example, a set of partial patch tensor data of dimensions "N x N x 3" ("N" is a predetermined natural number satisfying "N≦T"). For example, the motion patch data 9 may be a set of three-channel (r, g, b) partial patch image data of "N x N" pixels. That is, the size of the motion patch data 9 does not depend on the number of joints “J” included in the motion sequence data 8.
[0041] The motion patch data 9 may be, for example, a set of partial patch tensor data of dimension "N×L×3" ("N" is a predetermined natural number that satisfies "N≦T", and "L" is an arbitrary natural number). In other words, the motion patch data 9 may be a set of rectangular partial patch image data.
[0042] FIG. 1-2 is a flowchart showing an example of the flow of processing executed by the feature generating device 1 in this embodiment. Note that the processes described below are merely examples of processes for realizing the method of the present disclosure, and are not limited to these. Furthermore, other steps may be added to the processing described below, or some steps may be omitted (deleted) from the processing described below.
[0043] First, the motion sequence acquisition unit 4 of the feature generation device 1 executes a motion sequence acquisition process (P110). In the motion sequence acquisition process, for example, the motion sequence acquisition unit 4 acquires skeleton data of the first “N” frames from the motion sequence data 8. The number of frames acquired in the motion sequence acquisition process may be referred to as the first number of frames.
[0044] Here, "acquisition" of data can include not only acquisition (input) of data within the device itself, but also, for example, input of data (internal input) from other functional units within the device (e.g., a memory unit not shown), input of data (external input) or reception (external reception) from a device other than the device itself (external device), etc.
[0045] Then, the body part dividing unit 5 executes the body part dividing process (P120). In the body part division process, the body part division unit 5 classifies each joint point of the skeleton data into set parts (for example, five parts: torso including the head, left arm, right arm, left leg, and right leg) based on, for example, marker information recorded for each joint point.
[0046] It should be noted that the settable parts are not limited to the above five parts. For example, two parts, the right hand and the left hand, may be added to the above five parts. In other words, the settable parts may include parts that include many joint points that represent movements that are of interest. This makes it possible to generate partial patches that correspond to movements that are of particular interest, and to generate features that are effective for any part.
[0047] Furthermore, it may be possible to set a tail or the like, which does not exist in the human body, as a settable part. This makes it possible to generate effective partial patches for the motion models of animals other than the human body, and to generate motion patch data 9 that reflects the motion sequence data 8 of the animal.
[0048] Once each joint point has been classified, the joint point placement unit 6 executes a joint point unfolding process (P130). In the joint point unfolding process, the joint point placement unit 6, for example, one-dimensionally arranges joint points classified as a certain part (e.g., the right arm) in the skeleton data for the "first" frame, based on the distance from the torso, which is considered to be the center of gravity of the skeleton model. For example, in the case of an arm, the joint points are arranged in the following order: upper chest → shoulder → upper arm → lower arm → hand. When arranging in one dimension, for example, the relative positions of each joint point are the relative distances between joint points in one motion sequence for the same person, so by arranging each joint point based on the distance from the torso, where the center of gravity of the human body is located, the spatial structure of the skeleton can be maintained.
[0049] The joint point placement unit 6 determines whether or not the joint point unfolding process has been executed for a predetermined number of frames (for example, "N") of skeleton data (P140). If it is determined that there are any frames left to which the process has not been performed (P140: NO), the joint point placement unit 6 executes the joint point unfolding process on the skeleton data of the next frame (P130).
[0050] FIG. 1-3 shows an example of the joint point unfolding process. The upper left of the figure shows skeleton data for a certain frame "i." In this skeleton data, each joint point is assigned a number from "0" to "21" to identify it. Each joint point is also associated with position information (x0, y0, x0), etc., expressed in a Cartesian coordinate system.
[0051] For example, if the five joint points with joint point numbers "0," "1," "4," "7," and "10" are classified as the right leg, the right leg joint points in frame "i" are expanded one-dimensionally and arranged in the joint point expansion process, as shown in the upper right side of the figure. The arranged right leg joint points are called, for example, a "right side joint point series."
[0052] Below the skeleton data for frame "i," the skeleton data for frame "i+1" is shown. In this skeleton data, the position information of each joint point expressed in a Cartesian coordinate system changes as the human body moves, but the relative distance between joint points remains almost constant. Therefore, in the joint point unfolding process, the right-side joint point series for frame "i+1" is arranged as shown in the middle right of the figure.
[0053] By arranging the right-side joint point sequences in each of these frames in chronological order along an axis perpendicular to the axis on which the joint points are deployed, a two-dimensional right-side joint point sequence from frame “i” to frame “i+(N-1)” is generated.
[0054] Returning to FIG. 1-2, for example, when it is determined that the joint point unfolding process has been performed on the skeleton data of the "N" frame (P140: YES), the joint point placement unit 6 performs the joint point rearrangement process on the generated two-dimensional predetermined part joint point series (P150).
[0055] In the joint point rearrangement process, the joint point placement unit 6 standardizes the number of joint points in the predetermined part joint point series in each frame to a predetermined number "N". This standardization makes it possible to absorb differences in the placement of joint points due to differences in the target person or the measured joint points, which occur in the joint point deployment process, for example. This makes it possible to handle skeleton data, etc. of people with different physiques, etc., in a unified manner. The position information for each interpolated orthogonal coordinate system at each joint point may be obtained, for example, by linear interpolation based on the position information for each orthogonal coordinate system at each joint point before interpolation. In this case, the predetermined number may be the same as the number of first frames.
[0056] In the joint point rearrangement process, the joint point arrangement unit 6 may, for example, standardize the number of joint points to a predetermined number "L" (where "L≠N" is a natural number). In this case, the predetermined number may be a number different from the first number of frames.
[0057] After standardizing the number of joint points in the predetermined part joint point series in each frame to "N", the joint point placement unit 6 calculates the average and variance of the position information for each Cartesian coordinate system at each joint point for the entire two-dimensional predetermined part joint point series, and performs z-score normalization using these average and variance values, for example. As a result, the position information for each Cartesian coordinate system at each of the interpolated "N x N" joint points is normalized between "0" and "1".
[0058] The joint point arrangement unit 6 may normalize the position information for each of the interpolated "N×L" joint points in the Cartesian coordinate system.
[0059] Then, the joint point coordinate-color information conversion unit 7 executes, for example, a joint point position color information conversion process (P160). In the joint point position color information conversion process, the joint point coordinate-color information conversion unit 7 assigns (projects) each of the Cartesian coordinate systems (x, y, z) to each color in the RGB color space (r, g, b) for each of the joint points in the two-dimensional predetermined part joint point series. For example, if the Cartesian coordinate system position information for a certain joint point after interpolation is (x=0, y=0, z=1), the joint point is assigned color information of (r=0, g=0, b=255) in the joint point position color information conversion process.
[0060] In the joint point position color information conversion process, the Cartesian coordinate system position information is not limited to being projected onto the RGB color space, but may be projected onto a color space such as the HSV color space.
[0061] FIG. 1-4 shows an example of the joint point rearrangement process and the joint point position color information conversion process. The upper part of the figure shows the two-dimensional right-side joint point series from frame "i" to frame "i+(N-1)." In the joint point rearrangement process, the number of joint points for each right-side joint point series in this two-dimensional right-side joint point series is standardized from "5" to "N." In addition, the position information for each joint point in the Cartesian coordinate system is normalized between "0" and "1." That is, regardless of the sequence length "T" or the number of joint points "J" of the motion sequence data 8, for example, three-dimensional feature values (x, y, z) ("0≦x, y, z≦1") of "N×N" points are extracted for each specified part.
[0062] Then, by performing a joint point position color information conversion process on the two-dimensional right-side joint point series that has undergone the joint point rearrangement process, a partial patch, which is a feature that can be handled as an "N x N" pixel RGB image, is generated from the motion sequence data 8.
[0063] 1-2, for example, the partial patch generation unit 2 determines whether partial patches have been generated for all set portions (P170). If it determines that there are portions for which partial patches have not been generated (P170: NO), for example, the partial patch generation unit 2 changes the portion to be processed and executes steps P130 to P160 again.
[0064] If it is determined that partial patches have been generated for each set body part (P170: YES), for example, the motion patch generation unit 3 executes a motion patch generation process (P180). In the motion patch generation process, the motion patch generation unit 3 merges each partial patch in the x-axis direction (horizontal direction) of the image, for example, to generate motion patch data 9 corresponding to the first "N" frames of the motion sequence data 8. In the motion patch generation process, the motion patch generation unit 3 may, for example, merge only the partial patches relating to predetermined parts (for example, only the right leg, left leg, and torso) from among the partial patches, or may merge all partial patches corresponding to the entire skeleton.
[0065] Furthermore, when generating motion patch data corresponding to the first "N" frames of the motion sequence data 8, the motion patch generation unit 3 may, for example, cause the partial patch generation unit 2 to generate motion patch data corresponding to "N" frames following the "N+1" frame of the motion sequence data 8, according to steps P110 to P170. The generated motion patch data may then be merged in the y-axis direction of the image (vertical direction) to expand and output the motion patch data 9 in chronological order.
[0066] Furthermore, in the motion sequence acquisition process, the motion sequence acquisition unit 4 may read skeleton data for "N" frames, starting with any frame. That is, the frame corresponding to the y-axis direction of the image of the partial patch may be any frame within the time window "N" frames of the motion sequence data 8. Note that the motion sequence acquisition unit 4 may also read skeleton data for all frames of the motion sequence data 8 all at once.
[0067] In addition, when the joint points and frames are transposed and arranged in the joint point expansion process, in the motion patch generation process, the motion patch generation unit 3 may merge each partial patch in the y-axis direction of the image, and expand and output them in chronological order in the x-axis direction of the image.
[0068] Once motion patch data corresponding to an arbitrary frame length has been generated, the motion patch generation unit 3 outputs the generated motion patch data (P190). Note that the motion patch generation unit 3 may output the motion patch data as tensor data, or may output the motion patch data as image data. For example, visualized motion patch data may simply be called a "motion patch."
[0069] Here, "output" of data can include not only displaying data on the device itself (display output), but also, for example, outputting data to other functional units on the device itself (internal output), outputting data to a device other than the device itself (external device) (external output) or transmitting data (external transmission), etc.
[0070] <Example of generating a motion patch> Figures 1-5 and 1-6 show examples of motion patch data 9 visualized as RGB images. The left side of each figure shows the rendered motion sequence data 8 and its corresponding text label, and the right side shows the generated motion patch. In this example, the motion patch is an image that has been converted from an RGB image into grayscale. By comparing the generated motion patches, we can observe that different motion sequences result in different motion patches, and visually capture the unique characteristics of each motion in the form of motion patches.
[0071] <Effects of the First Embodiment> According to this embodiment, the control unit of a feature generation device (e.g., the feature generation device 1) acquires a three-dimensional skeletal partial time series (e.g., skeleton data of a first number of frames) from a three-dimensional skeletal motion sequence (e.g., motion sequence data 8), classifies the skeletal joints in the three-dimensional skeletal partial time series by skeletal part (e.g., five parts: torso including head, left arm, right arm, left leg, and right leg), rearranges the skeletal joints for each part based on their distance from a set joint (e.g., torso joint), interpolates the rearranged skeletal joints to a predetermined number (e.g., “N” points) of skeletal joints, standardizes position information (e.g., Cartesian coordinates) of the skeletal joints in the interpolated three-dimensional skeletal partial time series, generates part patches by converting the standardized position information into color information (e.g., RGB color information), and generates features (e.g., motion patches or motion patch data 9) by integrating the part patches generated for each part. With this configuration, the feature generation device can relocate a predetermined number of skeletal joints for each skeletal part from the 3D skeletal movement sequence, and then generate features consisting of partial patches that standardize the position information of the relocated skeletal joints and convert it into color information.
[0072] Furthermore, according to this embodiment, the control unit of the feature generation device generates first features (e.g., motion patches based on the HumanML 3D dataset) based on a first three-dimensional skeletal motion sequence (e.g., motion sequence data in the HumanML 3D dataset) expressed by skeletal joints with a first number of joints (e.g., 22). When the control unit of the feature generation device generates second features (e.g., motion patches based on the KIT-ML dataset) based on a second three-dimensional skeletal motion sequence (e.g., motion sequence data in the KIT-ML dataset) expressed by skeletal joints with a second number of joints different from the first number of joints (e.g., 21), the first features and the second features are features with the same number of dimensions (e.g., predetermined number "16" x predetermined number "16" x number of skeletal parts "5" x number of color information channels "3"). This allows the feature quantity generated by the feature quantity generating device to have a number of dimensions that is independent of the number of joints in the acquired three-dimensional skeletal movement sequence.
[0073] <Second Example> The second embodiment is an embodiment in which, by using the feature generation device 1 described in the first embodiment, cross-modal analysis between, for example, motion sequence data of a human body model and natural language is realized in an information processing device (which may also be called a motion sequence search device or a motion sequence creation device). By realizing cross-modal analysis of human body model motion sequence data and natural language, the information processing device can search for motion sequences based on input text descriptions and interpret actions in text based on specified motion sequences. The contents described in the second embodiment are similarly applicable to any of the other embodiments and other modified examples.
[0074] FIG. 2-1 is a block diagram showing an example of a functional configuration of an information processing device 10 according to an aspect of the present embodiment. The information processing device 10 includes, for example, a feature generating device 1, a text encoder unit 11, an image encoder unit 12, and a linear classification unit 13. These are, for example, functional units (functional blocks) included in a control unit (control device) (not shown) of the information processing device 10. The control unit may also be called a processing unit (processing device).
[0075] The configuration and basic operation of the information processing device 10 may be, for example, based on the CLIP framework, which is a cross-modal model of language and images. For CLIP, see Alec Radford, et al., "Learning transferable visual models from natural language supervision," 2021.
[0076] The text encoder unit 11 has a function of converting (encoding) the caption data 19 input to the information processing device 10 into "M" feature vectors, for example. The text encoder unit 11 is configured with a text encoder model such as a Transformer encoder or DistilBERT. For details about DistilBERT, see Victor Sanh, et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” 2019.
[0077] The image encoder unit 12 has a function of converting (encoding) the motion patch data 9 generated by the feature generator 1 based on the motion sequence data 8 input to the information processing device 10 into "M" feature vectors, for example. The image encoder unit 12 is configured with an image encoder model such as ViT (Vision Transformer). For details about ViT, see Alexey Dosovitskiy, et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” 2021. The image encoder unit 12 may also be a CNN model such as Residual Neural Networks (ResNet).
[0078] In the information processing device 10, the motion patch data 9 is used for input to the image encoder unit 12. Therefore, for example, the image encoder unit 12 can be trained in advance using an image database such as ImageNet.
[0079] For example, the scale of image and caption text pairs used to train the CLIP image language model is 400 million pairs, but the motion sequence and caption text pairs are only 14,616 motions in the relatively large HumanML3D dataset. This allows the image encoder unit 12 to be optimized more accurately and quickly than by directly calculating feature vectors based on the motion sequence data 8, which is difficult to collect in large quantities.
[0080] Furthermore, if the motion sequence data database changes, the number of joint points "J" of the skeleton model for each frame may change. Therefore, even when learning pairs of motion sequences and caption text across multiple databases, generating motion patch data 9 independent of the number of joint points "J" prevents the image encoder unit 12 from becoming complicated, such as by preparing an encoder corresponding to each number of joint points "J."
[0081] The linear classification unit 13 has a function of calculating the similarity relationship between, for example, an "M"-dimensional feature vector (referred to as a "caption feature vector") extracted based on the caption data 19 and an "M"-dimensional feature vector (referred to as a "motion feature vector") extracted based on the motion patch data 9. The linear classification unit 13 may also be called a similarity matrix unit. For example, when a set of motion sequences M and captions T is given, the linear classifier 13 classifies the caption feature vectors {t1, t2, . . . , t M} and the motion feature vectors {m1,m2,···,m M The text encoder 11 and the image encoder 12 are subjected to contrastive learning so that a high similarity is output between the caption feature vector and the motion feature vector for unrelated objects, and a low similarity is output between the caption feature vector and the motion feature vector for unrelated objects.
[0082] FIG. 2-2 is a flowchart showing an example of the flow of the contrastive learning process executed by the information processing device 10 in this embodiment.
[0083] First, the control unit (not shown) of the information processing device 10 acquires the caption data 19 and the motion patch data 9 associated (paired) for learning (P210-P220).
[0084] Then, for example, the feature generating device 1 of the information processing device 10 generates the motion patch data 9 in accordance with, for example, FIG. 1-2 based on the acquired motion patch data 9 (P230).
[0085] Then, for example, the image encoder unit 12 of the information processing device 10 calculates a motion feature vector based on the generated motion patch data 9 (P240).
[0086] For example, if the image encoder unit 12 is a ViT model, each partial patch ("N x N" pixel RGB image) of the motion patch data 9 can be input in correspondence with each patch obtained by dividing the input image in ViT. That is, for example, if ViT-B / 16, which has been pre-trained with a patch size of "16", is used for the image encoder unit 12, it is recommended to set the size of each partial patch to "16 x 16" in the motion patch data generation process (P230). This makes it possible to easily calculate an effective motion feature vector based on the results of pre-training ViT.
[0087] For example, information for identifying a class (a [class] token) may be embedded in the motion feature vector. As a result, the output of the [class] token may project a motion representation in a multimodal latent space.
[0088] Furthermore, for example, the text encoder unit 11 of the information processing device 10 calculates a caption feature vector based on the caption data 19 (P250).
[0089] For example, by repeating steps P210 to P250 to calculate "B" pairs (mini-batches) of motion feature vectors and caption feature vectors, the linear classifier 13 calculates the similarity between each motion feature vector and caption feature vector using, for example, cosine similarity so as to calculate the optimal similarity for the "B x B" pairs. Then, the linear classifier 13 calculates a symmetric cross-entropy loss so as to maximize the diagonal components of a similarity matrix based on each motion feature vector and caption feature vector, and optimizes the parameters (weights) of the text encoder 11 and the image encoder 12 based on the symmetric cross-entropy loss (P260).
[0090] Then, the processing unit of the information processing device 10 determines whether to end the contrastive learning (P270). For example, the processing unit of the information processing device 10 may determine to end the contrastive learning process when the contrastive cross-entropy loss is equal to or less than a threshold (P270: YES). It can be said that the linear classification unit 13 forms a multimodal latent space relating to language (caption T) and movement (motion sequence M) through the caption feature vectors and motion feature vectors calculated in the text encoder unit 11 and image encoder unit 12 of the contrastive learning process.
[0091] The processing unit of the information processing device 10 is capable of, for example, the following processes in the inference process after the contrastive learning process. ·(A): Retrieval of motion sequence data based on caption data. For example, a certain caption T is input to the text encoder unit 11, and a candidate set of [class] tokens is input to the image encoder unit 12. The linear classification unit 13 calculates the similarity between the resulting caption feature vector and each motion feature vector, and outputs the [class] token label with the maximum similarity as the inference result. ·(B): Retrieval of caption data based on motion sequence data. For example, a certain motion sequence M is converted into motion patch data and input to the image encoder unit 12. A template sentence and a candidate set of [object] tokens are also input to the text encoder unit 11. The linear classification unit 13 calculates the similarity between the resulting motion feature vector and each caption feature vector, and outputs the [object] token label with the maximum similarity as the inference result.
[0092] The information processing device 10 is based on the CLIP framework and therefore inherits the zero-shot learning and few-shot learning capabilities that are characteristic of the CLIP framework. That is, the information processing device 10 has the ability to generate an appropriate caption T for an unlearned motion sequence M, and the ability to generate an appropriate motion sequence M from an unlearned caption T.
[0093] <Example of searching for motion sequence data based on caption data> Figure 2-3 shows an example of a search result for motion sequence data based on caption data. In the upper part of the figure, when the caption T shown in the left caption box, "The person does 2 cartwheels," is input as caption data 19, the motion sequences M that are determined to have the top three similarities by the linear classification unit 13 are rendered and shown on the right side of the caption box. Note that below the rendered image of each motion sequence M, the caption T that was used as a pair in the contrastive learning is shown. For example, in the example at the top of the figure, the motion sequence M with the highest similarity matches the caption T in the caption box on the left, indicating that contrastive learning was performed appropriately.
[0094] Additionally, in the bottom row of the figure, as in the top row, the top two motion sequences M in terms of similarity match the caption T in the caption box on the left, demonstrating that contrastive learning was performed appropriately even for complex caption T.
[0095] Figure 2-4 shows an example of experimental results for zero-shot learning. The caption T shown in the left caption box, "A person is running in a circle," is data that was not included in the control training pair. Even for this untrained caption T, it was shown that motion sequences M similar to caption T can be appropriately found as the top three search results in terms of similarity.
[0096] <Quantitative performance evaluation> To evaluate the performance of the motion language model in the information processing device 10, benchmark evaluations were performed on the HumanML3D dataset and the KIT-ML dataset for text (caption T) to motion (motion sequence M) retrieval and motion to text retrieval. Figure 2-5 shows a list of the benchmark evaluation results for each dataset. The prior arts used for comparison were TEMOS (Mathis Petrovich, et al. “Temos: Generating diverse human motions from textual descriptions.”, 2022.), T2M (Chuan Guo, et al. “Generating diverse and natural 3D human motions from text.”, 2022.), and TMR (Mathis Petrovich, et al. “Tmr: Text-to-motion retrieval using contrastive 3D human motion synthesis.”, 2023.). We also compared our method with a scratch training method in the image encoder 12, which does not use pre-trained ViT weights.
[0097] In each table, the recall at rank "k" indicates the percentage of instances with the correct label in the top "k" similarity results, with higher values indicating better performance. Additionally, the median rank (MedR) is calculated, with lower values indicating better performance.
[0098] These tables show that the performance of the information processing device 10 consistently outperforms previous studies of varying difficulty across all evaluation sets. This demonstrates that the information processing device 10 is able to capture the nuanced nature of motion descriptions. This significant performance improvement is likely due to several factors. First, the motion sequence data is converted into motion patches to properly capture the spatiotemporal motion expression. Second, by utilizing ViT in the image encoder 12 and the pre-trained weights of ViT, transfer learning is performed via the motion patches, enabling the feature vectors of the motion sequence to be calculated effectively.
[0099] <Evaluation when using multiple datasets> As described above, in the information processing device 10, the motion sequence data 8 is not directly input to the image encoder unit 12, but is converted into motion patch data 9 by the feature generator 1. Therefore, by converting motion sequences made up of different skeletal structures into a unified representation (motion patch), it is possible to handle multiple data sets simultaneously. For example, the motion sequences in the HumanML3D dataset follow the skeletal structure of SMPL, which has 22 joints, while the poses in the KIT-ML dataset have 21 joints, which differs from SMPL in terms of the skeleton structure. Existing methods cannot simultaneously handle these two datasets because the dimensions of the input vector to the image encoder 12 depend on the skeleton structure, resulting in different dimensionality. However, our method converts motion sequences into motion patches that are independent of the skeleton structure (the number of joints "J"). Therefore, even if the skeleton structures are different, the image encoder 12 and text encoder 11 trained (optimized) on one dataset can be used on another dataset, or transfer learning can be performed.
[0100] Figures 2-6 show performance indicators for zero-shot learning, in which the text encoder 11 and image encoder 12 trained using the HumanML3D dataset are directly used for each search task in the KIT-ML dataset, and performance indicators for fine-tuning the KIT-ML dataset using the text encoder 11 and image encoder 12 trained using the HumanML3D dataset for transfer learning. For comparison, the performance indicators of an existing method trained using the KIT-ML dataset are also shown for reference.
[0101] Although the performance index of zero-shot training using the HumanML3D dataset is lower than that of training using the KIT-ML dataset, the text-to-motion search results show performance comparable to that of existing methods. In other words, it can be said that the information processing device 10 can operate appropriately even with datasets different from the dataset used for training.
[0102] Furthermore, the performance metrics obtained by fine-tuning using the KIT-ML dataset are superior to those of other methods, demonstrating the effectiveness of our approach in improving performance on small datasets (e.g., the KIT-ML dataset) by pre-training the text encoder 11 and image encoder 12 on large datasets (e.g., the HumanML3D dataset).
[0103] <Effects of the second embodiment> According to this embodiment, an information processing device capable of mutual conversion between a three-dimensional skeletal motion sequence (e.g., motion sequence data) and a sentence sequence (e.g., caption data) includes a feature generation device 1, an image encoding unit (e.g., image encoder 12), and a sentence encoding unit (e.g., text encoder 11). The image encoding unit calculates image features (e.g., motion feature vectors) based on features (e.g., motion patches) generated from the three-dimensional skeletal motion sequence by the feature extraction device, and the sentence encoding unit calculates sentence features (e.g., caption feature vectors) based on the sentence sequence. Furthermore, a control unit of the information processing device optimizes (e.g., contrastive learning) the image encoding unit and the sentence encoding unit so as to be able to calculate an appropriate combination of a three-dimensional skeletal motion sequence and a sentence sequence based on the image features and the sentence features. According to this, the image encoding unit of the information processing device can calculate image features from a 3D skeletal movement sequence using the features generated by the feature extraction device, thereby enabling the information processing device to more accurately calculate an appropriate combination of a 3D skeletal movement sequence and a sentence sequence through optimization.
[0104] <Second Modification Example (1)> In the above embodiment, an example has been described in which the information processing device 10 executes a search process between motion and text, such as searching for motion sequence data based on caption data, or searching for caption data based on motion sequence data, but the present invention is not limited to this. For example, the information processing device 10 after contrastive learning may generate motion sequence data based on arbitrary skeleton data and caption data, in which the skeleton data performs the action described in the caption data.
[0105] When generating motion sequence data, for example, the CLIP framework realized by the information processing device 10 can be combined with GLIDE (Alex Nichol, et al. “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models”, 2021.), a text-to-image generation model based on a diffusion model, to generate an image using motion patch data generated from skeleton data as a reference image and caption data as a prompt, and the generated image can be treated as motion patch data and reverse-converted into motion sequence data.
[0106] This makes it possible to easily generate motion sequence data for animals and the like that can be represented by any skeleton data.
[0107] <Third Example> The third embodiment is an embodiment relating to the use of a server 100 that includes the information processing device 10 described in the second embodiment. The contents described in the third embodiment can be similarly applied to any of the other embodiments and other modified examples.
[0108] In the following embodiments, a service for realizing a search for a motion sequence and a corresponding caption, or a search for a caption and a corresponding motion sequence, is referred to as a "motion search service" as an example.
[0109] In the following description, the user of terminal 200A communicating with server 100 will be referred to as "user AA," the user of terminal 200B as "user BB," the user of terminal 200C as "user CC," . . .
[0110] <System configuration> 3-1 is a diagram illustrating an example of a system configuration of a communication system 1000 according to this embodiment. In the communication system 1000, for example, a server 100 is connected to one or more terminals 200 (terminal 200A, terminal 200B, terminal 200C, etc.) via a network 300.
[0111] The server 100 has a function of providing, for example, a motion search service to a terminal 200 or the like owned by a user via a network 300 . The server 100 can also be expressed as a motion search service server, etc. In this embodiment, the user of the server 100 is assumed to be a company that provides a motion search service (a motion search service provider), for example. Note that the number of servers 100 and the number of terminals 200 connected to the network 300 are not limited to those described above. In this embodiment, the "motion search service" is a service provided by a company or other business that provides motion search services (server 100), and may be provided to a user (user's terminal 200), for example.
[0112] The terminal 200 (terminal 200A, terminal 200B, terminal 200C, etc.) may be any information processing terminal capable of implementing the functions described in each embodiment. Examples of the terminal 200 include a smartphone, a mobile phone (feature phone), a computer (including, but not limited to, a server, desktop, laptop, tablet, etc.), a media computer platform (including, but not limited to, a cable or satellite set-top box, or digital video recorder), a handheld computer device (including, but not limited to, a PDA (personal digital assistant), an email client, etc.), a wearable device (glasses-type device, watch-type device, etc.), a VR (Virtual Reality) terminal, a smart speaker (a device for voice recognition), or another type of computer or communication platform. The terminal 200 may also be referred to as an information processing terminal.
[0113] For example, the configurations of the terminals 200A, 200B, and 200C can be the same. Furthermore, as necessary, the terminal used by the user X may be expressed as the terminal 200X, and the user information in a predetermined service associated with the user X or the terminal 200X may or may not be expressed as the user information X. The user information is information of a user associated with an account used by the user in a predetermined service. The user information includes, but is not limited to, information associated with a user, such as the user's name, an icon image of the user, the user's age, the user's gender, the user's address, the user's hobbies and interests, and a user identifier, which is input by the user or assigned by the predetermined service, and may be any one of these, or a combination thereof, or may not be the same.
[0114] The network 300 serves to connect the devices constituting the communication system 1000. In other words, the network 300 refers to a communication network that provides connection paths so that the above-mentioned various devices can connect and transmit and receive data.
[0115] One or more portions of network 300 may or may not be a wired or wireless network. Network 300 may include, by way of example, an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular network, integrated service digital networks (ISDN), wireless LAN, long term evolution (LTE), code division multiple access (CDMA), Bluetooth, satellite communications, etc., or a combination of two or more thereof. Network 300 may include one or more networks 30.
[0116] The server 100 (which is not limited to an example of a server, information processing device, or information management device) has a function of providing a predetermined service (in this embodiment, a motion search service) to the terminal 200, etc. The server 100 may be any information processing device that can realize the functions described in each embodiment. The server 100 includes, for example, a server device, a computer (e.g., a desktop, laptop, tablet, etc.), a media computer platform (e.g., a cable or satellite set-top box, digital video recorder), a handheld computer device (e.g., a PDA, email client, etc.), or other types of computers or communication platforms. The server 100 may also be referred to as an information processing device. When there is no need to distinguish between the server 100 and the terminal 200, the server 100 and the terminal 200 may or may not each be referred to as an information processing device.
[0117] [Hardware (HW) configuration of each device] The hardware configuration of each device included in the communication system 1000 will be described.
[0118] (1) Device hardware configuration FIG. 3-1 shows an example of the hardware configuration of the terminal 200. The terminal 200 includes, for example, a control unit 210 (CPU: central processing unit), a storage unit 280, a communication I / F 220 (interface), an input / output unit 230, and a clock unit 290. The HW components of the terminal 200 are connected to each other, for example, via a bus B. It is not essential that the HW configuration of the terminal 200 includes all of the components. For example, the terminal 200 may or may not be configured such that individual components or multiple components are removable.
[0119] The communication I / F 220 transmits and receives various data via the network 300. The communication may be performed either wired or wirelessly, and any communication protocol may be used as long as mutual communication is possible. The communication I / F 220 has a function of communicating with various devices, such as the server 100, via the network 300. The communication I / F 220 transmits various data to various devices, such as the server 100, in accordance with instructions from the control unit 210. The communication I / F 220 also receives various data transmitted from various devices, such as the server 100, and transmits it to the control unit 210. The communication I / F 220 may also be simply referred to as a communication unit. When the communication I / F 220 is configured as a physically structured circuit, it may also be referred to as a communication circuit.
[0120] The input / output unit 230 includes a device for inputting various operations to the terminal 200, a device for outputting processing results processed by the terminal 200, etc. The input / output unit 230 may be an integrated unit of the input unit and the output unit, or may be separate units, or may not be the same.
[0121] The input unit is realized by any one or a combination of all types of devices that can accept input from a user and transmit information related to the input to the control unit 210. Examples of the input unit include hardware keys such as a touch panel, a touch display, and a keyboard, a pointing device such as a mouse, a camera (for inputting operations via moving images), and a microphone (for inputting operations by voice).
[0122] The output unit is realized by any one or a combination of all kinds of devices that can output the processing results processed by the control unit 210. Examples of the output unit include a touch panel, a touch display, a speaker (audio output), a lens (for example, 3D (three dimensions) output or hologram output), a printer, etc.
[0123] Although this is merely an example, the input / output unit 230 includes a display unit 240, a sound input unit 250, a sound output unit 260, and an imaging unit 270, for example.
[0124] The display unit 240 is realized by any one of all types of devices or a combination thereof that can display according to the display data written to the frame buffer. Examples of the display unit 240 include a touch panel, a touch display, a monitor (e.g., a liquid crystal display or an organic electroluminescence display (OLED)), a head mounted display (HDM), a projection mapping, a hologram, and a device that can display images, text information, etc. in air (which may or may not be a vacuum). Note that these display units 240 may or may not be capable of displaying display data in 3D.
[0125] The sound input unit 250 is used to input sound data (including voice data; the same applies below.) The sound input unit 250 includes a microphone and the like. The sound output unit 260 is used to output sound data and includes a speaker and the like. The imaging unit 270 is used to acquire image data (including still image data and moving image data; the same applies below.) The imaging unit 270 includes a camera and the like.
[0126] When the input / output unit 230 is a touch panel, the input / output unit 230 and the display unit 240 may be disposed facing each other and have approximately the same size and shape.
[0127] The clock unit 290 is a built-in clock of the terminal 200, and outputs time information (timekeeping information). The clock unit 290 is configured, for example, with a clock that uses a crystal oscillator. The clock unit 290 can also be expressed, for example, as a timekeeping unit or a time information detection unit.
[0128] The clock unit 290 may or may not have a clock that conforms to the NITZ (Network Identity and Time Zone) standard or the like.
[0129] The control unit 210 has a circuit physically structured to execute the functions realized by the code or instructions included in the program, and is realized by, for example, a data processing device built into hardware. Therefore, the control unit 210 may or may not be expressed as a control circuit.
[0130] The control unit 210 includes, for example, a central processing unit (CPU), a microprocessor, a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), or a field programmable gate array (FPGA).
[0131] The storage unit 280 has a function of storing various programs and various data required for the operation of the terminal 200. The storage unit 280 includes various storage media such as a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a random access memory (RAM), and a read only memory (ROM), for example. The storage unit 280 may or may not be expressed as a memory.
[0132] Terminal 200 stores program P in storage unit 280, and by executing this program P, control unit 210 executes the processing of each unit included in control unit 210. In other words, program P stored in storage unit 280 causes terminal 200 to realize each function executed by control unit 210. Furthermore, this program P may or may not be expressed as a program module.
[0133] (2) Server hardware configuration FIG. 3-1 shows an example of the hardware configuration of the server 100. The server 100 includes, for example, a control unit 110 (CPU), a storage unit 150, a communication I / F 140 (interface), an input / output unit 120, a display unit 130, and a clock unit 190. The components of the HW of the server 100 are connected to each other, for example, via a bus B. Note that the HW of the server 100 does not necessarily have to include all of the components as the configuration of the HW of the server 100. For example, the HW of the server 100 may or may not be configured so that individual components or multiple components can be removed.
[0134] The control unit 110 has circuits physically structured to execute functions realized by the codes or instructions contained in the program, and is realized, for example, by a data processing device built into hardware.
[0135] The control unit 110 is typically a central processing unit (CPU), but may also be a microprocessor, a processor core, a multiprocessor, an ASIC, or an FPGA. In the present disclosure, the control unit 110 is not limited to these.
[0136] The storage unit 150 has a function of storing various programs and various data required for the operation of the server 100. The storage unit 150 is realized by various storage media such as an HDD, an SSD, or a flash memory. However, in the present disclosure, the storage unit 150 is not limited to these. Furthermore, the storage unit 150 may or may not be expressed as a memory.
[0137] The communication I / F 140 transmits and receives various data via the network 300. The communication may be performed either wired or wirelessly, and any communication protocol may be used as long as mutual communication is possible. The communication I / F 140 has a function of communicating with various devices such as the terminal 200 via the network 300. The communication I / F 140 transmits various data to various devices such as the terminal 200 in accordance with instructions from the control unit 110. The communication I / F 140 also receives various data transmitted from various devices such as the terminal 200 and transmits it to the control unit 110. The communication I / F 140 may also be simply referred to as a communication unit. When the communication I / F 140 is configured as a physically structured circuit, it may also be referred to as a communication circuit.
[0138] The input / output unit 120 includes a device for inputting various operations to the server 100, a device for outputting processing results processed by the server 100, etc. The input / output unit 120 may be an integrated unit having an input unit and an output unit, or may be separate units having an input unit and an output unit, or may not be so.
[0139] The input unit is realized by any one of or a combination of all types of devices that can accept input from a user and transmit information related to the input to the control unit 110. The input unit is typically realized by hardware keys such as a keyboard or a pointing device such as a mouse. Note that the input unit may or may not include, for example, a touch panel, a camera (for operation input via moving images), or a microphone (for operation input by voice).
[0140] The output unit is realized by any one or a combination of all kinds of devices that can output the processing results processed by the control unit 110. Examples of the output unit include a touch panel, a touch display, a speaker (sound output), a lens (for example, 3D (three dimensions) output or hologram output), a printer, etc.
[0141] By way of example only, the input / output unit 120 includes a display unit 130, for example.
[0142] The display unit 130 is realized by a display or the like. The display is typically realized by a monitor (for example, a liquid crystal display or an organic electroluminescence display (OLED)). The display may or may not be a head-mounted display (HDM) or the like. These displays may or may not be capable of displaying display data in 3D. In the present disclosure, the display is not limited to these.
[0143] The clock unit 190 is a built-in clock of the server 100, and outputs time information (timekeeping information). The clock unit 190 is configured to include, for example, an RTC (Real Time Clock) as a hardware clock, a system clock, etc. The clock unit 190 can also be expressed as, for example, a timekeeping unit or a time information detection unit.
[0144] (3) Other Server 100 stores program P in storage unit 150, and by executing this program P, control unit 110 executes the processing of each unit included in control unit 110. In other words, program P stored in storage unit 150 causes server 100 to realize each function executed by control unit 110. This program P may or may not be expressed as a program module. The same applies to other devices.
[0145] In each embodiment of the present disclosure, the description will be given assuming that the CPU of the terminal 200 and / or the server 100 executes the program P to realize the present invention. The same applies to other devices.
[0146] The control unit 210 of the terminal 200 and / or the control unit 110 of the server 100 may or may not implement each process using not only a CPU having a control circuit but also a logic circuit (hardware) formed in an integrated circuit (IC (Integrated Circuit) chip, LSI (Large Scale Integration)), etc., or a dedicated circuit. These circuits may be implemented by one or more integrated circuits, and multiple processes shown in each embodiment may or may not be implemented by a single integrated circuit. LSIs may also be referred to as VLSIs, super LSIs, ultra LSIs, etc., depending on the degree of integration. Therefore, the control unit 21 may or may not be expressed as a control circuit. The same applies to other devices.
[0147] Furthermore, the program P (e.g., a software program, a computer program, or a program module) of each embodiment of the present disclosure may or may not be provided in a state stored in a computer-readable storage medium. The storage medium can store the program P in a "non-transitory tangible medium." The program P may or may not be intended to realize part of the functions of each embodiment of the present disclosure. Furthermore, the program P may or may not be a so-called differential file (differential program) that can realize the functions of each embodiment of the present disclosure in combination with a program P already recorded on a storage medium.
[0148] The storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards, or drives, any other suitable storage media, or any suitable combination of two or more of these. The storage medium may be volatile, nonvolatile, or a combination of volatile and nonvolatile, where appropriate. The storage medium is not limited to these examples and may be any device or medium capable of storing the program P. Furthermore, the storage medium may or may not be referred to as memory.
[0149] The server 100 and / or the terminal 200 can implement the functions of the multiple functional units shown in each embodiment by reading out the program P stored in a storage medium and executing the read out program P. The same applies to other devices.
[0150] Furthermore, the program P of the present disclosure may or may not be provided to the server 100 and / or the terminal 200 via any transmission medium capable of transmitting a program (such as a communication network or broadcast waves). The server 100 and / or the terminal 200 executes the program P downloaded via the Internet or the like, for example, to realize the functions of the multiple functional units shown in each embodiment. The same applies to other devices.
[0151] In addition, each embodiment of the present disclosure may also be realized in the form of a data signal in which the program P is embodied by electronic transmission. At least a part of the processing in the server 100 and / or the terminal 200 may or may not be realized by cloud computing configured by one or more computers. At least a part or all of the processing in terminal 200 may or may not be performed by server 100. In this case, at least a part or all of the processing of each functional unit of control unit 210 of terminal 200 may or may not be performed by server 100. At least a part or all of the processing in the server 100 may or may not be performed by the terminal 200. In this case, at least a part or all of the processing of each functional unit of the control unit 110 of the server 100 may or may not be performed by the terminal 200. Unless explicitly stated otherwise, the judgment configuration in the embodiments of the present disclosure is not essential, and a predetermined process may or may not be performed when the judgment condition is met, or when the judgment condition is not met.
[0152] The program of the present disclosure is implemented using, for example, a scripting language such as ActionScript or JavaScript (registered trademark), a compiler language such as Objective-C or Java (registered trademark), or a markup language such as HTML Living Standard.
[0153] [Functional configuration of each device] (1) Server functional configuration FIG. 3B is a diagram showing an example of functions realized by the control unit 110 of the server 100 in this embodiment. The control unit 110 includes, for example, the information processing device 10 and an application management processing unit 111 as functional units.
[0154] The information processing device 10 includes, for example, the functional units shown in FIG. 2-1. The information processing device 10 may also be called an information processing unit or a text-motion conversion unit. The server 100 may be a server system including the information processing device 10 separate from the server 100. The application management processing unit 111 also has a function of executing application management processing in accordance with an application management processing program 151 stored in the storage unit 150, for example.
[0155] FIG. 3-3 is a diagram illustrating an example of information stored in the storage unit 150 of the server 100 in this embodiment. The storage unit 15 stores, for example, an application management processing program 151 that is executed as application management processing, and token registration data 153.
[0156] The token registration data 153 is registration data relating to an access token (which may also be called an API token or an API key) for verifying authorization for use of the information processing device 10, and an example of the data configuration is shown in FIG. 3-4. The token registration data 153 stores, for example, an access token, an expiration date, and other token information in association with each other.
[0157] The access token is a token (for example, a character string) for providing authentication and authorization to the motion search service. For example, the server 100 sets a unique value (inherent value) for each access token and stores the token.
[0158] In the expiration date, for example, the expiration date of the token is set and stored.
[0159] Other token information may store, for example, information about the terminal 200 that instructed the generation of the token, or information about the user of the terminal 200. Furthermore, if the token has an authorization type (for example, it is possible to search for a motion sequence from text, but it is not possible to search for text from a motion sequence), the authorization type may be associated with the token and stored.
[0160] (2) Functional configuration of the terminal The control unit 210 of the terminal 200 includes, as a functional unit, an application processing unit 211 for executing application processing in accordance with an application processing program 281 stored in the storage unit 280, for example.
[0161] In addition, the storage unit 280 of the terminal 200 stores, for example, an application processing program 281 to be executed as application processing, and a terminal identification ID 283 (for example, an IMEI (International Mobile Equipment Identifier)) for identifying the terminal 20 or the user of the terminal 20.
[0162] <Processing> 3-5 is a flowchart showing an example of the flow of processing executed by each device in this embodiment. From the left, this figure shows an example of processing executed by the control unit 210 of the terminal 200A of user AA, and an example of processing executed by the control unit 110 of the server 100.
[0163] First, the server 100 executes a contrastive learning process (S110). In the contrast learning process, the server 100 causes the information processing device 10 to execute a contrast learning process similar to that shown in FIG. 2-2 based on, for example, a motion sequence database and a caption text database (not shown) stored in the storage unit 150.
[0164] The server 100 may read trained model information that has been optimized by performing a contrastive learning process in advance, and may set parameters (structures and weights) of each unit of the information processing device 10 based on the model information.
[0165] For example, based on an input to the input / output unit 230 of the terminal 200A (for example, a user input (such as an operation input or sound input by the user), the same may be applied hereinafter), the terminal 200A transmits to the server 100 access token generation request information for requesting an access token that enables a search for a motion sequence based on a caption in the information processing device 10 of the server 100 or a search for a caption based on a motion sequence (A110). The access token generation request information may include, for example, information for identifying the terminal 200A or the user of the terminal 200A.
[0166] When the server 100 receives the access token generation request information from the terminal 200A, the server 100 executes an access token generation process (S120). In the access token generation process, the server 100 calculates the expiration date of the access token to be issued, for example, based on the set access token validity period. Then, the server 100 generates an access token by hashing, for example, time information and information for identifying the terminal 200A or the user of the terminal 200A. Note that the method for generating the access token may be other methods (for example, generation by a random number generator using time information as an entropy source) as long as the method is a method that ensures that each access token has a unique value.
[0167] Note that, prior to the access token generation process, the server 100 may execute an access token payment process to settle the price for issuing the access token. Then, for example, if the payment process is performed in the terminal 200A and the payment is successful, the server 100 may execute the access token generation process.
[0168] In the access token payment process, for example, the higher the price for issuing the access token, the longer the validity period of the access token may be. Furthermore, a fee may be charged for enabling the search for a motion sequence based on a caption and the search for a caption based on a motion sequence. The number of times an access token can be used may be set. For example, the more times an access token can be used, the higher the price for issuing the access token may be.
[0169] Then, the server 100 stores the generated access token, the expiration date, and other token information in the token registration data 153 in association with each other, for example.
[0170] Then, the server 100 transmits access token information including the generated access token to the terminal 200A (S130).
[0171] When the terminal 200A receives the access token information from the server 100, the terminal 200A stores the received access token in the storage unit 280 (A120), for example.
[0172] For example, when the terminal 200A receives a caption indicating a motion based on a user input, the terminal 200A transmits motion search request information including the received caption data and an access token to the server 100 (A130).
[0173] When receiving motion search request information from terminal 200A, server 100 performs, for example, an access token matching process to match the access token included in the received motion search request information with the access token stored in token registration data 153 (S140).
[0174] If it is determined that a valid access token matching the received access token is stored in the token registration data 153 (S140: Approved), the server 100 executes a motion search process (S150). In the motion search process, the server 100 inputs the received caption data into the information processing device 10, and causes the information processing device 10 to execute an inference process for motion sequence data based on the caption data. Then, the server 100 transmits motion sequence information including the top k motion sequences in terms of similarity (where "k" is any natural number) to the terminal 200A (S160).
[0175] If it is determined that the token registration data 153 does not store an access token that is within its expiration date and matches the received access token, or that the access token has expired (S140: Disapproved), the server 100 skips steps S150 and S160, for example.
[0176] For example, when determining S140: Disapproval, the server 100 may transmit necessary token notification information to the terminal 200A notifying that an access token is required to execute the motion search process. For example, the terminal 200A may output the received necessary token notification information (for example, display it on the display unit 240).
[0177] When it is determined that the motion sequence information has been received from the server 100 (A140: YES), the terminal 200A outputs the received motion sequence information (for example, displays it on the display unit 240) (A150).
[0178] When it is determined that motion sequence information has not been received from the server 100 (A140: NO), the terminal 200A determines whether to end the process (A190). If it is determined that the process should be continued (A190: NO), the terminal 200A returns the process to A110, for example. On the other hand, if it is determined that the process should be ended (A190: YES), the terminal 200A ends the process.
[0179] The server 100 also determines whether to end the process (S190). If it is determined that the process should be continued (S190: NO), the server 100 waits for reception of, for example, access token generation request information, and returns the process to S120. On the other hand, if it is determined that the process should be ended (S190: YES), the server 100 ends the process.
[0180] In this example, the terminal 200A makes a search request for motion sequence data based on caption data, but the present invention is not limited to this. For example, the terminal 200A may transmit caption search request information including the motion sequence data and the access token to the server 100. If the server 100 determines that the access token is authorized, the server 100 may input the received motion sequence data to the information processing device 10 and cause the information processing device 10 to execute inference processing for caption data based on the motion sequence data. Then, for example, caption information including a caption with the highest similarity may be transmitted to the terminal 200A.
[0181] Furthermore, the access token is not limited to a token for executing a search process between motion and text in the information processing device 10 of the server 100. For example, as described in the second modified example (1), the access token may be an access token for causing the information processing device 10 of the server 100 to generate motion sequence data based on arbitrary skeleton data and caption data.
[0182] <Effects of the third embodiment> According to this embodiment, a server that communicates with at least a first terminal (e.g., terminal 200A) includes an information processing device 10 as a constituent element. The server also transmits first information (e.g., access token information or required token notification information) related to authorization to use the information processing device to the first terminal via a communication unit of the server, and receives second information (e.g., access token generation request information or motion search request information) related to a request to use the information processing device from the first terminal. When the second information includes the first information (e.g., when the access token can be verified), the server drives the information processing device (e.g., causes the information processing device to execute a motion search process). This shows an example of a configuration. According to this, when the server receives the second information including the first information transmitted in advance to the first terminal, the server can activate the information processing device based on the second information. That is, the server can authorize activation of the information processing device based on the second information transmitted from the first terminal based on the transmission of the first information to the first terminal.
[0183] <Other> At least a part of the processing that was to be performed by the server 100 in the above embodiment may be performed by the terminal 200. Conversely, at least a part of the processing that was to be performed by the terminal 200 in the above example may be performed by the server 100.
[0184] The operator of the server 100 may also be a motion search service provider in cooperation with a messaging service provider. In this case, the processing described in the above embodiments may be realized by one server, or the processing described in the above embodiments may be shared and realized by a server system consisting of multiple servers. A system configured with one or more servers may be defined as a server system, and the server of the present invention may be considered as a server system.
[0185] Furthermore, in the above-described embodiments, a server for distributing various applications (a server from which the terminal 200 downloads applications) may be configured as a server different from a server for providing the corresponding service (application), etc. In other words, a server for distributing applications and a server for performing application management processing, etc., described in the above-described embodiments, etc., may be configured as physically separated servers, or may be configured as a single server.
[0186] Furthermore, applications are not limited to various application programs, but may also include, for example, a program that provides the functionality of another service as one function of a base application (for example, a program that provides the functionality of a motion search service as one function of a messaging application, or vice versa), a program for updating the base application, etc. Data used in application programs (which may include data for updating applications, etc.) may also be included.
[0187] Furthermore, in the above embodiments, the present invention has been described as being implemented by a client-server system, but is not limited to this. As mentioned above, the present invention may be implemented by a system such as a distributed system in which the terminal 200 has the functions of a server or server system. For example, the processes described in the flowcharts of the above embodiments as being performed by a server or server system may be performed by a terminal.
[0188] Furthermore, as mentioned above, the contents described in the above-mentioned embodiments, modifications, other embodiments, etc. can be applied in combination with each other. [Explanation of symbols]
[0189] 1. Feature generation device 10. Information processing equipment 1000 Communication Systems 100 servers 200 devices 300 Network
Claims
1. A feature generation device, comprising: A control unit is provided, The control unit obtaining a three-dimensional skeletal sub-time series from the three-dimensional skeletal motion series; classifying the skeletal joints in the three-dimensional skeletal part time series into skeletal parts; rearrange the skeletal joints for each of the parts based on the distance from the set joint; Complementing the rearranged skeletal joints to a predetermined number of skeletal joints; standardizing the position information of the skeletal joints in the interpolated three-dimensional skeletal part time series; generating a region patch by converting the standardized position information into color information; generating a feature amount by integrating the part patches generated for each part; Feature generation device.
2. 2. The feature generating device according to claim 1, The control unit converts the feature into a color image and outputs the color image.
3. 2. The feature generating device according to claim 1, The body parts include at least a torso including a head, a left arm, a right arm, a left leg, and a right leg.
4. 4. The feature generating device according to claim 3, The feature generation device, wherein the set joint is a torso joint.
5. 5. The feature generating device according to claim 4, The three-dimensional skeletal motion sequence is a human body motion sequence.
6. 2. The feature generating device according to claim 1, The control unit generating a first feature amount, which is the feature amount, based on a first three-dimensional skeletal motion sequence represented by the skeletal joints having a first number of joints; generating a second feature amount that is the feature amount based on a second three-dimensional skeletal motion sequence expressed by the skeletal joints having a second number of joints different from the first number of joints; The feature generating device is characterized in that the first feature and the second feature have the same number of dimensions.
7. An information processing device capable of mutually converting the three-dimensional skeletal movement sequence and a sentence sequence, The information processing device includes the feature extraction device according to claim 1, an image encoding unit, and a text encoding unit, the image encoding unit calculates image features based on the features generated from the three-dimensional skeletal motion sequence by the feature extraction device; the text encoding unit calculates text features based on the text sequence; a control unit of the information processing device that optimizes the image encoding unit and the sentence encoding unit so as to be able to calculate an appropriate combination of the three-dimensional skeletal movement sequence and the sentence sequence based on the image feature amount and the sentence feature amount; Information processing device.
8. a server in communication with at least a first terminal, The control unit of the server includes the information processing device according to claim 7, The control unit of the server transmitting first information regarding authorization to use the information processing device to the first terminal by a communication unit of the server; receiving second information related to a request to use the information processing device from the first terminal by the communication unit; driving the information processing device when the second information includes the first information; server.
9. 9. The server of claim 8, the control unit of the server transmits the first information when it determines that a payment process related to the use authorization has been performed in the first terminal. server.
10. An information processing method in a feature generation device, comprising: obtaining a three-dimensional skeletal sub-time series from the three-dimensional skeletal motion series; classifying the skeletal joints in the three-dimensional skeletal partial time series into skeletal parts; rearranging the skeletal joints for each of the parts based on the distance from the set joint; Complementing the rearranged skeletal joints to a predetermined number of skeletal joints; normalizing the position information of the skeletal joints in the interpolated three-dimensional skeletal part time series; generating a region patch by converting the standardized position information into color information; generating a feature amount by integrating the part patches generated for each part; An information processing method including:
11. The feature generator obtaining a three-dimensional skeletal sub-time series from the three-dimensional skeletal motion series; classifying the skeletal joints in the three-dimensional skeletal partial time series into skeletal parts; rearranging the skeletal joints for each of the parts based on the distance from the set joint; Complementing the rearranged skeletal joints to a predetermined number of skeletal joints; normalizing the position information of the skeletal joints in the interpolated three-dimensional skeletal part time series; generating a region patch by converting the standardized position information into color information; generating a feature amount by integrating the part patches generated for each part; A program to execute.