Learning device, inference device, learning method, and program

By training a mathematical model with combined learning data sets to estimate modal data and feature vectors from different events, the method addresses the limitations of existing motion generation technologies, reducing data collection burdens and enhancing the diversity of generated actions.

JP2026019505APending Publication Date: 2026-02-05NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024121126
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing motion generation technologies, such as those using diffusion models, are limited by training data and struggle to generate actions that deviate from the training data, especially when generating motions for objects other than humans, leading to high data collection costs and burdens.

Method used

A control unit performs a first learning process using a combination of first and second learning data sets to train a mathematical model that estimates modal data and feature vectors, incorporating data from different events, and updates the model to reduce differences between estimated and actual feature vectors, allowing for the generation of diverse actions.

Benefits of technology

This approach reduces the burden of generating time-series data by enabling the generation of actions that deviate from training data, lowering costs and improving data collection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019505000001_ABST
    Figure 2026019505000001_ABST
Patent Text Reader

Abstract

To reduce a load required for generating time series data.SOLUTION: A control unit configured to perform first learning processing of performing learning based on a set of first modal data that is modal data related to a first event and a first feature vector indicating the first event, and a set of second modal data that is modal data related to a second event and is the same as or different from the first modal data, the first model that is a learning target of the first learning processing including a modal data estimation model configured to estimate the first modal data using the second modal data and a feature vector estimation model configured to estimate the first feature vector using the modal data, in a learning device, a control unit executes learning of a feature vector estimation model and learning of a modal data estimation model, and the control unit executes, after execution of first learning processing, processing for learning a second model that is a feature vector estimation model for estimating a first feature vector on the basis of second modal data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, an inference device, a learning method, and a program. [Background technology]

[0002] There are technologies aimed at generating motion. More specifically, there are technologies aimed at generating time-series data that represent motion. Motion is represented by time-series data of object features. For example, in the case of a human motion, it is represented by joint positions, joint velocities, joint rotation angles, whether the feet are on the ground, the positions of representative points of the entire body, the velocities of representative points of the entire body, and time-series data on photographed images of the person. Motion generation is a technology that has been researched in fields such as computer animation, and has a wide range of applications, including games, robotics, the metaverse, and content creation. Note that the definition of motion generation is the generation of motion. Therefore, motion generation means the generation of time-series data that represent motion.

[0003] As a technique for generating motion, for example, a motion generation technique using a diffusion model (see Non-Patent Document 1) has been proposed. This technique realizes the generation of motion according to text by training a diffusion model using training data consisting of paired data of text and motion. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. Human Motion Diffusion Model. In ICLR 2023. Summary of the Invention [Problem to be solved by the invention]

[0005] However, the actions that can be generated by the technology of Non-Patent Document 1 are greatly restricted by the training data, and it has been difficult to generate actions that deviate from the training data. Specifically, for example, when training is performed using training data that includes only human actions and text that describes them (e.g., "A man is walking"), when text that describes the movement of an object other than a human (e.g., "A pelican is catching fish from the sea") is given, it is difficult to generate human actions that match the text (e.g., "a motion like a pelican diving into the sea with its wings spread out to the side and its body submerged")

[0006] One solution to this problem is to increase the amount and variety of training data. Specifically, for example, by collecting paired data of text describing the movements of objects other than people and the corresponding human actions, and then performing training based on this data. By doing this, it becomes possible to generate appropriate human actions even for text describing the movements of objects other than people.

[0007] However, collecting motion data requires specialized equipment and software, and the cost of collecting motion data is higher than that of text data. Therefore, this solution results in high data collection costs and can impose a heavy burden on generating time-series data representing motion.

[0008] This situation is not limited to text, but also applies to generating motion from other modalities such as images, videos, and audio. Furthermore, it is not limited to generating motion, but also applies to generating any time-series data.

[0009] In view of the above circumstances, an object of the present invention is to provide a technique for reducing the burden required for generating time-series data. [Means for solving the problem]

[0010] One aspect of the present invention includes a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event. The first model, which is a mathematical model of a learning target in the first learning process, includes a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data. In the first model, input modal data, which is the modal data used for estimation by the feature vector estimation model, is an estimation result of the modal data estimation model, and the control unit performs a first learning process based on the first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event. and a second sub-learning process for learning the modal data estimation model using the first modal data and the first feature vector in the second training data set, and a second sub-learning process for learning the modal data estimation model using the first modal data and the second modal data in the second training data set, wherein after executing the first learning process, the control unit executes a second learning process for learning a second model, which is a feature vector estimation model that estimates a first feature vector based on second modal data, and in the second learning process, the second model is updated so as to reduce a difference between a first feature vector estimated by a model for second learning process, which is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, based on the second modal data in the second training data set, and a first feature vector estimated by the second model that uses the second modal data in the second training data set as the input modal data.

[0011] One aspect of the present invention is a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein the first model, which is a mathematical model to be learned in the first learning process, is a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a modal data and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data obtained by the feature vector estimation model, wherein in the first model, input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit executes, in the first learning process, a first sub-learning that is training of the feature vector estimation model using the first modal data and the first feature vector in the first training data set, and a second sub-learning that is training of the modal data estimation model using the first modal data and the second modal data in the second training data set.

[0012] One aspect of the present invention includes a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event. The first model that is a mathematical model of a learning target in the first learning process includes a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data. In the first model, input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit performs a first learning process based on the first modal data and the second modal data in the first learning data set. and a second sub-learning process that learns the modal data estimation model using the first modal data and the second modal data in the second training data set, wherein the control unit, after executing the first learning process, executes a second learning process that learns a second model that is a feature vector estimation model that estimates a first feature vector based on second modal data, and in the second learning process, a model for second learning process that is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, is updated so as to reduce a difference between a first feature vector estimated based on the second modal data in the second training data set and a first feature vector estimated by the second model that uses the second modal data in the second training data set as the input modal data.

[0013] One aspect of the present invention includes a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein the first model, which is a mathematical model to be learned in the first learning process, includes a modal data estimation model, which is a mathematical model that estimates the first modal data using the second modal data, and a modal data estimation model, which is a mathematical model that estimates the first feature vector using modal data that is modal data. and a feature vector estimation model which is a mathematical model for estimating a torque of a feature vector in the first model, wherein input modal data, which is the modal data used for estimation by the feature vector estimation model, is an estimation result of the modal data estimation model, and the control unit, in the first learning process, executes a first sub-learning which is training of the feature vector estimation model using the first modal data and the first feature vector in the first training data set, and a second sub-learning which is training of the modal data estimation model using the first modal data and the second modal data in the second training data set, and the inference device is provided with an inference unit which performs inference using the trained first model obtained by a learning device.

[0014] One aspect of the present invention is and a control unit that performs a first learning process that performs learning based on: a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event; and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein a first model that is a mathematical model to be learned in the first learning process includes: a modal data estimation model that is a mathematical model that estimates the first modal data using second modal data; and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data, and wherein in the first model, input modal data that is the modal data used in estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit performs a first learning process that performs learning based on the feature vector estimation model using the first modal data and the first feature vector in the first learning data set. a first sub-learning step in which the control unit executes the first learning process, which is training of the modal data estimation model using the first modal data and the second modal data in the second training data set; and a second sub-learning step in which the control unit executes the second learning process, which is training of a second model that is a feature vector estimation model that estimates a first feature vector based on second modal data. In the second learning process, a model for second learning process, which is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, is updated so as to reduce a difference between a first feature vector estimated based on the second modal data in the second training data set and a first feature vector estimated by the second model that uses the second modal data in the second training data set as the input modal data.

[0015] One aspect of the present invention is a control unit that performs a first learning process that performs learning based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein the first model, which is a mathematical model to be learned in the first learning process, is a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a mathematical model that estimates the first feature vector using the modal data that is modal data. and a feature vector estimation model which is a model of a first model, wherein in the first model, input modal data which is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit executes, in the first learning process, a first sub-learning which is training of the feature vector estimation model using the first modal data and the first feature vector in the first training data set, and a second sub-learning which is training of the modal data estimation model using the first modal data and the second modal data in the second training data set, the learning method being executed by a learning device, the learning method including a first learning step in which the control unit executes the first learning process.

[0016] One aspect of the present invention is a program for causing a computer to function as either the learning device or the inference device described above. [Effects of the Invention]

[0017] The present invention makes it possible to reduce the burden required for generating time-series data. [Brief explanation of the drawings]

[0018] [Figure 1]FIG. 1 is an explanatory diagram illustrating an information processing system according to an embodiment. [Figure 2] FIG. 10 is an explanatory diagram illustrating an example of conversion using a diffusion model of a feature vector estimation model according to an embodiment. [Figure 3] FIG. 10 is an explanatory diagram of an example of conversion using a large-scale language model as a text data estimation model in the embodiment. [Figure 4] FIG. 10 is an explanatory diagram of an example of learning using a large-scale language model as a text data estimation model in the embodiment. [Figure 5] FIG. 1 is an explanatory diagram illustrating an example of the flow of processing executed by a trained first model in an embodiment. [Figure 6] FIG. 6 is an explanatory diagram illustrating an example of a second learning process in the embodiment. [Figure 7] FIG. 10 is a diagram showing a first example of experimental results in the embodiment. [Figure 8] FIG. 10 is a diagram showing a second example of experimental results in the embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of a hardware configuration of a learning device according to an embodiment. [Figure 10] 5 is a flowchart showing a first example of the flow of processing executed by the learning device in the embodiment. [Figure 11] 10 is a flowchart showing a second example of the flow of processing executed by the learning device in the embodiment. [Figure 12] FIG. 2 is a diagram showing an example of the hardware configuration of an inference device according to an embodiment. [Figure 13] 10 is a flowchart showing an example of a flow of processing executed by the inference device of the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] (Embodiment) FIG. 1 is an explanatory diagram illustrating an information processing system 100 according to an embodiment. The information processing system 100 includes a learning device 1 and an inference device 2. For simplicity of the following explanation, we will first use an example in which the modal is text. Also, for simplicity of the following explanation, we will use an example in which the time-series data indicates behavior.

[0020] The learning device 1 includes a control unit 11 including a processor 91, such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit) or an NPU (Neural Network Processing Unit), and a memory 92, which are connected via a bus.

[0021] The control unit 11 performs a first learning process. The first learning process is a process in which learning of a mathematical model to be learned in this process (hereinafter referred to as the "first model") is performed until a predetermined condition regarding the end of the learning (hereinafter referred to as the "first learning end condition") is satisfied. The first learning end condition may be any condition regarding the end of learning of the first model, and may be, for example, a condition that the first model has been updated a predetermined number of times, a condition that the change in the first model due to the update is smaller than a predetermined change, or a condition that the score of a predetermined index for evaluating the performance of the first model satisfies a predetermined condition regarding the value.

[0022] It should be noted that for any learning, the state of the mathematical model or process at the point in time when a predetermined condition for the end of that learning is satisfied is the so-called "learned" state for that learning.

[0023] More specifically, the first learning process is a process of learning a first model based on a first learning data set and a second learning data set.

[0024] The first training data set is a set of first text data and a first feature vector. The first text data is text data of text related to a first action, which is an action of a first target. The text related to the first action is, for example, text explaining the first action. A specific example of text explaining the first action is, for example, "A man is walking."

[0025] The first feature vector is a feature vector that indicates a first action. Note that, hereinafter, a feature vector that indicates any action, not limited to the first action, is referred to as an action feature vector. Therefore, not only the first feature vector, but also a feature vector that indicates a second action, which is the action of a second object (hereinafter referred to as a "second feature vector"), is a type of action feature vector.

[0026] The second training data set is a set of first text data, which may be the same as or different from the first text data of the first training data set, and second text data. The first text data of the second training data set and the first text data of the first training data set are identical in that they are both text data relating to the motion of a first object. There can be various types of motion of the first object. Therefore, when the first text data of the second training data set and the first text data of the first training data set are said to be different, the difference may be, for example, in the content of the motion of the first object.

[0027] The second text data is text data of text related to a second action that is an action of a second object. The text related to the second action is, for example, text explaining the second action. The text explaining the second action is, for example, "A pelican is catching fish from the sea."

[0028] The first object is the first object, and the second object is the second object. The second object is an object different from the first object. The object may be a human, a non-human animal, or a machine such as a robot or a mechanical doll. Since the first object and the second object are different, for example, the first object may be a human and the second object may be a non-human animal, or the first object may be a non-human animal and the second object may be a human. The first object may be a non-human animal and the second object may be a machine, or the first object may be a machine and the second object may be a non-human animal. The first object may be a machine and the second object may be a human, or the first object may be a human and the second object may be a machine.

[0029] The first model includes a text data estimation model and a feature vector estimation model. The text data estimation model is a mathematical model that estimates the first text data using the second text data. The text data estimation model is, for example, a large language model (LLM).

[0030] The feature vector estimation model is a mathematical model that estimates a first feature vector using text data. In the first model, the input to the feature vector estimation model is the output of the text data estimation model. In other words, the text data used for estimation by the feature vector estimation model in the first model (hereinafter referred to as "input text data") is the result of estimation by the text data estimation model in the first model.

[0031] <More details on the first learning process> The processing executed by the control unit 11 in the first learning process will be described in more detail. In the first learning process, the control unit 11 performs first sub-learning. The first sub-learning is training of a feature vector estimation model using first text data and first feature vectors in the first training data set.

[0032] A more specific example of the first sub-learning will be described. In the first sub-learning, a feature vector estimation model is executed on first text data in a first training data set, and as a result, a first feature vector is estimated. Next, in the first sub-learning, the feature vector estimation model is updated so as to reduce the difference between the estimated first feature vector and the first feature vector in the first training data set. The update is performed until a predetermined condition for terminating the learning of the feature vector estimation model (hereinafter referred to as a "first sub-learning termination condition") is satisfied.

[0033] The first sub-learning termination condition may be any condition related to the termination of learning of the feature vector estimation model, such as a condition that the feature vector estimation model has been updated a predetermined number of times, a condition that the change in the feature vector estimation model due to the update is smaller than a predetermined change, or a condition that the score of a predetermined index used to evaluate the performance of the feature vector estimation model satisfies a predetermined condition related to the value.

[0034] In the first learning process, the control unit 11 performs second sub-learning, which is learning of a text data estimation model using the first text data and second text data in the second learning data set.

[0035] A more specific example of the second sub-learning will be described. In the second sub-learning, the text data estimation model is executed on the second text data in the second training data set, and as a result, the first text data is estimated. Next, in the second sub-learning, the text data estimation model is updated, for example, so as to reduce the difference between the estimated first text data and the first text data in the second training data set. The update is performed until a predetermined condition for terminating the learning of the text data estimation model (hereinafter referred to as the "second sub-learning termination condition") is satisfied.

[0036] Another more specific example of the second sub-learning will be described. In the second sub-learning, the internal state of the text data estimation model is updated based on examples of paired data of the second text data and the first text data in the second learning data set. The internal state of the text data estimation model is an abstract representation of the relationship between the example data. The update is performed until a second sub-learning termination condition, which is a predetermined condition for terminating the learning of the text data estimation model, is satisfied.

[0037] The second sub-learning termination condition may be any condition related to the termination of learning of the text data estimation model, such as a condition that pair data of the second text data and the first text data in the second learning data set have been illustrated for all learning data. The second sub-learning termination condition may also be a condition that the text data estimation model has been updated a predetermined number of times, that the change in the text data estimation model due to the update is smaller than a predetermined change, or that a predetermined condition related to the value of a predetermined index for evaluating the performance of the text data estimation model is satisfied.

[0038] Therefore, in the first learning process, the first learning end condition may be, for example, a condition that the first sub-learning end condition and the second sub-learning end condition are satisfied.

[0039] The trained first model obtained by executing such a first learning process estimates the first text data based on the second text data and estimates a first feature vector based on the estimated first text data. Therefore, the first feature vector obtained in this manner is a feature vector indicating a second action indicated explicitly or implicitly by the text indicated by the second text data. In other words, the first feature vector obtained by the trained first model obtained in this manner is an action feature vector that is less different from the second feature vector than the first feature vector obtained by the first model before learning.

[0040] The inference device 2 performs a first inference process. The first inference process is a process of making an inference using the trained first model obtained by the learning device 1. Therefore, by executing the first inference process, the inference device 2 estimates a feature vector indicating a second action based on the second text data using the trained first model.

[0041] <Effects of the first learning process> The trained first model obtained by the first learning process estimates a feature vector representing the second action based on the second text data. However, the feature vector representing the second action itself is not used in the first learning process. Therefore, the learning device 1 including the control unit 11 that executes such a first learning process can reduce the burden required to generate time-series data representing actions.

[0042] <More concrete examples of the first model and its learning> The training of the feature vector estimation model (i.e., the first sub-learning) may be, for example, a motion generation technique using a diffusion model as described in Non-Patent Document 1. The first sub-learning may use, for example, a feature vector estimation model trained using any deep learning model. Specifically, the training of the feature vector estimation model may be performed using any deep generative model, such as a variational autoencoder (VAE), a vector quantized variational autoencoder (VQ-VAE), a generative adversarial network (GAN), a flow-based model, an autoregressive model, a diffusion model, or any combination thereof.

[0043] 2 is an explanatory diagram illustrating an example of conversion of a feature vector estimation model using a diffusion model in an embodiment. For conversion of the feature vector estimation model, for example, a motion generation technique using a diffusion model as described in Non-Patent Document 1 or the like may be used.

[0044] Figure 2 shows the processing flow when a diffusion model is used as the feature vector estimation model. First, as a whole, when text describing a first action \hat{y} is given, the feature vector estimation model generates a corresponding action \hat{x}_0. Here, \hat{y} represents the symbol in the following equation (1), and \hat{x}_0 represents the symbol in the following equation (2).

[0045]

number

[0046]

number

[0047] In this case, as shown in Figure 2, before inputting text \hat{y} into the motion generator, \hat{y} may be converted into an embedded representation \hat{c} using multimodal embedding. \hat{c} represents the symbol in the following equation (3). Note that multimodal embedding may be, for example, a text-image embedding technique called Contrastive Language-Image Pre-training (CLIP). Note that the parameters of the multimodal embedding may be fixed when training the motion generator, or may be optimized simultaneously with the motion generator during training.

[0048]

number

[0049] While we have shown an example using multimodal embedding, any other embedding may be used. Furthermore, as shown in Figure 2, when a diffusion model is used as the feature vector estimation model (motion generation model), the sampled value t from the diffusion step sampler at diffusion step t may also be used as input to the motion generator. Note that t satisfies the relationship t~Uniform({1, . . . , T}). Uniform({1, . . . , T}) represents a uniform distribution of integers between 1 and T, and "~" indicates sampling from the distribution shown on the right-hand side. T is the total number of diffusion steps required to convert data into noise. For simplicity, Uniform({1, . . . , T}) is expressed as [1,T] below. Based on these inputs, the feature vector estimation model G generates a motion \hat{x}_0 corresponding to the text \hat{y}.

[0050] The process of generating the action \hat{x}_0 is expressed as the following formula (4). t is the data obtained by diffusing the motion data x for t steps. Note that x is an N×d real matrix. N represents the length in the time direction, and d represents the dimension of the feature of the motion data.

[0051]

number

[0052] To explain the data diffusion process in detail, for example, each diffusion step is assumed to follow a Markov process. That is, x t is the data of the previous step x t-1 It is assumed that the decision depends only on x. t-1 From x t Transition probability q(x t |x t-1 ) is expressed by a normal distribution, specifically, the mean is expressed by equation (5) and the variance is expressed by equation (6). Here, I is the unit matrix. Also, α t x t-1 From xt In the transition to x t-1 is a hyperparameter that indicates the survival rate, and after T steps of diffusion, x T has a value that follows a standard normal distribution (normal distribution with mean 0 and standard deviation I). With this definition, x approaches Gaussian noise as the diffusion step progresses, and after T steps of diffusion, x T represents Gaussian noise.

[0053]

number

[0054]

number

[0055] The feature vector estimation model G is trained so that x and \hat{x}_0 coincide. Specifically, for example, the function of the following equation (7) is used as the objective function.

[0056]

number

[0057] Here, (x,\hat{y})~q(x,\hat{y}) represents sampling of a pair (x,\hat{y}) of first feature data x and first text data \hat{y} from the first training data set. Note that in the example of Equation (7), L2 distance is used as the distance criterion between x and \hat{x}_0, but the distance criterion is not necessarily limited to this. The distance criterion may be any distance criterion, such as Lp distance (p is 1 or greater), hinge distance, Wasserstein distance, cosine distance, or a combination thereof. Also, in the example of Equation (7), the weight of the loss function is the same for all t values ​​between 1 and T, but the weight of the loss function may be changed depending on the value of t.

[0058] In addition to equation (7), one or more of the geometric constraint loss functions expressed by the following equations (8), (9), and (10) may be used.

[0059]

number

[0060]

number

[0061]

number

[0062] In equations (8) and (9), FK represents the forward kinematic function that converts the rotation of the joint into the position of the joint. i represents the data of the ith frame in the time direction of x (i is an integer from 1 to N). Similarly, \hat{x}_0^i represents the data of the ith frame in the time direction of \hat{x}_0. \hat{x}_0^i represents the symbol in the following equation (11). By minimizing the geometric constraint loss function shown in equation (8), the feature vector estimation model G is trained so that x and \hat{x}_0 match from the perspective of forward kinematics. Note that in the example of equation (8), the L2 loss function is used as the loss function, but the loss function is not necessarily limited to this. The loss function may be a loss function based on any distance criterion, such as a loss function based on the Lp distance (p is 1 or greater), a loss function based on the hinge distance, a loss function based on the Wasserstein distance, or a loss function based on the cosine distance, or a combination of these.

[0063]

number

[0064] In equation (9), fi is a binary mask representing whether or not the foot is in contact with the ground in the i-th frame. By minimizing the geometric constraint loss function shown in equation (8), the feature vector estimation model G is trained to prevent the foot from slipping when the foot is in contact with the ground. Note that, although the L2 loss function is used as the loss function in the example of equation (9), the loss function is not necessarily limited to this. The loss function may be a loss function based on any distance criterion, such as a loss function based on the Lp distance (p is 1 or greater), a loss function based on the hinge distance, a loss function based on the Wasserstein distance, or a loss function based on the cosine distance, or a combination of these.

[0065] In equation (10), the feature vector estimation model G is trained so that the change in x in each frame matches the change in \hat{x}_0 in each frame. Note that, although the L2 loss function is used as the loss function in the example of equation (10), the loss function is not necessarily limited to this. The loss function may be a loss function based on any distance criterion, such as a loss function based on the Lp distance (p is 1 or greater), a loss function based on the hinge distance, a loss function based on the Wasserstein distance, or a loss function based on the cosine distance, or a combination of these.

[0066] The overall loss function is expressed by equation (12).

[0067]

number

[0068] where λ pos , λ foot , λ vel represents the weight parameter of the loss function to be multiplied, and is a real number greater than or equal to 0. When it is 0, it means that the loss function to be multiplied is not used. The feature vector estimation model G is optimized by minimizing the loss function expressed by equation (12).

[0069] The conversion from the first text data to the first feature vector is performed using the feature vector estimation model that has been trained in the first sub-learning process. Specifically, first, the first text data \hat{y} is prepared as input data for the feature vector estimation model. Also, the diffusion step t is set to T. At this time, x t x T Specifically, as mentioned above, x T represents Gaussian noise, so it can be obtained by sampling from a standard normal distribution. t (=x T ) to generate the first feature vector \hat{x}_0.

[0070] The first feature vector \hat{x}_0 obtained in this way may be used as the final output data of the feature vector estimation model, or the diffusion and generation process may be repeated to further improve the estimation accuracy. Specifically, for t between 1 and T, x is diffused for t steps to obtain x t When \hat{y}, t, x are obtained, t By applying equation (4) to the first feature vector \hat{x}_0, the first feature vector \hat{x}_0 is generated. By performing the diffusion process for the step t-1 on the obtained \hat{x}_0, x t-1 Then, replace t with t-1 in equation (4) and calculate x t x t-1 The first feature vector \hat{x}_0 is generated by converting it using the result obtained by replacing \hat{x}_0 with \hat{x}_0. This process is repeated from t=T to t=1, and the resulting first feature vector \hat{x}_0 may be used as the final output data of the feature vector estimation model. Note that although the case where t is decreased by one step at a time during repeated calculations has been described here, the step update size may be set to any integer greater than or equal to 1.

[0071] Fig. 3 is an explanatory diagram of an example of conversion from second text data to first text data using a large-scale language model as a text data estimation model in an embodiment. Fig. 4 is an explanatory diagram of an example of learning for conversion from second text data to first text data using a large-scale language model as a text data estimation model in an embodiment.

[0072] For example, when a first action represents a human action and a second action represents a non-human object action, the large-scale language model converts text data of text explaining the non-human object action into text data of text explaining the human action. Note that the text explaining the human action is text explaining an action performed when a person imitates the non-human object action.

[0073] Specifically, for example, when the text describing the action of an object other than a person is "A pelican is catching fish from the sea," the large-scale language model converts the text data of that text into text data describing the action of a person: "A person dives into the sea with their arms and legs splayed out to the sides and their body submerging."

[0074] For example, In-Context Learning (ICL) may be used as a method for achieving this conversion. In-Context Learning is a technology that enables output tailored to a task to be solved by providing an example demonstration of the task to be solved when a pre-trained large-scale language model is available. A feature of ICL is that all that is required is the example demonstration, and there is no need to update the parameters of the model itself. Therefore, by using In-Context Learning (ICL), it is possible to quickly learn the conversion from the second text data to the first text data with a small amount of training data (example demonstrations).

[0075] The large-scale language model is configured, for example, by a deep neural network. More specifically, the large-scale language model is configured by a transformer neural network, a convolutional neural network, a recurrent neural network, a fully connected neural network, or a combination of these.

[0076] As a specific example, we will explain In-Context Learning, which aims to convert text data that describes the movements of objects other than people into text data that describes the movements of people. Below, we will explain using mathematical formulas and the example shown in Figure 4.

[0077] The pre-trained large-scale language model is represented by the symbol \mathcal{M}. The prompt template used in in-context learning is represented by the symbol \mathcal{T}. \mathcal{M} represents the symbol in the following equation (13). \mathcal{T} represents the symbol in the following equation (14). Then, \mathcal{T}={\mathcal{I}, \mathcal{D}}. \mathcal{I} represents the symbol in the following equation (15). \mathcal{D} represents the symbol in the following equation (16). Here, \mathcal{I} is an instruction, and \mathcal{D} is an illustrative demonstration.

[0078]

number

[0079]

number

[0080]

number

[0081]

number

[0082] Specifically, \mathcal{I} provides instructions for learning, and in the example in Figure 4, it provides the following three instructions. One of the instructions is to translate the sentence into a sentence that describes a person's motion ("Please translate the following sentence into human motion style"). Another instruction is that the output sentence must begin with a word related to a person and end with a period ("Sentence must begin with 'a person' or 'the person' or 'a man' or 'the man' and end with '."). The last instruction is that the output should be a single sentence ("Output one sentence only").

[0083] On the other hand, \mathcal{D} is a specific example, and is composed of pair data of text data y of input text and text data \hat{y} of output text. Let the number of pair data be k, and the input text data of the i-th pair data be y i Let the text data of the output text be \hat{y}_i. Note that \hat{y}_i represents the symbol in the following equation (17). In this case, \mathcal{D} is represented by a set of k paired data \mathcal{D} expressed by the following equation (18). Figure 4 shows an example where k=3, and shows an example of three input / output paired data.

[0084]

number

[0085]

number

[0086] Let f_{\mathcal{M}}(\hat{y},\mathcal{T},y) be the score function that calculates the score (relevance) between the text data y of the input text and the text data \hat{y} of the output text. f_{\mathcal{M}}(\hat{y},\mathcal{T},y) is expressed as the following equation (19). In this case, the likelihood P(\hat{y}|y) of the text data \hat{y} of the output text when the text data y of the input text is given is expressed as f_{\mathcal{M}}(\hat{y},\mathcal{T},y) based on the prompt template \mathcal{T} and the large-scale language model \mathcal{M}. In other words, P(\hat{y}|y) is defined as the following equation (20). In this case, the optimal output text \hat{y} when the text data y of the input text is given is obtained by the following equation (21). The definition of optimal here is the text data that is most suitable as the text data \hat{y} of the output text corresponding to the text data y of the input text when there is a large-scale language model \mathcal{M} and a prompt template \mathcal{T}. The likelihood P(\hat{y}|y) is expressed by the following equation (22). The definition of "most suitable" here is to maximize the likelihood P(\hat{y}|y).

[0087]

number

[0088]

number

[0089]

number

[0090]

number

[0091] Here, Y represents a set of candidate text data for the output text. This formula outputs the text data \hat{y} of the output text with the highest probability from Y.

[0092] Using a text data estimation model based on the large-scale language model obtained in this manner, text data y of text describing the second action is converted into text data \hat{y} of text describing the first action.

[0093] As shown in Figure 3, when converting the second text data into the first text data using a trained text data estimation model, a post-check of the text data obtained by the conversion may be performed. If the post-check determines that the conditions are not satisfied, the text data may be reconverted. This further improves the quality of the text conversion.

[0094] The post-check is a process of determining whether the text data obtained by conversion satisfies a predetermined condition. Therefore, the post-check may be, for example, a check of the length of the text indicated by the text data obtained by conversion, or a check of the words of the text indicated by the text data obtained by conversion.

[0095] The check of the length of the text indicated by the text data obtained by conversion is, for example, a process of determining whether the number of words constituting the text is less than a predetermined number. In this case, if it is determined that the number is less than the predetermined number, for example, the text data is reconverted.

[0096] Checking the words in the text indicated by the text data obtained by conversion may be, for example, a process of determining whether the text to be determined begins with a predetermined character string. In this case, if it is determined that the text does not begin with the predetermined character string, for example, the text data is reconverted. Therefore, checking the words in the text indicated by the text data obtained by conversion may be, for example, a process of determining whether the text to be determined begins with "A person," "The person," "A man," or "The man."

[0097] 5 is an explanatory diagram illustrating an example of the flow of processing executed by a trained first model in an embodiment, showing an example of processing executed by a text data estimation model and an example of processing executed by a feature vector estimation model.

[0098] Specifically, the contents shown in Figure 5 are as follows. That is, in the trained first model, the process shown in Figure 3, which is in a trained state with respect to the second sub-learning, is executed. Note that the process in a trained state with respect to the second sub-learning means the trained process obtained by the second sub-learning.

[0099] Next, the trained first model executes the process shown in Figure 2, which is in a trained state for the first sub-learning, on the execution result. Note that the process in a trained state for the first sub-learning means the trained process obtained by the first sub-learning. Figure 5 further shows that a feature vector indicating the second action is obtained from the second text data using the trained first model.

[0100] <Second Learning Process> Incidentally, the control unit 11 may further execute the second learning process after executing the first learning process. "After executing the first learning process" means after the first learning end condition is satisfied. Therefore, the second learning process is executed after the first sub-learning and the second sub-learning are executed. In other words, the second learning process is executed after the first sub-learning end condition and the second sub-learning end condition are satisfied.

[0101] The second learning process is a process in which a second model is used as a learning target. The second model is a feature vector estimation model that estimates a first feature vector based on second text data. The second learning process is a process in which the second model is obtained by distillation using the trained first model.

[0102] Therefore, in the second learning process, the second model is updated so as to reduce the difference between the first feature vector estimated by the model for the second learning process based on the second text data in the second learning data set and the first feature vector estimated by the second model that uses the second text data in the second learning data set as input text data.

[0103] The second learning process model is a first model including a trained text data estimation model obtained by the second sub-learning and a trained feature vector estimation model obtained by the first sub-learning.

[0104] The initial value of the second model may be the trained feature vector estimation model obtained by the first sub-learning. That is, the initial value of the second model may be the feature vector estimation model at the time when the first sub-learning termination condition is satisfied. The initial value of the second model may be a value sampled according to a predetermined distribution. The second model may be a mathematical model with fewer parameters than the trained feature vector estimation model obtained by the first sub-learning. The second model may be a mathematical model in which at least a portion of the parameters of the trained feature vector estimation model obtained by the first sub-learning are replaced with faster processing modules.

[0105] The second learning process is a process of updating the second model so as to reduce the difference between the first type first feature vector and the second type first feature vector. Therefore, if the feature vector estimation model at the time when the first sub-learning termination condition was satisfied is used as the initial value of the second model, the feature vector estimation model at the time when the first sub-learning termination condition was satisfied is further updated by the second learning process. Therefore, in this case, the second learning process can be said to be a process of relearning the second model.

[0106] The first-type first feature vector is the first feature vector estimated by the second learning process model based on the second text data in the second learning data set. The trained text data estimation model obtained by the second sub-learning refers to the text data estimation model at the time when the second sub-learning termination condition is satisfied.

[0107] The second type first feature vector is a first feature vector estimated by a second model that uses the second text data in the second training data set as input text data.

[0108] The second learning process is executed until a predetermined condition for terminating the second learning process (hereinafter referred to as the "second learning process termination condition") is satisfied. The second learning process termination condition may be any condition for terminating the second learning process, and may be, for example, a condition that the second model has been updated a predetermined number of times, a condition that the change due to the update of the second model is smaller than a predetermined change, or a condition that the score of a predetermined index for evaluating the performance of the second model satisfies a predetermined condition for the value.

[0109] Hereinafter, the second model at the point when the second learning process termination condition is satisfied will be referred to as a trained second model.

[0110] The mathematical model updated in the second learning process is the second model obtained by estimating the second type first feature vector. In the second learning process, the first model included in the model for the second learning process does not necessarily need to be updated.

[0111] When the second learning process is performed, the inference device 2 executes the second inference process instead of the first inference process. The second inference process is a process for making an inference using the trained second model obtained by the learning device 1. Therefore, by executing the second inference process, the inference device 2 estimates a feature vector indicating a second action based on the second text data by executing the trained second model.

[0112] <<Effects of Executing the Second Learning Process in addition to the First Learning Process>> If the first learning process is performed but the second learning process is not performed and the text data estimation model is a large-scale language model, the inference device 2 needs to perform text conversion using the large-scale language model, which is a heavy process, every time it attempts to generate an action. If the second learning process is performed, the inference device 2 can generate an action using the trained second model from the second text data without using a large-scale language model.

[0113] Furthermore, the trained second model estimates a feature vector representing the second action based on the second text data. However, the feature vector representing the second action itself is not used in either the first learning process or the second learning process, which are learning processes for obtaining the trained second model. Therefore, the learning device 1 including the control unit 11 that executes such a first learning process including the second learning process can reduce the burden required to generate time-series data representing actions.

[0114] Note that being able to generate a motion means generating time-series data representing the motion. That is, generating a motion means obtaining time-series data representing the motion. Therefore, obtaining a motion feature vector such as the first feature vector is an example of generating a motion.

[0115] 6 is an explanatory diagram illustrating an example of the second learning process in the embodiment, showing an example of the process executed by the text data estimation model, an example of the process executed by the second model, and an example of the process executed by the feature vector estimation model.

[0116] When the training data x of a movement is given, the movement data x that has been spread for t steps by executing the diffusion process is t The behavior generated from model G' that generates behavior from the second text data is expressed as formula (23) below. In FIG. 6, model G' is the mathematical model (i.e., the second model) described in area A101 in FIG. 6.

[0117]

number

[0118] The action generated from the first text data is expressed as the following formula (24).

[0119]

number

[0120] At this time, learning of G' is performed so that x0 and \hat{x}_0 coincide. Specifically, for example, the function of the following equation (25) is used as the objective function.

[0121]

number

[0122] In the example of equation (25), the L2 distance is used as the distance criterion between \hat{x}_0 and x0, but the distance criterion is not necessarily limited to this. The distance criterion may be any distance criterion, such as the Lp distance (p is 1 or greater), the hinge distance, the Wasserstein distance, the cosine distance, or a combination of these. Also, in the example of equation (25), the weight of the loss function is the same for all t values ​​between 1 and T, but the weight of the loss function may be changed depending on the value of t.

[0123] When training the feature vector estimation model G' as described above, the initial value of G' may be set to G. Furthermore, when training G', Low Rank Adaptation (LoRA) may be used to reduce training costs. In LoRA, when the weight of a certain layer of the model is W0 and the weight added to W0 when updating the weight through training is ΔW, low-rank decomposition is performed by substituting W0 + ΔW = W0 + BA. Hereinafter, the original model (a model with a weight of a certain layer W0) is referred to as the base model, and the model after applying LoRA (a model with a weight of a certain layer W0 + ΔW) is referred to as the trained model. Note that W0 and ΔW are j × k real matrices, B is a j × r real matrix, and A is an r × k real matrix. j, k, and r are all integers greater than or equal to 1. To perform low-rank decomposition, the rank r is sufficiently smaller than min(j, k).

[0124] The initialization of A and B is explained below. For example, A can be initialized with a normal distribution with standard deviation σ, and B can be initialized with B=0. When training a feature vector estimation model G' using LoRA, the weight W0 of the pre-trained model is fixed, and only the low-dimensional A and B are updated. With this definition, if the calculation of a certain layer is performed with h=W0x, applying LoRA makes it possible to calculate h as h=W0x+ΔWx=W0x+BAx.

[0125] By doing this, it is possible to reduce the dimension of the parameters to be learned in the feature vector estimation model G', resulting in more efficient learning. Furthermore, since the BA learned by LoRA is a linear mapping, it can also be merged with the base model. Merging is a process of replacing the weight W0 of the base model with a matrix of the same dimension, W0 + BA. Therefore, it is possible to make the inference speed of the learned model equivalent to that of the base model.

[0126] <Experimental Results> An example of the experimental results verifying the effectiveness of the first learning process will be described. In the experiment, a human motion was used as the first motion, and a motion of an object other than a human motion was used as the second motion. The technology described in Non-Patent Document 1 was used for the first sub-learning.

[0127] Specifically, the technology described in Non-Patent Document 1 is a model based on a diffusion model, and learning is performed using paired data of a first action and text data describing that action as learning data.

[0128] In the experiment, a text data estimation model was trained to convert text data of a text describing a second action into text data of a text describing a first action using the technology described with reference to Figure 6. In the second training process, LoRA was used to train the feature vector estimation model. In the experiment, the results of using the technology described in Non-Patent Document 1 as is were compared with the results of using the trained second model.

[0129] Fig. 7 is a diagram showing a first example of experimental results in the embodiment. Fig. 8 is a diagram showing a second example of experimental results in the embodiment. More specifically, Fig. 7 shows an example of the results using the trained second model, and Fig. 8 shows an example of the results using the technology described in Non-Patent Document 1 as is.

[0130] In the experiment that obtained the results of Figures 7 and 8, the text describing the second action was "A pelican is catching fish from the sea." Hereinafter, this text data will be referred to as experimental text data. Figure 7 shows an example of an action indicated by an action feature vector obtained from the experimental text data using the trained second model. Figure 8 shows an example of an action indicated by an action feature vector obtained from the experimental text data using the technology described in Non-Patent Document 1.

[0131] In an experiment using the trained second model, the feature vector estimation model of the second model was trained so that when the text describing the second action, "A pelican is catching fish from the sea," was converted using the trained text data estimation model into the text describing the first action, "A person dives into the sea with their arms and legs splayed out to the sides and their body submerging," the action generated by the feature vector estimation model of the trained first model would be generated by the second model.

[0132] In the experiment using the trained second model, the trained second model generated motion feature vectors directly from the experimental text data. The second model was trained so that the motion feature vectors generated from the text describing the second motion using the second model matched the motion feature vectors generated from the text describing the second motion using the first model.

[0133] Therefore, after executing the first sub-processing, the trained second model can generate a motion feature vector equivalent to the motion feature vector generated from the text describing the first action, "A person dives into the sea with their arms and legs splayed out to the sides and their body submerging," using the feature vector estimation model of the trained first model. Note that the first sub-processing is a process of converting the text describing the second action, "A pelican is catching fish from the sea," into the text describing the first action, "A person dives into the sea with their arms and legs splayed out to the sides and their body submerging," using the text data estimation model of the trained first model.

[0134] The results in Figure 7 and Figure 8 are compared. As shown by the actions indicated by symbols D101, D102, and D103 in Figure 7, when the trained second model is used, results are obtained that include actions such as spreading and closing the wings like a pelican, and bending down to bring the mouth closer to the water surface.

[0135] On the other hand, as shown by the movements indicated by symbols D201, D202, and D203 in Figure 8, when the technology described in Non-Patent Document 1 is used, the movement obtained is not one in which the wings are spread wide like a pelican, but rather one that resembles grabbing a fish with one's hands.

[0136] From the results of FIGS. 7 and 8, it can be seen that generating actions from text that deviates from the training data is easier with the trained second model than with the technique described in Non-Patent Document 1.

[0137] <Example of hardware configuration of learning device 1> 9 is a diagram showing an example of the hardware configuration of a learning device 1 according to an embodiment. The learning device 1 includes a control unit 11 having a processor 91, such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Network Processing Unit), and a memory 92, which are connected via a bus, and executes a program. By executing the program, the learning device 1 functions as a device including the control unit 11, an interface unit 12, and a storage unit 13.

[0138] More specifically, the processor 91 reads out a program stored in the storage unit 13 and stores the read out program in the memory 92. When the processor 91 executes the program stored in the memory 92, the learning device 1 functions as a device including the control unit 11, the interface unit 12, and the storage unit 13.

[0139] The control unit 11 controls the operation of each functional unit included in the learning device 1. The control unit 11 executes, for example, a first learning process. The control unit 11 may execute, for example, a second learning process. The control unit 11 acquires, for example, information stored in the memory unit 13. The process of acquiring the information stored in the memory unit 13 is specifically reading.

[0140] The interface unit 12 includes a communication interface for connecting the learning device 1 to an external device. The interface unit 12 communicates with the external device via wired or wireless communication. The external device is, for example, a device that has transmitted a first training data set. In such a case, the interface unit 12 acquires the first training data set by communicating with the device that has transmitted the first training data set. The first training data set acquired by the interface unit 12 is output to the control unit 11 or the storage unit 13. The external device is, for example, a device that has transmitted a second training data set. In such a case, the interface unit 12 acquires the second training data set by communicating with the device that has transmitted the second training data set. The second training data set acquired by the interface unit 12 is output to the control unit 11 or the storage unit 13.

[0141] The external device is, for example, the inference device 2. In such a case, when the control unit 11 executes the first learning process, the inference device 2 can execute a trained first model (hereinafter referred to as the "trained first model") obtained by executing the first learning process through communication via the interface unit 12. Furthermore, when the external device is the inference device 2 and the control unit 11 executes the first learning process and the second learning process, the inference device 2 can execute a trained second model (hereinafter referred to as the "trained second model") obtained by executing the first learning process and the second learning process through communication via the interface unit 12.

[0142] Interface unit 12 may be configured to include input devices such as a mouse, keyboard, or touch panel. Interface unit 12 may be configured as an interface that connects these input devices to learning device 1. In this way, the input device of interface unit 12 accepts input of various information to learning device 1 via wired or wireless connections. Note that various information that can be input to the communication interface of interface unit 12 does not necessarily have to be input to the communication interface of interface unit 12, but may also be input to the input device of interface unit 12.

[0143] Interface unit 12 outputs, for example, various types of information. Interface unit 12 is configured to include a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display, and a speaker. Interface unit 12 may be configured as an interface that connects these display devices or speakers to learning device 1. Therefore, interface unit 12 may output, for example, information input to an input device of interface unit 12 as an image or sound.

[0144] The storage unit 13 is configured using a computer-readable storage medium (non-transitory computer-readable recording medium) such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 13 stores various information related to the learning device 1. The storage unit 13 stores various information generated by the operation of the control unit 11, for example. The storage unit 13 may exist on a cloud, for example.

[0145] 10 is a flowchart showing a first example of the flow of processing executed by the learning device 1 in the embodiment. The control unit 11 of the learning device 1 acquires a first learning data set and a second learning data set (step S101). Next, the control unit 11 executes a first learning process (step S102).

[0146] 11 is a flowchart showing a second example of the flow of processing executed by the learning device 1 in the embodiment. The control unit 11 of the learning device 1 acquires a first learning data set and a second learning data set (step S101). Next, the control unit 11 executes a first learning process (step S102). Next, the control unit 11 executes a second learning process (step S103).

[0147] <An example of the hardware configuration of the inference device 2> 12 is a diagram showing an example of the hardware configuration of an inference device 2 in an embodiment. The inference device 2 is equipped with a control unit 21 having a processor 93 such as a CPU, GPU, or NPU, and a memory 94, which are connected by a bus, and executes a program. By executing the program, the inference device 2 functions as a device equipped with the control unit 21, an interface unit 22, and a memory unit 23.

[0148] More specifically, the processor 93 reads out a program stored in the storage unit 23 and stores the read out program in the memory 94. When the processor 93 executes the program stored in the memory 94, the inference device 2 functions as a device including the control unit 21, the interface unit 22, and the storage unit 23.

[0149] The control unit 21 controls the operation of each functional unit included in the inference device 2. When the control unit 11 of the learning device 1 executes the first learning process, the control unit 21 executes, for example, the first inference process. When the control unit 11 of the learning device 1 executes the first learning process and the second learning process, the control unit 21 executes, for example, the second inference process.

[0150] The control unit 21 acquires, for example, information stored in the storage unit 23. The process of acquiring information stored in the storage unit 23 is specifically a read process.

[0151] The interface unit 22 includes a communication interface for connecting the inference device 2 to an external device. The interface unit 22 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits second text data (hereinafter referred to as "inference target data") that is the inference target of the action feature vector obtained by the first inference process or the second inference process. In such a case, the interface unit 22 acquires the inference target data by communicating with the device that transmits the inference target data.

[0152] The external device may be, for example, the learning device 1. In such a case, the inference device 2 can execute the trained first model or trained second model obtained by the learning device 1 through communication via the interface unit 22.

[0153] The interface unit 22 may be configured to include input devices such as a mouse, keyboard, touch panel, etc. The interface unit 22 may be configured as an interface that connects these input devices to the inference device 2. In this way, the input device of the interface unit 22 accepts input of various information to the inference device 2 via wired or wireless connections. Note that the various information that can be input to the communication interface of the interface unit 22 does not necessarily have to be input to the communication interface of the interface unit 22, but may also be input to the input device of the interface unit 22.

[0154] The interface unit 22 outputs, for example, various types of information. The interface unit 22 is configured to include a display device such as a CRT display, a liquid crystal display, or an organic EL display, and a speaker. The interface unit 22 may be configured as an interface that connects these display devices or speakers to the inference device 2. Therefore, the interface unit 22 may output, for example, information input to an input device of the interface unit 22 as an image or sound.

[0155] The storage unit 23 is configured using a computer-readable storage medium device (non-transitory computer-readable recording medium) such as a magnetic hard disk device or a semiconductor storage device. The storage unit 23 stores various information related to the inference device 2. The storage unit 23 stores various information generated by the operation of the control unit 21, for example. The storage unit 23 may exist on a cloud, for example.

[0156] 13 is a flowchart showing an example of the flow of processing executed by the inference device 2 of the embodiment. The control unit 21 of the inference device 2 acquires inference target data (step S201). Next, the control unit 21 executes an inference process on the acquired inference target data, and infers an action feature vector representing the action indicated by the text of the inference target data (step S202). Note that the inference process refers to the second inference process when the control unit 11 of the learning device 1 executes the first learning process and the second learning process, and refers to the first inference process when the control unit 11 executes the first learning process but does not execute the second learning process.

[0157] The learning device 1 configured in this manner includes a control unit 11 that executes the first learning process. Therefore, when the second learning process is not performed in the first learning process, the burden required to generate time-series data indicating operations can be reduced, as described in <Effects of the First Learning Process>.

[0158] The learning device 1 configured in this manner may also include a control unit 11 that executes the first learning process and the second learning process. In this case, as described in <<Effects of executing the second learning process in addition to the first learning process>>, the burden required to generate time-series data indicating the behavior can be reduced.

[0159] Furthermore, the inference device 2 configured in this manner performs inference using the learning results obtained by the learning device 1. Therefore, the inference device 2 can reduce the burden required to generate time-series data indicating behavior.

[0160] Furthermore, the information processing system 100 configured in this manner includes the learning device 1. Therefore, the information processing system 100 can reduce the burden required to generate time-series data indicating actions.

[0161] (Variation) Although we have explained examples of generating actions based on text, actions may also be generated based on data of other modalities. Specifically, as shown in Figure 2, Figure 5, or Figure 6, when embeddings calculated by multimodal embedding are used as input to the action generator, these embeddings can also be calculated from modalities other than text, so actions may be generated from modalities other than text, such as images, videos, or sounds.

[0162] Furthermore, up to this point, the explanation has been given taking as an example the case where the generated time series data is time series data related to motion. However, as mentioned above, the generated time series data does not need to be related to motion, and may be time series data representing any predetermined type of event. Hereinafter, a predetermined type of event will be referred to as a target event. Motion is an example of a target event. Examples of target events include facial expressions (e.g., first event: human facial expression, second event: animal facial expression), vocalizations (e.g., first event: human utterance, second event: animal utterance), and musical instrument playing (e.g., first event: piano playing, second event: guitar playing).

[0163] Note that when the modal is text, the text data is image data when the modal is an image, video data when the modal is a video, and audio data when the modal is audio.

[0164] What has been described up to this point as a first action is a first type of target event (hereinafter referred to as a "first event") if the target event is not limited to an action. What has been described up to this point as a second action is a second type of target event (hereinafter referred to as a "second event") if the target event is not limited to an action. The second event is an event different from the first event.

[0165] When the target event is not limited to a movement, the first feature vector is a feature vector indicating the first event.When the target event is not limited to a movement, the second feature vector is a feature vector indicating the second event.

[0166] The text data estimation model is an example of a modal data estimation model. The modal data estimation model is a mathematical model that estimates first modal data using second modal data, which is modal data related to a second event. The first modal data is modal data related to the first event.

[0167] When the modal is an image, the feature vector estimation model estimates the first feature vector using image data. When the modal is a video, the feature vector estimation model estimates the first feature vector using video data. When the modal is sound, the feature vector estimation model estimates the first feature vector using sound data. In this way, the feature vector estimation model is a mathematical model that estimates the first feature vector using modal data, which is data of a modal.

[0168] When the modal is not limited to text and the target event is not limited to action, the first sub-learning is training of a feature vector estimation model using the first modal data and first feature vector in the first training data set, and the second sub-learning is training of a modal estimation model using the first modal data and second modal data in the second training data set.

[0169] When the modal is not limited to text and the target event is not limited to action, the model for the second learning process is a first model (i.e., a learned first model) that includes a learned modal data estimation model obtained by the second sub-learning and a learned feature vector estimation model obtained by the first sub-learning.

[0170] Note that the embedding calculated by multimodal embedding is the process represented by c or \hat{c} in Figure 2, Figure 5, or Figure 6.

[0171] The learning device 1 may be implemented using a plurality of information processing devices connected to each other via a network so that they can communicate with each other. In this case, the processes executed by the control unit 11 may be distributed among the plurality of information processing devices.

[0172] The inference device 2 may be implemented using a plurality of information processing devices connected to each other so as to be able to communicate via a network. In this case, each process executed by the control unit 21 may be distributed and executed by the plurality of information processing devices.

[0173] Learning device 1 and inference device 2 do not necessarily have to be mounted in different housings. Learning device 1 and inference device 2 may be mounted in a single housing. In such a case, control unit 11 and control unit 21 do not necessarily have to be different units, and a single control unit may execute the processes executed by control unit 11 and control unit 21.

[0174] All or part of the functions of the information processing system 100, the learning device 1, and the inference device 2 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via a telecommunications line.

[0175] The control unit 21 is an example of an inference unit.

[0176] Although an embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]

[0177] 100...information processing system, 1...learning device, 2...inference device, 11...control unit, 12...interface unit, 13...storage unit, 21...control unit, 22...interface unit, 23...storage unit, 91...processor, 92...memory, 93...processor, 94...memory

Claims

1. a control unit that performs a first learning process to perform learning based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event; Equipped with a first model that is a mathematical model of a learning target in the first learning process includes a modal data estimation model that is a mathematical model that estimates first modal data using second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data; in the first model, the input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, In the first learning process, the control unit a first sub-learning step for training the feature vector estimation model using the first modal data and the first feature vector in the first training data set; a second sub-learning step for training the modal data estimation model using the first modal data and the second modal data in the second training data set; Run the control unit, after executing the first learning process, executes a second learning process which is a process for learning a second model which is a feature vector estimation model that estimates a first feature vector based on second modal data; In the second learning process, the second model is updated so as to reduce a difference between a first feature vector estimated by a model for the second learning process, which is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, based on the second modal data in the second learning data set, and a first feature vector estimated by the second model that uses the second modal data in the second learning data set as the input modal data. Learning device.

2. the modal is text; The learning device according to claim 1 .

3. a control unit that performs a first learning process to perform learning based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event; Equipped with a first model that is a mathematical model of a learning target in the first learning process includes a modal data estimation model that is a mathematical model that estimates first modal data using second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data; in the first model, the input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, In the first learning process, the control unit a first sub-learning step for training the feature vector estimation model using the first modal data and the first feature vector in the first training data set; a second sub-learning step for training the modal data estimation model using the first modal data and the second modal data in the second training data set; To execute Learning device.

4. a control unit that performs a first learning process that performs learning based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein a first model that is a mathematical model of a learning target in the first learning process includes a modal data estimation model that is a mathematical model that estimates the first modal data using second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data, and in the first model, input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit performs a first learning process that performs learning based on the first modal data and a previous feature vector in the first learning data set. an inference unit that performs inference using the trained second model obtained by the learning device, the inference unit executing a first sub-learning process that is training of the feature vector estimation model using the first feature vector, and a second sub-learning process that is training of the modal data estimation model using the first modal data and the second modal data in the second training data set, wherein the control unit executes a second learning process that is a process of learning a second model that is a feature vector estimation model that estimates a first feature vector based on second modal data, after executing the first learning process, and the second model is updated in the second learning process so as to reduce a difference between a first feature vector estimated by a model for second learning process that is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, based on the second modal data in the second training data set, and a first feature vector estimated by the second model that uses the second modal data in the second training data set as the input modal data; An inference device comprising:

5. a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event; and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein the first model, which is a mathematical model of a learning target in the first learning process, is a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a first feature vector that estimates the first modal data using the modal data that is modal data. an inference unit that performs inference using the trained first model obtained by a learning device, the inference unit including a feature vector estimation model that is a mathematical model for estimating a vector of a feature vector, wherein in the first model, input modal data that is the modal data used in estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit performs, in the first learning process, a first sub-learning that is training of the feature vector estimation model using the first modal data and the first feature vector in the first training data set, and a second sub-learning that is training of the modal data estimation model using the first modal data and the second modal data in the second training data set; An inference device comprising:

6. a control unit that performs a first learning process based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event; and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, wherein a first model that is a mathematical model to be learned in the first learning process includes a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a feature vector estimation model that is a mathematical model that estimates a first feature vector using modal data that is modal data, and in the first model, input modal data that is the modal data used in estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit performs a first learning process based on the first modal data in the first learning data set. a first sub-learning process that is training of the feature vector estimation model using modal data and the first feature vector in the second training data set, and a second sub-learning process that is training of the modal data estimation model using the first modal data and the second modal data in the second training data set, wherein the control unit, after executing the first learning process, executes a second learning process that is a process of learning a second model that is a feature vector estimation model that estimates a first feature vector based on second modal data, and in the second learning process, the second model is updated so as to reduce a difference between a first feature vector estimated by a model for second learning process that is the first model including the trained modal data estimation model obtained by the second sub-learning and the trained feature vector estimation model obtained by the first sub-learning, based on the second modal data in the second training data set, and a first feature vector estimated by the second model that uses the second modal data in the second training data set as the input modal data, a first learning step in which the control unit executes the first learning process; a second learning step in which the control unit executes the second learning process; A learning method that has

7. a control unit that performs a first learning process that performs learning based on a first learning data set that is a combination of first modal data, which is modal data related to a first event that is a first type of event, and a first feature vector, which is a feature vector indicating the first event, and a second learning data set that is a combination of first modal data that is the same as or different from the first modal data, and second modal data, which is modal data related to a second event that is a type of event different from the first event, and the first model that is a mathematical model to be learned in the first learning process includes a modal data estimation model that is a mathematical model that estimates the first modal data using the second modal data, and a modal data estimation model that uses modal data that is modal data. a feature vector estimation model that is a mathematical model that estimates a first feature vector using the first feature vector estimation model, wherein in the first model, input modal data that is the modal data used for estimation by the feature vector estimation model is an estimation result of the modal data estimation model, and the control unit executes, in the first learning process, a first sub-learning that is training of the feature vector estimation model using the first modal data and the first feature vector in the first training data set, and a second sub-learning that is training of the modal data estimation model using the first modal data and the second modal data in the second training data set, a first learning step in which the control unit executes the first learning process; A learning method that has

8. A program for causing a computer to function as either the learning device according to any one of claims 1 to 3 or the inference device according to claim 4 or 5.