Learning device, learning method, and learning program

JPWO2025046718A5Pending Publication Date: 2026-05-11
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2023-08-29
Publication Date
2026-05-11

Smart Images

  • Figure 2025046718000001
    Figure 2025046718000001
Patent Text Reader

Abstract

This information processing device calculates a feature amount from training data including image data of time-series data, and determines parameters of a learning model so as to decrease an evaluation value related to a difference in the feature amount over a period.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, learning method, and learning program

[0001] The present disclosure relates to a learning device, a learning method, and a learning program.

[0002] A technology has been proposed that generates a control plan for a control object, such as a robot arm, using a trained model generated by machine learning. It has also been proposed that such a technology predicts future feature quantities related to the control object and uses the predicted feature quantities to increase the accuracy of the generated control plan. Patent Document 1 (JP-A-2011-102526) describes an example of a technology for predicting future feature quantities. Patent Document 1 describes acquiring multiple training data sets, each of which is composed of a combination of feature information included in a first sample of predetermined data at a first time and a second sample of the predetermined data at a second time that is future than the first time, and training a prediction model by machine learning to predict feature information at the second time from the first sample at the first time for each training data set.

[0003] Japanese Patent Application Publication No. 2020-068028

[0004] However, the technology described in Patent Document 1 has the problem that the transition of feature information may become complex or irregular, and when a learning model is trained using such feature information, the learning becomes unstable, making it impossible to improve the estimation accuracy of the learning model.

[0005] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology that can stabilize learning of a learning model using features.

[0006] A learning device according to an exemplary aspect of the present disclosure includes a calculation means for calculating features from training data including image data of time-series data, and a determination means for determining parameters of a learning model so as to reduce an evaluation value relating to the difference in the features over a period of time.

[0007] A learning method according to an exemplary aspect of the present disclosure includes a calculation process in which at least one processor calculates features from training data including image data of time-series data, and a determination process in which the at least one processor determines parameters of a learning model so as to reduce an evaluation value related to a difference in the features over a period of time.

[0008] A program according to an exemplary aspect of the present disclosure is a program that causes a computer to function as a learning device, and causes the computer to function as a calculation means that calculates features from training data that includes image data of time-series data, and a determination means that determines parameters of a learning model so as to reduce an evaluation value related to differences in features over a period of time.

[0009] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that learning of a learning model using feature quantities can be stabilized.

[0010] FIG. 1 is a block diagram showing a configuration of a learning device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of a learning method according to the present disclosure. FIG. 3 is a diagram for explaining an overview of a process for evaluating a behavior sequence using a learning model learned by an information processing device according to the present disclosure. FIG. 4 is a block diagram showing an example of a configuration of an information processing device according to the present disclosure. FIG. 5 is a block diagram showing an example of a functional configuration and a processing flow of an information processing device according to the present disclosure. FIG. 6 is a diagram showing an example of training data according to the present disclosure. FIG. 7 is a diagram showing an example of a configuration of a feature prediction unit according to the present disclosure. FIG. 8 is a diagram showing an example of a change in feature quantity over time according to the present disclosure. FIG. 9 is a diagram showing an example of a configuration of a regularization calculation unit according to the present disclosure. FIG. 10 is a diagram showing an example of a distance between directional vectors according to the present disclosure. FIG. 11 is a diagram showing an example of a graph showing a change in distance over time according to the present disclosure. FIG. 12 is a diagram showing an example of a configuration of a regularization calculation unit according to the present disclosure. FIG. 13 is a diagram showing a specific example of a reference directional vector according to the present disclosure. FIG. 14 is a diagram showing an example of a configuration of a regularization calculation unit according to the present disclosure. FIG. 15 is a diagram showing an example of a distribution of directional vectors according to the present disclosure. FIG. 16 is a block diagram showing an example of a configuration of an information processing device according to the present disclosure. FIG. 17 is a block diagram showing an example of a functional configuration and a processing flow of an information processing device according to the present disclosure.

[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0013] <Configuration of Learning Device> The configuration of the learning device 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the learning device 1. As shown in Fig. 1, the learning device 1 includes a calculation unit 11 and a determination unit 12. The calculation unit 11 calculates feature amounts from training data including image data of time-series data. The determination unit 12 determines parameters of a learning model so as to reduce an evaluation value related to differences in feature amounts over a period of time.

[0014] <Effects of the Learning Device> As described above, the learning device 1 employs a configuration including a calculation unit 11 that calculates feature quantities from training data including image data of time-series data, and a determination unit 12 that determines parameters of a learning model so as to reduce an evaluation value related to a difference in feature quantities over a period of time. Therefore, the learning device 1 has the effect of stabilizing learning of a learning model using feature quantities.

[0015] <Flow of Learning Method> The flow of learning method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of learning method S1. As shown in Fig. 2, learning method S1 includes calculation processing S11 and determination processing S12. In calculation processing S11, at least one processor calculates feature amounts from training data including image data of time-series data. In determination processing S12, at least one processor determines parameters of the learning model so as to reduce an evaluation value related to differences in feature amounts over a period of time.

[0016] <Effects of the Learning Method> As described above, the learning method S1 employs a configuration including a calculation process, performed by at least one processor, for calculating feature quantities from training data including image data of time-series data, and a determination process, performed by at least one processor, for determining parameters of a learning model so as to reduce an evaluation value related to a difference in feature quantities over a period of time. Therefore, the learning method S1 has the effect of stabilizing learning of a learning model using feature quantities.

[0017] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0018] <Overview of Information Processing Device> The information processing device 1A according to the present disclosure is a device that trains a learning model and is an example of a learning device according to the present disclosure. One example of the learning model trained by the information processing device 1A is a model for predicting future feature quantities from data related to a control object. Examples of the control object include, but are not limited to, moving objects such as a robot arm or an autonomous vehicle. One example of the learning model trained by the information processing device 1A is used to evaluate behavior information of the control object.

[0019] 3 is a diagram for explaining an overview of the evaluation process of behavioral information using a learning model trained by the information processing device 1A. In the example of FIG. 3, an extraction model LM1 and a prediction model LM2 are learning models trained by the information processing device 1A. In this example, data I 0 From feature e 0 The data I input to the extraction model LM1 is extracted. 0 is data relating to a control target, and is, for example, image data obtained by capturing an image of the control target or the environment of the control target.

[0020] Data I is extracted using the extraction model LM1. 0 The feature e extracted from 0 is input to the prediction model LM2. t-1 and the time series of the behavior information of the control target (a1, a 2, ...) and future feature value e t That is, the feature value e is predicted by the prediction model LM2. 0 and the time series of behavioral information (a 1, a 2, ...) and the time series of features (e 0, e 1, ..., e T ) is generated.

[0021] In addition, the generated time series (e 0, e 1, ..., e T ) and the time series of behavioral information (a 1, a 2, The success probability of the sequence of actions (a...) is predicted by the evaluation model LM3. 1, a 2, The evaluation results of the above (...) are referenced, for example, to generate control information that indicates what kind of control should be performed on the control target.

[0022] <Configuration of Information Processing Device> Fig. 4 is a block diagram showing an example of the configuration of the information processing device 1 A. As shown in Fig. 4, the information processing device 1 A includes a control unit 10 A, a storage unit 20 A, a communication unit 30 A, and an input / output unit 40 A.

[0023] (Communication Unit) The communication unit 30A communicates with devices external to the information processing device 1A via a communication line. While the specific configuration of the communication line does not limit the present exemplary embodiment, examples of the communication line include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices, and supplies data received from other devices to the control unit 10A.

[0024] (Input / Output Unit) Input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel are connected to the input / output unit 40A. The input / output unit 40A receives various types of information input to the information processing device 1A from the connected input devices. Furthermore, under the control of the control unit 10A, the input / output unit 40A outputs various types of information to connected output devices. Examples of the input / output unit 40A include an interface such as a USB (Universal Serial Bus).

[0025] (Control Unit) As shown in FIG. 4, the control unit 10A includes a data selection unit 11A, a feature extraction unit 12A, a feature prediction unit 13A, and an update unit 14A. The feature extraction unit 12A is an example of a calculation means according to the present disclosure. The update unit 14A is an example of a determination means according to the present disclosure. FIG. 5 is a block diagram showing an example of a functional configuration and processing flow of the information processing device 1A. Note that the unidirectional arrows in FIG. 5 simply indicate the direction of flow of a certain signal (data) or the direction of flow of processing, and do not exclude bidirectionality. The feature prediction unit 13A is an example of a prediction means according to the present disclosure.

[0026] (Data Selection Unit) The data selection unit 11A acquires training data to be used for learning. As an example, the data selection unit 11A randomly selects training data to be used for learning from the training data TD stored in the storage unit 20A. The data selection unit 11A may acquire training data input via the input / output unit 40A, or may acquire training data from a storage location specified by the user of the information processing device 1A (which may be within the storage device of the information processing device 1A or a storage device external to the information processing device 1A). The data selection unit 11A may also receive training data via the communication unit 30A.

[0027] (Training Data) The training data TD is data used for training a learning model. The training data TD includes image data of time-series data. The training data TD may further include behavioral information representing behavior during a period.

[0028] 6 is a diagram showing an example of training data TD. The training data TD shown in FIG. 6 is a time series I0 , I 1 , ... and the time series a of the behavior information of the control target 1 , a 2 , .... Image data I 0 , I 1 , . . . are, for example, moving image data obtained by capturing at least one of the controlled object and the environment in which the controlled object operates.

[0029] The behavior information includes at least one of an behavior category selected from a group of behavior categories including a plurality of behavior categories and an input value to a robot joint. Here, a behavior category is a class indicating a behavior of a robot arm, such as "grab," "release," or "move." The input value to a robot joint includes at least one of, for example, a joint angle, speed, acceleration, or torque value. To complete a certain behavior category, input values ​​to multiple joints are required. Hereinafter, a sequence of behavior information will also be referred to as an "behavior sequence."

[0030] (Feature Extraction Unit / Feature Amount) The feature extraction unit 12A calculates feature amounts from the training data TD using the extraction model LM1. The feature amounts are information representing the characteristics of the data, and include, for example, information indicating at least one of: (i) feature amounts obtained from an image; (ii) a vector listing the actions of the control object or feature amounts obtained from the vector; and (iii) the position and posture of an object in the image. For example, the feature amounts obtained from the image are feature amounts obtained by inputting images of the control object and / or the environment into a trained model such as a neural network. For example, the object in the image is an object that is the target of the robot arm's operation (grasping, moving, etc.). However, the feature amounts are not limited to the above example, and may include other information.

[0031] (Extraction Model) The extraction model LM1 is a model that extracts features from data. The extraction model LM1 is, for example, a neural network. The input of the extraction model LM1 is image data included in the training data TD, or image data and behavior included in the training data TD. The output of the extraction model LM1 includes features.

[0032] (Feature Prediction Unit) The feature prediction unit 13A uses a learning model to predict future feature quantities from the behavioral information included in the training data TD and the feature quantities extracted by the feature extraction unit 12 A. As an example, the feature prediction unit 13A uses a prediction model LM2 to predict future feature quantities from the feature quantities extracted by the feature extraction unit 12A and the behavioral information included in the training data TD.

[0033] (Prediction Model) The prediction model LM2 is a model for predicting future feature quantities from feature quantities and the behavior of the control target. The prediction model LM2 is, for example, a neural network. The prediction model LM2 may use a deterministic method or a probabilistic method. The prediction model LM2 includes, for example, at least one of a recurrent neural network and a graph neural network. Examples of recurrent neural networks include, but are not limited to, a recurrent neural network (RNN), a long short-term memory (LSTM), and a gated recurrent unit (GRU). The prediction model LM2 may also include, for example, a variational autoencoder (VAE).

[0034] The input of the prediction model LM2 is the action a at time t. t and the feature value e at time (t-1) t-1 The output of the prediction model LM2 is the feature value ^e at time t. t The feature ^e, which is the output of the prediction model LM2, is included. t As an example, the sequence of actions (a 1 , a 2 , ..., a T ) is used to calculate the success probability.

[0035] 7 is a diagram showing an example of the configuration of the feature prediction unit 13A. In the example of FIG. 7, the feature prediction unit 13A includes a prediction unit 131A and a state storage unit 132A. Note that the unidirectional arrows in FIG. 7 simply indicate the direction of flow of a certain signal (data) or the direction of flow of a process, and do not exclude bidirectionality. The prediction unit 131A predicts the feature amount e t-1 and the action sequence (a 1, a2, ...) and the prediction model LM2, the future feature value ^e t The state storage unit 132A stores the time series of future feature amounts (^e 1 , ^e 2 , ..., ^e T ) is stored and passed to the loss function calculation unit 15A and the regularization calculation unit 16A.

[0036] (Update Unit) The update unit 14A updates the model parameters of at least one of the extraction model LM1 and the prediction model LM2 so as to reduce the difference (deviation, distance, etc.) between the feature amount extracted by the feature extraction unit 12A and the feature amount predicted by the feature prediction unit 13A. More specifically, as an example, the parameter update unit 17A updates the model parameters using the mean square error between the feature amount extracted by the feature extraction unit 12A and the feature amount predicted by the feature prediction unit 13A.

[0037] At this time, the update unit 14A determines the parameters of the learning model so that the evaluation value relating to the difference in the feature amount over a predetermined period of time decreases.

[0038] As an example, the evaluation value represents the change between the difference in feature amounts in a first period and the difference in feature amounts in a second period different from the first period. In this case, a small evaluation value can be said to indicate that the time change in the difference in feature amounts before and after the transition is smooth. That is, in this case, the update unit 14A updates the model parameters so as to smooth the transition of feature amounts in at least one of the time series of feature amounts extracted by the feature extraction unit 12A and the time series of feature amounts predicted by the feature prediction unit 13A.

[0039] Fig. 8 is a diagram showing an example of changes over time in feature quantities. In the example of Fig. 8, the directions of directional vectors d101 to d105 indicating transitions in feature quantities vary and are inconsistent, whereas the directions of directional vectors d201 to d204 are consistent. In other words, it can be said that the changes over time in feature quantities transitioning along directional vectors d201 to d204 are consistent.

[0040] The evaluation value may represent, for example, a change between a difference in a feature amount in a first time period and a difference in a feature amount in a second time period different from the first time period, and the evaluation value may be larger as the behavioral information in the first time period and the behavioral information in the second time period are more similar. The evaluation value may be larger as the difference between the time periods deviates from a standard difference for each piece of behavioral information.

[0041] 5, the update unit 14A includes a loss function calculation unit 15A, a regularization calculation unit 16A, and a parameter update unit 17A. The loss function calculation unit 15A, the regularization calculation unit 16A, and the parameter update unit 17A are examples of a loss function calculation means, a regularization term calculation means, and a parameter update means, respectively, according to the present disclosure.

[0042] The loss function calculation unit 15A calculates a loss function using the feature amounts extracted by the feature extraction unit 12A and the feature amounts predicted by the feature prediction unit 13A. The regularization calculation unit 16A calculates a regularization term using at least one of the time series of the feature amounts extracted by the feature extraction unit 12A and the time series of the feature amounts predicted by the feature prediction unit 13A. The regularization term is an example of an evaluation value according to the present disclosure. Details of the calculation process performed by the regularization calculation unit 16A will be described later.

[0043] The parameter update unit 17A updates the parameters of the learning model using a second loss function determined by at least one of the time series of the feature amounts extracted by the feature extraction unit 12A and the time series of the feature amounts predicted by the feature prediction unit 13A and the loss function calculated by the loss function calculation unit 15A. More specifically, as an example, the parameter update unit 17A updates the parameters of the learning model using a second loss function that includes the regularization term calculated by the regularization calculation unit 16A in the loss function calculated by the loss function calculation unit 15A.

[0044] For example, the parameter update unit 17A outputs the model parameters by writing the model parameters to the storage unit 20A. However, the parameter update unit 17A may output the model parameters by transmitting the updated model parameters via the communication unit 30A, or may output the model parameters via the input / output unit 40A. Furthermore, the model parameters may be output by writing them to a storage destination designated by the user of the information processing device 1A (which may be within the storage device of the information processing device 1A or may be a storage device external to the information processing device 1A).

[0045] (Storage Unit) The storage unit 20A stores training data TD, feature values ​​FT, an extraction model LM1, and a prediction model LM2. The training data TD is data used for learning at least one of the extraction model LM1 and the prediction model LM2. The feature values ​​FT include feature values ​​extracted by the feature extraction unit 12A and feature values ​​predicted by the feature prediction unit 13A. Storing the extraction model LM1 and the prediction model LM2 in the storage unit 20A means that model parameters defining the extraction model LM1 and model parameters defining the prediction model LM2 are stored in the storage unit 20A.

[0046] <Configuration Examples of Regularization Calculation Unit> Next, configuration examples 1 to 4 will be described as configuration examples of the regularization calculation unit 16A. In the following examples, the regularization calculation unit 16A generates a time series of directional vectors representing differences in feature quantities at different times, using at least one of the time series of feature quantities extracted by the feature extraction unit 12A and the time series of feature quantities predicted by the feature prediction unit 13A. The regularization calculation unit 16A also calculates regularization terms including variables determined by the generated directional vectors.

[0047] (Configuration Example 1 of Regularization Calculation Unit) FIG. 9 is a diagram showing an example of the configuration of the regularization calculation unit 16A. In the example of FIG. 9, the regularization calculation unit 16A includes a direction vector calculation unit 161A, a distance calculation unit 162A, and a regularization term calculation unit 163A. Note that the unidirectional arrows in FIG. 9 simply indicate the direction of flow of a certain signal (data) or the direction of flow of a process, and do not exclude bidirectionality. The direction vector calculation unit 161A calculates a feature quantity e t-n(n is an integer of 1 or more) to obtain the feature value e t A direction vector d representing the transition to t If n=1, calculate the direction vector d t is the feature value e t and feature e t-1 Using and, -d t = e t -e t-1 where the direction vector d t is not limited to a vector calculated from continuous features. t may be calculated from features separated by two or more time points.

[0048] The distance calculation unit 162A calculates the direction vector d t and direction vector d t+n (n is an integer of 1 or more) t Calculate the distance θ t is, for example, the Euclidean distance or the angle between the vectors. t In the example of FIG. t is the direction vector d t and direction vector d t+1 The angle is the distance θ t is the distance between successive direction vectors (direction vector d t and direction vector d t+1 The distance θ t may be the distance between direction vectors that are two or more times apart.

[0049] In this example, the regularization term calculation unit 163A calculates the direction vector d t In other words, the regularization calculation unit 16A calculates a regularization term that increases at least one of the smoothness and sparsity of the time change of the distance between the direction vectors.

[0050] FIG. 11 shows the distance θ t 1 is a diagram showing an example of a graph showing the time change of the distance θ tWhen the graph 501 and the graph 502 are compared, the graph 502 has a larger distance θ t The time change of is smooth, and the distance θ t The frequency of time changes is low.

[0051] As an example, the regularization term calculation unit 163A calculates the distance θ t The time change of the distance θ t In this case, the regularization term calculation unit 163A calculates, as an example, a regularization term that minimizes the L2 norm of the gradient, which is expressed by the following equation:

[0052] In addition, as an example, the regularization term calculation unit 163A calculates the distance θ t The probability of occurrence of time change of the distance θ t In this case, the regularization term calculation unit 163A calculates, as an example, a regularization term that minimizes the L1 norm of the gradient, which is expressed by the following equation:

[0053] (Configuration Example 2 of Regularization Calculation Unit) FIG. 12 is a diagram showing a configuration example 2 of the regularization calculation unit 16A. Note that the unidirectional arrows in FIG. 12 simply indicate the direction of flow of a certain signal (data) or the direction of flow of a process, and do not exclude bidirectionality. In this example, the regularization calculation unit 16A calculates a regularization term so that the transition of the feature amount becomes smoother the higher the similarity between the actions before and after the transition. In the example of FIG. 12, the regularization calculation unit 16A includes a direction vector calculation unit 161A, a distance calculation unit 162A, a regularization term calculation unit 163b, and a behavioral similarity calculation unit 164b. Of these components, the direction vector calculation unit 161A and the distance calculation unit 162A are the same as the components included in FIG. 9 described above, and therefore, description thereof will not be repeated here.

[0054] The behavioral similarity calculation unit 164b calculates the behavior a t and Action A t+n Similarity with t Calculate the similarity w t As an example, (i) action a tBehavioral categories and behavior a t+n If the behavior categories are different, the value is 0, and if the behavior categories are the same, the value may be greater than 0 (for example, "1"). t For example, the similarity w may be the degree of similarity of the input values ​​to the robot joints (for example, the inner product of vectors in which the input values ​​are arranged as components). t may be a value obtained by integrating the value of (i) above and the value of (ii) above. Examples of the value obtained by integrating (i) and (ii) above include, but are not limited to, the multiplication value, addition value, or weighted sum of (i) and (ii) above.

[0055] The regularization term calculation unit 163b adjusts at least one of smoothness and sparsity according to the similarity of the transition behavior. t At least one of the smoothness and sparsity of the time change of is determined by the similarity of the behavior before and after the transition w t The higher the value of , the larger the regularization term is calculated.

[0056] As an example, the regularization term calculation unit 163b calculates a regularization term that smooths the change in distance over time as the behavior before and after the transition becomes more similar. In this case, as an example, the regularization term calculation unit 163b calculates a regularization term expressed by the following formula: t is the similarity calculated by the behavioral similarity calculation unit 164b.

[0057] In addition, as an example, the regularization term calculation unit 163b calculates a distance θ t Calculate a regularization term that reduces the probability of occurrence of time changes in .

[0058] By using the regularization term calculated by the regularization term calculation unit 163b, the parameter update unit 17A updates the model parameters so that the transition of the feature in at least one of the time series of the feature extracted by the feature extraction unit 12A and the time series of the feature predicted by the feature prediction unit 13A becomes smoother the higher the similarity between the behavior before and after the transition.

[0059] (Configuration Example 3 of Regularization Calculation Unit) FIG. 13 is a diagram showing Configuration Example 3 of the regularization calculation unit 16A. Note that the unidirectional arrows in FIG. 13 simply indicate the direction of flow of a certain signal (data) or the direction of flow of a process, and do not exclude bidirectionality. In this example, the regularization calculation unit 16A calculates a regularization term so that directional vectors of the same behavior category are similar. In FIG. 13, the regularization calculation unit 16A includes a directional vector calculation unit 161A, a distance calculation unit 162c, a regularization term calculation unit 163c, and a reference vector calculation unit 165c. The processing performed by the directional vector calculation unit 161A is similar to the processing performed by the directional vector calculation unit 161A shown in FIG. 9 above, and therefore, description thereof will not be repeated here.

[0060] The reference vector calculation unit 165c generates a reference direction vector from the direction vector and the time series of behavior information. The reference direction vector is a direction vector that serves as a reference for each behavior category. Examples of the reference direction vector include an average vector of direction vectors corresponding to behavior information belonging to a certain behavior category, or a vector representing the difference between the first feature value and the last feature value in a time series of feature values ​​belonging to the same behavior category. However, the reference direction vector is not limited to the above example. As an example, the reference direction vector may be a vector determined by the user for each behavior category.

[0061] 14 is a diagram showing a specific example of the reference direction vector bc. In the example of FIG. 14, the direction vector d 11 , d 12 , d 13 indicates the transition of the feature amount in the behavior belonging to the behavior category c1. Also, the direction vector d 21 , d 22 , d 23 indicates the transition of the feature amount in the behavior belonging to the behavior category c2. 31 , d 32 indicates the transition of the feature amount in the behavior belonging to the behavior category c3. In this case, the reference direction vector b c1 As an example, the feature value e 11 ~e 14The reference direction vector b c2 As an example, the feature value e 14 ~e 17 The reference direction vector b of the behavior category c3 is a vector representing the difference between the c3 As an example, the feature value e 17 ~e 19 is a vector representing the difference between

[0062] The distance calculation unit 162c calculates the direction vector d i (0≦i≦t) and the reference direction vector b c Calculate the distance between the direction vector d and i and the reference direction vector b c The distance between is, for example, the direction vector d i and the reference direction vector b c is the Euclidean distance, or the angle between the vectors.

[0063] The regularization term calculation unit 163c calculates the reference direction vector b c In other words, the regularization term calculation unit 163c calculates a regularization term that minimizes the distance between the direction vector d i and the reference direction vector b c Calculate the regularization term determined by the distance between

[0064] As an example, the regularization term calculation unit 163c calculates the regularization term expressed by the following equation: c is a set of times that belong to the behavior category c.

[0065] By using the regularization term calculated by the regularization term calculation unit 163c, the parameter update unit 17A updates the model parameters so that the transition patterns of the features in at least one of the time series of the features extracted by the feature extraction unit 12A and the time series of the features predicted by the feature prediction unit 13A become more similar as the similarity of the behaviors increases.

[0066] (Configuration Example 4 of Regularization Calculation Unit) FIG. 15 is a diagram showing Configuration Example 4 of the regularization calculation unit 16A. Note that the unidirectional arrows in FIG. 15 simply indicate the direction of flow of a certain signal (data) or the direction of flow of a process, and do not exclude bidirectionality. In FIG. 15, the regularization calculation unit 16A calculates regularization terms so that directional vectors of different behavior categories are not similar. In the example of FIG. 15, the regularization calculation unit 16A includes a direction vector calculation unit 161A, a distance calculation unit 162d, a regularization term calculation unit 163d, a reference vector calculation unit 165d, and a behavior similarity calculation unit 166d.

[0067] The reference vector calculation unit 165d generates a reference direction vector for each behavior category using the time series of the direction vector. The distance calculation unit 162d calculates the distance between the reference direction vectors of different behavior categories. The behavior similarity calculation unit 166d calculates the reference direction vector b i Behavioral category and reference direction vector b j The similarity δ(i, j) between the behavior category of i and j is calculated.

[0068] The regularization term calculation unit 163d calculates a regularization term that increases the distance between the reference direction vectors as the similarity between the actions decreases. In other words, the regularization term calculation unit 163d calculates a regularization term that separates clusters of direction vectors belonging to different action categories. As an example, the regularization term calculation unit 163d calculates a regularization term that increases the distance between the reference direction vectors b 1 and b 2 A regularization term that maximizes the distance between is calculated. The regularization term is expressed by the following formula, for example:

[0069] Here, w i,j may be a constant determined by the user, or may be the similarity of the input values ​​to the robot joints (for example, the inner product of a vector in which the input values ​​are arranged as components). i Behavioral category and reference direction vector b j The similarity to the behavior category.

[0070] 16 is a diagram illustrating an example of the distribution of directional vectors. In the example of Fig. 16, the regularization term calculation unit 163b calculates a regularization term so as to separate a directional vector group G11 of a first behavior category from a directional vector group G12 of a second behavior category in a predetermined vector space.

[0071] (Effects of Information Processing Device) As described above, the information processing device 1A determines the parameters of the learning model so that the transition of the feature quantities becomes smoother. If the transition of the feature quantities is complex, the learning of the learning model may become unstable, which may result in a decrease in the estimation accuracy of the learning model. In contrast, according to the information processing device 1A, by performing learning so that the transition of the feature quantities becomes smoother, it is possible to stabilize the learning and increase the accuracy of future predictions.

[0072] Furthermore, the information processing device 1A is configured such that the evaluation value represents the change between the difference in feature quantity in a first period and the difference in feature quantity in a second period different from the first period. Therefore, according to the information processing device 1A, by determining parameters such that the change between the difference in feature quantity in the first period and the difference in feature quantity in the second period is reduced, learning can be performed so that the transition of feature quantity becomes smooth.

[0073] Furthermore, in the information processing device 1A, the training data further includes behavioral information representing behavior over a period of time, the evaluation value represents the change between the difference in feature values ​​over a first period of time and the difference in feature values ​​over a second period of time that is different from the first period of time, and the evaluation value increases as the behavioral information over the first period of time and the behavioral information over the second period of time become more similar. Therefore, according to the information processing device 1A, by determining parameters that reduce the change between the difference in feature values ​​over the first period of time and the difference in feature values ​​over the second period of time, learning can be performed to smooth the transition of feature values.

[0074] Furthermore, in the information processing device 1A, the training data further includes behavioral information representing behavior during a period, and the evaluation value increases as the difference between the period and the standard difference for each behavioral information deviates. Therefore, the information processing device 1A can train a learning model so that the transition patterns of features are similar when the behavioral information is the same.

[0075] In addition, in the information processing device 1A, the training data further includes behavioral information representing behavior over a period of time, and the information processing device 1A further includes a feature prediction unit 13A that predicts future feature quantities using the learning model from the behavioral information and the feature quantities calculated by the feature extraction unit 12A, and the update unit 14A is configured to update the parameters using a loss function calculation unit 15A that calculates a loss function using the feature quantities calculated by the feature extraction unit 12A and the feature quantities predicted by the feature prediction unit 13A, and a second loss function determined by at least one of the time series of the feature quantities calculated by the feature extraction unit 12A and the time series of the feature quantities predicted by the feature prediction unit 13A and the loss function calculated by the loss function calculation unit 15A. As a result, the information processing device 1A can train a learning model taking into account constraints on transitions of the feature quantities.

[0076] Furthermore, in the information processing device 1A, the update unit 14A further includes a regularization calculation unit 16A that calculates a regularization term using at least one of the time series of feature quantities calculated by the feature extraction unit 12A and the time series of feature quantities predicted by the feature prediction unit 13A, and the parameter update unit 17A updates the parameters using a second loss function that includes the regularization term in the loss function calculated by the loss function calculation unit 15A. As a result, the information processing device 1A can train a learning model using a loss function that includes a regularization term related to transitions in the feature quantities.

[0077] Furthermore, in the information processing device 1A, the regularization calculation unit 16A uses at least one of the time series of the feature amounts calculated by the feature extraction unit 12A and the time series of the feature amounts predicted by the feature prediction unit 13A to generate a time series of directional vectors representing differences in the feature amounts at different times, and calculates a regularization term including a variable determined by the generated directional vector. This makes it possible to train a learning model using a loss function including a regularization term related to the transition of the feature amounts.

[0078] Furthermore, the information processing device 1A employs a configuration in which the regularization calculation unit 16A calculates a regularization term that increases at least one of the smoothness and sparsity of the time change in the distance between the direction vectors. As a result, the information processing device 1A can smooth the transition of the feature values ​​used in learning and improve the estimation accuracy of the learning model.

[0079] Furthermore, the information processing device 1A employs a configuration in which the regularization calculation unit 16A calculates a regularization term that increases at least one of the smoothness and sparsity of the time change in the distance between the directional vectors as the similarity between the actions before and after the transition increases. When the actions are similar, the transition of the feature values ​​is likely to remain constant. Therefore, with this configuration, the learning model is trained so that the transition of the feature values ​​becomes smoother as the similarity between the actions increases, thereby improving the estimation accuracy of the learning model.

[0080] Furthermore, in the information processing device 1A, the regularization calculation unit 16A generates a reference direction vector from the direction vector and the time series of the behavioral information, and calculates a regularization term determined by the distance between the direction vector and the reference direction vector. Similar behaviors are likely to result in similar transitions of feature quantities. For example, by training a learning model so that the distance between the reference direction vector and the direction vector is small, the more similar the behaviors, the more similar the transitions of feature quantities can be, thereby improving the estimation accuracy of the learning model.

[0081] Furthermore, the information processing device 1A employs a configuration in which the regularization calculation unit 16A generates a reference direction vector for each behavior category using the time series of the direction vectors, and calculates the regularization term such that the distance between the reference direction vectors increases as the similarity between the behaviors decreases. When the behaviors are dissimilar, the transition patterns of the feature quantities are likely to be dissimilar as well. Therefore, according to the information processing device 1A, the learning model is trained so that the transition patterns of the feature quantities are dissimilar when the behaviors are dissimilar, thereby improving the estimation accuracy of the learning model.

[0082] Furthermore, the information processing device 1A employs a configuration in which the feature amounts include information indicating at least one of feature amounts obtained from an image, a vector listing the actions of the control target or feature amounts obtained from the vector, and the position and orientation of an object in the image. Therefore, the information processing device 1A can stabilize learning of a learning model using information indicating at least one of feature amounts obtained from an image, a vector listing the actions of the control target or feature amounts obtained from the vector, and the position and orientation of an object in the image.

[0083] <Another Configuration Example of an Information Processing Device> Fig. 17 is a block diagram showing the configuration of an information processing device 1B according to the present disclosure. A control unit 10B of the information processing device 1B includes a learning phase execution unit 110B and an estimation phase execution unit 120B. The learning phase execution unit 110B includes the data selection unit 11A, feature extraction unit 12A, feature prediction unit 13A, and update unit 14A shown in Fig. 4. The estimation phase execution unit 120B includes a data acquisition unit 21B, a second extraction unit 22B, a second prediction unit 23B, a calculation unit 24B, and an output unit 25B. The data acquisition unit 21B is an example of a second acquisition means according to the present disclosure. The calculation unit 24B is an example of a success probability calculation means according to the present disclosure.

[0084] The data acquisition unit 21B acquires data including a time series of image data and behavioral information representing behavior over a period of time. The second extraction unit 22B inputs the data acquired by the data acquisition unit 21B into an extraction model LM1, thereby extracting feature quantities from the data. The second prediction unit 23B inputs the feature quantities extracted by the second extraction unit 22B and the behavior acquired by the data acquisition unit 21B into a prediction model LM2, thereby predicting future feature quantities from the feature quantities.

[0085] The calculation unit 24B calculates the success probability of the behavior information acquired by the data acquisition unit 21B using future feature amounts obtained by inputting the data acquired by the data acquisition unit 21B into a learning model. More specifically, the calculation unit 24B calculates the success probability of the time series of behavior acquired by the data acquisition unit 21B using the future feature amounts predicted by the second prediction unit 23B. The calculation unit 24B outputs the calculated success probability to the output unit 25B. As an example, the calculation unit 24B calculates the success probability using the evaluation model LM3.

[0086] The evaluation model LM3 is a model used by the success / failure prediction unit 18A to calculate the success probability of the plan sequence PS, and is a trained model constructed by machine learning using training data. The machine learning method for the evaluation model LM3 is not limited, and as an example, a decision tree-based, linear regression, or neural network method may be used, or two or more of these methods may be used.

[0087] The success probability calculated by the calculation unit 24B is used to create a plan sequence for operating a control target such as a robot arm. For example, the information processing device 1B determines a control plan to instruct the robot arm by performing an optimization process for the plan sequence using the calculated success probability. However, the use of the success probability calculated by the information processing device 1B is not limited to the above-mentioned example.

[0088] The output unit 25B generates control information for the control target based on the success probability calculated by the calculation unit 24B, and outputs the generated control information. The output unit 25B may output the generated control information to an output device connected via the input / output unit 40A, or may transmit the control information via the communication unit 30A. Furthermore, the output unit 25B may output the control information by writing the control information to a storage destination designated by the user of the information processing device 1B (which may be within a storage device of the information processing device 1B or may be a storage device external to the information processing device 1B).

[0089] As described above, the information processing device 1B employs a configuration further including a data acquisition unit 21B that acquires image data of time-series data and behavioral information representing behavior over a period of time, and a calculation unit 24B that calculates the success probability of the behavioral information acquired by the data acquisition unit 21B using future feature quantities obtained by inputting the data acquired by the data acquisition unit 21B into a learning model. Thus, the information processing device 1B can more accurately calculate the success probability of the behavioral information.

[0090] Furthermore, the information processing device 1B is configured to further include an output unit 25B that generates control information for the control target based on the success probability calculated by the calculation unit 24B and outputs the generated control information. Therefore, the information processing device 1B can perform stable control of the control target.

[0091] <Another configuration example of an information processing device> Fig. 18 is a block diagram showing an example of the functional configuration and processing flow of an information processing device 1C according to the present disclosure. Note that the unidirectional arrows in Fig. 18 simply indicate the direction of flow of a certain signal (data) or the direction of flow of processing, and do not exclude bidirectionality.

[0092] 18, an information processing device 1C includes an image restoration unit 18C in addition to the data selection unit 11A, feature extraction unit 12A, feature prediction unit 13A, loss function calculation unit 15A, regularization calculation unit 16A, and parameter update unit 17A shown in FIG. 5.

[0093] The image restoration unit 18C restores the image using at least one of the feature amounts extracted by the feature extraction unit 12A and the feature amounts predicted by the feature prediction unit 13A, and outputs the restored image to the loss function calculation unit 15A. In this case, the loss function calculation unit 15A calculates the loss function using the feature amounts extracted by the feature extraction unit 12A and the feature amounts predicted by the feature prediction, as well as the image data included in the training data selected by the data selection unit 11A and the restored image data restored by the image restoration unit 18C.

[0094] [Example of implementation using software] Some or all of the functions of the learning device 1 and the information processing devices 1A, 1B, and 1C (hereinafter also referred to as "each of the above devices") may be implemented using hardware such as an integrated circuit (IC chip), or may be implemented using software.

[0095] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 19. Figure 19 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0096] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0097] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0098] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0099] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0100] [Appendix 1] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0101] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0102] (Appendix A1) A learning device comprising: a calculation means for calculating a feature from training data including image data of time-series data; and a determination means for determining parameters of a learning model so as to reduce an evaluation value relating to a difference in the feature over a period of time.

[0103] (Supplementary Note A2) The learning device according to Supplementary Note A1, wherein the evaluation value represents a change between a difference in a feature amount in a first period and a difference in a feature amount in a second period different from the first period.

[0104] (Appendix A3) The learning device described in Appendix A1, wherein the training data further includes behavioral information representing behavior during a period, the evaluation value represents a change between a difference in a feature value during a first period and a difference in a feature value during a second period different from the first period, and the evaluation value is larger the more similar the behavioral information during the first period and the behavioral information during the second period are.

[0105] (Appendix A4) The learning device according to Appendix A1, wherein the training data further includes behavioral information representing behavior during a period, and the evaluation value is a value that increases as the difference between the period and a standard of difference for each of the behavioral information deviates.

[0106] (Appendix A5) The learning device according to any one of Appendices A1 to A4, wherein the training data further includes behavioral information representing behavior during a period, and the device further comprises a prediction means for predicting future feature quantities using the learning model from the behavioral information and the feature quantities calculated by the calculation means, and the determination means comprises: a loss function calculation means for calculating a loss function using the feature quantities calculated by the calculation means and the feature quantities predicted by the prediction means, and a parameter update means for updating the parameters using a second loss function determined by at least one of a time series of the feature quantities calculated by the calculation means and a time series of the feature quantities predicted by the prediction means, and a loss function calculated by the loss function calculation means.

[0107] (Appendix A6) The learning device according to Appendix A5, wherein the determining means further comprises: a regularization term calculating means for calculating a regularization term using at least one of the time series of the feature amounts calculated by the calculating means and the time series of the feature amounts predicted by the predicting means; and the parameter updating means updates the parameters using the second loss function obtained by including the regularization term in the loss function calculated by the loss function calculating means.

[0108] (Appendix A7) The learning device according to Appendix A6, wherein the regularization term calculation means uses at least one of the time series of the feature values ​​calculated by the calculation means and the time series of the feature values ​​predicted by the prediction means to generate a time series of directional vectors representing differences between the feature values ​​at different times, and calculates a regularization term including a variable determined by the generated directional vector.

[0109] (Supplementary Note A8) The learning device according to Supplementary Note A7, wherein the regularization term calculation means calculates the regularization term that increases at least one of smoothness and sparsity of a change in the distance between the direction vectors over time.

[0110] (Supplementary Note A9) The learning device according to Supplementary Note A7, wherein the regularization term calculation means calculates the regularization term that increases at least one of smoothness and sparsity of a change in time of the distance between the direction vectors as the similarity between actions before and after a transition increases.

[0111] (Supplementary Note A10) The learning device according to Supplementary Note A7, wherein the regularization term calculation means generates a reference direction vector from the direction vector and the time series of the behavior information, and calculates the regularization term determined by a distance between the direction vector and the reference direction vector.

[0112] (Appendix A11) The learning device according to Appendix A7, wherein the regularization term calculation means generates a reference direction vector for each behavior category using the time series of the direction vectors, and calculates the regularization term such that the distance between the reference direction vectors increases as the similarity between the behaviors decreases.

[0113] (Appendix A12) The learning device according to any one of Appendices A1 to A11, wherein the feature amount includes information indicating at least one of a feature amount obtained from an image, a vector listing actions of a control target or a feature amount obtained from the vector, and a position and posture of an object in the image.

[0114] (Supplementary Note A13) The learning device according to any one of Supplementary Notes A1 to A12, further comprising: a second acquisition means for acquiring image data of time-series data and behavioral information representing behavior during a period; and a success probability calculation means for calculating a success probability of the behavioral information acquired by the second acquisition means using future features obtained by inputting the data acquired by the second acquisition means into the learning model.

[0115] (Supplementary Note A14) The learning device according to Supplementary Note A13, further comprising: an output unit that generates control information for a control object based on the success probability calculated by the success probability calculation unit, and outputs the generated control information.

[0116] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0117] (Appendix B1) A learning method including: a calculation process in which at least one processor calculates features from training data including image data of time series data; and a determination process in which the at least one processor determines parameters of a learning model so as to reduce an evaluation value related to a difference in features over a period of time.

[0118] (Supplementary Note B2) The learning method according to Supplementary Note B1, wherein the evaluation value represents a change between a difference in a feature amount in a first period and a difference in a feature amount in a second period different from the first period.

[0119] (Appendix B3) The learning method described in Appendix B1, wherein the training data further includes behavioral information representing behavior during a period, the evaluation value represents the change between the difference in feature values ​​during a first period and the difference in feature values ​​during a second period different from the first period, and the evaluation value is larger the more similar the behavioral information during the first period and the behavioral information during the second period are.

[0120] (Appendix B4) The learning method according to Appendix B1, wherein the training data further includes behavioral information representing behavior during a period, and the evaluation value is a value that increases as the difference between the period and a standard difference for each piece of behavioral information diverges.

[0121] (Supplementary Note B5) The learning method according to any one of Supplementary Notes B1 to B4, wherein the training data further includes behavioral information representing behavior during a period, and the at least one processor further includes a prediction process that predicts future feature quantities using the learning model from the behavioral information and the feature quantities calculated by the calculation process, and wherein in the determination process, the at least one processor executes: a loss function calculation process that calculates a loss function using the feature quantities calculated by the calculation process and the feature quantities predicted by the prediction process; and a parameter update process that updates the parameters using a second loss function determined by at least one of a time series of the feature quantities calculated by the calculation process and a time series of the feature quantities predicted by the prediction process, and the loss function calculated by the loss function calculation process.

[0122] (Supplementary Note B6) The learning method according to Supplementary Note B5, wherein in the determination process, the at least one processor further executes a regularization term calculation process that calculates a regularization term using at least one of the time series of the feature amounts calculated in the calculation process and the time series of the feature amounts predicted in the prediction process, and in the parameter update process, the at least one processor updates the parameters using the second loss function obtained by including the regularization term in the loss function calculated in the loss function calculation process.

[0123] (Appendix B7) The learning method according to Appendix B6, wherein in the regularization term calculation process, the at least one processor generates a time series of directional vectors representing differences in feature quantities at different times using at least one of the time series of feature quantities calculated in the calculation process and the time series of feature quantities predicted in the prediction process, and calculates regularization terms including variables determined by the generated directional vectors.

[0124] (Supplementary Note B8) The learning method according to Supplementary Note B7, wherein in the regularization term calculation process, the at least one processor calculates the regularization term that increases at least one of smoothness and sparsity of a change in time of the distance between the direction vectors.

[0125] (Appendix B9) The learning method according to Appendix B7, wherein in the regularization term calculation process, the at least one processor calculates the regularization term that increases at least one of smoothness and sparsity of the change in time of the distance between the direction vectors as the similarity between the actions before and after the transition increases.

[0126] (Appendix B10) The learning method according to Appendix B7, wherein in the regularization term calculation process, the at least one processor generates a reference direction vector from the direction vector and the time series of the behavior information, and calculates the regularization term determined by a distance between the direction vector and the reference direction vector.

[0127] (Appendix B11) The learning method according to Appendix B7, wherein in the regularization term calculation process, the at least one processor generates a reference direction vector for each behavior category using the time series of the direction vectors, and calculates the regularization term such that the distance between the reference direction vectors increases as the similarity between the behaviors decreases.

[0128] (Appendix B12) The learning method according to any one of Appendices B1 to B11, wherein the feature amount includes information indicating at least one of a feature amount obtained from an image, a vector listing actions of a control target or a feature amount obtained from the vector, and a position and a posture of an object in the image.

[0129] (Appendix B13) The learning method described in any one of Appendices B1 to B12, further including: a second acquisition process in which the at least one processor acquires image data of time-series data and behavioral information representing behavior during a period; and a success probability calculation process in which the at least one processor calculates a success probability of the behavioral information acquired by the second acquisition process using future features obtained by inputting the data acquired by the second acquisition process into the learning model.

[0130] (Appendix B14) The learning method according to Appendix B13, further comprising an output process in which the at least one processor generates control information for a control object based on the success probability calculated by the success probability calculation process, and outputs the generated control information.

[0131] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0132] (Appendix C1) A program that causes a computer to function as a learning device, the program causing the computer to function as: a calculation means that calculates features from training data including image data of time-series data; and a determination means that determines parameters of a learning model so as to reduce an evaluation value related to differences in features over a period of time.

[0133] (Supplementary Note C2) The learning program according to Supplementary Note C1, wherein the evaluation value represents a change between a difference in a feature amount in a first period and a difference in a feature amount in a second period different from the first period.

[0134] (Appendix C3) The learning program described in Appendix C1, wherein the training data further includes behavioral information representing behavior during a period, the evaluation value represents the change between the difference in feature values ​​during a first period and the difference in feature values ​​during a second period different from the first period, and the evaluation value is larger the more similar the behavioral information during the first period and the behavioral information during the second period are.

[0135] (Appendix C4) The learning program according to Appendix C1, wherein the training data further includes behavioral information representing behavior over a period of time, and the evaluation value is a value that increases as the difference between the period of time and a standard difference for each of the behavioral information deviates.

[0136] (Appendix C5) The learning program according to any one of Appendices C1 to C4, wherein the training data further includes behavioral information representing behavior during a period, and the computer is further caused to function as a prediction means that predicts future feature quantities using the learning model from the behavioral information and the feature quantities calculated by the calculation means, and the determination means includes: a loss function calculation means that calculates a loss function using the feature quantities calculated by the calculation means and the feature quantities predicted by the prediction means, and a parameter update process that updates the parameters using a second loss function determined by at least one of a time series of the feature quantities calculated by the calculation means and a time series of the feature quantities predicted by the prediction means, and a loss function calculated by the loss function calculation means.

[0137] (Appendix C6) The learning program according to Appendix C5, wherein the determining means further causes the computer to function as a regularization term calculating means that calculates a regularization term using at least one of the time series of the feature amounts calculated by the calculating means and the time series of the feature amounts predicted by the predicting means, and the parameter updating means updates the parameters using the second loss function that includes the regularization term in the loss function calculated by the loss function calculating means.

[0138] (Appendix C7) The learning program according to Appendix C6, wherein the regularization term calculation means uses at least one of the time series of the feature values ​​calculated by the calculation means and the time series of the feature values ​​predicted by the prediction means to generate a time series of directional vectors representing differences between the feature values ​​at different times, and calculates a regularization term including a variable determined by the generated directional vector.

[0139] (Supplementary Note C8) The learning program according to Supplementary Note C7, wherein the regularization term calculation means calculates the regularization term that increases at least one of smoothness and sparsity of a change over time in the distance between the direction vectors.

[0140] (Appendix C9) The learning program according to Appendix C7, wherein the regularization term calculation means calculates the regularization term that increases at least one of smoothness and sparsity of the change in time of the distance between the direction vectors as the similarity between actions before and after a transition increases.

[0141] (Appendix C10) The learning program according to Appendix C7, wherein the regularization term calculation means generates a reference direction vector from the direction vector and the time series of the behavior information, and calculates the regularization term determined by a distance between the direction vector and the reference direction vector.

[0142] (Appendix C11) The learning program according to Appendix C7, wherein the regularization term calculation means generates a reference direction vector for each behavior category using the time series of the direction vectors, and calculates the regularization term such that the distance between the reference direction vectors increases as the similarity between the behaviors decreases.

[0143] (Supplementary Note C12) The learning program according to any one of Supplementary Notes C1 to C11, wherein the feature amount includes information indicating at least one of a feature amount obtained from an image, a vector listing actions of a control target or a feature amount obtained from the vector, and a position and posture of an object in the image.

[0144] (Appendix C13) The learning program described in any one of Appendices C1 to C12, wherein the computer is further made to function as: a second acquisition means for acquiring image data of time-series data and behavioral information representing behavior over a period of time; and a success probability calculation means for calculating the success probability of the behavioral information acquired by the second acquisition means using future features obtained by inputting the data acquired by the second acquisition means into the learning model.

[0145] (Appendix C14) The learning program according to Appendix C13, further causing the computer to function as an output means for generating control information for a control object based on the success probability calculated by the success probability calculation means, and outputting the generated control information.

[0146] REFERENCE SIGNS LIST 1 Learning device 1A, 1B, 1C Information processing device 11 Calculation unit 12 Determination unit S1 Learning method S11 Calculation process S12 Determination process

Claims

1. A calculation method for calculating features from training data that includes time-series image data, A decision-making means for determining the parameters of a learning model so that the evaluation value regarding the difference in features over a period of time decreases, A learning device equipped with the following features.

2. The aforementioned evaluation value represents the change between the difference in feature quantities during the first period and the difference in feature quantities during a second period that is different from the first period. The learning device according to claim 1.

3. The aforementioned training data further includes behavioral information representing actions during the period, The aforementioned evaluation value represents the change between the difference in feature quantities during the first period and the difference in feature quantities during a second period that is different from the first period. The evaluation value is larger the more similar the behavioral information in the first period and the behavioral information in the second period are. The learning device according to claim 1.

4. The aforementioned training data further includes behavioral information representing actions during the period, The evaluation value is a value that becomes larger as the difference between the standard for the difference for each behavioral information and the difference over the period. The learning device according to claim 1.

5. The aforementioned training data further includes behavioral information representing actions during the period, The system further comprises a prediction means that uses the learning model to predict future features from the behavioral information and the features calculated by the calculation means, The aforementioned determination means is A loss function calculation means calculates a loss function using the features calculated by the calculation means and the features predicted by the prediction means, A parameter update means updates the parameters using a second loss function determined by at least one of the time series of the feature quantities calculated by the calculation means and the time series of the feature quantities predicted by the prediction means, and the loss function calculated by the loss function calculation means. A learning device according to any one of claims 1 to 4, comprising:

6. The aforementioned determination means is The system further comprises a regularization term calculation means that calculates a regularization term using at least one of the time series of feature quantities calculated by the calculation means and the time series of feature quantities predicted by the prediction means, The parameter update means is The parameters are updated using the second loss function, which includes the regularization term in the loss function calculated by the loss function calculation means. The learning device according to claim 5.

7. The regularization term calculation means is, Using at least one of the time series of feature quantities calculated by the calculation means and the time series of feature quantities predicted by the prediction means, a time series of direction vectors representing the difference between feature quantities at different time points is generated. Calculate the regularization term which includes a variable determined by the generated direction vector. The learning device according to claim 6.

8. The regularization term calculation means is, The regularization term is calculated to increase at least one of the smoothness and sparsity of the time change of the distance between the direction vectors. The learning device according to claim 7.

9. At least one processor performs a computation process that calculates features from training data, which includes time-series image data, The at least one processor performs a decision process to determine the parameters of the learning model such that the evaluation value for the difference in features over a period of time decreases, Learning methods that include this.

10. A learning program that makes a computer function as a learning device, The aforementioned computer, A calculation method for calculating features from training data that includes time-series image data, A decision-making means for determining the parameters of a learning model so that the evaluation value regarding the difference in features over a period of time decreases, A learning program designed to function as such.