Learning device, trained model generation method, trained model, and estimating method
The neural network system with multiple branches and attention layers, trained to minimize a loss function with cross-entropy and similarity penalties, addresses the issue of inconsistent feature focus in classical machine learning, improving the accuracy and interpretability of time-series data analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-04-02
AI Technical Summary
Existing methods for comparative analysis of time-series data, such as animal behavior or manufacturing machinery, often overlook important features due to reliance on classical machine learning techniques that manually analyze pre-formulated hypotheses and lack effective constraints on attention mechanisms, leading to inconsistent focus on relevant features.
A neural network system with multiple branches, each containing an attention layer and an output layer, trained to minimize a loss function that includes cross-entropy losses and similarity penalties, ensuring each branch focuses on distinct features for accurate classification.
The system effectively discovers and highlights multiple label-dependent features by training to minimize a loss function that penalizes similarity and variation, enhancing the accuracy and interpretability of feature extraction and classification.
Smart Images

Figure JP2025033243_02042026_PF_FP_ABST
Abstract
Description
Learning device, method for generating a trained model, trained model, and estimation method
[0001] This invention relates to a learning device that uses a neural network system.
[0002] Currently, comparative analysis using time-series data acquired by sensors regarding the behavior of animals or humans, or the operating status of manufacturing machinery, is widely practiced. For example, in animal behavior analysis, comparative analysis comparing two groups, such as an experimental group versus a control group or a male group versus a female group, is one of the most fundamental approaches. In the classical knowledge-driven approach, biologists generally identify (discover) behaviors that characterize one group, such as sex-specific locomotion strategies, by visually comparing vast amounts of time-series data. Based on these discoveries, biologists design statistical values (features) that distinguish the two groups from the behavioral data. Subsequently, biologists verify their discoveries using statistical tests (e.g., significance tests for feature differences between male and female groups) with the calculated features. However, this approach carries the potential risk of researchers overlooking important features. Although time-series data analysis based on classical machine learning is being studied, it still relies on methods that involve manually analyzing features designed based on pre-formulated hypotheses.
[0003] Non-patent documents 1-3 disclose neural networks having multiple attention mechanisms.
[0004] Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang, “Multi-head attention with disagreement regularization”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2897-2903, 2018. Jian Li, Xing Wang, Zhaopeng Tu, and Michael R Lyu, “On the diversity of multi-head attention”, Neurocomputing, 454:14-24, 2021. Takuya Maekawa, Kazuya Ohara, Yizhe Zhang, Matasaburo Fukutomi, Sakiko Matsumoto, Kentarou Matsumura, Hisashi Shidara, Shuhei J Yamazaki, Ryusuke Fujisawa, Kaoru Ide, et al, “Deep learning-assisted comparative analysis of animal trajectories with DeepHL”, Nature Communications, 11(1):5316, 2020.
[0005] However, the technology described in Non-Patent Documents 1-2 is a language processing technique. The system described in Non-Patent Documents 1-2 includes multiple attention heads. However, each of the multiple attention heads does not necessarily focus on features that contribute to group classification. Therefore, it is inconvenient for data analysis.
[0006] In the system described in Non-Patent Document 3, there are no constraints on multiple attention mechanisms focusing on the same feature. Therefore, through learning, many attention mechanisms end up focusing on the same feature. This can lead to missing features that differ between groups.
[0007] One aspect of the present invention aims to realize a learning device that contributes to data analysis by analysts.
[0008] A learning device according to one aspect of the present invention comprises a neural network system implemented by a computer and a learning unit, wherein the neural network system includes a plurality of branches into which feature data is input, each of the plurality of branches includes an attention layer that outputs attention data focusing on a portion of the feature data and an output layer that outputs estimation results based on attention-weighted feature data reflecting the attention data, and the learning unit is configured to train the neural network system to reduce a loss function that includes a plurality of losses in the plurality of output layers of the plurality of branches.
[0009] A method for generating a trained model according to one aspect of the present invention is a method for generating a trained model according to a neural network system, which involves the steps of: inputting feature data into a plurality of branches included in a trained model provided by a neural network system; an attention step in which attention data focusing on a portion of the feature data is generated in each of the plurality of branches; and an output step in which estimation results are output based on attention-weighted feature data that reflects the attention data, and then training the neural network system to reduce a loss function that includes a plurality of losses in the outputs of the plurality of branches.
[0010] A trained model according to one aspect of the present invention comprises a plurality of branches into which feature data is input, each of which includes an attention layer that outputs attention data focusing on a portion of the feature data, and an output layer that outputs estimation results based on attention-weighted feature data reflecting the attention data, thereby causing the computer to function to perform estimation by focusing on a plurality of different features.
[0011] According to one aspect of the present invention, a learning device that contributes to data analysis by analysts can be realized.
[0012] This is a block diagram showing the configuration of a learning device according to one embodiment of the present invention. This is a diagram showing the processing flow of the learning device. This is a diagram showing an example of input data for the experimental group and the control group. This is a diagram showing the change in the loss function L for the validation data during the learning process. This is the cross-entropy loss L for the validation data during the learning process. CE This figure shows the average trend of the cross-entropy loss L for the validation data during the learning process. CE This figure shows the number of times a branch appeared in which the coefficient did not decrease sufficiently. This figure shows an example of input data and attention data for the experimental group. This figure shows the Pearson correlation coefficients between the attention data and the low-frequency, high-frequency, and medium-frequency sequences. This figure shows the results of principal component analysis of the attention data obtained from each branch of the trained model in the experimental example. This figure shows a graph (left) of the nematode's behavioral trajectory (x-coordinate, y-coordinate) with emphasis on the period focused on by the attention layer of the first branch, and a graph (right) showing the cumulative distance traveled and the attention data. This figure shows the time course of the cumulative distance traveled by multiple individuals under the naive condition and the cumulative distance traveled by multiple individuals under the preexposed condition. This figure shows a graph (left) of the nematode's behavioral trajectory (x-coordinate, y-coordinate) with emphasis on the period focused on by the attention layer of the fifth branch, and a graph (right) showing the movement variance of the turning angle and the attention data. This figure shows the time course of the movement variance of the turning angle for multiple individuals under the naive condition and the movement variance of the turning angle for multiple individuals under the preexposed condition.
[0013] [Embodiment 1] (Configuration of the learning device) Figure 1 is a block diagram showing the configuration of the learning device 1 of this embodiment. The learning device 1 comprises a neural network system 2 and a learning unit 5. The neural network system 2 comprises an acquisition unit 3, a presentation unit 4, and a learning model 10. The neural network system 2 is implemented by a computer equipped with a processing unit and a memory device.
[0014] The acquisition unit 3 acquires input data from a storage device or an external device (not shown). The input data is the data to be analyzed. In this case, the input data is time-series data. For example, the input data may be time-series data showing the acceleration, velocity, or position of a sensed object. The acquisition unit 3 outputs the input data to the learning model 10.
[0015] The learning model 10 includes a feature extraction layer 11 and multiple branches 12. Although three branches 12 are shown here, the number of branches 12 is not limited to this; there may be two, four or more, or any number of branches 12.
[0016] The feature extraction layer 11 receives input data and generates feature data by extracting features from the input data. The feature extraction layer 11 includes, for example, a one-dimensional convolutional layer and / or a recurrent layer (LSTM layer). The length of the input data and the length of the feature data are the same. The feature data is a multidimensional vector. The feature extraction layer 11 outputs the feature data, which is time-series data, to multiple branches 12.
[0017] Each of the multiple branches 12 includes an attention layer 13, a linear layer 14, and an output layer 15. The multiple branches 12 are configured in parallel. Here, the results of calculations in each branch 12 are not input to the other branches 12. Each branch 12 performs calculations independently based on the input feature data.
[0018] The attention layer 13 receives feature data and generates attention data that focuses on a portion of the feature data. For example, the attention layer 13 calculates the importance (attention weight) of the feature data at each time step. The attention data represents the attention weight at each time step. The length of the input data and the length of the attention data are the same. The attention data is normalized so that the sum of the values of the attention data at all time steps equals 1. Multiple attention layers 13 in multiple branches 12 generate different attention data from each other. The attention layer 13 outputs the attention data.
[0019] Branch 12 obtains attention-weighted feature data, which is feature data that reflects the attention data, by multiplying the feature data and the attention data. Specifically, branch 12 generates attention-weighted feature data by multiplying the importance (attention weight) at each time step by the value of the feature data at each time step. This generates attention-weighted feature data that focuses on a part (a part of the period) of the feature data. Branch 12 inputs the attention-weighted feature data into the linear layer 14 of the same branch 12.
[0020] The linear layer 14 performs a linear transformation on the input attention-weighted feature data. For example, the linear layer 14 includes one or more layers that perform weighted linear sums. The linear layer 14 outputs the transformed data, after performing the linear transformation, to the output layer 15.
[0021] The output layer 15 receives the transformed data output by the linear layer 14 and performs estimation based on this transformed data. That is, the output layer 15 performs estimation based on attention-weighted feature data via the linear layer 14. For example, the output layer 15 estimates which of several classes the input data belongs to by performing a weighted linear sum on the output of the linear layer 14. For example, the output layer 15 estimates the probability that the input data belongs to each of several classes. Multiple output layers 15 of multiple branches 12 have the same classification targets (multiple classes to classify, such as male and female). The output layer 15 outputs the estimation results to the presentation unit 4 and the learning unit 5.
[0022] The presentation unit 4 presents the user (analyst) with information regarding the range (interval) that the attention layer 13 of the trained learning model 10 was focusing on. The presentation unit 4 also presents the user with the output of the trained learning model 10 (the estimation results of each output layer 15). For example, the presentation unit 4 presents the user with the estimation results of each output layer 15, associated with the attention data of the attention layer 13 corresponding to each output layer 15. The presentation unit 4 displays, for example, the estimation results and the corresponding attention data side by side on the display device. For example, the presentation unit 4 may select the attention data of the attention layer 13 corresponding to the output layer 15 whose estimation results are better than the standard and present it to the user.
[0023] The learning unit 5 causes the neural network system 2, which includes the learning model 10, to perform training. The learning unit 5 trains the learning model 10 using training data that includes multiple input data to which multiple labels are associated. For example, the feature extraction layer 11, multiple attention layers 13, multiple linear layers 14, and multiple output layers 15 include learning parameters that are adjusted during training. The learning unit 5 adjusts each learning parameter based on the estimation results of the output layer 15 for the training data. The learning parameters differ for each branch 12. During training, the learning unit 5 trains the learning model 10 by adjusting each learning parameter to reduce (minimize) the loss function. The loss function includes multiple losses in the multiple output layers 15 of the multiple branches 12. The losses in the output layers 15 may be, for example, cross-entropy loss or squared error. By training, the learning unit 5 generates a trained learning model 10 (trained model) with optimized learning parameters.
[0024] (Processing flow of the learning device) Figure 2 is a diagram showing the processing flow of the learning device 1. The acquisition unit 3 of the learning device 1 acquires input data, which is the training data, from a storage device or an external device (S1).
[0025] The feature extraction layer 11 generates feature data by extracting features from the input data (S2).
[0026] Each attention layer 13 generates attention data that focuses on a portion of the feature data (S3). Each branch 12 obtains attention-weighted feature data, which is feature data that reflects the attention data.
[0027] Each linear layer 14 performs a linear transformation on the attention-weighted feature data (S4).
[0028] The output layer 15 performs estimation based on the data output by the linear layer 14 and outputs the estimation result (S5).
[0029] The learning unit 5 performs training on the learning model 10. The learning unit 5 determines whether a predetermined number of training sessions have been performed (S6).
[0030] If learning for a predetermined number of times has not been performed yet (No in S6), the process proceeds to S7. The learning unit 5 adjusts each learning parameter in the feature extraction layer 11, the plurality of attention layers 13, the plurality of linear layers 14, and the plurality of output layers 15 based on the estimation results of the output layer 15 for a plurality of teacher data (S7). Then, the process returns to S1 and the learning process is repeated.
[0031] If learning for a predetermined number of times has been performed (Yes in S6), the process proceeds to the process of S8. The presentation unit 4 presents information regarding the section that the attention layer 13 of the learned learning model 10 has focused on, to the user (S8).
[0032] (Learning Process) The learning unit 5 calculates one loss function L for the learning model 10 including a plurality of branches 12. An example of the loss function L is shown in the following formula. ・・・(1) Here, B is the number of branches 12, L i CE is the cross-entropy loss of the i-th branch 12, S i is the similarity of the attention data of the i-th branch 12, C CE and C S are constants. λ is the weight (positive value) of the variation index term, TV() is the total variation function, a i is a vector representing the attention data of the i-th branch 12. w i is the penalty for the value of the cross-entropy loss, T CE is the threshold of the cross-entropy loss, C h is a value greater than 1. δ ij is the Kronecker delta (the value is 1 when i and j are equal, and the value is 0 otherwise).
[0033] The loss function L includes a plurality of cross-entropy losses L i CE in the plurality of output layers 15, and the similarity S i of the plurality of attention data. The cross-entropy loss L i CE represents the error of the estimation result of the output layer 15 of the i-th branch 12 with respect to the label of the teacher data.
[0034] Similarity S i This is the attention data a of the i-th branch 12. i and attention data a of other branch 12 j This is an index of similarity between two things, and the larger the value (≧0), the more similar the two things are. Here, the similarity S i This is the attention data a of the i-th branch 12. i and attention data a of other branch 12 j This is the average of the cosine similarity between the two values, normalized to a range of 0 to 1 (above 0). Similarity S i The value will be in the range of 0 to 1. Attention data a of the i-th branch 12 i Attention data a of other branch 12 j If it is similar to, then the similarity score is S. i The similarity S of the attention layers 13 increases. In other words, when multiple attention layers 13 focus on the same period (same feature) of the feature data, the similarity S of the attention layers 13 increases. i It will become larger. Note that the attention data a of the i-th branch 12 i and attention data a of other branch 12 j A value based on the distance to, similarity S i This may also be done. For example, the similarity S may be calculated by normalizing the result of subtracting the distance from a constant to a range of 0 or greater (0 to 1). i That is also acceptable.
[0035] The loss function L is the cross-entropy loss L i CE And the similarity S of the attention data i It includes an intermediate function with and as variables (inputs). This intermediate function is a cross-entropy loss L i CE And the similarity S of the attention data i It is a function whose value decreases non-linearly as both decrease. For example, L i CE and S i When it is within a specific range, the cross-entropy loss L i CE And the similarity S of the attention data iWhen both decrease to 1 / k, the value of the intermediate function decreases to less than 1 / k (k > 1). Here, the intermediate function is the cross-entropy loss L i CE And the similarity S of the attention data i It includes the product of with. In equation (1), the intermediate function is (w i L i CE +C CE ) (S i +C S ) . As can be seen when expanded, the intermediate function is L i CE and S i It includes a product term of (w i L i CE +C CE ) and (S i +C S When both of the values decrease to 1 / k, the value of the intermediate function decreases (non-linearly) to less than 1 / k (k > 1). The loss function L includes the sum of the intermediate functions of multiple branches 12.
[0036] Cross-entropy loss L i CE A small similarity value means that the output layer 15 of the i-th branch 12 outputs an appropriate estimation result. i A small value means that the attention layer 13 of the i-th branch 12 is focusing on a different period (i.e., a different feature) than the attention layers 13 of the other branches 12. The loss function L is the cross-entropy loss L i CE And the similarity S of the attention data iThis includes the following. Therefore, by training the learning unit 5 to minimize the loss function L, each output layer 15 can be appropriately estimated, and each attention layer 13 can be trained to focus on different features. Multiple attention layers 13 focusing on different features means that multiple label-dependent features inherent in the input data can be discovered. For example, if the input data is behavioral data of male and female animals, this means that the learning model 10 can focus on multiple features that differentiate the behavior of males and females, and appropriately classify them based on each feature.
[0037] For example, the loss function L is the cross-entropy loss L i CE And, similarity S i It may also include a sum rather than a product. In this case as well, each cross-entropy loss L i CE And each similarity S i If both decrease, the loss function L decreases. Therefore, it is expected that the learning model 10 will learn so that each output layer 15 makes appropriate estimations and each attention layer 13 focuses on different features.
[0038] However, in the case of summation, learning may not be possible depending on the task. For example, cross-entropy loss L i CE Even if the similarity S increases, i If the decrease is even greater, the loss function L may decrease as a result. In such a case, multiple attention layers 13 of multiple branches 12 may focus on the same feature that does not contribute to proper estimation. The output layer 15 of such a branch 12 cannot output proper estimation results. Furthermore, it cannot provide the user, who is the analyst, with useful information for discovering various features inherent in the input data.
[0039] On the other hand, in equation (1), the loss function L is the cross-entropy loss L i CE And, similarity S i As both decrease, it includes an intermediate function whose value decreases non-linearly. Therefore, the cross-entropy loss Li CE and the similarity S i both decrease, the value of the loss function L decreases significantly. As a result, the cross-entropy loss L i CE and the similarity S i learning in which one of them increases is less likely to occur.
[0040] The first constant C CE and the second constant C S are terms to prevent the value of the intermediate function (w i L i CE +C CE )(S i +C S ) from taking an extremely small value. C CE and C S If there are no such terms, the following problem may occur. For example, during the learning process, if the L i CE and S i of a certain branch 12 become very small, the change in the other will contribute little to the change in the loss function L. Therefore, after that, it becomes difficult to perform learning in which both the L i CE and S i of the branch 12 decrease. This means that appropriate learning is not performed for some branches 12.
[0041] To avoid this problem, the intermediate function is in a form that includes the product of the sum of the cross-entropy loss L i CE and the first constant C CE and the sum of the similarity S i and the second constant C S . The values of C CE and C S are arbitrary, but from the above perspective, they may be set to values greater than, for example, 1. Also, C CE and C S may be different values. C CE and C S are, respectively, S i and L i CEIt can be thought of as the weight of C. CE Compared to C S If L is large, i CE The change in this value is greatly influenced by the value of the loss function L.
[0042] The learning unit 5 has a cross-entropy loss L i CE A penalty based on this is applied to the loss function L. i L is the cross-entropy loss. i CE This is a coefficient that determines the penalty for the value of L. i CE The threshold T CE If it's smaller, there's no penalty, i.e., lol i = 1. L i CE The threshold T CE If it's more than that, there's a penalty, i.e., lol i = C h > 1. As a result, the learning unit 5 calculates the cross-entropy loss L of a certain branch. i CE If the value is too large, a penalty is applied to the loss function L such that the loss function L becomes larger. This penalty results in "low estimation performance (cross-entropy loss L)". i CE (is large), but similarity S i This has the effect of suppressing the appearance of "small branches".
[0043] The loss function L includes a variation index term. The variation index term includes the variation index for each of the multiple attention data from multiple branches 12. The variation index represents the magnitude of the variation in the entire attention data (over all time points). Here, the variation index term is λΣTV(a i ) / B. The variation indicator term is the total variation TV (a) of each of the multiple attention data from multiple branches 12. i Includes the sum of ). Total Variation TV (a i ) is attention data a iThis is the sum of the absolute values of the fluctuations in the value over the entire time period (from start to finish). Since the sum of the attention data values is normalized to 1, attention data with sharp peaks will have a larger total variation than attention data with gradually spreading peaks. Therefore, the variation index term has the effect of suppressing learning in which the attention layer 13 focuses only on a very limited narrow period. In addition, the variation index term is the total variation TV (a) of multiple branches 12. i This includes the sum of ) . Therefore, multiple attention layers 13 of multiple branches 12 generate attention data that focuses on different periods (different features), and it is possible to avoid the attention data of each branch 12 having large values only in a very limited narrow period. The variation index term works in training to penalize large fluctuations in the attention data. Therefore, the variation index term suppresses fluctuations and has the effect of smoothing out the attention data. For example, the variation index is the attention data a i This could also be the sum of the squares of the fluctuations in the value over all time points.
[0044] (Experimental Example 1) Figure 3 shows an example of input data (time series data) for the experimental group and the control group. The vertical axis represents amplitude, and the horizontal axis represents the number of time steps. Reference numeral 201 indicates the input data for the experimental group. Reference numeral 202 indicates the input data for the control group. Each input data is an artificially generated signal according to the following rules.
[0045] The input data for the experimental group includes intervals of low-frequency (1 Hz) sine waves, intervals of multiple medium-frequency (6 Hz, 7 Hz, 8 Hz, or 9 Hz) sine waves, and intervals of high-frequency (14 Hz) sine waves. The length of each interval is random, ranging from 5% to 10% of the total time series length, and the order in which the intervals appear is also random. However, the input data for the experimental group includes at least one low-frequency interval and at least one high-frequency interval. Gaussian noise is superimposed on the sine waves.
[0046] The input data for the control group consists only of intervals of multiple medium-frequency (6 Hz, 7 Hz, 8 Hz, or 9 Hz) sine waves. The length of each interval is random, ranging from 5% to 10% of the total time series length, and the order in which the intervals appear is also random. Gaussian noise is superimposed on the sine waves.
[0047] The low-frequency (1 Hz) and high-frequency (14 Hz) ranges are characteristic features of the input data for the experimental groups. The dataset consists of input data for 500 experimental groups and input data for 500 control groups.
[0048] Using these experimental and control groups, comparative experiments were conducted on learning using learning device 1 for multiple loss functions. The network configuration of learning device 1 was identical except for the loss function. Specifically, the feature extraction layer 11 has two one-dimensional convolutional layers. Here, the linear layer 14 was omitted, and attention-weighted feature data was input to the output layer 15. The number of branches 12 B was set to 10. The number of epochs during learning was set to 30. C CE = 1, C S =3, λ=20, T CE = 0.5, and C h The value was set to 10. The seed was changed, and each condition was tested 10 times. 60% of the dataset was used for training, 20% for validation, and 20% for testing.
[0049] Under condition 1, the above equation (1) is used as the loss function.
[0050] Under condition 2, the loss function used is the same as equation (1) but with the variation index term removed. The loss function for condition 2 is shown below. ... (2) In condition 3, the loss function is the variation index term and the penalty w of the cross-entropy loss from equation (1). i We will use the one that excludes [the specified element]. The loss function for condition 3 is shown below. ... (3) In condition 4, the sum of the cross-entropy loss and the similarity of the attention data is used as the loss function. The loss function for condition 4 is shown below. ... (4) In condition 5, the sum of cross-entropy losses is used as the loss function. The loss function for condition 5 is shown below. ... (5) Learning was performed using the learning device 1 under these conditions 1 to 5.
[0051] Figure 4 shows the change in the loss function L for validation data during the learning process. The vertical axis represents the value of the loss function L, and the horizontal axis represents the number of epochs. Although there are differences in the range of values because the loss function itself differs depending on the conditions, in all conditions, the value of the loss function L steadily decreased as learning progressed.
[0052] Figure 5 shows the cross-entropy loss L for validation data during the learning process. CE This figure shows the average trend of the cross-entropy loss L under condition 5. CE The average becomes the lowest, followed by the cross-entropy loss L of condition 1. CE The average of the cross-entropy loss L was lower. However, under all conditions, CE The average value is less than 0.2, indicating that each output layer 15 can appropriately classify the experimental group and the control group.
[0053] Figure 6 shows the trend of the average similarity S of the attention data to the validation data during the learning process. In conditions 2 to 4, the average similarity S tends to decrease and is kept low. In condition 1, the similarity S also decreased as learning progressed. In condition 1, the loss function L includes a variation index term, so the attention data has a broad peak. On the other hand, in conditions 2 to 5, the loss function L does not include a variation index term, so the attention data has a sharp peak. Therefore, in condition 1, the average similarity S tends to be calculated to be larger than in conditions 2 to 4. On the other hand, in condition 5, the average similarity S increased as learning progressed. That is, in condition 5, it is thought that in order to reduce the loss function L, learning progressed so that multiple attention layers 13, with some exceptions, focused on the same features.
[0054] Figure 7 shows the cross-entropy loss L for the validation data over 10 trials. CEThis figure shows the number of times a branch 12 appeared in which the cross-entropy loss did not decrease sufficiently. A branch 12 in which the cross-entropy loss does not decrease sufficiently is a branch with poor estimation performance that remains at the chance level. For example, some branches 12 can reduce similarity and keep the loss function L low by generating random attention data.
[0055] In condition 4, the loss function was a simple sum of the cross-entropy loss and the similarity score. When using a sum, learning tends to occur where the decrease in similarity score compensates for the increase in the cross-entropy loss, thereby reducing the value of the loss function. On the other hand, when the product of the cross-entropy loss and the similarity score is used as the loss function L, the loss function L decreases significantly when both the cross-entropy loss and the similarity score decrease. Therefore, using a product makes it less likely for learning to compensate for the increase in the cross-entropy loss with the decrease in similarity score to occur.
[0056] Under condition 3, there is no penalty to the value of the cross-entropy loss. Therefore, in training, it is more acceptable for the cross-entropy loss to increase in some branches 12 in exchange for a decrease in similarity, compared to conditions 1 and 2.
[0057] Under conditions 1 and 2, there is a penalty applied to the value of the cross-entropy loss. Therefore, it is thought that the emergence of branches that remain at the chance level was effectively suppressed.
[0058] In both Condition 1 and Condition 2, out of 10 branches in each of the 10 trials (100 branches for each condition), five branches had a classification accuracy < 0.9. Excluding these branches, the classification accuracy was > 0.99 in Condition 1 and > 0.96 in Condition 2. Very high classification performance was also achieved in Conditions 1 and 2.
[0059] Figure 8 shows an example of the input data and attention data of the experimental group. The horizontal axis represents the number of time steps, which is consistent across all graphs. Graph 301 shows the input data of the experimental group. Graph 302 shows the frequency of the input data before the application of Gaussian noise. Graph 303 shows the spectrogram of the short-time Fourier transform for the input data. Graph 304 shows the attention data output by an attention layer 13 that mainly focuses on the low-frequency range in the trained model under condition 2. Graph 305 shows the attention data output by a certain attention layer 13 that mainly focuses on the low-frequency range in the trained model under condition 1. Graph 306 shows the attention data output by another attention layer 13 that mainly focuses on the low-frequency range in the trained model under condition 1. Graph 307 shows the attention data output by an attention layer 13 that mainly focuses on the high-frequency range in the trained model under condition 2. Graph 308 shows the attention data output by a particular attention layer 13 that focuses mainly on high-frequency ranges in the trained model under condition 1. Graph 309 shows the attention data output by another attention layer 13 that focuses mainly on high-frequency ranges in the trained model under condition 1.
[0060] In condition 2, the loss function L does not include a variation index term. Therefore, in condition 2, the attention values of the attention data were high only in very limited intervals within the low-frequency or high-frequency ranges (Graphs 304, 307). Consequently, in condition 2, it is difficult for a user referring to the attention data to determine, for example, whether the entire low-frequency range is important for classifying the experimental group from the control group, or whether it is a different feature within the low-frequency range that is important for classifying the experimental group from the control group. This problem can become more pronounced when the input data is the complex trajectory of animals.
[0061] On the other hand, under condition 1, the loss function L includes a variation index term. Therefore, under condition 1, the interval with high attention values extended to correspond to either low-frequency or high-frequency intervals (Graphs 305, 306, 308, 309). Thus, under condition 1, users who refer to the attention data can infer that the entire low-frequency or high-frequency interval represents a feature that differs between the experimental group and the control group.
[0062] Figure 9 shows the Pearson correlation coefficients between attention data and low-frequency, high-frequency, and medium-frequency sequences. The low-frequency sequence consists of data where the value in the low-frequency interval is 1 and the value in the other intervals is 0. The high-frequency and medium-frequency sequences are defined similarly. In other words, the correlation coefficients shown in the figure indicate the correlation between multiple features (frequencies) of the input data and multiple attention data.
[0063] Under condition 1, the correlation coefficient between the attention data from the low-frequency-focused attention layer 13 and the low-frequency sequence was a high value of 0.42. Similarly, the correlation coefficient between the attention data from the high-frequency-focused attention layer 13 and the high-frequency sequence was also a high value of 0.52. In other combinations, the correlation was low.
[0064] Under condition 2, the correlation coefficient between the attention data of the attention layer 13, which focuses on low frequencies, and the low-frequency sequence was 0.14, indicating a certain degree of correlation. Similarly, the correlation coefficient between the attention data of the attention layer 13, which focuses on high frequencies, and the high-frequency sequence was 0.10, also indicating a certain degree of correlation. In other combinations, the correlation was low. Note that the period during which the attention value is greater than 1 / T is defined as the period that the attention layer 13 is focusing on. The sum of the attention values in one set of attention data is 1, and T is the time series length. Under condition 2, it can be said that the attention layer 13 does not focus on irrelevant intervals, but only on intervals of interest.
[0065] (Example of the display unit's operation) The display unit 4 presents the estimation results of each output layer 15 to the user, associating them with the attention data of the attention layer 13 corresponding to each output layer 15. The attention data can be said to be information about the range (interval) that the attention layer 13 focused on. Furthermore, the display unit 4 may also present the cross-entropy loss of each branch 12. In either learning device 1 of conditions 1 or 2, the user can understand which part of the input data each branch focused on to classify the experimental group and the control group. For example, in this example, the user can understand that the branch that focused on the low-frequency interval and the branch that focused on the high-frequency interval made appropriate classifications. Also, multiple attention layers 13 focus on different features from each other. Therefore, the learning device 1 extracts multiple features that differ between the experimental group and the control group, and the user can pay attention to multiple features. Therefore, the learning device 1 can also focus on features that might be overlooked in verification based on the user's hypothesis, and make the user aware of those features.
[0066] The presentation unit 4 may extract the input data for the period that branch 12 focused on, or highlight the input data for that period, and present it to the user. For example, the period of focus may be a period in which the attention value is greater than a threshold (e.g., 1 / T). By examining the input data for the period that branch 12 of the learning model 10 focused on, the user can investigate which features of the input data contribute to accurate estimation (classification). For example, the presentation unit 4 may calculate the Hellinger distance between groups of explanatory variables for the period in which the attention value is greater than a threshold and present it to the user.
[0067] The presentation unit 4 may present to the user the correlation coefficient between the attention data of each branch 12 and the processed data obtained by processing the input data. The processed data is, for example, data corresponding to features that may contribute to estimation, such as the low-frequency series, high-frequency series, and medium-frequency series mentioned above. For example, the low-frequency series is data that has values corresponding to the frequencies of the input data. For example, the processed data may be data divided into intervals based on the magnitude of the input data values, or data divided into intervals based on the magnitude of the first derivative (change) or second derivative of the input data values. The presentation unit 4 may present the processed data to the user along with the correlation coefficient. The presentation unit 4 may select and present to the user processed data with a relatively high correlation coefficient (higher than a threshold or higher than others), or input data corresponding to said processed data. For example, if the input data is position coordinates, velocity, acceleration, or turning angle, the presentation unit 4 may generate their moving average or moving variance as processed data.
[0068] The presentation unit 4 may present to the user an index that shows the relationship between attention data and specific features, not limited to the correlation coefficient. For example, the presentation unit 4 may obtain the index by correlation analysis or coherence analysis using input data, processed data, or data related to the input data. Data related to the input data may not be data that has been input into the learning model 10. Data related to the input data may be, for example, the original data of the input data. For example, acceleration data obtained by measurement using an acceleration sensor may be the original data, and velocity data or position data calculated from the original data may become the input data. Data related to the input data is data that corresponds to features that may contribute to estimation. The presentation unit 4 may also present an index that shows the relationship between attention data and multiple features. For example, the presentation unit 4 may use regression to determine how much each feature contributes to the attention data as the index.
[0069] The presentation unit 4 may highlight or extract portions of the processed data that correspond to the period (region) that the attention layer 13 of branch 12 focused on, and present them to the user. The processed data may be data in which the features inherent in the input data are expressed in a more easily understandable way. This makes it easier for the user to understand the relationship between features and differences between groups by examining the processed data for the period that branch 12 of the learning model 10 focused on.
[0070] The presentation unit 4 may cluster the multiple attention data from multiple branches 12 using principal component analysis or the like. Attention data and corresponding branches 12 that focus on similar features are classified into the same cluster. The presentation unit 4 presents the clustering results to the user. For example, the presentation unit 4 presents the user with information on the branches 12 belonging to each cluster and the corresponding attention data as clustering results. The presentation unit 4 may also present the above information to the user grouped by cluster. In this way, branches 12 that focus on similar features and their corresponding information are presented to the user as clusters. This eliminates the need for the user to fine-tune the number of branches, for example, even if the number of branches is greater than the number of features in the input data. Therefore, the interpretability of the presented results is improved for the user. The presentation unit 4 may also present the user with the number of clusters in the clustering results. The number of clusters may correspond to the number of features discovered.
[0071] (Variations) The input data to be analyzed may, for example, be time-series data showing the trajectory of the coordinates of an object. The object may be a living organism, a device, or a product, etc. The input data may also be time-series data measuring anything other than position.
[0072] The input data is not limited to time-series data; it can be any data. The input data may be a two-dimensional or three-dimensional image, or data showing the spatial distribution of measurements. For example, the input data may be data (a matrix) containing a parameter that replaces time, and one or more measurements corresponding to that parameter.
[0073] The input data may be, for example, a video of a worker. The feature extraction layer 11 may be pre-trained to extract the trajectories of multiple joints of a person from the video as feature data. The learning device 1 uses the input data, which is the training data, to train multiple branches 12 to focus on different features (different time periods or different joints). The presentation unit 4 may extract the portion of the video corresponding to the range (focused time period or region) that the branches 12 focused on and present it to the user. The presentation unit 4 may generate a video that emphasizes the range (focused time period or region) that the branches 12 focused on and present it to the user. By examining the presented video, estimation results, and / or attention data, the user can discover where the differences lie between the movements of a skilled worker and an inexperienced worker. This allows the tacit knowledge of a skilled worker's technical know-how to be conveyed to an inexperienced worker. Instead of a worker, video or measurement data of a work device or robot may be used as input data. In this case, for example, the user can identify the cause of any abnormalities in the work device, etc.
[0074] The learning model 10 may include multiple feature extraction layers 11. The multiple feature extraction layers 11 generate different feature data from the same input data. Each of the multiple feature extraction layers 11 outputs the feature data to its corresponding branch 12. The multiple branches 12 generate different attention data based on different feature quantities. In this case, each branch 12 can be said to include its own feature extraction layer 11. With this configuration, the learning device 1 can focus on various features inherent in the input data and contribute to the discovery of important features.
[0075] Each branch 12 of the learning model 10 may have additional feature extraction layers. For example, each branch 12 may have additional feature extraction layers immediately before and / or immediately after the attention layer 13.
[0076] Each branch 12 of the learning model 10 does not necessarily have a linear layer 14. In this case, the output layer 15 performs estimation based on the attention-weighted feature data input from the branch 12. For example, the output layer 15 performs estimation by calculating a weighted linear sum on the attention-weighted feature data. The output layer 15 may include two or more layers that calculate a weighted linear sum on the attention-weighted feature data (the linear layer 14 may be part of the output layer 15).
[0077] The learning parameters of the feature extraction layer 11 may be adjusted simultaneously with the learning parameters of the attention layer 13, the linear layer 14, and the output layer 15 through learning, or they may be adjusted at different times. For example, the learning parameters of the feature extraction layer 11 may be pre-adjusted through learning using separate training data.
[0078] The intermediate function of the loss function L is a two-variable (cross-entropy loss L) i CE and similarity S i The product is not limited to the product of (e^L), but may also be the product of exponential functions of two variables, for example. For example, the intermediate function is (e^L i CE )・(e^S i It can also be in the form of ). Here, "^" represents exponentiation. For example, the intermediate function is (L i CE ^2)・(S i It may also be in the form of ^2). The constants added to each term can be arbitrarily added. As an intermediate function, the cross-entropy loss L i CE And the similarity S of the attention data i Any function can be used in which the value decreases non-linearly as both decrease.
[0079] In the example above, cross-entropy loss was used as an indicator representing the error of the estimation result of the output layer 15 with respect to the labels of the training data. However, other methods such as mean squared error may also be used.
[0080] The learning device 1 is not limited to classifying data; it may also be used to solve regression problems. The loss in the output layer 15 may be, for example, the squared error.
[0081] The penalty given to the loss function L based on the error is the coefficient w of the cross-entropy loss. i However, this is not the only way; a penalty can also be applied by adding it to the loss function L. The penalty should be such that it further increases the loss function L when the error (cross-entropy loss) is relatively large.
[0082] Furthermore, as shown in equation (5) above, the loss function L is the cross-entropy loss L of the multiple branches 12. i CE It may also be the sum of these. Some attention layers 13 are expected to focus on features different from those of other attention layers 13. The presentation unit 4 may present multiple attention data from multiple attention layers 13 that focus on different ranges from each other. This allows the user to know of multiple candidate features that contribute to estimation.
[0083] (Example of preprocessing) The acquisition unit 3 may perform preprocessing on the input data. For example, the acquisition unit 3 may acquire time-series data of the animal's movement trajectory (x coordinate, y coordinate) from an external source as input data, and perform preprocessing to generate time-series data of one or more other explanatory variables (features) from the time-series data of the movement trajectory. For example, the acquisition unit 3 may calculate time-series data such as velocity, acceleration, direction of travel, turning angle, distance from the initial position, azimuth angle from the initial position, cumulative distance traveled, and / or the cumulative difference in the direction of travel at each time from the time-series data of the movement trajectory. For example, the acquisition unit 3 generates a matrix containing the time-series data of these multiple explanatory variables. The acquisition unit 3 may input a matrix containing the time-series data of multiple explanatory variables as input data to be input to the learning model 10.
[0084] The acquisition unit 3 may present the user with options for explanatory variables to be included in the input data to be input to the learning model 10, and generate a matrix containing the time-series data of the explanatory variables selected by the user as input data to be input to the learning model 10. In addition, the acquisition unit 3 may acquire additional data from the user to be added to the input data, separate from the behavioral trajectory. The additional data may be, for example, time-series data representing the animal's posture or time-series data of neural activity. The acquisition unit 3 also adds the additional data to the matrix as input data to be input to the learning model 10.
[0085] The feature extraction layer 11 generates feature data by extracting features from input data (matrix) containing multiple explanatory variables. In this case, the feature data is a matrix. Each branch 12 performs calculations using this feature data. The attention data generated by the attention layer 13 is a vector whose elements are the attention weights at each time step. The processing in each branch 12 and the learning unit 5 is performed in the same manner as in the embodiment described above, except that the input data is a matrix.
[0086] (Example of multi-class classification) In the example above, the neural network system 2 performed two-class classification, but it may also perform three or more-class classification. For example, the i-th branch 12 may be trained to classify whether the input data belongs to the i-th class or another class. In this way, the classification targets (classes to be classified) of the multiple output layers 15 may differ from one another. In this case, each branch 12 is equivalent to performing two-class classification. In training data containing three or more labels, data obtained by merging labels other than the i-th class can be used to train the i-th branch 12. That is, the true labels used in calculating the cross-entropy loss will differ for each branch 12.
[0087] The loss function L described above can be used as the loss function. Since the labels of the training data differ depending on the branch 12, it can be expected that each attention layer 13 will learn to focus on different features. Therefore, the loss function L is the similarity S. i It does not have to be included.
[0088] Alternatively, each of the multiple branches 12 may perform multi-class classification. In this case, the multiple output layers 15 will have the same classification targets (multiple classes to be classified). That is, each branch 12 will be trained using the same training data, for example, with the loss function L of equation (1).
[0089] The learning model 10 may include multiple branches 12 that classify the i-th class. In this case, to prevent the multiple branches 12 from focusing on the same feature, the loss function L of the multiple branches 12 that classify the i-th class is the similarity S. i It is preferable to include it.
[0090] (Experimental Example 2) In the drug discovery process, animal experiments using animals such as mice are conducted to screen drug candidates. Conventional evaluation methods in animal experiments rely on predetermined specific behavioral indicators or observations at the end of the experiment. Therefore, it has been difficult to evaluate the effects of drugs from multiple perspectives and efficiently. In particular, it has been difficult to evaluate cases where behavioral changes caused by drugs appear only during the experiment, or to evaluate complex behavioral patterns.
[0091] Learning device 1 can also be used, for example, for screening drug candidates in the drug discovery process. An example of applying learning device 1 to nematode behavioral trajectory data will be described. Here, a dataset from an odor avoidance experiment of the nematode (C. elegans) described in Non-Patent Literature 3 was used. This nematode is known to dislike the odor substance 2-nonanone and to take avoidance behavior from odor sources. This dataset includes data on the behavioral trajectories of nematodes that were previously exposed to 2-nonanone (preexposed condition) and nematodes that were not exposed (naive condition) when 2-nonanone was presented. In the experiment, nematodes in the preexposed or naive condition were placed on an experimental plate, and 2-nonanone was dropped onto the plate to obtain the nematode's behavior in response to the odor. Each data point is time-series data of the nematode's positional information (x, y coordinates). The number of data points (number of individuals) for the naive condition was 34, and the number of data points for the preexposed condition was 33.
[0092] In Experimental Example 2, the number of branches in the learning device 1 was set to 5. The loss function L used was equation (1) described above. The input data to the learning device 1 was time-series data of the nematode's position information (x, y coordinates). The acquisition unit 3 calculated time-series data of velocity, direction of travel, turning angle, distance from the initial position, azimuth angle from the initial position, cumulative distance traveled, and the cumulative difference in direction of travel at each time point from the time-series data of the position information. In addition to the position information, the acquisition unit 3 used input data (matrix) containing these multiple explanatory variables as input to the learning model 10.
[0093] Figure 10 shows the results of principal component analysis of the attention data obtained from each branch 12 of the trained learning model 10 in Experimental Example 2. Specifically, the presentation unit 4 performed principal component analysis on a matrix that integrated multiple attention data, including both naive and preexposed conditions, for each branch 12. Attention1 to 5 in the figure correspond to the attention data of the first to fifth branches 12. In the space of principal components PC1, PC2, and PC3, the attention data of the first branch and the third branch are located close to each other. Therefore, it is possible that the first branch and the third branch are focusing on similar features. The attention data of the other branches are located far apart from each other. Therefore, for example, it is highly likely that the first, second, fourth, and fifth branches are focusing on different features.
[0094] Figure 11 shows a graph (left) of the nematode's behavioral trajectory (x-coordinate, y-coordinate) with the portion corresponding to the period of focus of the attention layer 13 of the first branch highlighted, and a graph (right) showing the cumulative distance traveled and attention data. In the graph on the left, the behavioral trajectory for periods when the attention data of the first branch is greater than the threshold is drawn darker, and the behavioral trajectory for other periods is drawn lighter. In the graph on the right, the dotted line represents the cumulative distance traveled, and the solid line represents the attention data. Here, the behavioral trajectory, cumulative distance traveled, and attention data are shown as examples for one individual each under the naive condition (upper panel) and the preexposed condition (lower panel). Note that 2-nonanone was dropped at a predetermined position on the positive y-axis side of the initial position (start).
[0095] The accuracy rate for the first branch was 0.89. The attention data for the first branch correlated best with the explanatory variable, cumulative distance traveled.
[0096] Figure 12 shows the time course of the cumulative distance traveled by multiple individuals under the naive condition and the cumulative distance traveled by multiple individuals under the preexposed condition. The dark lines represent the average cumulative distance traveled under the naive condition and the average cumulative distance traveled under the preexposed condition. The lighter areas represent the standard error. The distortion of the plot at the end of the time series is due to missing input data for some individuals during that time period.
[0097] Previous studies have shown that in escape behavior experiments, nematodes in the preexposed condition reach a greater final destination (distance from the initial position after a predetermined period) than nematodes in the naive condition.
[0098] On the other hand, in this experimental example, behavioral analysis was performed using the cumulative distance traveled, which had the highest correlation with the attention data of the first branch. As a result, the nematodes in the preexposed condition traveled significantly less cumulative distance than those in the naive condition. In other words, it was found that the nematodes in the preexposed condition moved more efficiently than those in the naive condition. As shown in Figure 12, for example, a significant difference in cumulative distance traveled between the preexposed and naive conditions was observed after 100 seconds. It is possible that a significant difference in cumulative distance traveled may occur sooner than a significant difference in the final destination. For example, by using cumulative distance traveled as a judgment metric, accurate judgments can be made in a shorter amount of time.
[0099] Figure 13 shows a graph (left) of the nematode's behavioral trajectory (x-coordinate, y-coordinate) with the portion corresponding to the period focused on by the attention layer 13 of the fifth branch highlighted, and a graph (right) showing the displacement variance of the turning angle and attention data. In the graph on the left, the behavioral trajectory for periods when the attention data of the fifth branch is greater than the threshold is drawn darker, and the behavioral trajectory for other periods is drawn lighter. In the graph on the right, the dotted line shows the displacement variance of the turning angle, and the solid line shows the attention data. Here, the behavioral trajectory, displacement variance of the turning angle, and attention data are illustrated for one individual each under the naive condition (upper panel) and the preexposed condition (lower panel).
[0100] The accuracy rate for the fifth branch was 0.91. The attention data for the fifth branch showed a high correlation with the moving average of velocity, and a high inverse correlation (low correlation coefficient) with the moving variance of turning angle and moving variance of velocity. The turning angle indicates the change in direction of movement over time. A low moving variance of turning angle indicates that the nematode is moving at a constant turning angle. Looking at the trajectory of the parts with high attention values (darker parts) in the behavioral trajectory, it can be seen that the nematode is paying attention to the segment in which it is turning.
[0101] Figure 14 shows the time course of the moving variance of turning angles for multiple individuals under the naive condition and for multiple individuals under the preexposed condition. The dark lines represent the average value of the moving variance of turning angles under the naive condition and under the preexposed condition. The lighter areas represent the standard error.
[0102] In this experimental example, behavioral analysis was performed using the movement variance of turning angles, which showed a high inverse correlation with the attention data of the fifth branch. As a result, it was found that the movement variance of turning angles was significantly smaller in the nematodes under the preexposed condition compared with those under the naive condition. In other words, it was found that the nematodes under the preexposed condition turned and moved more smoothly than those under the naive condition. The nematodes under the naive condition may have been unsure of which direction to escape.
[0103] Thus, the presentation unit 4 may present to the user explanatory variables that have a high inverse correlation (strong negative correlation) with the attention data of each branch 12. A high inverse correlation means that there is a high correlation with the reciprocal of the explanatory variable or with an explanatory variable that has a negative sign.
[0104] The presentation unit 4 may present the user with information showing the correlation between the input data, processed data obtained by processing the input data, or data related to the input data and the attention data of one or more branches 12 (such as the correlation coefficient or information showing the degree of correlation). For example, the time series data of the turning angle is part of the input data input to the neural network system 2. The time series data of the moving variance of the turning angle is processed data obtained by processing the input data. The presentation unit 4 may present the user with information showing which explanatory variables had a relatively high correlation with the attention data of one or more branches 12.
[0105] When screening drug candidates, experimental data from administering a drug candidate (a specific substance) is input to a neural network system 2 that has been trained using data from preexposed and naive conditions, as in the experimental example above. The data from drug administration is, for example, data on the behavioral trajectory of animals in the preexposed condition that have been administered the drug and subjected to a similar escape behavior experiment. Drug administration may be performed before or after prior exposure to the odor substance (2-nonanone). The difference between the preexposed and naive conditions in the escape behavior experiment is related to dopamine receptors, which are involved in memory.
[0106] For example, multiple behavioral data from animals in a preexposed condition after being administered a candidate drug are input into a trained neural network system 2, and each branch 12 makes an estimation. If each branch 12 estimates "preexposed" with a high probability, it is considered that there is no significant change in escape behavior, and therefore it can be concluded that the candidate drug does not affect dopamine receptors, i.e., it has no effect.
[0107] If each branch 12 is estimated to be "naive" with a high probability, it means that no effect of prior exposure was observed despite the preexposed condition. In this case, it can be concluded that the candidate drug is likely to affect dopamine receptors, i.e., to be effective.
[0108] If the accuracy rate for some branches 12 is moderate (close to 0.5), it means that the behavioral results obtained cannot be classified as either preexposed or naive. In this case, it is possible that the drug candidate is partially affecting dopamine receptors, or that an unexpected other effect (side effect) is occurring.
[0109] In this way, the learning device 1 can be used to screen for drug candidates that have a specific effect. By selecting an experimental model according to the effect to be investigated, it is possible to screen for drug candidates that will produce the desired effect. The experimental animals may be other animals such as mice.
[0110] Thus, the presentation unit 4 of the trained neural network system 2, which has been input with behavioral data from animals administered with the drug candidate, may present the user with the estimated results or accuracy rates of one or more branches 12.
[0111] The presentation unit 4 of the trained neural network system 2, which has received behavioral data from animals administered with the drug candidate, may present the user with explanatory variables that have a high correlation with the attention data of each branch 12. The user can then learn what changes occurred due to the drug candidate. This allows the user to examine the effects of the drug candidate from multiple perspectives.
[0112] Furthermore, the learning device 1 may be retrained using behavioral data from animals administered with the drug candidate. Because multiple branches 12 focus on different features, the learning device 1 can help the user discover unexpected effects of the drug candidate.
[0113] For example, in experiments such as the escape behavior experiment with nematodes, it may be possible to gain insight into the effectiveness of a specific judgment metric (cumulative distance traveled in Experiment Example 2) in an experimental model. If this is known in advance, the user may have the acquisition unit 3 generate time-series data of the specific judgment metric (cumulative distance traveled) from the input data (data on the behavioral trajectory). This allows for accurate judgment with input data from a shorter period of time. Therefore, it may be possible to significantly reduce the required experimental time.
[0114] Thus, a trained model, which has been trained using animal data from the first experimental condition and animal data from a second experimental condition different from the first, may be further trained by inputting animal data from the second experimental condition in which a specific treatment was applied, and obtaining the output of the trained model. Based on the output of the trained model (estimated results or attention data), it is possible to estimate whether or not the specific treatment has an effect. The specific treatment may be the administration of a specific substance, a change in the environment, the application / reduction of a specific stimulus, or genetic manipulation. The animal data may be biological data or behavioral data, etc.
[0115] [Example of implementation using software] The function of the learning device 1 (hereinafter referred to as "device") is a program that causes a computer to function as the device, and can be realized by a program that causes a computer to function as each block of the device (acquisition unit 3, feature extraction layer 11, attention layer 13, linear layer 14, output layer 15, presentation unit 4, and learning unit 5).
[0116] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0117] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0118] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.
[0119] [Summary] The learning device according to aspect 1 of the present invention comprises a neural network system implemented by a computer and a learning unit, wherein the neural network system comprises a plurality of branches into which feature data is input, each of the plurality of branches includes an attention layer that outputs attention data focusing on a part of the feature data and an output layer that outputs estimation results based on attention-weighted feature data reflecting the attention data, and the learning unit is configured to train the neural network system to reduce a loss function that includes a plurality of losses in the plurality of output layers of the plurality of branches.
[0120] In the learning device according to aspect 2 of the present invention, in aspect 1 described above, the loss function may be configured to include a plurality of cross-entropy losses in a plurality of output layers and a plurality of similarity scores of the attention data.
[0121] In the learning device according to embodiment 3 of the present invention, in embodiment 2 described above, the loss function may include an intermediate function with the cross-entropy loss and the similarity as variables, and the value of the intermediate function may decrease non-linearly as the cross-entropy loss and the similarity decrease.
[0122] In the learning device according to embodiment 4 of the present invention, in embodiment 2 or 3 described above, the loss function may be configured to include the product of the cross-entropy loss and the similarity.
[0123] In the learning device according to embodiment 5 of the present invention, in any of embodiments 2 to 4 described above, the learning unit may be configured to apply a penalty based on the cross-entropy loss to the loss function.
[0124] In the learning device according to embodiment 6 of the present invention, in any of embodiments 1 to 5 described above, the loss function may be configured to include a variation index for each of the multiple attention data.
[0125] In the learning device according to embodiment 7 of the present invention, in embodiment 2 or 6 above, the loss function may be configured to include the product of the sum of the cross-entropy loss and the first constant and the sum of the similarity and the second constant.
[0126] In the learning device according to embodiment 8 of the present invention, in any of embodiments 1 to 7 described above, the plurality of output layers may be configured such that the classification targets are the same to each other.
[0127] In the learning device according to embodiment 9 of the present invention, the neural network system may be configured to perform multi-class classification in any of embodiments 1 to 8 described above.
[0128] In the learning device according to embodiment 10 of the present invention, in any of embodiments 1 to 7 described above, the neural network system performs multi-class classification, and the multiple output layers may be configured to classify objects that are different from each other.
[0129] The learning device according to embodiment 11 of the present invention may be configured to include a presentation unit that presents to the user information regarding the range that the attention layer has focused on, in any of embodiments 1 to 10 described above.
[0130] In the learning device according to embodiment 12 of the present invention, in embodiment 11 described above, the presentation unit may be configured to cluster a plurality of attention data and present the clustered results to the user.
[0131] A learning device according to embodiment 13 of the present invention includes a feature extraction layer that outputs feature data by extracting features from input data, as in embodiment 11 or 12 above, and the presentation unit may be configured to highlight or extract and present portions of the input data, processed data obtained by processing the input data, or data related to the input data that correspond to the range that the attention layer has focused on.
[0132] A learning device according to embodiment 14 of the present invention may be configured to include, in any of embodiments 1 to 10 above, a feature extraction layer that outputs feature data by extracting features from input data, and a presentation unit that presents to the user information showing the correlation between the input data, processed data obtained by processing the input data, or data related to the input data, and the attention data.
[0133] A method for generating a trained model according to aspect 15 of the present invention is a method for generating a trained model according to a neural network system, which involves inputting feature data into a plurality of branches included in a trained model provided by a neural network system, performing an attention step in which attention data focusing on a part of the feature data is generated in each of the plurality of branches, and performing an output step in which estimation results are output based on attention-weighted feature data that reflects the attention data, and then performing a training step in which the neural network system is trained to reduce a loss function that includes a plurality of losses in the outputs of the plurality of branches.
[0134] A trained model according to aspect 16 of the present invention comprises a plurality of branches into which feature data is input, each of which includes an attention layer that outputs attention data focusing on a portion of the feature data, and an output layer that outputs estimation results based on attention-weighted feature data reflecting the attention data, thereby causing the computer to function to perform estimation by focusing on a plurality of different features.
[0135] An estimation method according to aspect 17 of the present invention is an estimation method using the trained model of aspect 16 described above, wherein the trained model, which has been trained using animal data under first experimental conditions and animal data under second experimental conditions, is input with data of animals that have been subjected to a specific treatment under the second experimental conditions, and the output of the trained model is obtained.
[0136] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0137] 1. Learning device 2. Neural network system 3. Acquisition unit 4. Presentation unit 5. Learning unit 10. Learning model 11. Feature extraction layer 12. Branch 13. Attention layer 14. Linear layer 15. Output layer
Claims
1. A learning device comprising a neural network system implemented by a computer and a learning unit, wherein the neural network system comprises a plurality of branches into which feature data is input, each of the plurality of branches includes an attention layer that outputs attention data focusing on a portion of the feature data, and an output layer that outputs estimation results based on attention-weighted feature data reflecting the attention data, and the learning unit causes the neural network system to learn to reduce a loss function that includes a plurality of losses in the plurality of output layers of the plurality of branches.
2. The learning apparatus according to claim 1, wherein the loss function includes a plurality of cross-entropy losses in a plurality of output layers and a plurality of similarity scores of the attention data.
3. The learning device according to claim 2, wherein the loss function includes an intermediate function with the cross-entropy loss and the similarity as variables, and the value of the intermediate function decreases non-linearly as the cross-entropy loss and the similarity decrease.
4. The learning device according to claim 2, wherein the loss function includes the product of the cross-entropy loss and the similarity.
5. The learning device according to any one of claims 2 to 4, wherein the learning unit applies a penalty based on the cross-entropy loss to the loss function.
6. The learning device according to any one of claims 1 to 4, wherein the loss function includes a variation index for each of the plurality of attention data.
7. The learning device according to claim 2, wherein the loss function includes the product of the sum of the cross-entropy loss and a first constant and the sum of the similarity and a second constant.
8. The learning device according to claim 1, further comprising a display unit that presents to the user information regarding the range that the attention layer has focused on.
9. The learning device according to claim 8, wherein the display unit clusters a plurality of attention data and presents the clustered results to the user.
10. A learning device according to claim 8 or 9, comprising a feature extraction layer that outputs feature data by extracting features from input data, wherein the presentation unit highlights or extracts and presents portions of the input data, processed data obtained by processing the input data, or data related to the input data that correspond to the range that the attention layer has focused on.
11. A learning device according to any one of claims 1 to 4, comprising: a feature extraction layer that outputs feature data by extracting features from input data; and a presentation unit that presents to the user information showing the correlation between the input data, processed data obtained by processing the input data, or data related to the input data, and the attention data.
12. A method for generating a trained model, comprising: inputting feature data into multiple branches included in a learning model provided by a neural network system; an attention step in each of the multiple branches, generating attention data that focuses on a portion of the feature data; and an output step, outputting estimation results based on attention-weighted feature data that reflects the attention data; and a learning step that trains the neural network system to reduce a loss function that includes multiple losses in the outputs of the multiple branches.
13. A trained model for causing a computer to perform estimation by focusing on multiple different features, each of which includes a branch into which feature data is input, an attention layer that outputs attention data focusing on a portion of the feature data, and an output layer that outputs estimation results based on attention-weighted feature data that reflects the attention data.
14. An estimation method using a trained model according to claim 13, wherein the trained model has been trained using animal data under a first experimental condition and animal data under a second experimental condition, and data of animals that have been further subjected to a specific treatment under the second experimental condition is input into the trained model to obtain the output of the trained model.
Citation Information
Patent Citations
Gesture recognition method and system based on space-time diagram convolution alternating converter
CN117079345A
Health monitoring method and device based on voiceprint recognition, equipment and storage medium
CN117198339A
Sleep staging model construction method based on multi-scale time residual shrinkage network
CN118000664A
Railway station guiding operation information closed-loop detection method, system and device
CN118597231A
Attention Neural Network with Sparse Attention Mechanism
JP2023529801A