Learning device, estimation device, and learning method
The learning device improves mental state estimation accuracy for unknown users by using similarity in feature spaces as auxiliary information, addressing the challenges of high storage and computational loads in existing methods.
Patent Information
- Application Number
- PCT/JP2023/039147
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for estimating mental states from non-verbal and paralinguistic information face challenges in improving accuracy for unknown users, as they require multiple user-specific models, leading to high storage and computational loads, and do not effectively account for individual differences in emotional recognition.
A learning device that includes a feature extraction unit, a similarity calculation unit, an estimation unit, and a learning unit, which extracts feature quantities, calculates similarity between them, uses similarity as auxiliary information for mental state estimation, and learns parameters to improve estimation accuracy across different users.
This approach efficiently enhances the accuracy of mental state estimation for both known and unknown users by leveraging similarity in feature spaces, reducing the computational and storage burdens associated with multiple user models.
Smart Images

Figure JP2023039147_08052025_PF_FP_ABST
Abstract
Description
Learning device, estimation device, and learning method
[0001] The present invention relates to a learning device, an estimation device, and a learning method.
[0002] Research and development has been conducted on technologies that attempt to automatically estimate mental states expressed in non-verbal and paralinguistic information such as human voice, facial expression, and gestures. For example, it is expected that this technology can be used to reflect the mental state of the person interacting with an agent or robot when generating responses, to utilize the results of this estimation as part of mental health care, or to quantify the state of participants in web conferences to make it easier to understand.
[0003] It is generally known that the estimation accuracy for unknown users (not included in the training data) is worse than that for known users (included in the training data). This is thought to be because the mental state expression patterns differ for each user.
[0004] As a method for solving these problems and improving estimation accuracy for unknown users, for example, in estimating confidence from body movements, a method has been proposed in which an estimation model is trained for each user, the center distance of the feature space between users is regarded as the similarity, and a weighted sum using the similarity is used (Non-Patent Document 1).
[0005] Furthermore, in emotion recognition from speech, a method has been proposed in which individual differences between speakers are normalized by using a speaker vector representing the individuality of the speaker as auxiliary information (Non-Patent Document 2).
[0006] A. Matsufuji, E. Sato-Shimokawara, and T. Yamaguchi, "Adaptive Personalized Multiple Machine Learning Architecture for Estimating Human Emotional States", Journal of Advanced Computational Intelligence and Intelligent Informatics, Vol.24 No.5, 2020C. L. Moine and N. Obin and A. Roebel, "Speaker Attentive Emotion Recognition", INTERSPEECH 2021、30 August - 3 September, 2021, Brno, Czechia
[0007] However, in the method of Non-Patent Document 1, although it is expected that the estimation accuracy will improve if users with similar feature spaces exist in the training data, since there are as many models as there are users, the storage and calculation load is large, making it less practical.
[0008] Furthermore, with regard to the method of Non-Patent Document 2, the speaker vector is merely used to identify individuals, and does not necessarily indicate individual differences in emotion recognition or the like.
[0009] The present invention has been made in view of the above points, and has an object to efficiently improve the accuracy of estimating a person's mental state.
[0010] In order to solve the above problem, the learning device has a feature extraction unit configured to extract features from each of a plurality of data items related to different people; a similarity calculation unit configured to calculate the similarity between the features; an estimation unit configured to estimate, for each of the data items, the mental state of the person related to the data from the features based on learnable parameters, using the similarity between the person related to the data and another person as auxiliary information; and a learning unit configured to learn the parameters so that the mental state assigned as a correct answer for each of the plurality of data items approaches the estimation result by the estimation unit for each of the plurality of data items.
[0011] The accuracy of estimating a person's mental state can be improved efficiently.
[0012] FIG. 1 is a diagram illustrating an example of a hardware configuration of an estimation device 10 according to an embodiment of the present invention. FIG. 2 is a diagram illustrating an example of a functional configuration of the estimation device 10 according to an embodiment of the present invention. FIG. 3 is a diagram illustrating an example of a configuration of a training data storage unit 13. FIG. 4 is a diagram illustrating a specific example of a latent representation calculation unit 112. FIG. 5 is a diagram illustrating equation (2). FIG. 6 is a diagram illustrating equation (3) and learning of a k-means model. FIG. 7 is a diagram illustrating equation (4). FIG. 8 is a diagram illustrating a specific example of a posterior probability estimation unit 114.
[0013]
[0023] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of the hardware configuration of an estimation device 10 according to an embodiment of the present invention. The estimation device 10 in Fig. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, all of which are interconnected via a bus B.
[0014] A program that realizes the processing in the estimation device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.
[0015] When an instruction to start the program is received, the memory device 103 reads and stores the program from the auxiliary storage device 102. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the estimation device 10 in accordance with the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.
[0016] Fig. 2 is a diagram showing an example of the functional configuration of an estimation device 10 according to an embodiment of the present invention. In Fig. 2, the estimation device 10 includes a non-linguistic / paralinguistic information estimation unit 11 and a learning unit 12. These units are realized by a processor 104 executing one or more programs installed in the estimation device 10. The estimation device 10 also uses a training data storage unit 13 and a model parameter storage unit 14. These storage units can be realized using, for example, an auxiliary storage device 102, a storage device connectable to the estimation device 10 via a network, or the like.
[0017] In this embodiment, a concrete example is a five-level estimation of the level of understanding using a video showing the upper body of the subject using a neural network. The level of understanding is defined as follows, with a higher number indicating a higher level of understanding: 1: Not understood 2: Somewhat not understood 3: Normal state 4: Somewhat understood 5: Understood Each part will be explained below.
[0018] [Learning Data Storage Unit 13] The learning data storage unit 13 stores learning data used to train the model as the non-verbal / paralinguistic information estimation unit 11 during the learning phase. The learning data includes a data ID, a correct answer label (correct mental state) for the non-verbal / paralinguistic information to be estimated, and the ID of the recorded person. The learning data may also include labels such as person attributes (personality, age, gender, etc.). Furthermore, division into learning, development, and evaluation sets and data expansion may be performed as necessary, and pre-processing such as contrast normalization or use of only certain regions using face detection may be performed. Furthermore, the codec of the input information is not particularly specified.
[0019] For example, when estimating comprehension levels from videos, the training data storage unit 13 may have a configuration as shown in FIG. 3 . In FIG. 3 , the videos used for the training data are recorded by recording multiple different people participating in a 30-minute web conference or the like. Each person's video may relate to the same event (e.g., a web conference) or different events (e.g., a web conference). The video data is H264 format video recorded at 30 frames per second (FPS) with a webcam and may be resized to 224 pixels on each side. The video data is assigned a personal ID identifying each person recorded in the video data and a comprehension label as a correct label for each individual's comprehension level. There are a total of X videos, recorded by S participants. Comprehension labels may be assigned at either a frame-by-frame or segment-by-segment level of any length. Here, labels are assigned to segments every 5 seconds. That is, each row in FIG. 3 corresponds to a 5-second segment of a 30-minute video. Approximately 30 minutes of video data divided into 5-second intervals is stored for each person.
[0020] [Non-verbal / Paralinguistic Information Estimation Unit 11] Here, the non-verbal / paralinguistic information estimation unit 11 will be described when estimating the above-mentioned five-level comprehension level from a video. While the present embodiment illustrates an example in which the non-verbal / paralinguistic information estimation unit 11 uses an estimation method based on a neural network, other known techniques, such as batch normalization, dropout, or L1 / L2 regularization, may be applied to any location. Furthermore, any location in the model structure may be pre-trained using any task or method, and it may also be determined arbitrarily whether pre-trained model parameters are trainable or not.
[0021] In FIG. 2, the non-linguistic / paralinguistic information estimation unit 11 includes a feature extraction unit 111 , a latent expression calculation unit 112 , a similarity calculation unit 113 , and a posterior probability estimation unit 114 .
[0022] [Feature Extraction Unit 111] The feature extraction unit 111 extracts features from input data. Known features can be used. For example, when a video is input, the input video itself may be used as a feature, or a feature such as a histogram of oriented gradients (HOG) for each frame or an embedded representation of a model pre-trained in an arbitrary task may be used as a feature. Furthermore, the feature extraction unit 111 may perform preprocessing such as noise removal, contrast normalization, extraction of a face peripheral region, and normalization / standardization of features as necessary. Before calculating the features, the feature extraction unit 111 may perform processing such as video rotation or noise addition as data augmentation.
[0023] A specific example of estimating the level of understanding from a video will be described. The feature extraction unit 111 extracts an embedded representation X of a time series length T obtained from an input video by a pre-trained Transformer Encoder. (u,i_u) is calculated as a feature. Note that the subscript x_u (x is an arbitrary symbol) corresponds to the symbol where u is a subscript of x in the figures and formulas described later. u indicates a personal ID (personal identification information), and i u indicates a data ID within a person u.
[0024] [Latent Representation Calculation Unit 112] Based on learnable parameters, the latent representation calculation unit 112 calculates latent representations from the features output by the feature extraction unit 111. For example, it is preferable to acquire, as the latent representation, an embedded representation that takes time-series changes into consideration using a model structure that can handle time series, such as an RNN or a Transformer.
[0025] A specific example of estimating the level of understanding from a video will be described. (u,i_u) Using Enc(·), a Transformer Encoder of any structure, we can generate a time series latent representation Z (u,i_u) Calculate.
[0026] This specific example is shown in FIG.
[0027] [Similarity Calculation Unit 113] The similarity calculation unit 113 calculates the similarity of the feature space for each person. The "similarity of the feature space for each person" refers to the similarity between the feature space of a certain person and the feature space of another person, when focusing on that person. Any known method such as clustering can be used to calculate the similarity.
[0028] For example, as in Non-Patent Document 1, the center point of the average values of each feature value for each person may be calculated, and the Euclidean distance between the center points of each person may be used as the similarity, or any distance function such as cosine similarity may be used. Here, not only the average value but also a statistical quantity such as variance may be used. Furthermore, clusters may be created using a clustering method such as k-means or a Gaussian mixture model, and the distance from the center point of each cluster or the membership probability may be used instead of the distance function. Any feature value may be used to calculate the similarity. For example, the input feature value itself may be used, or a latent representation of any model may be used. The feature value may be preprocessed, such as by using principal component analysis to reduce the dimension. When time series feature values are used as feature values, time series clustering such as time series k-means may be used instead of using statistical quantities.
[0029] A specific example of estimating the level of understanding from a video will be described. In this case, the similarity calculation unit 113 uses an encoder Enc b (・), MHA that performs multi-head attentive pooling using self-attention b (·), decoder Dec b (·) and the output layer Out b In this embodiment, a trained model (hereinafter referred to as a "baseline model") is constructed from (·) and is not used as a training target. In this model, the latent representation vector e obtained by attentive pooling is (i,u) b The similarity calculation unit 113 uses the baseline model and its parameters to acquire latent expressions from all the training data, and generates a similarity feature vector x u s is calculated (the following formula (3)). u s The similarity calculation unit 113 employs the k-means method as the clustering method, specifies an arbitrary number of clusters N, and calculates x u s In the k-means method, the center point c of each cluster is used to train the model. n (n is the cluster number). The similarity calculation unit 113 calculates the similarity feature vector x u s is calculated, and the cosine similarity between the center point of each cluster and the similarity feature vector is calculated. u (Equation (4) below).
[0030] Equation (2) is shown in FIG. 5, equation (3) and the learning of the k-means model are shown in FIG. 6, and equation (4) is shown in FIG.
[0031] [Posterior Probability Estimation Unit 114] The posterior probability estimation unit 114 estimates the posterior probability of non-verbal / paralinguistic information (a person's mental state) from the latent expressions calculated by the latent expression calculation unit 112, based on learnable parameters, using the similarity in each person's feature space as an auxiliary feature. The similarity can be input at any point in the estimation model serving as the non-verbal / paralinguistic information estimation unit 11. For example, a similarity vector extended by the time series may be combined with the time-series latent expression (i.e., data equal to the similarity vector multiplied by the time series length may be combined after the time-series latent expression), or a similarity vector may be combined with the latent expression vector after pooling. Alternatively, when the latent expression calculation unit 112 calculates the time-series latent expression, a similarity vector extended by the time series may be combined with the feature. These methods may be implemented independently or in combination. Furthermore, the similarity vector may be subjected to any transformation or normalization.
[0032] A specific example of estimating the level of understanding from a video will be described. First, the posterior probability estimation unit 114 extracts S u is generated (Equation (5) below). Next, the posterior probability estimation unit 114 acquires attention weights in the time direction using a multi-head attention mechanism, and calculates the sum of weights in the time direction using MHA(·) (Equation (6) below). MHA(·) is expanded as shown in Equation (7) below. For simplification, the number of heads in multi-head attention is set to 1 in Equations (5) and (6). MHA(·) is given a concatenation of the key and value, which are the extensions of the time-series latent representation and the similarity vector, and the time-series latent representation is given as a query. W {q,k,v} are learnable weights. The posterior probability estimation unit 114 calculates the latent representation vector obtained by MHA(·) as e (i,u) This is then used to generate the posterior probability P(l|X (u,i_u) , s u, Ω) is obtained (Equation (8)). Note that Ω is a model parameter set (a set of parameters Enc(·) in Equation (1), MHA(·) in Equation (6), Dec(·) and Out(·) in Equation (8)). In Equation (8), l indicates the class ID of the comprehension level. The decoder is composed of any function, and may be, for example, a combination of several fully connected layers and an activation function, or an identity function. The output layer may be composed of a fully connected layer and a Softmax function.
[0033] This specific example is shown in FIG.
[0034] [Learning Unit 12] To learn the model parameters, the learning unit 12 updates the model parameter set Ω for each piece of training data so that the estimation result by the posterior probability estimation unit 114 approaches the understanding level label as the correct answer for that training data, and acquires learned parameters Ω'. Any known technology may be used for the loss function and update method. As described above, the model parameter set Ω may be a mixture of parameters previously trained in any other task, or initial values may be generated using any random number. Also, some parameters may not need to be updated.
[0035] A specific example of estimating comprehension from a video will be described. As a specific example, the model parameter set Ω is updated using the stochastic gradient descent method. At this time, any value is used for the hyperparameters such as the learning rate. For example, the posterior probability P(l|X (u,i_u) ), the Kullback-Leibler information loss L may be used as a loss function for the correct labels of the training data corresponding to u and i. The loss function is defined as follows:
[0036] At this time, A(l│X (u,i_u)) is the correct distribution of the input features, and may be expressed as a one-hot vector in which the probability of the correct class is 1.0 and all other classes are 0.0, or any distribution, such as a correct distribution approximating a normal distribution centered on the correct class, may be used. Furthermore, loss functions corresponding to other probability density functions, not limited to Kullback-Leibler divergence, may also be used. Additionally, any task may be simultaneously learned as an auxiliary multi-task learning task. For example, if a class has a continuous relationship between labels, such as comprehension, a regression task that takes this into account may be added to the multi-task learning.
[0037] [Model Parameter Storage Unit 14] The model parameter storage unit 14 stores a learned parameter set Ω'. During inference, the Ω' stored in the model parameter storage unit 14 is used. The model parameter storage unit 14 also stores a parameter set of a baseline model and parameters of a k-means model obtained during learning (in the above example, the central points c of each cluster obtained by the k-means method) in order to acquire a similarity vector from the feature quantities of input data. n (n is the cluster number) is also stored in the same way. [During inference] During inference, the non-verbal / para-linguistic information estimation unit 11 inputs a video of a certain person as input data and outputs a certain piece of non-verbal / para-linguistic information (mental state) estimated from the video.
[0038] Specifically, the feature extraction unit 111 extracts feature amounts from the input data as described above.
[0039] As described in FIG. 4, the latent expression calculation unit 112 calculates a latent expression from the feature quantity output by the feature quantity extraction unit 111 based on the learned parameters stored in the model parameter storage unit 14.
[0040] The similarity calculation unit 113 calculates the similarity using the parameters of the k-means model stored in the model parameter storage unit 14 (the central point c of each cluster). n (n is the cluster number))) is used to calculate the similarity vector for the person. That is, the similarity vector in the feature space of the person to be inferred indicates the similarity with the feature space of each person to be learned.
[0041] The posterior probability estimation unit 114 estimates the posterior probability of non-linguistic and paralinguistic information from the latent expression calculated by the latent expression calculation unit 112, using the similarity vector calculated by the similarity calculation unit 113 as an auxiliary feature.
[0042] Although the above example shows that the input data is a video, the data that can be input data is not limited to video. For example, any data related to a person that records events that change depending on the person's mental state, such as still images, audio data, or sensor information such as brain waves, can be used as input data.
[0043] As described above, according to this embodiment, it is possible to estimate a person's mental state (non-linguistic and paralinguistic information) by clearly using the similarity in feature space as auxiliary information, rather than information that identifies an individual, such as a speaker vector. Since a single model is used for estimation, it is possible to suppress increases in the amount of calculation and storage load. When making an estimation using data of an unknown person as input, it is necessary to calculate the similarity in feature space, but there is no need to retrain the model itself. Therefore, it is possible to efficiently improve the accuracy of estimating a person's mental state.
[0044] In this embodiment, the latent expression calculation unit 112 and the posterior probability estimation unit 114 are an example of an estimation unit. The estimation device 10 in the learning phase is an example of a learning device.
[0045] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0046] REFERENCE SIGNS LIST 10 Estimation device 11 Non-linguistic / paralinguistic information estimation unit 12 Learning unit 13 Learning data storage unit 14 Model parameter storage unit 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 Processor 105 Interface device 111 Feature extraction unit 112 Latent expression calculation unit 113 Similarity calculation unit 114 Posterior probability estimation unit B Bus
Claims
1. A learning device comprising: a feature extraction unit configured to extract features from each of a plurality of data relating to different people; a similarity calculation unit configured to calculate a similarity between the features; an estimation unit configured to estimate, for each of the data, a state of mind of the person from the features based on learnable parameters, using the similarity between the person related to the data and another person as auxiliary information; and a learning unit configured to learn the parameters so that the state of mind assigned as a correct answer for each of the plurality of data approaches the estimation result by the estimation unit for each of the plurality of data.
2. The learning device described in claim 1, characterized in that the estimation unit includes: a latent expression calculation unit configured to calculate a latent expression from the feature amount; and a posterior probability estimation unit configured to estimate the posterior probability of the mental state from the latent expression using the similarity as auxiliary information.
3. An estimation device comprising: a feature extraction unit configured to extract features from data related to a certain person; a similarity calculation unit configured to calculate a similarity between each feature extracted by the feature extraction unit from each of a plurality of data related to different people and a feature extracted by the feature extraction unit from the data related to the certain person; and an estimation unit configured to estimate the state of mind of the certain person from the features using the similarity as auxiliary information based on learnable parameters, wherein the parameters are learned so that a state of mind assigned as a correct answer for each of the plurality of data approaches an estimation result by the estimation unit for each of the plurality of data.
4. A learning method characterized in that a computer executes the following steps: a feature extraction step for extracting features from each of a plurality of data relating to different people; a similarity calculation step for calculating the similarity between the features; an estimation step for estimating, for each of the data, the state of mind of a person from the features using the similarity between the person related to that data and other people as auxiliary information based on learnable parameters; and a learning step for learning the parameters so that the state of mind assigned as a correct answer for each of the plurality of data approaches the estimation result obtained by the estimation step for each of the plurality of data.
Citation Information
Patent Citations
Information processing method and apparatus
JP2005346471A
Emotion information estimation device, emotion information estimation method and emotion information estimation program
JP2016106689A
Computer system
JP2019032591A
Systems and methods for prediction of user affect within SAAS applications
US20210109607A1