Estimation method, estimation device, and estimation program

The estimation device accurately estimates mental states by calculating embedded representations and comparing them using neural networks, addressing the challenges of individual differences and data requirements in existing technologies.

JP7700863B2Active Publication Date: 2025-07-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023544819
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-30
Publication Date
2025-07-01
Estimated Expiration
2041-08-30

AI Technical Summary

Technical Problem

Existing methods struggle to accurately estimate mental states from non-verbal and para-verbal information due to individual differences and the difficulty in absorbing or normalizing these variations, especially in cases with high label dispersion, and require substantial data for model retraining.

Method used

An estimation method using an estimation device that calculates embedded representations of mental states through feature amounts and compares them using neural networks, employing techniques like 2D CNN and RNN, with multi-head self-attention mechanisms to estimate mental states accurately.

Benefits of technology

The method enables precise estimation of mental states despite high label dispersion and individual differences, reducing the need for extensive data collection and model retraining, while maintaining accuracy and resource efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700863000008
    Figure 0007700863000008
  • Figure 0007700863000009
    Figure 0007700863000009
  • Figure 0007700863000010
    Figure 0007700863000010
Patent Text Reader

Abstract

An acquisition unit (15a) acquires learning data that includes nonverbal information or paralinguistic information and correct labels representing the states of mind indicated in the nonverbal information or paralinguistic information. A calculation unit (15b) uses a feature quantity of each of input data (14b) and reference data (14c) in the acquired learning data (14a) to calculate an embedded representation of the state of mind indicated by the input data (14b) and an embedded representation of the state of mind indicated by the reference data (14c). An estimation unit (15c) estimates the state of mind indicated by the input data (14b), using the result of a comparison between the embedded representation calculated from the input data (14b) and the embedded representation calculated from the reference data (14c).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an estimation method, an estimation device, and an estimation program.

Background Art

[0002] Conventionally, research and development have been conducted on technologies for automatically estimating the mental state expressed in non-verbal and para-verbal information such as human voice, face, gestures, etc. For example, in conversations with agents or robots, it is expected to reflect the mental state of the conversation partner when generating their responses, utilize the estimation results as part of mental health care, or quantify the mental state of participants in web conferences, etc., to make it easier to grasp.

[0003] Estimation of the mental state expressed in such non-verbal and para-verbal information is generally defined as supervised learning that outputs the posterior probability, etc. of each label representing the defined mental state for inputs such as feature quantities and the data itself extracted from voice and moving images (see Non-Patent Document 1).

[0004] Here, it is said that there are individual differences in mental states and their expressions. In contrast, generally, a method of collecting data from a large number of people and absorbing individual differences into the machine learning model is used. Also, methods of normalizing or absorbing individual differences using pre-registered reference voices and moving images are known (see Non-Patent Documents 2 and 3).

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

[0006] However, in the prior art, it has been difficult to accurately estimate the labels representing the mental states expressed in non-verbal and para-verbal information. For example, when the dispersion of the corresponding labels increases due to the subdivision of general emotion recognition such as normal, joy, sadness, surprise, fear, hatred, anger, contempt, etc., or the level of specific indicators such as the degree of understanding, it may be difficult to sufficiently absorb or normalize individual differences only in the normal state. Also, in the retraining of the model, it is difficult to secure a sufficient amount of data to stably adapt the model parameters to the evaluation subject and to surely learn.

[0007] The present invention has been made in view of the above, and an object thereof is to accurately estimate the labels representing the mental states expressed in non-verbal and para-verbal information. [Means for Solving the Problems]

[0008] In order to solve the above-described problems and achieve the object, an estimation method according to the present invention is an estimation method executed by an estimation device, including: an acquisition step of acquiring learning data including non-verbal information or para-verbal information and a correct label representing a mental state represented by the non-verbal information or para-verbal information; a calculation step of calculating an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using respective feature amounts of the input data and the reference data among the acquired learning data; and an estimation step of estimating the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data. It is characterized by including the above.

Effect of the Invention

[0009] According to the present invention, it is possible to accurately estimate a label representing a mental state represented by non-verbal / para-verbal information.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Mode for Carrying Out the Invention

[0011] Hereinafter, with reference to the drawings, an embodiment of the present invention will be described in detail. Note that the present invention is not limited by this embodiment. Also, in the description of the drawings, the same parts are denoted by the same reference numerals.

[0012] [Configuration of the estimation device] FIG. 1 is a schematic diagram illustrating the schematic configuration of the estimation device. Also, FIG. 2 is a diagram for explaining the processing of the estimation device. The estimation device 10 of the present embodiment estimates, using a neural network, the state of mind represented by non-verbal and para-verbal information as the degree of understanding in five stages for a video in which the upper body of the subject, which is non-verbal and para-verbal information, is reflected. The degree of understanding is defined, for example, as 1. Not understood, 2. Slightly not understood, 3. Normal state, 4. Slightly understood, 5. Understood, and the larger the number, the more understood it is.

[0013] First, as illustrated in FIG. 1, the estimation device 10 of the present embodiment is realized by a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.

[0014] The input unit 11 is realized using an input device such as a keyboard or a mouse, and inputs various instruction information such as a processing start to the control unit 15 in response to an input operation by the operator. The output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, an information communication device, or the like. The communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication between the control unit 15 and an external device such as a server or a device that manages learning data via a network.

[0015] The storage unit 14 is implemented by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. Note that the storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13. In the present embodiment, the storage unit 14 stores, for example, learning data 14a used in the estimation process described later, model parameters 14d generated and updated in the estimation process, and the like.

[0016] Here, as shown in FIG. 1, the learning data 14a of the present embodiment includes input data 14b and reference data 14c, and the data configurations are the same. FIG. 3 is a diagram illustrating the data configuration of the learning data. As shown in FIG. 3, the learning data 14a includes at least video data in which the upper body of the subject appears as non-verbal and para-verbal information, a data ID for identifying each video data, a personal ID for identifying the subject, and a correct label indicating a mental state such as the degree of understanding appearing in each video data. The learning data 14a may include labels representing attributes of a person such as age and gender. Further, if necessary, learning, development, or splitting into evaluation sets or data augmentation of the learning data 14a may be performed.

[0017] Note that preprocessing such as contrast normalization and face detection may be performed so that only a certain area of the video data is used. Also, the codec or the like of the input data (video data) is not particularly limited.

[0018] Specifically, when estimating the degree of understanding from video data in the estimation process described later, for example, video data in the H264 format recorded at 30 frames per second with a web camera may be resized so that one side becomes 224 pixels. Each of the X video data is assigned a personal ID of S subjects and a correct label of the degree of understanding.

[0019] Also, the input data and the reference data only need not to have the same data mixed together. During the estimation process described later, the input data 14b and the reference data 14c may be generated from any combination of the learning data 14a so as to avoid the mixing of the same data.

[0020] The control unit 15 is implemented using a CPU (Central Processing Unit), an NP (Network Processor), an FPGA (Field Programmable Gate Array), etc., and executes a processing program stored in the memory. Thereby, as illustrated in FIG. 1, the control unit 15 functions as an acquisition unit 15a, a calculation unit 15b, an estimation unit 15c, and a learning unit 15d. Note that these functional units may be implemented on different hardware. For example, the acquisition unit 15a may be implemented on hardware different from other functional units. Also, the control unit 15 may include other functional units.

[0021] The acquisition unit 15a acquires learning data including non-verbal information or para-verbal information and a correct label representing the mental state represented by the non-verbal information or para-verbal information. Specifically, the acquisition unit 15a acquires, via the input unit 11 or via the communication control unit 13 from a device that generates learning data or the like, learning data 14a including video data showing the upper body of a subject as non-verbal / para-verbal information, a data ID for identifying each video data, and a correct label representing a mental state such as the degree of understanding represented in each video data.

[0022] Also, the acquisition unit 15a stores the learning data 14a acquired in advance prior to the following processing in the storage unit 14. Note that the acquisition unit 15a may transfer the acquired learning data 14a to the estimation unit 15c shown below without storing it in the storage unit 14.

[0023] Returning to the description of FIG. 1, the calculation unit 15b calculates an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the respective feature amounts of the acquired learning data 14a for the input data 14b and the reference data 14c.

[0024] Note that the processing using the neural network described below is not limited to this embodiment. For example, elements of well-known techniques such as Batch Normalization, dropout, and L1 / L2 regularization may be added at any location.

[0025] Specifically, the calculation unit 15b first extracts respective feature amounts from the input data 14b and the reference data 14c for the same subject. For example, as the feature amounts, the calculation unit 15b extracts log mel-filterbank of voice, MFCC (Mel Frequency Cepstral Coefficients), HOG (Histogram of Oriented Gradients) for each frame of video, HOF (Histogram of Optical Flow), etc. The video itself may be used as the feature amount.

[0026] Also, the calculation unit 15b may perform preprocessing such as voice enhancement, noise removal, contrast normalization, extraction of the area around the face, and normalization of the feature amounts as necessary. Further, the calculation unit 15b may perform data augmentation processing such as superposition of noise and reverberation, rotation of video, and addition of noise before extracting the feature amounts.

[0027] Specifically, the calculation unit 15b extracts the video data x of the input data 14b with a frame length of T 1:T and the video data y of N reference data 14c 1:T (1,…,N) For these, only the area around the face is cut out into a square and resized again to be 224 pixels on each side. Further, the calculation unit 15b normalizes so that the value of each pixel becomes 0.0 to 1.0. The correct label of understanding level 3 may be assigned to all of the video data of the N reference data 14c, or labels of each of understanding levels 1 to 5 may be mixed.

[0028] Also, it is assumed that the input data 14b and the reference data 14c are selected so that the same video data does not coexist. For the video data of the reference data 14c, preprocessing such as burning out or deforming the metadata may be performed to avoid the coexistence of the same video data.

[0029] Next, the calculation unit 15b calculates an embedded representation from the feature amounts of the input data 14b and the reference data 14c. For example, the calculation unit 15b calculates an embedded representation H for each time using a 2D CNN (Convolutional Neural Network) or an RNN (Recurrent Neural Network). Note that the calculation unit 15b may replace the 2D CNN with a 3D CNN, or may replace the RNN with a Transformer.

[0030] Also, the model parameter 14d may include those pre-learned in any other task, or the initial value may be generated with any random number. Further, when the learned model parameter 14d is used, whether or not to update the model parameter 14d may be arbitrarily determined.

[0031] Specifically, as shown in FIG. 4, the calculation unit 15b uses a 2D CNN and an RNN with a D-dimensional output dimension, and as shown in the following equation (1), for the input video data x 1:T to calculate the embedded representation tensor H x . Here, θ is a set of parameters of the CNN, and φ is a set of parameters of the RNN.

[0032]

Equation

[0033] Also, as shown in the following equation (2), the calculation unit 15b calculates the embedded representation tensor H 1:T (1,…,N) from the reference video data y y .

[0034]

Equation

[0035] Next, the calculation unit 15b calculates the embedded representation from the input video data x 1:T and the embedded representation calculated from the reference video data y 1:T (1,…,N) and compares them. Specifically, as shown in FIG. 4, the calculation unit 15b compares the embedded representation tensor H x calculated from the input data 14b with the embedded representation tensor H y calculated from the reference data 14c, and calculates e (1,…,N) .

[0036] For example, the calculation unit 15b compares using a source - target attention mechanism. In this case, the calculation unit 15b calculates the comparison result vector e (1,…,N) as shown in the following formula (3).

[0037]

Equation

[0038] In the above formula (3), the calculation unit 15b calculates the attention weight from query Q i (1,…,N) and key K i , applies it to value V i , and finally calculates the sum in the time direction.

[0039] Here, d1 is the number of attention heads, i is each attention head, and W i Q , W i K , W i V represent the weights for Query, key, and value in each attention head, respectively.

[0040] Note that the calculation unit 15b is not limited to the source-target attention mechanism, and may perform comparison using a predetermined arithmetic operation or combination between the embedding representations.

[0041] Next, as shown in FIG. 4, the calculation unit 15b compares e (1,…,N) with each other to calculate a comparison result vector v. At this time, the calculation unit 15b may add arbitrary information such as metadata of y (1,…,N) to e 1:T (1,…,N) .

[0042] For example, the calculation unit 15b uses the multi-head self attention mechanism to combine m (1,…,N) of y 1:T (1,…,N) with e (1,…,N) as shown in the following formula (4), and further generates a tensor E 1:T by combining them. Here, the metadata m (1,…,N) represents the understanding degree label of the C-th stage (C = 5) in the form of a one-hot vector.

[0043]

Number

[0044] Also, the calculation unit 15b calculates v from E 1:T as shown in the following formula (5) using the multi-head self attention mechanism.

[0045]

Number

[0046] Here, d2 is the number of attention heads, j is each attention head, W j Q , W j K , W j Vrepresent the weights for Query, Key, and Value in each attention head, respectively.

[0047] Returning to the description of FIG. 1, the estimation unit 15c estimates the mental state of the input data 14b using the comparison result between the embedding representation calculated from the input data 14b and the embedding representation calculated from the reference data 14c.

[0048] For example, as described above, the estimation unit 15c estimates the mental state of the input data 14b using the result of comparing the embedding representation calculated from the input data 14b by the calculation unit 15b and the embedding representation calculated from the reference data 14c using the multi-head self-attention mechanism.

[0049] Specifically, as shown in FIG. 4, the estimation unit 15c estimates the mental state of the input data 14b from the comparison result vector v calculated by the calculation unit 15b. At this time, the estimation unit 15c may calculate the posterior probability for each class as a classification problem with an arbitrary number of classes and estimate the mental state. That is, the estimation unit 15c may calculate the posterior probability for each class of the mental state classification and estimate the mental state. Alternatively, the estimation unit 15c may estimate a numerical value representing the mental state as a regression problem.

[0050] For example, as shown in the following formula (6), the estimation unit 15c uses a two-layer fully connected layer to calculate the posterior probability p(C|x 1:T , y 1:T (1,…,N) ) for each of the five levels of understanding.

[0051]

Equation

[0052] Here, W1 FC , W2 FC represent the weights of the two-layer fully connected layer, and D FCrepresents the output dimension of the first fully-connected layer, and C represents the number of predicted labels (in this embodiment, C = 5). Also, the ReLU function is used as the activation function of the first fully-connected layer.

[0053] Returning to the description of FIG. 1. The learning unit 15d learns the model parameters 14d of a model that estimates the mental state represented by the input non-verbal information or para-verbal information using the input data 14b and the estimated mental state of the input data 14b.

[0054] Specifically, the learning unit 15d updates the model parameter set Ω and obtains the learned model parameter set Ω'. The learning unit 15d can apply well-known loss functions and update methods. For example, the model parameter set Ω may include those pre-trained in any other task, or the initial values may be generated by arbitrary random numbers, or some model parameters may not be updated.

[0055] For example, the learning unit 15d uses the stochastic gradient descent (SGD) method to update the model parameter set Ω with the cross-entropy L shown in the following formula (7) as the loss function.

[0056]

Equation

[0057] Here, m x is the correct distribution of the input video data x 1:T The expression method of the correct distribution is not particularly limited. For example, it may be expressed as a one-hot vector. Alternatively, the correct distribution may be represented by approximating a normal distribution centered on the correct class.

[0058] Note that the learning unit 15d stores the obtained learned model parameter set Ω' as the model parameters 14d in the storage unit 14.

[0059] In this case, as described above, the calculation unit 15b calculates the embedded representation of the mental state of the input data 14b and the embedded representation of the mental state of the reference data 14c using the learned model parameters 14b.

[0060] [Estimation Process] Next, the estimation process by the estimation device 10 will be described. FIG. 5 is a flowchart showing the estimation process procedure. The flowchart of FIG. 5 starts, for example, at the timing when an input instructing the start of the estimation process is received.

[0061] First, the acquisition unit 15a acquires learning data 14a including non-verbal information or para-verbal information and a correct label representing the mental state represented by the non-verbal information or para-verbal information (step S1).

[0062] Next, the calculation unit 15b calculates the embedded representation of the mental state of the input data and the embedded representation of the mental state of the reference data using the respective feature amounts of the input data 14b and the reference data 14c in the acquired learning data 14a (step S2).

[0063] Also, the calculation unit 15b compares the embedded representation calculated from the input data 14b with the embedded representation calculated from the reference data 14c (step S3).

[0064] Then, the estimation unit 15c estimates the mental state of the input data 14b using the comparison result between the embedded representation calculated from the input data 14b and the embedded representation calculated from the reference data 14c (step S4). Thereby, a series of estimation processes ends.

[0065] [Effect] As described above, in the estimation device 10 of the present embodiment, the acquisition unit 15a acquires learning data including non-verbal information or para-verbal information and a correct label representing the mental state represented by the non-verbal information or para-verbal information. Further, the calculation unit 15b calculates the embedded representation of the mental state of the input data 14b and the embedded representation of the mental state of the reference data 14c using the respective feature amounts of the acquired learning data 14a between the input data 14b and the reference data 14c. Further, the estimation unit 15c estimates the mental state of the input data 14b using the comparison result between the embedded representation calculated from the input data 14b and the embedded representation calculated from the reference data 14c.

[0066] Specifically, the estimation unit 15c compares the embedded representation calculated from the input data 14b and the embedded representation calculated from the reference data 14c using a multi-head self-attention mechanism. Further, the estimation unit 15c calculates the posterior probability for each class of the mental state classification and estimates the mental state. Alternatively, the estimation unit 15c estimates a numerical value representing the mental state as a regression problem.

[0067] In this way, the estimation device 10 estimates the mental state using a plurality of correct labels other than the normal state registered in advance as reference information. As a result, even if the dispersion of the labels is so large that individual differences cannot be fully normalized or absorbed, it is possible to accurately estimate the label representing the mental state represented by the non-verbal / para-verbal information.

[0068] Further, since it is not necessary to re-learn the model, it is not necessary to collect an appropriate amount of data well-balancedly for each class of classification or to monitor the learning, and processing can be performed with low resources.

[0069] Further, the learning unit 15d learns the model parameters 14d of a model that estimates the mental state represented by the input non-verbal information or paralinguistic information using the input data 14b and the estimated mental state of the input data 14b. In this case, the calculation unit 15b calculates the embedded representation of the mental state of the input data 14b and the embedded representation of the mental state of the reference data 14c using the learned model parameters 14b. As a result, it becomes possible to estimate more accurately the label representing the mental state represented by the non-verbal / paralinguistic information.

[0070] [Embodiment] FIG. 6 is a diagram for explaining an embodiment. FIG. 6 shows the accuracy in each method when the degree of understanding of unknown video data of the same person is estimated in three methods including the present invention at five levels. Here, a general method that does not use reference data (none), a method that absorbs individual differences using only reference data in a normal state (only degree of understanding 3), and a method to which the present invention is applied (degrees of understanding 2, 3, 4) were applied. In "none", the degree of understanding was estimated using the self-attention mechanism and the fully connected layer for the embedded representation H x without using reference data. Also, in "only degree of understanding 3", the degree of understanding was estimated using only the reference data of degree of understanding 3 (N = 3). Also, in "degrees of understanding 2, 3, 4", the degree of understanding was estimated using the reference data of degrees of understanding 2 to 4 (N = 3).

[0071] As shown in FIG. 6, compared with a general method that does not use reference data (none) or a method that uses only reference data in a normal state (only degree of understanding 3), in the method of the present invention that uses reference data of a plurality of states (degrees of understanding 2, 3, 4), it was confirmed that the accuracy of estimating the degree of understanding represented by the video data is stably improved.

[0072] [Program] It is also possible to create a program that describes the processing executed by the estimation device 10 according to the above embodiment in a language executable by a computer. As one embodiment, the estimation device 10 can be implemented by installing an estimation program that executes the above-described estimation processing as package software or online software on a desired computer. For example, by causing the information processing device to execute the above-described estimation program, the information processing device can function as the estimation device 10. In addition, other information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and further slate terminals such as PDAs (Personal Digital Assistant). Further, the functions of the estimation device 10 may be implemented on a cloud server.

[0073] FIG. 7 is a diagram showing an example of a computer that executes an estimation program. The computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0074] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System), for example. The hard disk drive interface 1030 is connected to the hard disk drive 1031. The disk drive interface 1040 is connected to the disk drive 1041. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1041, for example. A mouse 1051 and a keyboard 1052 are connected to the serial port interface 1050, for example. A display 1061 is connected to the video adapter 1060, for example.

[0075] Here, the hard disk drive 1031 stores, for example, the OS 1091, application programs 1092, program modules 1093, and program data 1094. Each piece of information described in the above embodiment is stored, for example, in the hard disk drive 1031 or the memory 1010.

[0076] Also, the estimation program is stored in the hard disk drive 1031 as a program module 1093 in which instructions executed by the computer 1000 are described, for example. Specifically, a program module 1093 in which each process executed by the estimation device 10 described in the above embodiment is described is stored in the hard disk drive 1031.

[0077] Also, the data used for the information processing by the estimation program is stored in the hard disk drive 1031 as program data 1094, for example. Then, the CPU 1020 reads out the program module 1093 and program data 1094 stored in the hard disk drive 1031 to the RAM 1012 as needed and executes each of the above-described procedures.

[0078] Note that the program module 1093 and program data 1094 related to the estimation program are not limited to being stored in the hard disk drive 1031. For example, they may be stored in a removable storage medium and read out by the CPU 1020 via a disk drive 1041 or the like. Alternatively, the program module 1093 and program data 1094 related to the estimation program may be stored in another computer connected via a network such as a LAN (Local Area Network) or WAN (Wide Area Network) and read out by the CPU 1020 via the network interface 1070.

[0079] The embodiments to which the invention made by the present inventor has been applied have been described above. However, the present invention is not limited by the description and drawings that form part of the disclosure of the present invention according to the present embodiment. That is, all other embodiments, examples, operation techniques, etc. made by those skilled in the art based on the present embodiment are included in the scope of the present invention.

Explanation of Signs

[0080] 10 Estimation device 11 Input unit 12 Output unit 13 Communication control unit 14 Storage unit 14a Learning data 14b Input data 14c Reference data 14d Model parameters 15 Control unit 15a Acquisition unit 15b Calculation unit 15c Estimation unit 15d Learning unit

Claims

1. An estimation method executed by an estimation device, comprising: an acquisition step of acquiring learning data including non-verbal information or para-verbal information and a correct label representing a mental state represented by the non-verbal information or para-verbal information; a calculation step of calculating an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using respective feature amounts of the input data and the reference data among the acquired learning data; an estimation step of estimating the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data; wherein the estimation step is characterized in that the embedded representation calculated from the input data and the embedded representation calculated from the reference data are compared using a multi-head self-attention mechanism.

2. An estimation method executed by an estimation device, comprising: an acquisition step of acquiring learning data including non-verbal information or para-verbal information and a correct label representing a mental state represented by the non-verbal information or para-verbal information; a calculation step of calculating an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using respective feature amounts of the input data and the reference data among the acquired learning data; an estimation step of estimating the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data; a learning step of learning model parameters of a model that estimates the mental state represented by the input non-verbal information or para-verbal information using the input data and the estimated mental state of the input data; wherein the calculation step is characterized in that the embedded representation of the mental state of the input data and the embedded representation of the mental state of the reference data are calculated using the learned model parameters.

3. The estimation step is characterized in that the mental state is estimated by calculating a posterior probability for each class of the classification of the mental state, according to the estimation method of Claim 1 or 2.

4. The estimation step is characterized in that a numerical value representing the mental state is estimated as a regression problem, according to the estimation method of Claim 1 or 2.

5. An acquisition unit that acquires learning data including non-verbal information or paralinguistic information and a correct label representing the mental state represented by the non-verbal information or paralinguistic information; A calculation unit that calculates an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the respective feature amounts of the input data and the reference data among the acquired learning data; An estimation unit that estimates the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data; having; The estimation unit is characterized in that it compares the embedded representation calculated from the input data and the embedded representation calculated from the reference data using a multi-head self-attention mechanism.

6. An acquisition unit that acquires learning data including non-verbal information or paralinguistic information and a correct label representing the mental state represented by the non-verbal information or paralinguistic information; A calculation unit that calculates an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the respective feature amounts of the input data and the reference data among the acquired learning data; An estimation unit that estimates the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data; A learning unit that learns model parameters of a model that estimates the mental state represented by the input non-verbal information or paralinguistic information using the input data and the estimated mental state of the input data; having; The calculation unit is characterized in that it calculates an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the learned model parameters.

7. An acquisition step of acquiring learning data including non-verbal information or paralinguistic information and a correct label representing the mental state represented by the non-verbal information or paralinguistic information; A calculation step of calculating an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the respective feature amounts of the input data and the reference data among the acquired learning data; An estimation step of estimating the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data; causing a computer to execute, The estimation step compares an embedded representation calculated from the input data and an embedded representation calculated from the reference data using a multi-head self-attention mechanism. An estimation program characterized by this.

8. An acquisition step of acquiring learning data including non-verbal information or para-verbal information and a correct label representing a mental state represented by the non-verbal information or para-verbal information, A calculation step of calculating an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the respective feature amounts of the input data and the reference data among the acquired learning data, An estimation step of estimating the mental state of the input data using a comparison result between the embedded representation calculated from the input data and the embedded representation calculated from the reference data, A learning step of learning model parameters of a model that estimates the mental state represented by the input non-verbal information or para-verbal information using the input data and the estimated mental state of the input data, causing a computer to execute, The calculation step calculates an embedded representation of the mental state of the input data and an embedded representation of the mental state of the reference data using the learned model parameters. An estimation program characterized by this.

Citation Information

Patent Citations

  • Method and apparatus for recognizing facial expression

    US20180114057A1

  • Recognition device, learning device, method for same, and program

    WO2021166207A1