Behavior estimation device, behavior estimation method, and behavior estimation program
The behavior estimation device predicts feature quantities and their occurrence times with fine granularity by employing SHAP analysis and kernel density estimation, addressing the coarseness of existing methods to identify impactful behaviors and scenes in conversations.
Patent Information
- Application Number
- JP2022080134
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-12-15
- Estimated Expiration
- 2042-05-16
AI Technical Summary
Existing technologies fail to predict the occurrence time of feature quantities that significantly influence interlocutors during a conversation with fine granularity, as they segment data coarsely, preventing accurate data preparation.
A behavior estimation device using a model to extract feature quantities with significant influence on interlocutors, calculated through SHAP analysis and kernel density estimation to identify time points of maximum impact, enabling fine-grained prediction.
Enables precise prediction of feature quantities and their occurrence times that affect interlocutors during conversations, allowing for the detection of behaviors and scenes with significant impact on impressions.
Smart Images

Figure 0007785290000018 
Figure 0007785290000019 
Figure 0007785290000020
Abstract
Description
[Technical Field]
[0001] The present invention relates to a behavior estimation device, a behavior estimation method, and a behavior estimation program. [Background technology]
[0002] Among the nonverbal behaviors that occur during human interaction, head movements are known to play a variety of roles. For example, speakers use head movements to emphasize what they are saying or to confirm a response, while listeners use head movements as interjections, responses, or signs of agreement. As such, head movements have multiple functions, and it is known that a single head movement can simultaneously have multiple meanings.
[0003] Taking note of the diversity and ambiguity of the functions of head movements, there are known techniques for extracting the function and meaning of a user's head movements during a conversation, or predicting the user's subjective impression (see Non-Patent Documents 1 and 2).
[0004] In addition to head movements, the characteristics of the people engaged in the conversation and their subjective impressions should be reflected in features that have a large influence at certain points during the conversation. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] K. Otsuka and M. Tsumori, “Analyzing Multifunctionality of Head Movements in Face-to-Face Conversations Using Deep Convolutional Neural Networks”, IEEE Access, 2020, vol.8, pp.217169-217195 [Non-patent document 2] Shumepi Otsuchi, et al., “Prediction of Interlocutors’ Subjective Impressions Based on Functional Head-Movement Features in Group Meetings” 、[online]、2021年、in Proceedings of ACM International Conference on Multimodal Interaction (ICMI2021), pp.352-360, [2022年4月13日検索]、インターネット<URL:https: / / doi.org / 10.1145 / 3462244.<3479930>
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, the prior art has a problem that it cannot predict the occurrence time of feature quantities that have a great influence on interlocutors during the conversation with a fine granularity. For example, if the feature quantities that have a great influence on the interlocutors during the conversation and their occurrence times are used as correct data, it becomes possible to predict the occurrence times of the time series of the feature quantities that have a great influence. However, in the prior art, data is segmented every two minutes, and it is impossible to identify the feature quantities that have influenced the interlocutors with a finer granularity than that, so correct data cannot be prepared.
[0007] The present invention has been made in view of the above, and an object thereof is to predict, with a fine granularity, the feature quantities that have a great influence on the interlocutors during the conversation and the occurrence times thereof.
Means for Solving the Problems
[0008] In order to solve the above-mentioned problems and achieve the object, the behavior estimation device of the present invention is characterized by having an extraction unit that uses a model that outputs a predicted value of a value related to a conversation partner in a conversation for a predetermined input feature value, and extracts feature values from data in which the conversation partners are in a conversation and that include the predetermined feature value, the magnitude of the influence on the predicted value of the value related to the conversation partner being equal to or greater than a predetermined threshold; a calculation unit that calculates the time distribution of the extracted feature values; and an identification unit that identifies the time point at which the magnitude of the influence of the feature values reaches a maximum value. [Effects of the Invention]
[0009] According to the present invention, it is possible to predict, with fine granularity, feature quantities that have a large influence on interlocutors during a conversation and the time points at which they occur. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram illustrating a schematic configuration of a behavior inference device. [Figure 2] FIG. 2 is a diagram for explaining the processing of the calculation unit. [Figure 3] FIG. 3 is a diagram for explaining the processing of the identification unit. [Figure 4] FIG. 4 is a flowchart showing the procedure of the behavior estimation process. [Figure 5] FIG. 5 is a schematic diagram illustrating a schematic configuration of a behavior inference device according to the second embodiment. [Figure 6] FIG. 6 is a diagram illustrating a computer that executes a behavior estimation program. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.
[0012] [Outline of the behavior estimation device] The behavior inference device of this embodiment predicts, at a fine level of granularity, feature quantities that influence values related to interlocutors during a dialogue and the time points at which those feature quantities occur. Here, the values related to interlocutors during a dialogue may be, for example, numerical values representing the impressions that an interlocutor has of the dialogue in progress or other interlocutors, or values representing the personality traits of the interlocutors themselves.
[0013] Specifically, the behavior estimation device uses a model that outputs a predicted value related to an interlocutor in response to input of a plurality of feature quantities to extract feature quantities that affect the predicted value and the time points at which they occur. This enables the behavior estimation device to predict feature quantities that affect the predicted value with fine granularity even if the granularity of the time intervals in the learning data of the model is not fine.
[0014] In this embodiment, the behavior estimation device, for example, takes the impression (hereinafter also referred to as subjective impression) that a conversational participant has of a conversation during the conversation as a value related to the conversational participant, and predicts feature quantities that have a large influence on this subjective impression and the time points at which they occur. For example, the behavior estimation device applies SHAP analysis to a trained model to calculate a contribution degree that indicates the magnitude of influence that each feature quantity has on the predicted value of the impression. The behavior estimation device also extracts a set of one or more feature quantities with the highest contribution degrees, approximates the time distribution of each feature quantity with the distribution of the occurrence probability of each feature quantity, and estimates it using a kernel density estimation method. The behavior estimation device then calculates the time distribution of the contribution degree by taking the sum of the products of the time distribution and the contribution degree, thereby identifying feature quantities that have a large influence on the conversational participant's impression and the time points at which they occur.
[0015] [Configuration of behavior estimation device] Fig. 1 is a schematic diagram illustrating the overall configuration of a behavior inference device. As illustrated in Fig. 1, the behavior inference device 10 is realized by a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.
[0016] The input unit 11 is realized using input devices such as a keyboard and a mouse, and inputs various instruction information such as a command to start processing to the control unit 15 in response to input operations by an operator. The output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, or the like.
[0017] The communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication between an external device such as a server via a network and the control unit 15. For example, the communication control unit 13 controls communication between the control unit 15 and a management device or the like that manages data used for behavior estimation, which will be described later.
[0018] The storage unit 14 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 14 stores in advance a processing program for operating the behavior estimation device 10, data used during execution of the processing program, and the like, or temporarily stores the data each time processing is performed. For example, the storage unit 14 stores a model 14a and the like used in the behavior estimation process described below. The storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13.
[0019] Here, the model 14a outputs a predicted value of a value related to the interlocutors during a dialogue for a predetermined input feature amount. The model 14a of this embodiment is trained to output a predicted value of the subjective impression that the interlocutors have of the dialogue during the dialogue.
[0020] Specifically, the model 14a calculates the introspection score y for the impression item I input by the interlocutor j. i,j (i∈I) is used as the correct answer data. Model 14a is a regression model that trains each item. When feature quantity xj is input, the predicted value y i ^(x j ) is output.
[0021] The feature quantity is, for example, the motion duration Hrate , 3-DOF head posture angle θ azi,t , θ ele,t , θ roll,t , and their respective variances δ 2 azi , δ 2 ele , δ 2 roll Alternatively, it may be one or more of the function content rate, the functional division composition ratio, the function appearance rate, etc.
[0022] The control unit 15 is realized using a CPU (Central Processing Unit) or the like, and executes a processing program stored in a memory. As a result, the control unit 15 functions as an extraction unit 15a, a calculation unit 15b, an identification unit 15c, and an estimation unit 15d, as exemplified in FIG. 1, and executes a behavior estimation process described below. Note that these functional units may be implemented individually or in part in different hardware. For example, the estimation unit 15d may be implemented as a device separate from the other functional units. The control unit 15 may also include other functional units.
[0023] The extraction unit 15a uses the model 14a that outputs a predicted value of a value related to a conversation partner during a conversation for an input predetermined feature quantity to extract, from data of the conversation partner during a conversation that includes the predetermined feature quantity, a feature quantity whose magnitude of influence on the predicted value of the value related to the conversation partner is equal to or greater than a predetermined threshold. For example, the extraction unit 15a extracts a feature quantity whose magnitude of influence on the subjective impression that the conversation partner has of the conversation during the conversation is equal to or greater than a predetermined threshold.
[0024] Specifically, the extraction unit 15a extracts the features by calculating a SHAP value that reflects the contribution of each feature to a predicted value of the interlocutor obtained by using the feature. Here, the SHAP value is calculated as the average of the marginal contributions of each feature for all permutations of the features, and for each feature f∈x in the dataset, j represents the magnitude of the influence on the final impression prediction result.
[0025] The marginal contribution is the expected contribution to the predicted value obtained by using feature f. Given the impression i to be predicted and feature x obtained from interlocutor j, j The following equations (1) and (2) hold between the prediction result for , the expected value of the prediction result of the model, and the SHAP value of the feature value f.
[0026]
number
number
[0027] The above formula (2) means that for all permutations of features, the difference in prediction results when feature f is used is calculated, and then the average is calculated. The magnitude of the absolute value of the SHAP value obtained here represents the degree of influence on the impression. In other words, the larger this value, the greater the influence the feature had on the impression. Furthermore, if the sign of this value is positive, it means that it had an influence on improving the impression, and if it is negative, it means that it had an influence on worsening the impression.
[0028] Therefore, in this embodiment, the extraction unit 15a selects and extracts R feature quantities in descending order of the absolute value of the SHAP value. Each feature quantity selected here has a clear correlation with an action. For example, a feature quantity called "the frequency of nodding in a conversation" is associated with the action of "nodding." If, as a result of the SHAP analysis, the set F^ of R feature quantities extracted as having a significant impact on the impression i of a certain person j includes "the frequency of nodding in a conversation," the action of "nodding" corresponding to this feature quantity is identified as having contributed to forming the impression of this person. Furthermore, the sign of the SHAP value makes it possible to infer whether the action contributed to improving or worsening the impression.
[0029] The calculation unit 15b calculates the time distribution of the extracted feature amount. Specifically, the calculation unit 15b performs time expansion of the feature amount in order to estimate the influence of an action at a certain time point in the conversation on the formation of an impression.
[0030] At this time, the calculation unit 15b calculates the time distribution of the feature using a temporal addition method in which feature values obtained from partial time intervals are added for the entire time interval. The temporal addition method of feature values refers to the property that feature values obtained from partial time intervals of a dialogue become equivalent to feature values for the entire dialogue. In other words, by calculating the distribution in which the feature value f∈F^ is expanded over time, it becomes possible to understand, based on the level of the distribution, which point in time contributed to the composition of the feature value and to what extent.
[0031] To calculate the distribution of time-expanded features, we focus on the function appearance rate, which indicates the proportion of frames in which each function, which represents behavior such as head movement function, appears. The function appearance rate can be calculated by counting the number of times a function appears in each frame from the start to the end of the dialogue and dividing this by the duration of the dialogue, which shows that the temporal summation method is valid.
[0032] Each feature is detected as a discrete value of 0 / 1 at each time. Therefore, we consider this process to be a type of stochastic process, and assume that features are generated and detected probabilistically. Furthermore, the time distribution of the probability that a feature is generated or detected at each time is called the occurrence probability distribution of the feature, and is considered to be the distribution of the feature expanded over time.
[0033] Then, the calculation unit 15b uses kernel density estimation to estimate the probability distribution of occurrence of a function corresponding to the feature, thereby calculating the time distribution of the feature. That is, the calculation unit 15b uses kernel density estimation to approximately estimate the occurrence probability distribution of the feature. When a finite number of sample points are given, the kernel density estimation estimates a continuous distribution that is the basis of the sample points. In this case, if a Gaussian function is used as the kernel function, the occurrence probability distribution of the function appearance rate at time t for each person j is expressed as in the following equations (3) to (5).
[0034]
number
number
number
[0035] The bandwidth h is a parameter that controls the degree of temporal smoothing of the occurrence probability distribution. In the above formula (3), the value of the occurrence probability distribution at time t represents the probability that a function such as head movement corresponding to feature f at that time will occur, i.e., an estimated value of the occurrence rate per unit frame. In time periods when that function frequently appears, the occurrence probability distribution shows a high value, and is considered to contribute more significantly to the construction of feature f.
[0036] Since this feature is the feature f that contributed to the formation of the impression identified by the SHAP analysis, it is thought that behavior that occurred at times when the occurrence probability distribution showed higher values had a greater impact on the formation of the impression.
[0037] The following equation (6) holds between the occurrence probability distribution and the feature amount.
[0038]
number
[0039] The above formula (6) means that the sum of the occurrence probability distribution over all intervals is equal to the value of the feature, suggesting that the feature is additive over time.
[0040] Here, Fig. 2 is a diagram for explaining the processing of the calculation unit. Fig. 2 illustrates an example of an estimation result of an occurrence probability distribution for a time series of detection results of a certain function, using the kernel density estimation method. Specifically, Fig. 2 illustrates a curve showing the occurrence probability distribution and the time at which a head movement function is detected.
[0041] The occurrence probability distribution for the function content is expressed as the following equations (7) and (8) using the occurrence probability distribution for the function appearance rate.
[0042]
number
number
[0043] Similarly, the occurrence probability distribution regarding the functional division composition ratio is expressed by the following equations (9) to (11).
[0044]
number
number
number
[0045] Similarly, the occurrence probability distribution of the variance of the movement duration and head posture angle, which are kinematic features, is calculated as the head movement detection result d t Using the above, it is defined as in the following equations (12) to (14).
[0046]
number
number
number
[0047] The calculation unit 15b uses the above definition to calculate, for each feature quantity that contributes greatly to the predicted value of the impression, an occurrence probability distribution p j,fThe occurrence probability distribution may be calculated for all feature quantities, or for any one of the feature quantities. In this case, any one of the feature quantities may be selected manually or a predetermined number n of feature quantities may be selected randomly.
[0048] Returning to the explanation of FIG. 1, the identification unit 15c identifies the time point at which the magnitude of the influence of the feature amount reaches a maximum value. Here, FIG. 3 is a diagram for explaining the processing of the identification unit. The identification unit 15c receives from the calculation unit 15b an occurrence probability distribution corresponding to the time expansion of the feature amount obtained for each feature amount, and expands a SHAP value corresponding to the degree to which these feature amounts contributed to the predicted value of the impression on the time axis. Then, as exemplified in FIG. 3, the identification unit 15c identifies the time point at which the distribution of the SHAP value reaches a maximum value, and estimates this time point as the time that contributed most to the formation of the impression.
[0049] Specifically, the specification unit 15c calculates the occurrence probability distribution p j,f Normalize (t).
[0050]
number
[0051] The normalized occurrence probability distribution of the above formula (15) indicates the proportion of contribution of the feature f to the formation of the impression at each time. The identification unit 15c calculates the absolute value of the sum of the products of this normalized occurrence probability distribution and the SHAP value for the feature included in the feature set, and defines the time distribution of the contribution to the impression (hereinafter also referred to as the contribution distribution) as shown in the following formula (16).
[0052]
number
[0053] Here, the degree of contribution of a set of multiple feature quantities to the prediction result is calculated as the sum of the SHAP values of the feature quantities, and the additivity of the SHAP values is utilized. When calculating the sum of the SHAP values, the specifying unit 15c uses a predetermined weight w n (n is the number of features) and calculate the SHAP value of each feature using w n Alternatively, the specification unit 15c may select any feature amount and calculate the sum of the SHAP values of the selected feature amount. Alternatively, the specification unit 15c may select any feature amount and calculate the sum of the SHAP values of the selected feature amount. Alternatively, the specification unit 15c may select any feature amount multiple times, or may assign any weight w to each SHAP value of each selected feature amount. n The sum of the SHAP values may be calculated after integrating the above.
[0054] The time point (maximum time point) t^ showing the maximum value of the contribution distribution obtained in this way is identified as shown in the following equation (17).
[0055]
number
[0056] In this embodiment, it is assumed that the impression of a conversation partner is significantly influenced by their behavior at a specific time point, and the time point at which the obtained SHAP value is at its maximum and the behavior suggested by the feature values at that time are identified as the behavior that had the greatest impact on the impression of the entire conversation.
[0057] In this case, if a plurality of sums of SHAP values are calculated, the same number of maximum time points as the number of sums of SHAP values can be obtained. The identification unit 15c outputs the maximum time points and the feature amounts used to identify the maximum time points.
[0058] Returning to the explanation of Figure 1, the estimation unit 15d estimates the behavior of the interlocutor corresponding to the feature at the time of the maximum value. Specifically, the estimation unit 15d receives the maximum time output by the identification unit 15c and the feature used to identify the maximum time, and estimates what behavior the feature corresponding to the maximum time corresponds to.
[0059] Specifically, by registering in advance the behaviors corresponding to the combination of feature amounts and the time of the maximum time point, the estimation unit 15d identifies the behavior corresponding to the information output from the identification unit 15c. For example, assume that "behavior A" is registered corresponding to the case where the feature amount of "head posture angle" exists from time hh:mm:00 to hh:mm:59. Then, if the maximum time point falls within the above time period and the feature amount corresponding to the maximum time point is the feature amount of "head posture angle," the estimation unit 15d estimates "behavior A." This makes it possible to detect behaviors that have a large impact on the impression, etc., of interlocutors during a conversation.
[0060] [Behavior estimation processing] Next, the behavior inference process performed by the behavior inference device 10 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the behavior inference process procedure. The flowchart in Fig. 4 starts, for example, when an operation input is made to instruct the start of the behavior inference process.
[0061] First, the extraction unit 15a uses the model 14a that outputs a predicted value of a value related to a conversation partner during a conversation for an input predetermined feature quantity to extract a feature quantity whose magnitude of influence on the predicted value of a value related to the conversation partner is equal to or greater than a predetermined threshold from data of the conversation partner during a conversation that includes the predetermined feature quantity (step S1). For example, the extraction unit 15a extracts a feature quantity whose magnitude of influence on the subjective impression that the conversation partner has of the conversation during the conversation is equal to or greater than a predetermined threshold.
[0062] Next, the calculation unit 15b calculates the time distribution of the extracted feature amounts (step S2). Specifically, the calculation unit 15b performs time expansion of the feature amounts in order to estimate the extent to which actions at which points in the dialogue influenced the formation of an impression. In this case, the calculation unit 15b calculates the time distribution of the feature amounts using a temporal addition method in which feature amounts obtained from partial time intervals are added for the entire time interval.
[0063] Furthermore, the calculation unit 15b calculates the time distribution of the feature by estimating the probability distribution of the occurrence of a function corresponding to the feature using a kernel density estimation method. That is, the calculation unit 15b approximately estimates the occurrence probability distribution of the feature using the kernel density estimation method.
[0064] Next, the identifying unit 15c identifies the time point at which the magnitude of the influence of the feature amount reaches a maximum value (step S3). Specifically, the identifying unit 15c uses the occurrence probability distribution of the feature amount to develop, on the time axis, a SHAP value corresponding to the degree to which these feature amounts contributed to the predicted value of the impression. Then, the identifying unit 15c identifies the time point at which the distribution of the SHAP value reaches a maximum value, and estimates this time point as the time that most contributed to the formation of the impression.
[0065] Then, the estimation unit 15d estimates the behavior of the interlocutor corresponding to the feature amount at the time of the maximum value, and outputs the estimated behavior via, for example, the output unit 12. This completes a series of behavior estimation processes.
[0066] [Second embodiment] 5 is a schematic diagram illustrating a schematic configuration of a behavior inference device according to the second embodiment. In the following, only differences from the behavior inference process of the behavior inference device 10 of the above-described embodiment will be described, and a description of commonalities will be omitted.
[0067] As shown in FIG. 5, the behavior inference device 10 of the second embodiment differs from the behavior inference device 10 of the above embodiment in that it has a topic extraction unit 15e and related data 14b instead of the estimation unit 15d of the behavior inference device 10 of the above embodiment.
[0068] The related data 14b is time-series data related to the dialogue, such as video data of images. The time-series data may be audio data or point cloud data. The related data 14b is acquired in advance from, for example, an external management device via the communication control unit 13 and stored in the storage unit 14.
[0069] In the behavior inference device 10 of the second embodiment, the identification unit 15c identifies n (≧1) maximum time points t n Identify the ^.
[0070] Then, the topic extraction unit 15e extracts a topic corresponding to the time point of the maximum value from the data related to the conversation divided into predetermined topics. Specifically, the topic extraction unit 15e receives, for example, video data as data related to the conversation and divides it into topics.
[0071] The segmentation into topic units may be performed manually. Alternatively, the topic extraction unit 15e may segment a video when, for example, the sound pressure level is below a predetermined threshold based on audio features, or may segment the video by setting multiple thresholds. Alternatively, the topic extraction unit 15e may extract optical flow from video features and segment the video when a vector equal to or greater than a predetermined threshold is detected, or may segment the video by setting multiple thresholds based on other video features. Alternatively, the topic extraction unit 15e may build a machine learning model by learning using data that has already been segmented as correct answer data, and predict segmentation locations. Other commonly used methods for segmenting videos may also be used.
[0072] Then, the topic extraction unit 15e extracts a topic that includes a maximum time point from among the divided topics. In this case, the topic extraction unit 15e may extract the topic from the same time-series data as the time-series data used for topic division, or may extract the topic from other time-series data other than video data that has a common time series with the time-series data used for topic division, such as audio data or point cloud data.
[0073] Furthermore, the topic extraction unit 15e may extract topics corresponding to all of the maximum time points, or may extract a topic corresponding to any selected maximum time point. The topic extraction unit 15e outputs the extracted topics, for example, via the output unit 12. This allows detection of scenes related to the dialogue that have a large impact on the impressions of the interlocutors during the dialogue.
[0074] [effect] As described above, in the behavior estimation device 10, the extraction unit 15a uses the model 14a that outputs a predicted value of a value related to a conversation partner in a conversation for an input predetermined feature quantity to extract a feature quantity whose magnitude of influence on the predicted value of a value related to the conversation partner is equal to or greater than a predetermined threshold from data in which the conversation partner is in a conversation and which includes the predetermined feature quantity. Furthermore, the calculation unit 15b calculates the time distribution of the extracted feature quantity. Furthermore, the identification unit 15c identifies a time point at which the magnitude of the influence of the feature quantity reaches a maximum value.
[0075] Specifically, the extraction unit 15a extracts features whose influence is equal to or greater than a predetermined threshold by calculating a SHAP value that reflects the contribution of each feature to the predicted value of the value related to the interlocutor obtained by using the feature.
[0076] Furthermore, the calculation unit 15b calculates the time distribution of the feature quantity using a temporal addition method in which feature quantities obtained from partial time intervals are added for the entire time interval. In this case, the calculation unit 15b calculates the time distribution of the feature quantity by estimating the probability distribution of the occurrence of a function corresponding to the feature quantity using a kernel density estimation method.
[0077] This allows the behavior estimation device 10 to predict, with fine granularity, feature amounts that affect the predicted value and the time points at which they occur, even if the granularity of the time intervals in the learning data of the model 14a is not fine.
[0078] Furthermore, the estimation unit 15d estimates the behavior of the interlocutor corresponding to the feature amount at the time of the maximum value, thereby enabling the behavior estimation device 10 to detect behavior that has a large influence on the interlocutor during a conversation.
[0079] Furthermore, the topic extraction unit 15e extracts a topic corresponding to the time point of the maximum value from the data on the conversation divided into predetermined topics, which enables the behavior estimation device 10 to detect a scene related to the conversation that has a large influence on the impression of the conversation participants during the conversation.
[0080] [program] A program describing the processing executed by the behavior estimation device 10 according to the above embodiment in a computer-executable language can also be created. As one embodiment, the behavior estimation device 10 can be implemented by installing a behavior estimation program that executes the above behavior estimation processing as package software or online software on a desired computer. For example, by causing an information processing device to execute the behavior estimation program, the information processing device can function as the behavior estimation device 10. Other examples of information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHSs (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants). The functions of the behavior estimation device 10 may also be implemented on a cloud server.
[0081] 6 is a diagram showing an example of a computer that executes a behavior estimation program. The computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0082] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1031. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to, for example, a mouse 1051 and a keyboard 1052. The video adapter 1060 is connected to, for example, a display 1061.
[0083] Here, the hard disk drive 1031 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. Each piece of information described in the above embodiment is stored in the hard disk drive 1031 or memory 1010, for example.
[0084] The behavior estimation program is stored in the hard disk drive 1031, for example, as a program module 1093 in which instructions to be executed by the computer 1000 are written. Specifically, the hard disk drive 1031 stores the program module 1093 in which each process executed by the behavior estimation device 10 described in the above embodiment is written.
[0085] Furthermore, data used for information processing by the behavior estimation program is stored as program data 1094, for example, in the hard disk drive 1031. Then, the CPU 1020 reads the program module 1093 and the program data 1094 stored in the hard disk drive 1031 into the RAM 1012 as necessary, and executes each of the above-described procedures.
[0086] The program module 1093 and program data 1094 related to the behavior estimation program are not limited to being stored in the hard disk drive 1031, and may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1041 or the like. Alternatively, the program module 1093 and program data 1094 related to the behavior estimation program may be stored in another computer connected via a network such as a LAN (Local Area Network) or a WAN (Wide Area Network), and read by the CPU 1020 via the network interface 1070.
[0087] Although the present invention has been described above as an embodiment, the present invention is not limited to the description and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]
[0088] 10 Behavior estimation device 11 Input section 12 Output section 13 Communication control section 14 Storage section 14a model 14b Related Data 15 Control Unit 15a Extraction part 15b Calculation part 15c Specific part 15d Estimation part 15e Topic Extraction
Claims
1. an extraction unit that uses a model that outputs a predicted value representing a subjective impression that a given interlocutor has toward a given dialogue during the dialogue, for a given feature value related to the given interlocutor during the dialogue, and extracts, from data during the dialogue of the interlocutor that includes the given feature value, a feature value whose magnitude of influence on the predicted value representing the subjective impression is equal to or greater than a given threshold; a calculation unit that calculates a time distribution of the extracted feature amount; an identification unit that identifies a time point at which the magnitude of the influence of the feature amount reaches a maximum value; A behavior estimation device comprising:
2. 2. The behavior inference device according to claim 1, wherein the extraction unit extracts the feature amounts by calculating a SHAP value that reflects a contribution of each feature amount to a predicted value of the value representing the subjective impression obtained by using each feature amount.
3. The behavior inference device according to claim 1 , wherein the calculation unit calculates the time distribution of the feature amounts using a temporal addition method in which the feature amounts obtained from partial time intervals are added for the entire time interval.
4. The behavior inference device according to claim 3 , wherein the calculation unit calculates the time distribution of the feature by estimating a probability distribution of occurrence of a function corresponding to the feature using a kernel density estimation method.
5. The behavior estimation device according to claim 1 , further comprising an estimation unit that estimates the behavior of the interlocutor corresponding to the feature amount at the time of the maximum value.
6. The behavior estimation device according to claim 1 , further comprising a topic extraction unit that extracts a topic corresponding to a time point of the maximum value from data relating to conversations divided into predetermined topics.
7. A behavior estimation method executed by a behavior estimation device, an extraction step of extracting, from data during a conversation by an interlocutor that includes a predetermined feature quantity, a feature quantity whose magnitude of influence on the predicted value of the value representing the subjective impression is equal to or greater than a predetermined threshold, using a model that outputs a predicted value representing a subjective impression that the interlocutor has toward the conversation during the conversation, for the predetermined feature quantity related to the interlocutor that is input; a calculation step of calculating a time distribution of the extracted feature amount; a step of identifying a time point at which the magnitude of the influence of the feature amount reaches a maximum value; A behavior estimation method comprising:
8. A behavior estimation program for causing a computer to function as the behavior estimation device according to any one of claims 1 to 6.
Citation Information
Patent Citations
JP1145346224A
Information processing device, information processing method and program
JP2010061450A
Customer service data processing device and customer service data processing method
JP2016206736A
Call recording text analysis system and method
JP2021189817A
Information processing unit, information processing method and program
JP2022015687A