Communication Prediction Device and Communication Prediction Method

The communication prediction system uses machine learning on social networking service data to derive linguistic and relational features, addressing inefficiencies in data collection and enhancing speaker prediction accuracy.

JP7697400B2Active Publication Date: 2025-06-24TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022068321
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-06-24
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

Existing communication prediction systems require significant labor for data collection to accurately estimate the next speaker, which is inefficient and time-consuming.

Method used

A communication prediction apparatus and method that utilizes machine learning to derive linguistic and relational feature amounts from social networking service data to predict the next speaker, reducing the need for extensive data collection by leveraging information from posted texts, reactions, and user interactions.

Benefits of technology

Accurately predicts the next speaker while significantly reducing the labor required for data collection, enabling more efficient and precise communication prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697400000001
    Figure 0007697400000001
  • Figure 0007697400000002
    Figure 0007697400000002
  • Figure 0007697400000003
    Figure 0007697400000003
Patent Text Reader

Abstract

To provide a technology which can accurately predict a next speaker, while reducing the labor of collection of learning data.SOLUTION: In a communication prediction device 10, a first acquisition unit 20 acquires information inputted by a plurality of users in social networking service. A learning unit 24 carries out machine learning of a model for predicting a next speaker from among a plurality of conversation participants, on the basis of the acquired information.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a communication prediction apparatus and a communication prediction method.

Background Art

[0002] Patent Document 1 discloses a conversation support system that prompts a participant in a conversation who has missed the appropriate timing of speaking during the conversation to speak. In this system, based on the measurement results of the non-verbal behavior of each participant during the conversation, the probability of the next speaker, which is the probability that each participant will be the next speaker at an arbitrary time, is estimated, and based on the probability of the next speaker of each participant, the predicted next speaker, who is the participant who should speak next, is estimated.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technology of Patent Document 1, since data acquired in past conversations is used as learning data and a model for estimating the probability of the next speaker is learned, a great deal of labor is required for data collection.

[0005] An object of the present invention is to provide a technology that can accurately predict the next speaker while reducing the labor of collecting learning data.

Means for Solving the Problems

[0006] In order to solve the above problems, a communication prediction apparatus according to an aspect of the present invention includes an acquisition unit that acquires information input by a plurality of users in a social networking service, and a learning unit that machine-learns a model for predicting the next speaker from among a plurality of conversation participants based on the acquired information. Derive the linguistic feature amounts of a plurality of users based on the posted texts input by the plurality of users included in the acquired information, and based on the information regarding the reactions of the target user to the posted texts of other users input by the target user included in the acquired information, derive a relational feature amount representing the relationship between the target user and other users, and information representing the number of speaker alternations from the target user to other users. A derivation unit comprises. The learning unit takes as input the linguistic feature amount of the target user, the linguistic feature amounts of other users, and the relational feature amount between the target user and other users, and learns a model using the information representing the number of speaker alternations from the target user to other users as the correct label.

[0007] Another aspect of the present invention is also a communication prediction device. This device includes an acquisition unit that acquires information regarding the communication between each conversation participant and other conversation participants during a conversation among a plurality of conversation participants, and a prediction unit that predicts the next speaker from among the plurality of conversation participants based on the acquired communication-related information using a learned model that has been machine-learned based on information input by a plurality of users in a social networking service. A derivation unit that derives, based on the acquired information regarding communication, the respective second feature amount vectors of the plurality of conversation participants comprises. The prediction unit acquires the respective first feature amount vectors of the plurality of conversation participants from a storage unit that stores the first feature amount vectors of each user derived from the information input by the plurality of users in the social networking service, and uses the learned model to predict the next speaker based on the plurality of acquired first feature amount vectors and the plurality of derived second feature amount vectors.

[0008] Yet another aspect of the present invention is a communication prediction method. This method A communication prediction method executed by a computer, includes a step of acquiring information input by a plurality of users in a social networking service, and a step of machine-learning a model for predicting the next speaker from among a plurality of conversation participants based on the acquired information. Derive the linguistic feature amounts of a plurality of users based on the posted texts input by the plurality of users included in the acquired information, and based on the information regarding the reactions of the target user to the posted texts of other users input by the target user included in the acquired information, derive a relational feature amount representing the relationship between the target user and other users, and information representing the number of speaker alternations from the target user to other users. A step comprises. In the step of machine learning, take as input the linguistic feature amount of the target user, the linguistic feature amounts of other users, and the relational feature amount between the target user and other users, and learn a model using the information representing the number of speaker alternations from the target user to other users as the correct label.

[0009] Yet another aspect of the present invention is also a communication prediction method. This method A communication prediction method executed by a computer, includes a step of acquiring information regarding the communication between each conversation participant and other conversation participants during a conversation among a plurality of conversation participants, and a step of predicting the next speaker from among the plurality of conversation participants based on the acquired communication-related information using a learned model that has been machine-learned based on information input by a plurality of users in a social networking service. Derive, based on the acquired information regarding communication, the respective second feature amount vectors of the plurality of conversation participants. A step comprises. In the step of predicting, from a storage unit that stores the first feature vector of each user derived from information input by a plurality of users in a social networking service, the first feature vector of each of the plurality of conversation participants is obtained, and using the learned model, the next speaker is predicted based on the plurality of obtained first feature vectors and the plurality of derived second feature vectors.

Advantages of the Invention

[0010] According to the present invention, it is possible to provide a technique for accurately predicting the next speaker while reducing the labor of collecting learning data.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0012] (First Embodiment) FIG. 1 is a diagram for explaining the outline of the processing of the communication prediction system according to the first embodiment. The communication prediction system predicts information regarding conversations of a plurality of conversation participants in the physical space using communication information in the cyber space. Hereinafter, an example of three conversation participants u1, u2, and u3 will be described, but the number of people is not particularly limited.

[0013] During a conversation by a plurality of conversation participants u1, u2, and u3, the communication prediction system acquires voice data, video data, etc. of the conversation participants, and based on the acquired data, regularly acquires information regarding the communication between each conversation participant and other conversation participants. The information regarding communication includes information regarding conversation content, eye line, nodding, etc.

[0014] The learned model 50 is a model for predicting the next speaker from among a plurality of conversation participants, and is pre-trained by machine learning based on information such as post texts and reactions input by a plurality of users in a social networking service (hereinafter also referred to as SNS). It is assumed that each conversation participant u1, u2, u3 uses SNS, and the plurality of users to be learned includes the conversation participants u1, u2, u3. However, as will be described later, the conversation participants do not necessarily have to use SNS.

[0015] The communication prediction system uses the learned model 50 to periodically predict, for each conversation participant, the number of speaker alternations per unit time from the respective conversation participant to each other conversation participant, based on information regarding the communication of the conversation participants u1, u2, u3 and information input by the conversation participants u1, u2, u3 to the SNS. The prediction result can be represented in a tabular form as shown in the figure. The unit time and the prediction frequency can be appropriately determined by experiments or simulations. For example, the unit time may be 1 minute, and the prediction frequency may be once per minute.

[0016] In the example shown in the figure, the number of speaker alternations per unit time from the conversation participant u1, who is the speaker, to the conversation participant u2, who is the next speaker, is predicted to be "2", and the number of speaker alternations per unit time to the conversation participant u3 is predicted to be "20".

[0017] The number of speaker alternations per unit time from the conversation participant u2, who is the speaker, to the conversation participant u1, who is the next speaker, is predicted to be "5", and the number of speaker alternations per unit time to the conversation participant u3 is predicted to be "10".

[0018] The number of speaker alternations per unit time from the conversation participant u3, who is the speaker, to the conversation participant u1, who is the next speaker, is predicted to be "10", and the number of speaker alternations per unit time to the conversation participant u2 is predicted to be "3".

[0019] Note that in this example, transitions where the speaker and the next speaker are the same are excluded, and the number of speaker alternations is fixed at "0". However, depending on the application, transitions where the speaker and the next speaker are the same may also be defined as speaker alternations, and the number of such speaker alternations may also be predicted.

[0020] The number of speaker alternations per unit time represents the frequency of speaker alternation and also represents the transition probability of the speaker. For example, when the speaker is conversation participant u1, the probability that the next speaker is conversation participant u2 is 100×2 / (2 + 20) [%], and the probability that the next speaker is conversation participant u3 is 100×20 / (2 + 20) [%]. It can be predicted that the next speaker is conversation participant u3. That is, the communication prediction system predicts the next speaker from among a plurality of conversation participants.

[0021] Compared with the case of collecting the data of actual conversations as learning data, the collection of SNS data requires less effort, and more data regarding more users can be collected in a shorter time. By using a prediction model learned using such a large amount of learning data, the next speaker can be predicted accurately.

[0022] By predicting the next speaker, for example, the predicted next speaker can be notified to a plurality of conversation participants. Depending on whether the predicted next speaker matches the actual next speaker, it is also possible to confirm whether the utterance was made at an appropriate timing.

[0023] Also, the number of speaker alternations per unit time can be said to represent the amount of conversation per unit time. In the example of FIG. 1, the amount of conversation of conversation participant u2 following the utterance of conversation participant u1 can be expressed as "2", and the amount of conversation of conversation participant u3 following the utterance of conversation participant u1 can also be expressed as "20". That is, the communication prediction system predicts, for each conversation participant, the amount of conversation per unit time of each other conversation participant following the utterance of that conversation participant. Also, for example, the amount of conversation of conversation participant u2 following the utterance of conversation participant u1 and the amount of conversation of conversation participant u1 following the utterance of conversation participant u2 are each regarded as the amount of conversation between conversation participant u1 and conversation participant u2. Therefore, the amount of conversation between conversation participant u1 and conversation participant u2 is predicted to be "2" + "5" = "7". That is, it can be said that the communication prediction system predicts the amount of conversation per unit time for each combination of two conversation participants among a plurality of conversation participants. In this way, the amount of conversation can be predicted in advance. That is, it is possible to predict in advance who will talk and how much.

[0024] The predicted amount of conversation can be used, for example, as the amount of conversation per unit time between two members in the communication support system described in Japanese Patent Application No. 2021-163089 previously filed by the present applicant. For example, the future stress state value of conversation participant u1 can be predicted from the current stress state values of each of conversation participants u1, u2, u3, the predicted amount of conversation between conversation participant u1 and conversation participant u2, and the predicted amount of conversation between conversation participant u1 and conversation participant u3. That is, it is possible to predict the future stress state value of an individual considering the influence of conversation with others.

[0025] FIG. 2 shows the configuration of the communication prediction system 1 according to the first embodiment. The communication prediction system 1 includes a microphone 2, a camera 4, a sensor 6, and a communication prediction device 10. Although not shown, the communication prediction system 1 has a plurality of microphones 2, a plurality of cameras 4, and a plurality of sensors 6.

[0026] The microphone 2 acquires the voices of a plurality of conversation participants and supplies the acquired voice data to the communication prediction device 10. The camera 4 photographs a plurality of conversation participants and supplies the photographed image data to the communication prediction device 10. The sensor 6 is attached to the body of each conversation participant, detects the posture of each conversation participant, and supplies the detected sensor data to the communication prediction device 10. The sensor 6 may detect, for example, the heart rate of each conversation participant.

[0027] The communication prediction device 10 includes a processing unit 12, an SNS data storage unit 14, a feature dictionary storage unit 16, and a learned model storage unit 18. The processing unit 12 includes a first acquisition unit 20, a first derivation unit 22, a learning unit 24, a second acquisition unit 26, a voice recognition unit 28, a second derivation unit 30, a prediction unit 32, and an output unit 34.

[0028] In terms of hardware, the configuration of the processing unit 12 can be realized by the CPU, memory, and other LSIs of an arbitrary computer, and in terms of software, it can be realized by a program loaded into the memory, etc. Here, however, functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.

[0029] First, the processing related to learning will be described. The first acquisition unit 20 acquires information (hereinafter also referred to as SNS data) input by a plurality of users on the SNS and stores the acquired SNS data in the SNS data storage unit 14. The SNS data includes information such as the content of the posted text, reactions such as emojis to the posted text, replies to the posted text, and the image or profile image of the user's account.

[0030] The first acquisition unit 20 performs preprocessing on the acquired SNS data as necessary, creates learning data from the SNS data, and supplies the learning data to the first derivation unit 22. The preprocessing is, for example, text formatting processing of the posted text, and includes processing for unifying full-width characters such as alphabets and katakana into half-width characters, deleting emoticons, etc., and performing word segmentation.

[0031] The first derivation unit 22 derives the first feature vector of each user from the learning data. The first feature vector includes a linguistic feature, an image feature, and a relational feature.

[0032] The first derivation unit 22 derives the linguistic features of a plurality of users based on the posted texts input by the plurality of users included in the information acquired by the first acquisition unit 20, the reactions to the posted texts, and the replies to the posted texts. The linguistic features represent topics that the user is good at, topics that the user is interested in, the user's expertise, etc.

[0033] The first derivation unit 22 derives the linguistic features of the target user based on the frequency of occurrence of each word in the posted text input by the target user and the frequency of occurrence of each word in the posted texts of other users that are the targets of the reactions or replies input by the target user. The frequency of occurrence of each word in the reply input by the target user may be used. That is, the first derivation unit 22 derives the linguistic features of the target user based on the posted text input by the target user and the reactions and replies, which are information regarding the reactions of the target user to the posted texts of other users.

[0034] FIG. 3 shows an example of SNS information. FIG. 3 shows an example of a screen on which a posted text 60 posted by user A, a reply 62 to the posted text 60 by user B, and reactions 64 to the post by user A by a plurality of users are displayed. The account image 66 of user A and the account image 68 of user B are also displayed.

[0035] Figure 4 shows an example of the linguistic features of a plurality of users u1 to um (m is an integer of 2 or more). w1 to wn (n is an integer of 2 or more) represent words. Each element of the matrix represents the frequency of occurrence of a word in a posted text. One row is a vector representing the linguistic features of a certain user. In the illustrated example, in the vector of the linguistic features of user u1, the frequency of occurrence of word w1 is "0", and the frequency of occurrence of word w2 is "10". Since the vocabulary of words is very large, the linguistic features can be, for example, tens of thousands of dimensions, and the matrix also becomes sparse. Therefore, dimensionality reduction may be performed using a low-rank approximation method (NMF) or the like. The dimensionality-reduced linguistic features of each user may be, for example, 200 dimensions.

[0036] The first derivation unit 22 also derives the image feature amount of each user based on an image for identifying each user in the SNS. The image for identifying a user is an image of the user's account or a profile image. The first derivation unit 22 derives the image feature amount based on the RGB values of a plurality of pixels such as an account image. When there is a model for estimating personality with an account image or the like as an input, the output value of the model may be used as the image feature amount. The image feature amount is, for example, one-dimensional and represents personal characteristics and internal properties such as the personality of each user. For example, the larger the numerical value of the image feature amount, the more extroverted the personality may be. By using the image feature amount, the personality of the user can also be reflected in the model, and the prediction accuracy can be improved.

[0037] The first derivation unit 22 derives a relational feature amount representing the relationship between the target user and other users, and information representing the number of speaker alternations from the target user to other users, based on the reaction and reply information of the target user to the posted texts of other users input by the target user included in the acquired information.

[0038] The relational feature amount can also be called a network feature amount and represents the relationship between users. Regarding each user as a node, when there is a reaction or a reply, an edge is drawn between the original poster and the user who made the reaction or the reply. If there are multiple reactions or the like, the weight of the edge is added.

[0039] The value of the relational feature quantity may be the difference between the average of the whole network such as the shortest path length, the assortativity, the degree, etc. and each value between certain specific persons. The relational feature quantity includes, for example, a value based on the shortest path length, a value based on the assortativity, and a value based on the degree, and is three-dimensional.

[0040] The number of speaker alternations from the target user to other users is derived to be larger as the number of reactions, etc. to the posted texts of other users input by the target user is larger, based on a predetermined relational expression. The predetermined relational expression can be appropriately determined by experiments so as to approximate the relationship between the number of reactions, etc. to the posted texts of other users input by the target user and the number of speaker alternations from the target user to other users in an actual conversation. The number of speaker alternations from the target user A to other user B and the number of speaker alternations from other user B to the target user A are distinguished.

[0041] The first derivation unit 22 concatenates the linguistic feature quantity and the image feature quantity for each user to create a feature quantity vector unique to the user, creates a feature quantity dictionary including this feature quantity vector, and stores the created dictionary in the feature quantity dictionary storage unit 16. The feature quantity vector unique to one user is, for example, 200 + 1 = 201 - dimensional.

[0042] The first derivation unit 22 also stores in the feature quantity dictionary storage unit 16 the relational feature quantity of each combination of two users among a plurality of users and the information representing the number of speaker alternations. That is, it can be said that the feature quantity dictionary storage unit 16 stores the first feature quantity vectors of each user.

[0043] Based on the information acquired by the first acquisition unit 20, the learning unit 24 performs machine learning on a model for predicting the next speaker from among a plurality of conversation participants, and stores the learned model in the learned model storage unit 18. Specifically, the learning unit 24 acquires from the feature dictionary storage unit 16 the first feature vector of the target user, the first feature vector of other users, and information representing the number of speaker alternations from the target user to other users, and learns the model based on these. The first feature vector of the target user includes the linguistic feature amount and image feature amount of the target user, and the relational feature amount between the target user and other users, and is, for example, 204-dimensional. The first feature vector of other users includes the linguistic feature amount and image feature amount of other users, and the relational feature amount between other users and the target user, and is, for example, 204-dimensional. The relational feature amount between the target user and other users may be the same as the relational feature amount between other users and the target user.

[0044] More specifically, the learning unit 24 takes as input the linguistic feature amount and image feature amount of the target user, the relational feature amount between the target user and other users, the linguistic feature amount and image feature amount of other users, and the relational feature amount between other users and the target user, and learns the model using the information representing the number of speaker alternations from the target user to other users as the correct label. The learning unit 24 repeats learning the model for every two users out of a plurality of users. For two users, learning can be performed twice, once with the information representing the number of speaker alternations from the target user to other users as the correct label, and once with the information representing the number of speaker alternations from other users to the target user as the correct label. By learning for every two users, the number of learning samples can be increased, and a model robust to the number of conversation participants can be created. For example, in the case of 100 users, 100×100 - 100 = 9900 samples can be secured.

[0045] The model is, for example, a multi-layer neural network model. The model may be a regression model or a classification model. In the case of a regression model, the learning unit 24 learns the numerical value of the number of speaker alternations as the correct label. For example, when targeting a conversation from user A to user B, the 204-dimensional feature amount of user A and the 204-dimensional feature amount of user B are input, and "54", which is the number of speaker alternations from user A to user B derived from SNS data, becomes the correct label.

[0046] On the other hand, for example, when targeting a conversation from user B to user A, the 204-dimensional feature amount of user A and the 204-dimensional feature amount of user B are input, and "30", which is the number of speaker alternations from user B to user A derived from SNS data, becomes the correct label.

[0047] In order to distinguish and learn the number of speaker alternations from user A to user B and the number of speaker alternations from user B to user A, a vector for specifying the order of the first feature vectors of the two users may also be supplied to the learning unit 24.

[0048] In the case of a classification model, the learning unit 24 learns the quantized value of the number of speaker alternations as the correct label. The quantized value of the number of speaker alternations is information representing the number of speaker alternations.

[0049] FIG. 5 shows an example of the relationship between the number of speaker alternations and the frequency. When the number of speaker alternations is less than the threshold th1, the number of speaker alternations is defined as "small". When the number of speaker alternations is greater than or equal to the threshold th1 and less than the threshold th2, the number of speaker alternations is defined as "medium". When the number of speaker alternations is greater than or equal to the threshold th2, the number of speaker alternations is defined as "large". The thresholds th1 and th2 are set in advance by a person.

[0050] For example, when targeting a conversation from user A to user B, the 204-dimensional feature amount of user A and the 204-dimensional feature amount of user B are input, and "large", which is the number of speaker alternations from user A to user B derived from SNS data, becomes the correct label.

[0051] When targeting the conversation from user B to user A, the 204-dimensional feature vectors of user A and the 204-dimensional feature vectors of user B are input, and the "medium", which is the number of speaker alternations from user B to user A derived from SNS data, becomes the correct label.

[0052] By the way, it is possible that the actual conversation participants do not use SNS and the data of those conversation participants is not included in the SNS data. In this case, the following processing is executed.

[0053] Figure 6 schematically shows the distribution of the unique feature vectors of multiple users. The dimension of the unique feature vector is, for example, 201 dimensions as described above, but in Figure 6, it is simplified to 2 dimensions.

[0054] The first derivation unit 22 classifies the unique feature vectors of multiple users into multiple clusters by unsupervised learning using k-means or the like, and derives the unique feature vectors of virtual representative users representing each cluster. The first derivation unit 22 derives, for each cluster, the unique feature vector of a representative user representing the centroid of the multiple feature vectors of that cluster. In the example of Figure 6, it is classified into five clusters C1 to C5, and five unique feature vectors V1 to V5 of five virtual representative users are derived. That is, when there are 100 users, feature amounts for 105 people are created.

[0055] When the first derivation unit 22 also targets users for whom there is no information input to the SNS for learning, it sets the unique feature vector of a representative user of any cluster to the unique feature vector of that user and stores it in the feature dictionary storage unit 16. For example, an administrator of the system or the like may check whether a user for whom there is no SNS data is extroverted or introverted, etc., and select the unique feature vector of a representative user of the cluster to which a user with a personality relatively similar to that of that user belongs.

[0056] The first derivation unit 22 derives the relational feature amounts and correct labels of users who do not have information input to the SNS based on the users classified into each cluster, and stores them in the feature amount dictionary storage unit 16.

[0057] For example, when it is known that the number of speaker alternations from user A to user B is "54", it is regarded that the number of speaker alternations from cluster C1 in which user A is classified to cluster C3 in which user B is classified is "54". Also, the relational feature amount between cluster C1 in which user A is classified and cluster C3 in which user B is classified is regarded as the relational feature amount between user A and user B.

[0058] That is, when a user (referred to as user X) who does not have information input to the SNS is classified into cluster C1 and the conversation from user X to user B is targeted, the unique feature amount vector V1 of user X, the relational feature amount between user A and user B, the unique feature amount vector 70 of user B, and the relational feature amount between user B and user A are input, and "54", which is the number of speaker alternations from user A to user B, becomes the correct label.

[0059] In this way, even when the actual conversation participants do not use the SNS, learning can be performed using the feature amount vector that may represent the characteristics of the conversation participants.

[0060] Also, when there are missing values in the SNS data for some reason, the first derivation unit 22 may complement the missing values using a learned missing value estimation model, and derive the feature amount vector of each user based on the SNS data for which the complementation of the missing values has been completed. For the complementation of the missing values, for example, the technology described in Japanese Patent Application No. 2021-165372 previously filed by the applicant can be used.

[0061] Next, after learning, the process regarding the prediction executed during the conversation by the conversation participants will be described.

[0062] During the conversation among a plurality of conversation participants, the second acquisition unit 26 acquires the audio data supplied from the microphone 2, the image data supplied from the camera 4, and the sensor data supplied from the sensor 6, supplies the audio data to the speech recognition unit 28, and supplies the image data and the sensor data to the second derivation unit 30. This process corresponds to the second acquisition unit 26 acquiring information regarding the communication between each conversation participant and other conversation participants.

[0063] The speech recognition unit 28 performs speech recognition on the audio data acquired by the second acquisition unit 26, executes preprocessing on the speech recognition result as necessary, creates text data, and supplies the text data to the second derivation unit 30. The preprocessing includes, for example, processing to unify full-width characters into half-width characters, anonymize the proper nouns of people, and perform word segmentation.

[0064] The second derivation unit 30 derives language feature amounts for each conversation participant based on the text data. The second derivation unit 30 derives the language feature amount of the target conversation participant from the frequency of occurrence of each word in the text data representing the conversation content of the target conversation participant. This language feature amount is also a vector with the same structure and the same dimension as the language feature amount derived by the first derivation unit 22 shown in FIG. 4.

[0065] The second derivation unit 30 derives image feature amounts for each conversation participant based on the sensor data. The posture of the conversation participant during the conversation identified by the sensor data is considered to represent the personality of the conversation participant. Therefore, the second derivation unit 30 derives the image feature amount of the target conversation participant from the posture derived from the sensor data of the target conversation participant. This image feature amount is also a one-dimensional vector same as the image feature amount derived by the first derivation unit 22.

[0066] The second derivation unit 30 derives relational feature amounts for each pair of two conversation participants based on the image data. The second derivation unit 30, with each conversation participant as a node, increases the weight of the edge between two conversation participants when it is specified by the image data that they have looked at each other or nodded. Similar to the processing by the first derivation unit 22, the second derivation unit 30 derives three-dimensional relational feature amounts based on the nodes and edges of a plurality of conversation participants. That is, the second derivation unit 30 derives the relational feature amount between a target conversation participant and other conversation participants based on information regarding the actions taken by the target conversation participant with respect to other conversation participants.

[0067] These processes correspond to the second derivation unit 30 deriving the second feature amount vectors of each of the plurality of conversation participants based on the acquired information regarding the communication. The second feature amount vectors include linguistic feature amounts, image feature amounts, and relational feature amounts. The number of dimensions of the second feature amount vectors is the same as the number of dimensions of the first feature amount vectors.

[0068] The prediction unit 32 predicts the next speaker from among a plurality of conversation participants based on the acquired information regarding the communication, using the learned model stored in the learned model storage unit 18 that has been machine-learned based on the information input by a plurality of users on the SNS. The prediction unit 32 predicts the conversation amount per unit time of each of the other conversation participants following the speech of each conversation participant.

[0069] The prediction unit 32 acquires the first feature amount vectors of each of the plurality of conversation participants from the feature amount dictionary storage unit 16 that stores the first feature amount vectors of each user derived from the information input by a plurality of users on the SNS. Even when there are conversation participants who do not use the SNS, as described above, the first feature amount vectors of those conversation participants are also stored in the feature amount dictionary storage unit 16.

[0070] The prediction unit 32 predicts the next speaker based on the acquired plurality of first feature vectors and the derived plurality of second feature vectors using the learned model. The prediction unit 32 mixes the acquired first feature vectors and the derived second feature vectors for each conversation participant. Specifically, for each conversation participant, the prediction unit 32 derives a third feature vector by weighted addition of corresponding features in the acquired first feature vector and the derived second feature vector, inputs the derived third feature vector into the learned model, and predicts the next speaker.

[0071] The weights may be different for each of the linguistic feature, the image feature, and the relational feature. The weight for a specific feature of the first feature vector may be zero, in which case the specific feature of the first feature vector is replaced with the feature of the second feature vector. The weights may be obtained separately by machine learning.

[0072] More specifically, the prediction unit 32 inputs the third feature vector of the target conversation participant and the third feature vectors of other conversation participants into the learned model, and outputs the number of speaker alternations per unit time from the target conversation participant to other conversation participants. The third feature vector of the target conversation participant includes the linguistic feature and the image feature of the target conversation participant and the relational feature between the target conversation participant and other conversation participants. The third feature vector of other conversation participants includes the linguistic feature and the image feature of other conversation participants and the relational feature between other conversation participants and the target conversation participant. The prediction unit 32 repeats the above prediction process for every two of the plurality of conversation participants. For two conversation participants, prediction can be made twice. A vector for specifying the order of the third feature vectors of the two conversation participants may also be supplied to the prediction unit 32.

[0073] The output unit 34 outputs the result predicted by the prediction unit 32. The output unit 34 may output the matrix of FIG. 1 output from the model learned by the prediction unit 32.

[0074] By inputting both the information entered into the SNS and the information acquired during the conversation into the learned model, information about the physical space during the conversation can also be reflected, enabling more accurate prediction.

[0075] FIG. 7 is a flowchart showing the learning process of the communication prediction system 1 in FIG. 2. The first acquisition unit 20 acquires SNS data of a plurality of users (S10), and creates learning data from the SNS data (S12). The first derivation unit 22 creates a feature dictionary from the learning data and stores it in the feature dictionary storage unit 16 (S14). The learning unit 24 performs machine learning on the prediction model (S16), stores the learned model (S18), and ends the process.

[0076] FIG. 8 is a flowchart showing the prediction process of the communication prediction system 1 in FIG. 2. This process is executed after the learning is completed. It is determined whether a conversation has started (S30). If not (N in S30), the process returns to S30.

[0077] If a conversation has started (Y in S30), the second acquisition unit 26 acquires data from the microphone 2, the camera 4, and the sensor 6 (S32), and the speech recognition unit 28 performs speech recognition on the speech data and creates text data (S34).

[0078] The second derivation unit 30 derives features for each conversation participant based on the text data, image data, and sensor data (S36). The second derivation unit 30 mixes the features acquired from the feature dictionary and the derived features for each conversation participant to create a feature vector (S38). The prediction unit 32 inputs the created feature vector into the learned model and makes a prediction (S40). If it is during the conversation (Y in S42), the process returns to S32. If it is not during the conversation (N in S42), the process ends.

[0079] Next, a specific example of the process by the communication prediction system 1 will be described. [Learning] Assume that the frequency of occurrence of the word "camp" in the SNS post of user A is 100, and the frequencies of occurrence of "fishing" and "politics" are 0.

[0080] Assume that the frequency of appearance of "fishing" in the SNS post of User B is 200, and the frequencies of appearance of "camping" and "politics" are 0.

[0081] Assume that the frequencies of appearance of "camping" and "fishing" in the SNS post of User C are 0, and the frequency of appearance of "politics" is 50. Hereinafter, User A, User B, and User C are also simply represented as A, B, and C.

[0082] Assume that the relational feature amount indicating the relationship between A and B is 10, and the relational feature amounts indicating the relationships between A and C, and between B and C are 0.

[0083] Assume that the number of speaker alternations from A to B is 50, the number of speaker alternations from B to A is 40, and the numbers of speaker alternations from A to C, from C to A, from B to C, and from C to B are 0.

[0084] The model is trained using the SNS data of 100 users including A, B, and C.

[0085] [Prediction] During the conversation among the three users A, B, and C, a matrix is output from the trained model at the current time t1.

[0086] Assume that the frequency of appearance of "camping" in the speech content of A until time t1 is 10, and the frequencies of appearance of "fishing" and "politics" are 1.

[0087] Assume that the frequency of appearance of "camping" in the speech content of B is 4, and the frequencies of appearance of "fishing" and "politics" are 1.

[0088] Assume that the frequency of appearance of "camping" in the speech content of C is 1, and the frequencies of appearance of "fishing" and "politics" are 0.

[0089] Assume that the relational feature amount (for example, the frequency of eye contact, the same hereinafter) between A and B obtained by Camera 4 is 8, the relational feature amount between A and C is 1, and the relational feature amount between B and C is 0.

[0090] In the output matrix, assume that when the current speaker is A, the number of speaker switches to the next speaker B is 40, and the number of speaker switches to the next speaker C is 1.

[0091] Assume that when the current speaker is B, the number of speaker switches to the next speaker A is 30, and the number of speaker switches to the next speaker C is 0.

[0092] Assume that when the current speaker is C, the number of speaker switches to the next speaker A is 30, and the number of speaker switches to the next speaker B is 20.

[0093] Actually, if the speaker is A at time t1, the next speaker is B, and the conversation volume of B speaking after A is predicted to be "40".

[0094] Next, after time t1, the matrix is output at time t2 and the matrix is output at time t3. The description of the process at time t2 is omitted.

[0095] Assume that from time t2 to t3, the appearance frequencies of "camp", "fishing", and "politics" in A's speech content are 0.

[0096] Assume that the appearance frequencies of "camp", "fishing", and "politics" in B's speech content are 0, the appearance frequencies of "camp" and "fishing" in C's speech content are 0, and the appearance frequency of "politics" is 10.

[0097] Assume that the relational feature quantity between A and B obtained by the sensor is 0, the relational feature quantity between A and C is 0, and the relational feature quantity between B and C is 5.

[0098] In the output matrix, assume that when the current speaker is A, the number of speaker switches to the next speaker B is 10, and the number of speaker switches to the next speaker C is 20.

[0099] Assume that when the current speaker is B, the number of speaker switches to the next speaker A is 10, and the number of speaker switches to the next speaker C is 20.

[0100] Assume that when the current speaker is C, the number of speaker alternations to the next speaker A = 0, and the number of speaker alternations to the next speaker B = 20.

[0101] Actually, when the speaker was C at time t3, the next speaker is B, and the conversation volume of B speaking after C is predicted to be "20".

[0102] (Second Embodiment) In the second embodiment, it is different from the first embodiment that the feature amounts obtained from the SNS and the feature amounts obtained during the conversation for each user are treated in parallel as separate feature amounts. Hereinafter, the description will focus on the differences from the first embodiment.

[0103] In this embodiment, additional learning is performed on the learned model learned from the SNS data according to the processing of the first embodiment. A conversation experiment is performed in advance by a plurality of conversation participants to be predicted, and the model is additionally learned using the data obtained in the conversation experiment.

[0104] First, the learning will be described. The second derivation unit 30 derives the second feature amount vectors of each of the plurality of conversation participants based on the information on the communication obtained in the conversation experiment. The second feature amount vector can include feature amounts different from the first feature amount vector, and may have a different number of dimensions from the first feature amount vector. Therefore, the degree of freedom in designing the feature amounts can be improved.

[0105] For example, the second feature amount vector may include linguistic feature amounts, prosodic feature amounts, and feature amounts obtained from image data and sensor data. The linguistic feature amount may have the same structure as in the first embodiment, for example, 200 dimensions.

[0106] The second derivation unit 30 derives prosodic feature amounts for each conversation participant based on the voice data. The prosodic feature amount is, for example, 2-dimensional and may include a feature amount related to the pitch of the voice and a feature amount related to the power of the voice.

[0107] The feature amounts obtained from the image data and the sensor data may be two-dimensional, for example, and may include a feature amount related to the line of sight and a feature amount related to the posture.

[0108] The second derivation unit 30 derives a feature amount related to the line of sight for each pair of two conversation participants based on the image data. For example, the second derivation unit 30 may derive a larger feature amount related to the line of sight between the two conversation participants as the frequency of directing the line of sight between the two conversation participants specified by the image data is higher.

[0109] The second derivation unit 30 derives a feature amount related to the posture for each conversation participant based on the sensor data.

[0110] The learning unit 24 acquires the first feature amount vectors of the respective multiple conversation participants from the feature amount dictionary storage unit 16, and concatenates the acquired first feature amount vectors and the derived second feature amount vectors for each conversation participant to derive a third feature amount vector. The third feature amount vector is, for example, 204 + 204 = 408 dimensions.

[0111] The learning unit 24 inputs the third feature amount vector of the target conversation participant and the third feature amount vectors of the other conversation participants, and learns the model using the information representing the number of speaker alternations from the target conversation participant to the other conversation participants as the correct label. The information representing the number of speaker alternations may be the same as in the first embodiment or may be derived from the results of the conversation experiment.

[0112] Next, prediction will be described. During a conversation by a plurality of conversation participants performed after learning, the second derivation unit 30 derives the second feature amount vectors of the respective multiple conversation participants based on the acquired information regarding the communication.

[0113] The prediction unit 32 acquires, from the feature dictionary storage unit 16, the first feature vectors of the plurality of conversation participants, and for each conversation participant, concatenates the acquired first feature vector and the derived second feature vector to derive a third feature vector. The prediction unit 32 inputs the derived third feature vector into a learned model to predict the next speaker. Specifically, the prediction unit 32 inputs the third feature vector of the target conversation participant and the third feature vectors of other conversation participants into the learned model, and outputs the number of speaker alternations per unit time from the target conversation participant to other conversation participants.

[0114] Also in this embodiment, the effects of the first embodiment can be obtained.

[0115] Note that since the data that can be secured in an actual application scenario may be limited, the input feature amount may be created as follows.

[0116] For example, when there is a user who can acquire SNS data but cannot acquire the data of the conversation experiment, for the second feature vector of that user, sampling may be performed from a uniform distribution or the like to create dummy data. Alternatively, based on the features of the SNS data, similar users with conversation data may be extracted and the features of the similar users may be substituted. Similar users can be extracted by calculating the inner product between users using the first feature vector based on the SNS data.

[0117] Also, when there is a user who cannot acquire SNS data but can acquire the data of the conversation experiment, the process of acquiring the feature vector of the representative user of the cluster in the first embodiment can be applied. When there is a user who can acquire neither SNS data nor the data of the conversation experiment, the above two processes can be used in combination.

[0118] The present invention has been described based on the embodiments. It should be understood by those skilled in the art that the embodiments are merely examples, and various modifications are possible for the combination of each component and each processing process, and such modifications are also within the scope of the present invention.

[0119] For example, the communication prediction device 10 may be configured not to execute learning. In this case, the communication prediction device 10 does not necessarily include the first acquisition unit 20, the first derivation unit 22, the learning unit 24, and the SNS data storage unit 14. The feature dictionary and the learned model are acquired from the outside. In this case, a machine learning device that only executes machine learning may be configured as a device separate from the communication prediction device 10. The machine learning device that executes learning in the first embodiment includes the first acquisition unit 20, the first derivation unit 22, the learning unit 24, the SNS data storage unit 14, the feature dictionary storage unit 16, and the learned model storage unit 18. The machine learning device that executes learning in the second embodiment includes the first acquisition unit 20, the first derivation unit 22, the learning unit 24, the second acquisition unit 26, the speech recognition unit 28, the second derivation unit 30, the SNS data storage unit 14, the feature dictionary storage unit 16, and the learned model storage unit 18.

Explanation of Signs

[0120] 1... Communication prediction system, 10... Communication prediction device, 12... Processing unit, 14... SNS data storage unit, 16... Feature dictionary storage unit, 18... Learned model storage unit, 20... First acquisition unit, 22... First derivation unit, 24... Learning unit, 26... Second acquisition unit, 28... Speech recognition unit, 30... Second derivation unit, 32... Prediction unit, 34... Output unit.

Claims

1. An acquisition unit that acquires information input by a plurality of users in a social networking service; A learning unit that machine-learns a model for predicting the next speaker from among a plurality of conversation participants based on the acquired information; Based on the posted texts input by a plurality of users included in the acquired information, a plurality of language feature amounts of the plurality of users are derived, and based on the information regarding the reaction of the target user to the posted texts of other users input by the target user included in the acquired information, a relational feature amount representing the relationship between the target user and other users, and information representing the number of speaker alternations from the target user to other users are derived; a derivation unit; Comprising: The learning unit inputs the language feature amount of the target user, the language feature amount of other users, and the relational feature amount between the target user and other users, and learns the model using the information representing the number of speaker alternations from the target user to other users as the correct label. A communication prediction device characterized by the above.

2. The derivation unit derives respective image feature amounts of the target user and other users based on an image for identifying the target user included in the acquired information and an image for identifying other users; The learning unit also inputs the respective image feature amounts of the target user and other users and learns the model. The communication prediction device according to claim 1, characterized by the above.

3. The derivation unit: Classifies the language feature amounts of a plurality of users into clusters; Derives the language feature amounts of virtual representative users representing each cluster; When a user who has no information input to the social networking service is also a learning target, sets the language feature amount of a representative user of any one of the clusters as the language feature amount of the user. The communication prediction device according to claim 1, characterized by the above.

4. An acquisition unit that acquires information regarding the communication between each conversation participant and other conversation participants during a conversation by a plurality of conversation participants; A prediction unit that predicts the next speaker from among the plurality of conversation participants based on the acquired communication-related information using a learned model machine-learned based on information input by a plurality of users in a social networking service; A derivation unit that derives respective second feature amount vectors of the plurality of conversation participants based on the acquired communication-related information; Comprising: The prediction unit: Obtain the first feature vector of each of the plurality of conversation participants from a storage unit that stores the first feature vector of each user derived from information input by a plurality of users in a social networking service. Predict the next speaker based on the plurality of obtained first feature vectors and the plurality of derived second feature vectors using the learned model. A communication prediction apparatus characterized by the above.

5. The prediction unit: For each conversation participant, weight and add corresponding feature amounts in the obtained first feature vector and the derived second feature vector to derive a third feature vector. Input the derived third feature vector into the learned model to predict the next speaker. The communication prediction apparatus according to claim 4, characterized by the above.

6. The prediction unit: For each conversation participant, concatenate the obtained first feature vector and the derived second feature vector to derive a third feature vector. Input the derived third feature vector into the learned model to predict the next speaker. The communication prediction apparatus according to claim 4, characterized by the above.

7. A communication prediction method executed by a computer, comprising: Obtaining information input by a plurality of users in a social networking service; Machine learning a model for predicting the next speaker from among a plurality of conversation participants based on the obtained information; Deriving linguistic feature amounts of a plurality of users based on posted texts input by the plurality of users included in the obtained information, and based on information regarding the reaction of the target user to the posted texts of other users input by the target user included in the obtained information, deriving a relational feature amount representing the relationship between the target user and other users, and information representing the number of speaker alternations from the target user to other users; Comprising: In the step of machine learning, using the linguistic feature amount of the target user, the linguistic feature amount of other users, and the relational feature amount between the target user and other users as inputs, and using the information representing the number of speaker alternations from the target user to other users as the correct label to learn the model. A communication prediction method characterized by the above.

8. A communication prediction method executed by a computer, comprising: During a conversation among a plurality of conversation participants, obtaining information regarding the communication between each conversation participant and other conversation participants; Using a learned model that has been machine-learned based on information input by a plurality of users in a social networking service, predicting the next speaker from among the plurality of conversation participants based on the obtained information regarding the communication; Deriving, based on the obtained information regarding the communication, a second feature vector for each of the plurality of conversation participants; comprising; In the step of predicting, Obtaining, from a storage unit that stores a first feature vector of each user derived from information input by a plurality of users in a social networking service, a first feature vector of each of the plurality of conversation participants; Predicting the next speaker using the learned model based on the plurality of obtained first feature vectors and the plurality of derived second feature vectors; A communication prediction method characterized by the above.

Citation Information

Patent Citations

  • Conversation support system, conversation support device, and conversation support program

    JP2017123027A

  • Speaker determination system, speaker determination method and speaker determination program

    JP2018146844A

  • Speech analyzer, speech analyzing method, speech analyzing program, and speech analyzing system

    JP2020035467A

  • System and method

    JP2021045568A

  • Communication skill evaluation system, device, method, and program

    WO2019093392A1