Device and method
By employing a two-tiered interest estimation and a knowledge graph to enhance topic candidates, the system addresses monotonous conversations by providing more diverse and personalized responses in user interactions.
Patent Information
- Application Number
- PCT/JP2024/000481
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-17
AI Technical Summary
Existing systems for estimating user interest in casual conversations often result in monotonous interactions due to insufficient diversity in topic candidates, limiting the variety of responses.
An apparatus and method that estimate user interest using a first estimation unit for initial interest information and a second estimation unit for relationship-based interest information, incorporating a knowledge graph to expand topic candidates, thereby enhancing conversation diversity.
The approach allows for more diverse and relevant responses by increasing the number of topic candidates, making conversations with users more engaging and personalized.
Smart Images

Figure JP2024000481_17072025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] The present invention relates to an apparatus and a method.
[0002] Patent Literature 1 describes a device that can obtain information indicating a user's interests based on information indicating words included in the user's utterance. Specifically, the device estimates the user's level of interest in each of a plurality of predetermined topics based on the user's utterance. In this way, the device of Patent Literature 1 estimates the user's topics of interest from the user's utterances during casual conversation.
[0003] Japanese Patent Application Laid-Open No. 2020-53015
[0004] However, in the device described in Patent Document 1, a topic of interest to the user is estimated from a plurality of predetermined topics. In this case, the number of topics estimated to be of interest to the user is insufficient, and therefore, if a response to a user's utterance is made based on the topic of interest to the user, the conversation with the user becomes monotonous.
[0005] In view of the above, an object of the present invention is to provide an apparatus and method that can provide a wider variety of replies to a user when replying to a user's utterance based on a topic of interest to the user.
[0006] In order to achieve the above object, the device of the present invention includes a first estimation unit that estimates first interest information that indicates a user's interest based on information related to the content of the user's utterance, a second estimation unit that estimates second interest information for estimating topics that indicate the user's interest based on the first interest information and relationship information that indicates the relationship between a plurality of words, and a topic estimation unit that estimates topic candidate information that indicates topic candidates for the user based on the second interest information.
[0007] In the device and method according to the present invention, second interest information is estimated based on first interest information and relationship information, and topic candidate information indicating topic candidates for the user is estimated based on the second interest information. Here, if topic candidate information is estimated based only on the first interest information, the number of topics included in the topic candidate information is insufficient. Therefore, if a response to a user's utterance is based on the topic candidate information, the conversation with the user becomes monotonous. On the other hand, with the above configuration, the topic candidate information estimated based on the second interest information includes more topics than the topic candidate information estimated based only on the first interest information, resulting in more diverse replies based on the topic candidate information. This allows for more diverse replies to be returned to the user when a response to a user's utterance is based on topics of the user's interest.
[0008] The device and method of the present disclosure make it possible to return a wider variety of responses to user utterances.
[0009] 7(a) and 7(b) are diagrams for explaining an overview of an apparatus and a method according to an embodiment of the present invention.
[0033] FIG. 7(a) is a block diagram showing the functional configuration of an apparatus according to an embodiment of the present invention.
[0034] FIG. 7(b) is a block diagram showing the functional configuration of a first estimation unit of the apparatus shown in FIG. 2.
[0035] FIG. 7(a) is a diagram showing an example of a vector indicating the meaning of a user's characteristics or words, and FIG. 7(b) is a diagram showing an example of value information.
[0036] FIG. 7(a) is a diagram showing an example of an embedded representation of a user and an embedded representation of multiple feature words, and FIG. 7(b) is a diagram showing an example of information indicating whether or not there is a relationship between the user and each feature word, and FIG. 7(c) is a diagram showing an example of a knowledge graph, and FIG. 7(d) is a diagram showing an example of multiple pieces of second semantic information.
[0037] FIGS. 8(a) and 8(b) are diagrams for explaining a knowledge graph, and FIG. 9 is a diagram showing an example of second interest information.
[0038] FIG. 10 is a diagram for explaining a process for generating second interest information, and FIG. 11 is a flowchart showing an example of a process for generating utterances for a user, and FIG. 12 is a flowchart showing an example of a process for estimating first interest information.
[0039] FIG. 7(b) is a diagram showing the hardware configuration of a terminal and a server included in an apparatus according to an embodiment of the present invention.
[0010] Hereinafter, an embodiment of the device and method according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same elements are designated by the same reference numerals, and duplicated explanations will be omitted.
[0011] FIG. 1 is a diagram illustrating an overview of a device and a method according to the present disclosure. The device and method of the present disclosure are a topic estimation device and a topic estimation method for estimating a topic for a user. Specifically, in the device and method of the present disclosure, first, information related to the content of a user's utterance is acquired from a server, a terminal, or the like. Next, information indicating a relationship between the user and a topic related to the content of the utterance is estimated based on the content of the user's utterance. Next, information indicating a topic candidate for the user is estimated based on the information indicating the relationship. Finally, a response to the user's utterance is generated based on the information indicating the topic candidate, and the generated response is output to a server, a terminal, or the like. In this way, the device of the present disclosure enables, for example, a character in a virtual space managed by a server to converse with the user in real time while providing the user with new topics.
[0012] In the example shown in FIG. 1 , first, the content of a conversation between user A and user B is acquired. Next, information indicating the relationship between user A or user B and topics related to the content of the conversation of user A or user B is generated. For example, based on the content of user A's conversation, "The Nadeshiko game was exciting!", a topic "Nadeshiko" that is directly related to the content of the conversation is generated. Based on the content of user B's conversation, "I finally bought a PS5," a topic "PS5" that is directly related to the content of the conversation is generated. Then, "soccer" and "Japan national team" are generated as topics indirectly related to the content of the conversation via the topic "Nadeshiko." A topic "game software F2" is generated as a topic indirectly related to the content of the conversation via the topic "PS5." A topic "soccer game W11" is generated as a topic indirectly related to the content of the conversation via the topics "Nadeshiko" and "PS5."
[0013] Next, a plurality of topics including "PS5" and "soccer game W11" are added to a list of chat topics for user A and user B. Finally, based on the list of chat topics, an utterance is generated with the content "Have you heard of the recently released soccer game W11 for PS5? It seems to be popular overseas!", and an NPC (Non Player Character) utters this content to user A and user B in the virtual space. In this way, the device 10 of the present disclosure allows the NPC to converse with user A and user B in real time while providing the new topic "soccer game W11" to user A and user B.
[0014] Next, the function of generating utterances will be described in detail. Fig. 2 is a block diagram showing the functional configuration of the device according to this embodiment. As shown in Fig. 2, the device 10 includes an input unit 1, a first estimation unit 2, a knowledge graph construction unit 3, a second estimation unit 4, a topic estimation unit 5, a dialogue history storage unit 6, an utterance generation unit 7, and a construction unit 8.
[0015] The input unit 1 inputs information related to the user's utterance content to the first estimation unit 2 and the dialogue history storage unit 6. Specifically, the input unit 1 acquires the user's utterance content in at least one of the real space and the virtual space. The utterance may be an utterance included in a conversation between users, an utterance included in a conversation between an NPC and the user in the virtual space, or an utterance spoken by the user alone. The information related to the user's utterance content includes, for example, information indicating the user's utterance content and information indicating the user's emotion in the utterance content. The information indicating the user's utterance content may be, for example, text indicating the user's utterance content or vector information indicating the character string. The information indicating the user's emotion may be, for example, information indicating a positive emotion or a negative emotion, or vector information indicating these emotions. FIG. 3 is a diagram illustrating an example of information related to the user's utterance content. The example illustrated in FIG. 3 shows that User A uttered the content "I watched a soccer game today" with a positive emotion. Furthermore, User B uttered the content "Tomorrow is Halloween" with a negative emotion.
[0016] The input unit 1 inputs information about a plurality of words to the knowledge graph construction unit 3. The information about a plurality of words is information indicating the meaning and usage of each of the plurality of words. For example, the input unit 1 acquires a sentence posted on a website as information about a plurality of words. As an example, the input unit 1 acquires a sentence posted on the Japanese version of Wikipedia as information about a plurality of words.
[0017] The input unit 1 may input information indicating topics of users, information indicating relationships between users, and information indicating locations visited by users to the first estimation unit 2. The information indicating topics of users is, for example, keywords indicating topics of conversation between users. The information indicating relationships between users is, for example, information indicating whether or not the users have had conversations with each other. The information indicating locations visited by the user may be, for example, places visited by the user in real space, worlds visited by the user or coordinates visited by the user in virtual space, or stores visited by the user in real space or virtual space.
[0018] The first estimation unit 2 estimates first interest information indicating the user's interests based on information about the user's utterance content input from the input unit 1. Fig. 4 is a block diagram showing the functional configuration of the first estimation unit 2. As shown in Fig. 4, the first estimation unit 2 includes a value estimation unit 21 and an interest estimation unit 22.
[0019] The value estimation unit 21 estimates value information indicating the user's values based on information related to the content of the user's utterance. The value estimation unit 21 has a language understanding unit 21a and a relationship understanding unit 21b. The value information is information indicating the user's values regarding multiple words. For example, the value information includes information indicating the user's level of interest in multiple words and information indicating whether or not there is a relationship between the user and each word. As an example, the information indicating the user's level of interest in multiple words is the user's embedded representation and the word's embedded representation, which will be described later.
[0020] The language understanding unit 21a inputs information about the user's utterance into the embedding model and obtains multiple embeddings, including the user's embedding and multiple word embeddings related to the user's utterance. The user's embedding may be a random real vector, or a real vector consisting of features that reflect some characteristic of the user. The word embedding is represented by a real vector. The word embedding is a real vector consisting of features that reflect the meaning of the word. In the device of this embodiment, the method for obtaining the user's embedding and the word embedding is not limited and may be any well-known method.
[0021] Specifically, the language understanding unit 21a inputs information indicating the user's utterance content and information indicating the user's emotions in the utterance content into an embedding model, and generates multiple embedded representations including user embedded representations indicating the user's characteristics and word embedded representations indicating the meanings of words related to the user's utterance content. The embedding model is generated in advance using machine learning. FIG. 5(a) is a diagram showing an example of user embedded representations and word embedded representations. In the example shown in FIG. 5(a), the process of generating the multiple embedded representations described above is realized using publicly known technology.
[0022] The relationship understanding unit 21b estimates value information indicating the user's values based on information about the user's utterance content input from the input unit 1 and the multiple embedded expressions generated by the language understanding unit 21a. Specifically, the relationship understanding unit 21b generates information indicating the presence or absence of a relationship between the user and each word based on the information about the user's utterance content. Similar to the second estimation unit 4 described below, the relationship understanding unit 21b generates the value information by solving a problem called Link Prediction.
[0023] FIG. 5B is a diagram showing an example of value information. In FIG. 5B, embedded expressions are shown as nodes, and relationships between embedded expressions are shown as edges. In the example shown in FIG. 5B, the relationship understanding unit 21b refers to information related to the content of user utterances. When a conversation takes place between user A and user B, the relationship understanding unit 21b creates an edge E between the node NA of user A and the node NB of user B. When the content of user A's utterance includes the words "Halloween" and "soccer," the relationship understanding unit 21b creates an edge E between the node NA of user A and the node NC of "Halloween" and the node ND of "soccer."
[0024] The relationship understanding unit 21b may acquire information indicating users' topics, information indicating relationships between users, or information indicating locations visited by users from the input unit 1, and may generate information indicating whether or not there is a relationship between the users and each word based on the acquired information. Information indicating users' topics may be converted into embedded representations of words and displayed as nodes. Information indicating locations visited by users may be converted into embedded representations of words and displayed as nodes. In the example shown in FIG. 5(b), the relationship understanding unit 21b may add an edge E between the node NA of user A and the node NB of user B based on the information indicating the relationships between users. When user A visits a location related to "Tomica," the relationship understanding unit 21b may generate a node NE called "Tomica" and add an edge E between the node NA of user A and the node NE called "Tomica." If user B visits a place related to "goldfish," the relationship understanding unit 21b may generate a node NF called "goldfish" and create an edge E between user B's node NB and the node NF called "goldfish."
[0025] In this way, the value estimation unit 21 is a value engine that understands the inner thoughts of humans by converting the user's utterances and emotions into real vectors, and estimates embedded representations of people, topics, and places. For example, the value estimation unit 21 acquires embedded representations of people, topics, and places using a graph neural network (GNN) based on information about language and information about relationships.
[0026] The interest estimation unit 22 estimates the first interest information based on the value information estimated by the value estimation unit 21. Specifically, the interest estimation unit 22 estimates the first interest information for each user based on information indicating the degree of interest of the user in a plurality of words and information indicating the presence or absence of a relationship between the user and each word.
[0027] The first interest information includes a plurality of first semantic information and first semantic relationship information. The plurality of first semantic information is information indicating the characteristics of the user and the meaning of at least one or more characteristic words related to the content of the user's utterance. The first semantic information is first vector information that represents the meaning of the user's characteristics or characteristic words as a real vector. For example, the first vector information is an embedded expression of the user or an embedded expression of a word. The first semantic relationship information is information indicating the relationship between the first semantic information. For example, the first semantic relationship information is information indicating the relationship between the embedded expressions. As an example, the first semantic relationship information is information indicating whether or not there is a relationship between the user and each word.
[0028] First, the interest estimation unit 22 extracts a plurality of feature words from a plurality of words related to the content of the utterance based on information indicating the user's degree of interest in the plurality of words estimated by the value estimation unit 21. As an example, the interest estimation unit 22 extracts a plurality of feature words based on the distance between the user's embedded expression and the word's embedded expression acquired by the language understanding unit 21 a.
[0029] FIG. 6 is a diagram illustrating the processing performed by the interest estimation unit 22. In FIG. 6, the user's embedded representation and the word's embedded representation are represented in the same space. As shown in FIG. 6, when extracting feature words for user A, the interest estimation unit 22 performs the following processing. Since the distance between user A's node NA and the "Halloween" node NC is less than a predetermined distance, the interest estimation unit 22 extracts "Halloween" as a feature word. Since the distance between user A's node NA and the "soccer" node ND is less than a predetermined distance, the interest estimation unit 22 extracts "soccer" as a feature word. Since the distance between user A's node NA and the "Tomica" node NE is greater than the predetermined distance, the interest estimation unit 22 does not extract "Tomica" as a feature word. In this case, the interest estimation unit 22 evaluates the distance between the embedded representations by calculating the inner product of the user's embedded representation and the word's embedded representation. The interest estimation unit 22 extracts feature words in descending order of similarity to the user's embedded expressions, since the larger the inner product, the greater the similarity between the embedded expressions (for example, the two words "soccer" and "Halloween" are extracted as feature words for user A). The interest estimation unit 22 may also evaluate the distance between embedded expressions using other known calculation methods.
[0030] Next, the interest estimation unit 22 acquires information indicating whether or not there is a relationship between the user and each word from the relationship understanding unit 21b and extracts information indicating whether or not there is a relationship between the user and each feature word. Finally, the interest estimation unit 22 defines the user's embedded expression and the embedded expressions of multiple feature words (multiple first semantic information), and information indicating whether or not there is a relationship between the user and each feature word (first semantic relationship information), as first interest information. FIG. 7 is a diagram showing an example of information input and output to the second estimation unit 4 described below. FIG. 7(a) is a diagram showing an example of a user's embedded expression and an example of embedded expressions of multiple feature words. In FIG. 7(a), a user's embedded expression indicating the characteristics of user A is shown as second semantic information indicating the user's characteristics. As second semantic information indicating the meaning of the feature word, embedded expressions of words indicating the meaning of "soccer" and "Halloween" are shown. FIG. 7(b) is a diagram showing an example of information indicating whether or not there is a relationship between the user and each feature word. In Figure 7 (b), the first semantic relationship information indicates that there is a relationship between User A's node and the ``Soccer'' node, and that there is a relationship between User B's node and the ``Halloween'' node.
[0031] In this way, the interest estimation unit 22 calculates the distance between embedded expressions based on the embedded expressions calculated by the value estimation unit 21, and acquires multiple topics (characteristic words) for each user.
[0032] The knowledge graph construction unit 3 generates relationship information indicating the relationships between multiple words based on information about the multiple words input by the input unit 1. The knowledge graph construction unit 3 stores the generated relationship information. Specifically, the knowledge graph construction unit 3 generates the relationship information based on information indicating the meaning and use of each of the multiple words. The relationship information is a knowledge graph including multiple entities each representing the multiple words and multiple relations each representing the relationships between the multiple entities. Knowledge graphs are used to share human knowledge with AI and the like. FIGS. 8( a) and (b) are diagrams for explaining knowledge graphs. As shown in FIG. 8( a), in a knowledge graph, entities represent things or concepts in the real world. In the knowledge graph, the relationship between a head entity h and a tail entity t is represented as a relation r.
[0033] In the example shown in Fig. 8(b), the knowledge graph construction unit 3 represents information indicating that cats are carnivores, that dogs are carnivores, that dogs eat meat, and that carnivores eat meat, using a head entity h, a tail entity t, and a relation r. Fig. 7(c) is a diagram showing an example of a knowledge graph constructed by the knowledge graph construction unit 3.
[0034] In this way, the knowledge graph construction unit 3 constructs graph data consisting of triples (head entity, relation and tail entity) from an external source.
[0035] The second estimation unit 4 estimates second interest information for estimating topics that the user is interested in, based on the first interest information and relationship information that indicates relationships between multiple words. For example, the second estimation unit 4 estimates second interest information for each user, based on the first interest information and the relationship information. Specifically, the second estimation unit 4 calculates the second interest information using an estimation model trained by machine learning, with the first interest information and the relationship information as input.
[0036] The second interest information includes a plurality of second semantic information and second semantic relationship information. The plurality of second semantic information is information indicating the user's characteristics, the meaning of a feature word, and the meanings of a plurality of words. The second semantic information is second vector information that represents the user's characteristics, the meaning of a feature word, or the meaning of one of the plurality of words as a real number vector. The second semantic relationship information is information indicating the relationship between the second semantic information.
[0037] More specifically, the second estimation unit 4 estimates multiple pieces of second semantic information for each user based on the embedded representation of the user and the embedded representations of multiple feature words ( FIG. 7( a) ), information indicating whether or not there is a relationship between the user and each feature word ( FIG. 7( b) ), and relationship information indicating the relationship between multiple words ( FIG. 7( c) ). FIG. 7( d ) is a diagram showing an example of multiple pieces of second semantic information corresponding to user A. In the example shown in FIG. 7( d ), an embedded representation indicating user A's node is shown as second semantic information indicating the feature of user A. As second semantic information indicating the meaning of feature words, embedded representations indicating the meanings of "soccer" and "Halloween" are shown. As second semantic information indicating the meaning of words, an embedded representation indicating the meaning of "FIFA World Cup" is shown.
[0038] By inputting relationship information into the estimation model in addition to the embedded representations and information indicating the presence or absence of a relationship, the number of pieces of second semantic information included in the plurality of pieces of second semantic information increases. In the example shown in Fig. 7, the embedded representations of the user indicating the characteristics of user A and the embedded representation of the word "soccer" (Fig. 7(a)), as well as information indicating that the word "soccer" has a relationship with user A (Fig. 7(b)), and relationship information indicating that the word "soccer" has a relationship with at least the word "FIFA World Cup" (Fig. 7(c)) are input into the estimation model. As a result, the plurality of pieces of second semantic information corresponding to user A includes the embedded representation of the word "FIFA World Cup" in addition to the embedded representations of the words "soccer" and "Halloween" (Fig. 7(d)).
[0039] Next, the second estimation unit 4 estimates second semantic relationship information based on the user's embedded representation and the embedded representations of multiple feature words, information indicating the presence or absence of a relationship between the user and each feature word, and relationship information indicating the relationship between multiple words. Figure 9 is a diagram showing an example of second interest information corresponding to user A. In the example shown in Figure 9, node N1 of user A, node N2 indicating "soccer," node N3 indicating "FIFA World Cup," and node N4 indicating "Halloween" are shown. Edges E12, E23, and E14, which indicate the presence or absence of a relationship between node N1 and node N2, node N2 and node N3, and node N1 and node N4, respectively, are shown as second semantic relationship information. Edges E12, E23, and E14 are attached according to the information indicating the presence or absence of a relationship between the user and each word and the relationship information indicating the relationship between multiple words.
[0040] For a word that is not considered to have a relationship with a user, the second estimation unit 4 may determine whether or not the word has a relationship with the user based on the distance between the embedded representation of the word and the embedded representation of the user (or may predict a link between nodes). In the example shown in Figure 9, the second estimation unit 4 may add an edge E13 based on the result of calculating the distance between node N3 representing "FIFA World Cup" and node N1 of user A.
[0041] The following describes the process for generating second information of interest in the second estimator 4. The estimation model generates a relationship graph, which is second information of interest, by solving a problem called Link Prediction. Note that the process for generating a relationship graph, which is value information, in the relationship understanding unit 21b of the first estimator 2 is also performed in a similar manner.
[0042] FIG. 10 is a diagram showing an example of a relationship graph and an example of extracting positive examples and negative examples from the relationship graph. In the example shown in FIG. 10, the relationship graph gn includes nodes n1 to n5 corresponding to users or words. The second estimation unit 4 randomly samples a node of interest. In the example shown in FIG. 10, it is assumed that node n2 is sampled as the node of interest.
[0043] The second estimation unit 4 extracts a positive example graph g1 and a negative example graph g2 from the relationship graph gn. The positive example graph g1 includes node n2, which is a node of interest, and nodes n1 and n5, which are connected to node n2 by edges. The negative example graph g2 includes node n2, which is a node of interest, and nodes n3 and n4, which are not connected to node n2 by edges. Note that the negative example graph g2 does not need to include all nodes that are not connected to the node of interest by edges. An example of learning the relationship graph gn will be described below, but since the learning process of a graph neural network is a well-known technique, a brief description will be given.
[0044] First, learning in the positive example graph g1 will be described. Based on the positive example graph g1, the second estimation unit 4 extracts an adjacency matrix Y in which the nodes included in the graph are represented as rows and columns, and the connection relationships via edges with the node n2, which is the node of interest, are represented as elements. The second estimation unit 4 also extracts a diagonal matrix I in which the nodes included in the graph are represented as rows and columns, and the self-loops of the nodes are represented as elements. Then, if the real number vector representing the feature amount of a node is the feature amount X of the node, the feature amount of each node is expressed as the sum (convolution) of the feature amount X of the node with which it is connected, represented by the adjacency matrix Y, and the feature amount of the node itself, represented by the diagonal matrix I, as shown in the following equation (1): (Y+I)·X (1)
[0045] The second estimation unit 4 multiplies the convolved feature (Y+I)·X of each node by a weight W, as expressed by the following formula, and then inputs the result into an activation function f to obtain an output H: H (positive example) = f ((Y+I)·X·W).The relationship learning unit 19 then learns the weight W and the feature X so that the output H (positive example) obtained based on the positive example graph g1 becomes 1.
[0046] The second estimation unit 4 similarly obtains an output H (negative example) based on the negative example graph g2. The second estimation unit 4 then learns the weights W and the feature quantities X so that the output H (negative example) obtained based on the negative example graph g2 becomes 0.
[0047] In this way, the second estimation unit 4 uses a GNN to link the graph data of the knowledge graph generated by the knowledge graph construction unit 3 with the characteristic words extracted for each user by the interest estimation unit 22, thereby re-acquiring multiple embedded expressions that respectively indicate the user's characteristics and the meanings of words.
[0048] The topic estimation unit 5 estimates topic candidate information indicating topic candidates for the user based on the second interest information estimated by the second estimation unit 4. For example, the topic candidate information may be a list indicating topic candidates for the user, or may be data such as a matrix indicating topic candidates for the user.
[0049] Specifically, when a word embedding is located within a predetermined processing range from the user's embedding and has a relationship with the user's embedding, the topic estimation unit 5 extracts the word as a topic candidate for the user. In this way, the topic estimation unit 5 calculates the distance between multiple embedded expressions and selects embedded expressions of words that are close to the user's embedded expression, thereby acquiring topics close to the user (topics in which the user is interested).
[0050] The dialogue history storage unit 6 stores information related to the content of user utterances input from the input unit 1. For example, the dialogue history storage unit 6 summarizes and stores older utterances from the user. The dialogue history storage unit 6 stores newer utterances from the user (the most recent utterances) as they are.
[0051] The utterance generation unit 7 generates an utterance for the user based on information about the user's utterance content and topic candidate information. Specifically, the utterance generation unit 7 generates information for generating an utterance based on information about the user's utterance content and topic candidate information. For example, the utterance generation unit 7 generates a prompt for causing the language model L to generate an utterance based on the information about the user's utterance content and the topic candidate information. The utterance generation unit 7 inputs the generated prompt into the language model L to generate an utterance for the user. The language model L is provided outside the device 10, but may also be provided inside the device 10. The language model L may be a known model, such as a large language model (LLM). Furthermore, the utterance generation unit 7 provides the language model with topics that match the user's interests and preferences in accordance with the user's utterance content, thereby allowing the language model to generate an utterance that matches the user's interests and preferences.
[0052] The construction unit 8 constructs an embedding model used in the first estimation unit 2 and an estimation model used in the second estimation unit by machine learning. The estimation model is trained using the first interest information and the related information as explanatory variables and the second interest information as a target variable.
[0053] As described above, the device 10 acquires information about the content of a user's utterance and generates topic candidates based on the acquired information. The device 10 then generates an utterance to respond to the user's utterance based on the topic candidates and outputs the utterance to, for example, a virtual space. The device 10 automatically performs this process every time the user speaks, thereby enabling the user to have a conversation with, for example, a character in a virtual space. The device 10 outputs the generated utterance to a server and a terminal.
[0054] Next, a process executed by the device 10 according to this embodiment will be described with reference to the flowcharts of Figures 11 and 12. Figure 11 is a flowchart showing an example of a process for generating an utterance to a user. Before executing the process for generating an utterance to a user, the device 10 generates relationship information indicating the relationship between a plurality of words based on information about the plurality of words.
[0055] First, the device 10 acquires information about the user's utterance (step S1). The device 10 acquires information about a plurality of words (step S2). The device 10 acquires relationship information indicating the relationships between the plurality of words (step S3). The device 10 estimates first interest information indicating the user's interests based on the information about the user's utterance (step S4: first estimation step). The device 10 estimates second interest information for estimating topics that indicate the user's interest based on the first interest information and the relationship information (step S5: second estimation step). The device 10 estimates topic candidate information based on the second interest information (step S6: topic estimation step). A response to the user's utterance is generated based on the topic candidate information (step S7).
[0056] 12 is a flowchart showing an example of the process of estimating first interest information in step S4. In the process of step S4, device 10 estimates value information indicating the user's values based on information about the user's utterance content (step S41). Then, first interest information is estimated based on the value information (step S42).
[0057] Next, the effects of the device 10 according to this embodiment will be described.
[0058] The device 10 according to this embodiment includes a first estimation unit 2 that estimates first interest information indicating a user's interest based on information related to the content of the user's utterance, a second estimation unit 4 that estimates second interest information for estimating topics that indicate the user's interest based on the first interest information and relationship information indicating the relationship between a plurality of words, and a topic estimation unit 5 that estimates topic candidate information indicating topic candidates for the user based on the second interest information.
[0059] In this embodiment, second interest information is estimated based on first interest information and related information, and topic candidate information indicating topic candidates for the user is estimated based on the second interest information. Here, if topic candidate information is estimated based only on the first interest information, the number of topics included in the topic candidate information is insufficient. Therefore, if a response to a user's utterance is based on the topic candidate information, the conversation with the user will become monotonous. On the other hand, according to the device 10, topic candidate information estimated based on the second interest information includes more topics than topic candidate information estimated based only on the first interest information, resulting in more diverse responses based on the topic candidate information. This allows for more diverse responses to be returned to the user when a response to a user's utterance is based on a topic of interest to the user. In other words, by estimating second interest information based not only on the first interest information but also on related information, knowledge related to the user's utterance can be inferred using not only pre-learned information but also externally acquired information. This prevents the conversation with the user from becoming monotonous.
[0060] For example, the number of embedded expressions of words having vertices within a predetermined distance from the vertex of the user's embedded expression in the second interest information can be increased compared to the number of embedded expressions in the first interest information. Since the number of topics included in the topic candidate information estimated based on the second interest information increases, replies based on the topic candidate information become more diverse. This allows for a more diverse response to the user's utterances based on topics of the user's interest.
[0061] Furthermore, in the past, it was difficult to generate utterances for each user by having the language model L learn the individual characteristics of the user. In contrast, the device 10 can estimate topics that the user is likely to be interested in for each user based on the content of the user's past utterances, and input the estimated information into the language model L. This makes the content of utterances generated by the language model L diverse and suitable for each user.
[0062] The first estimation unit 2 may include a value estimation unit that estimates value information indicating the user's values based on information related to the content of the user's utterance, and an interest estimation unit that estimates first interest information based on the value information. In this case, the first interest information, the second interest information, and the topic candidate information are information that reflects the user's values. This allows responses to the user's utterance to be related to topics that correspond to the user's values, making it possible to more reliably generate utterances related to topics that interest the user.
[0063] The first interest information may include a plurality of first semantic information pieces indicating the user's characteristics and the meanings of at least one or more feature words related to the user's speech content, and the second interest information may include a plurality of second semantic information pieces indicating the user's characteristics, the meanings of at least one or more feature words, and the meanings of multiple words. In this case, the first interest information and the second interest information represent the user's level of interest in each feature word or word as a difference between the meaning of each feature word or word and the user's characteristics. This allows the second interest information estimated based on the first interest information to more appropriately reflect the user's interests. As a result, the topics included in the topic candidate information estimated based on the second interest information are more likely to be topics in which the user is interested.
[0064] The first semantic information may be first vector information that represents the user's characteristics or the meaning of a characteristic word as a real number vector, and the second semantic information may be second vector information that represents the user's characteristics, the meaning of the characteristic word, and the meaning of one of the multiple words as a real number vector. In this case, in the first interest information and the second interest information, the user's degree of interest in each characteristic word or word is represented as a distance between a vector indicating the meaning of each characteristic word or word and a vector indicating the user's characteristics. This makes the first interest information and the second interest information information that more appropriately reflects the user's interests. As a result, the topic included in the topic candidate information estimated based on the second interest information is more likely to be a topic in which the user is interested.
[0065] The first interest information may include first semantic relationship information indicating a relationship between the first semantic information. In this case, the second interest information can be more accurately estimated based on not only the meanings of the first semantic information but also the relationship between the first semantic information. For example, in the second interest information, the distance between the node "Nadeshiko" and the node "Soccer Game W11" can be set closer, taking into consideration that the meanings of the words "Nadeshiko" and "Soccer Game W11" are different but each have a relationship with the word "Soccer."
[0066] The second interest information may include second semantic relationship information indicating the relationship between the second semantic information. In this case, the second estimation unit 4 can more accurately estimate the difference in meaning between the second semantic information based on the relationship between the second semantic information. For example, the second estimation unit 4 can estimate the second interest information using link prediction. This allows the second interest information to be more accurately estimated.
[0067] The second estimation unit 4 may use an estimation model trained by machine learning to calculate the second interest information using the first interest information and the related information as input. The estimation model may be trained using the first interest information and the related information as explanatory variables and the second interest information as a target variable. In this case, the second interest information is more accurately estimated by the estimation model constructed by machine learning. This increases the number of topics that interest the user in the topic candidate information. As a result, when responding to a user's utterance based on the user's topics of interest, a more diverse range of responses can be returned to the user.
[0068] The relationship information may be a knowledge graph including multiple entities each representing multiple words and multiple relations each representing a relationship between the multiple entities. In this case, it is possible to include words related to multiple feature words in the topic candidate information based on the multiple relations each representing a relationship between the multiple entities. This increases the number of topics that interest the user in the topic candidate information. As a result, when responding to a user's utterance based on the user's topics of interest, it is possible to return a more diverse response to the user.
[0069] [About the Present Disclosure] The device 10 of the present disclosure has the following configuration.
[0070] [1] A device comprising: a first estimation unit that estimates first interest information that indicates an interest of a user based on information about the content of the user's utterance; a second estimation unit that estimates second interest information for estimating topics that indicate an interest of the user based on the first interest information and relationship information that indicates relationships between a plurality of words; and a topic estimation unit that estimates topic candidate information that indicates topic candidates for the user based on the second interest information.
[0071] [2] The device described in [1], wherein the first estimation unit includes: a value estimation unit that estimates value information indicating the user's values based on information about the content of the user's utterance; and an interest estimation unit that estimates the first interest information based on the value information.
[0072] [3] The device described in [1] or [2], wherein the first interest information has a plurality of first semantic information indicating the characteristics of the user and the meanings of at least one or more characteristic words related to the user's speech content, and the second interest information has a plurality of second semantic information indicating the characteristics of the user, the meanings of the at least one or more characteristic words, and the meanings of the plurality of words, respectively.
[0073] [4] The device described in [3], wherein the first semantic information is first vector information that represents the user's characteristics or the meaning of the characteristic word as a real number vector, and the second semantic information is second vector information that represents the user's characteristics, the meaning of the characteristic word, or the meaning of one of the multiple words as a real number vector.
[0074] [5] The device according to [3] or [4], wherein the first interest information further includes first semantic relationship information indicating a relationship between the first semantic information.
[0075] [6] The device according to any one of [3] to [5], wherein the second interest information further includes second semantic relationship information indicating a relationship between the second semantic information.
[0076] [7] The device described in any one of [1] to [6], wherein the second estimation unit calculates the second interest information using an estimation model trained by machine learning, with the first interest information and the related information as input, and the estimation model is trained using the first interest information and the related information as explanatory variables and the second interest information as a target variable.
[0077] [8] The device according to any one of [1] to [7], wherein the relationship information is a knowledge graph including a plurality of entities each representing a plurality of words and a plurality of relations each representing a relationship between the plurality of entities.
[0078] [9] A method comprising: a first estimation step of estimating first interest information indicating an interest of a user based on information about the content of the user's utterance; a second estimation step of estimating second interest information for estimating topics indicating an interest of the user based on the first interest information and relationship information indicating relationships between a plurality of words; and a topic estimation step of estimating topic candidate information indicating topic candidates for the user based on the second interest information.
[0079] [Definition of Terms, etc.] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.
[0080] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0081] For example, the device 10 according to an embodiment of the present disclosure may function as a computer that performs processing of the virtual space providing method of the present disclosure. Fig. 13 is a diagram illustrating an example of the hardware configuration of the device 10 according to an embodiment of the present disclosure. The device 10 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0082] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of apparatus 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.
[0083] Each function of the device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0084] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned first estimator 2 and the like may be realized by the processor 1001.
[0085] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the first estimation unit 2 and the like may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may also be made for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0086] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a virtual space providing method according to one embodiment of the present disclosure.
[0087] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0088] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, or a communication module. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the input unit 1 described above may be realized by the communication device 1004. The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated from each other.
[0089] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0090] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0091] The device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0092] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0093] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0094] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0095] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0096] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0097] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0098] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0099] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0100] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0101] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0102] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0103] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0104] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.
[0105] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
[0106] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0107] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0108] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0109] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0110] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0111] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0112] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0113] 10...device, 2...first estimation unit, 4...second estimation unit, 5...topic estimation unit, 1001...processor, 1002...memory, 1003...storage, 1004...communication device, 1005...input device, 1006...output device.
Claims
1. A device comprising: a first estimation unit that estimates first interest information indicating the user's interest based on information regarding the user's utterance content; a second estimation unit that estimates second interest information for estimating a topic indicating the user's interest based on the first interest information and relationship information indicating the relationship between a plurality of words; and a topic estimation unit that estimates topic candidate information indicating a topic candidate for the user based on the second interest information.
2. The device according to claim 1, wherein the first estimation unit includes: a values estimation unit that estimates values information indicating the user's values based on information regarding the user's utterance content; and an interest estimation unit that estimates the first interest information based on the values information.
3. The device according to claim 1, wherein the first interest information has a plurality of first meaning information respectively indicating the user's characteristics and the meanings of at least one or more characteristic words regarding the user's utterance content, and the second interest information has a plurality of second meaning information respectively indicating the user's characteristics, the meanings of the at least one or more characteristic words, and the meanings of the at least one or more words.
4. The device according to claim 3, wherein the first meaning information is first vector information representing the user's characteristics or the meaning of the characteristic word as a real number vector, and the second meaning information is second vector information representing the user's characteristics, the meaning of the characteristic word, or the meaning of one of the plurality of words as a real number vector.
5. The device according to claim 3, wherein the first interest information further has first meaning relationship information indicating the relationship between the first meaning information.
6. The device according to claim 3, wherein the second interest information further has second meaning relationship information indicating the relationship between the second meaning information.
7. The device according to claim 1, wherein the second estimation unit calculates the second interest information using an estimation model learned by machine learning, with the first interest information and the relationship information as inputs, and the estimation model is learned with the first interest information and the relationship information as explanatory variables and the second interest information as the target variable.
8. The device according to claim 1, wherein the relationship information is a knowledge graph including a plurality of entities respectively representing a plurality of words and a plurality of relations respectively representing the relationships between the plurality of entities.
9. A method comprising: a first estimation step of estimating first interest information indicating the user's interest based on information regarding the user's utterance content; a second estimation step of estimating second interest information for estimating a topic indicating the user's interest based on the first interest information and relationship information indicating the relationship between a plurality of words; and a topic estimation step of estimating topic candidate information indicating a topic candidate for the user based on the second interest information.
Citation Information
Patent Citations
Speech sentence generation device, and method and program for the same
JP2015045833A
Interest determination device, interest determination method, and program
JP2018190136A
Generation of data related to underrepresented data based on received data input
JP2020057365A