Interactive intention information extraction program, device, and method

The interactive intention information extraction program addresses the challenge of processing free-form responses in multi-stage questioning by utilizing Japanese language processing, enabling accurate user intent extraction and analysis for surveys.

JP7774793B2Active Publication Date: 2025-11-25DENTSU INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021096018
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-08
Publication Date
2025-11-25
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

Existing AI technologies for multi-stage questioning in surveys lack effective Japanese language analysis, making it difficult to process free-form user responses and extract user intent accurately, especially in fields like marketing research.

Method used

An interactive intention information extraction program using Japanese language processing technology for multi-stage questioning, which includes a comparative expression database and dialogue control to analyze and manage user responses, allowing for the extraction of user intent through free-form text inputs.

Benefits of technology

Enables automatic extraction and analysis of user intent from free-form responses, providing easy-to-read charts and supporting dialogue on a wide range of topics, overcoming the limitations of Yes/No or multiple-choice questionnaires.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007774793000001
    Figure 0007774793000001
  • Figure 0007774793000002
    Figure 0007774793000002
  • Figure 0007774793000003
    Figure 0007774793000003
Patent Text Reader

Abstract

To provide an interactive intention information extraction program, apparatus, and method capable of extracting intention information of a user from dialogue with the user in a chat format.SOLUTION: A program implemented in a computer causes the computer to operate as: text input means 11 for interactively capturing a free answer from a user; intention information extraction means 12 for extracting intention information expressing user's intention contained in the user answer using a comparative expression database 21; dialogue control means 13 for controlling progress of the dialogue with the user by referring to a classification in the comparative expression database 21 to which the extracted intention information belongs and a dialogue scenario database 23; and log analysis means 14 for analyzing and aggregating dialogue history.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an interactive intention information extraction program, device, and method that can extract user intention information from a chat-style dialogue with a user. [Background technology]

[0002] Multi-stage questioning is frequently used in various surveys and market research, but traditionally, this has been done manually. This has required time, effort, and expense, as well as complex processes from creating the questions to asking the subjects and compiling the collected data. Meanwhile, with the recent advances in information processing technology, various AI (artificial intelligence) technologies have been introduced into various fields of society, and AI is now being used in surveys and other areas. For example, Patent Document 1 discloses a technology aimed at extracting appropriate information from diverse dialogues with users. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-193533 Summary of the Invention [Problem to be solved by the invention]

[0004] Patent Document 1 aims to reduce labor by using AI, but does not provide any details on Japanese language analysis, which is essential for processing and affects the success of the system. Therefore, even if multi-stage questions using a Yes / No type or menu format are feasible, it seems difficult to realize a type of question that allows the user to respond freely. In order to solve this problem, the present invention aims to realize multi-stage queries that are easy for users to understand and manage, without using keyword matching that is prone to errors or complicated menu formats, by using Japanese language processing technology that utilizes structural comparison of sentences. This multi-step questioning technique is expected to be used in a variety of fields, one of which is laddering, a marketing research method in which survey respondents are asked a series of questions about products and brands, in order to clarify the benefits and value of a product, as if climbing a ladder. In this invention, this laddering survey is conducted using AI, with the aim not only to extract the characteristics and elements of the targeted products from the free responses of the survey subjects, but also to clarify how these elements are connected to create value. [Means for solving the problem]

[0005] In order to solve the above problems, the interactive intention information extraction program of the present invention is A program for extracting user intent information through multi-stage questioning, Computer, a text input means for interactively capturing free responses from users; an intention information extraction means for extracting intention information that expresses the user's intention contained in the user's reply sentence using a comparative expression database; As a dialogue control means for controlling the progress of a dialogue with a user by referring to the classification in the comparative expression database to which the extracted intention information belongs and a dialogue scenario database. It is characterized by operating Furthermore, it is also preferable that the system be operated as a log analysis means for analyzing and aggregating the dialogue history. The user's response is spoken in Japanese and can be a sentence (simple, compound, complex, or overlapping), a word, or a combination of an adjective and a noun.

[0006] Here, intention information refers to information contained in a user's utterance that a listener wants to extract for a specific purpose. For example, when asked about beer, a listener might respond at length, "I drink any alcoholic beverage, including wine and beer. I always drink wine with lasagna. Beer is best enjoyed with yakitori on a hot day." If the listener is conducting market research on beer, the information "yakitori and beer on a hot day" is of interest to the listener, and the listener will simply ignore the mention of wine. In this way, intention information is the information that the listener wants to extract from the utterance. The speaker intends to talk about their favorite drinks, wine and beer, but only the mention of beer, which is of interest to the listener, matches their intention. Thus, the intention information in this invention is interpreted from the listener's perspective and differs from the general meaning of "intention"—"what the speaker is thinking" (Kojien). For example, if the topic is market research on beer, the comparative expression database must contain a large number of expressions related to beer. These expressions (which can be sentences or words) are called "factors" in this invention.

[0007] The comparative expression database preferably has one or more areas of interest classified into any number of categories for each purpose of the dialogue, with comparison sentences registered in each area of ​​interest, and when a match is established between all or part of the user's response sentence and one of the comparison sentences, the comparison sentence is considered to be intention information. In the following embodiments, the "area of ​​interest" corresponds to the "layer," and the "comparison sentence" corresponds to the "factor." The comparison sentences are also written in Japanese and can be any sentence (simple, compound, complex or overlapping), a word or a combination of an adjective and a noun.

[0008] The dialogue scenario database preferably stores rules that define the progression of the dialogue and the individual probing questions that make up the multi-stage questioning of the user. As a basic mode of proceeding with the dialogue, it is preferable to start with a primary response to a primary question posed to the user, execute multi-stage questions equal to the number of pieces of intention information extracted from the primary response, and extract intention information from the responses to the in-depth questions that make up the multi-stage questions.

[0009] The comparative expression database preferably has an interest area to which a comparative sentence expressing the user's indifference belongs, and when a comparative sentence belonging to this area is extracted, the dialogue control means preferably terminates the dialogue.

[0010] The present invention can be realized by implementing an interactive intention information extraction program in a computer and causing it to operate as an interactive intention information extraction device, or as a method for operating an interactive intention information extraction device. [Effects of the Invention]

[0011] According to this invention, AI can automatically extract intent information from dialogue with users, so that responses can be freely entered in text rather than being limited to Yes / No or multiple choice types in questionnaires, etc. Also, while free-entry type questionnaires are not easy to tally manually, this invention can also perform this automatically with AI, and the analysis results can be displayed in easy-to-read charts. Furthermore, simply by switching the comparative expression database, it opens the door to dialogue on topics from a wide range of fields. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating an outline of processing according to an embodiment of the present invention. [Figure 2] FIG. 10 is a state transition diagram of intention information extraction according to an embodiment of the present invention. [Figure 3] FIG. 3 is a diagram showing comparison target sentences (factors) registered in a factor list according to the embodiment of the present invention. [Figure 4] 1 is a functional block diagram of an intention information extraction device according to an embodiment of the present invention. [Figure 5]FIG. 10 is a diagram illustrating a structure of a factor list according to an embodiment of the present invention. [Figure 6] FIG. 10 is a diagram illustrating a probing question determination rule according to an embodiment of the present invention. [Figure 7] 1A to 1C are diagrams illustrating basic ladder movement states (up, down, slide) according to an embodiment of the present invention. [Figure 8] 10A to 10C are diagrams illustrating an example of laddering operation according to an embodiment of the present invention. [Figure 9] FIG. 10 is a diagram illustrating a factor extraction process according to an embodiment of the present invention. [Figure 10] FIG. 2 is an overall processing flow diagram of the embodiment of the present invention. [Figure 11] FIG. 10 is a process flow diagram of the primary response phase according to the embodiment of the present invention. [Figure 12] FIG. 10 is a process flow diagram of the laddering phase according to an embodiment of the present invention. [Figure 13] FIG. 10 is a diagram illustrating a factor restriction of a transition destination layer according to an embodiment of the present invention. [Figure 14] FIG. 1 is a chain diagram of extracted factors according to an embodiment of the present invention. [Figure 15] 1 is a context map of extracted factors of an embodiment of the present invention; [Figure 16] FIG. 10 is a diagram visualizing clustering of extracted factors according to an embodiment of the present invention. [Figure 17] FIG. 10 is a diagram visualizing the extracted factors according to the embodiment of the present invention, analyzed based on the ladder connection rate and the number of appearances of the factors. [Figure 18] FIG. 10 is a visualization of the extracted factors of an embodiment of the present invention compared between different brands. [Figure 19] FIG. 10 is a diagram illustrating normalized visualization of transitions in users' extraction factors according to an embodiment of the present invention. [Figure 20] FIG. 10 is a diagram illustrating an example in which the present invention is applied to a field other than marketing, such as healthcare. DETAILED DESCRIPTION OF THE INVENTION

[0013] An embodiment of the present invention will be described below with reference to the drawings. In the following description, the comparative expression database will be called a "factor list," and each of the comparison sentences, i.e., the words and sentences registered in the factor list, will be called a "factor." "User" refers to the speaker, but can also mean "subject" or "survey subject." Furthermore, the entity that asks questions to the user and receives answers is called the "system," but can also be called the "listener" or "reader."

[0014] ≪1. System Overview≫ First, referring to Figure 1, we will explain the overview of a processing system (hereinafter referred to as "this system") using an interactive intention information extraction device (hereinafter referred to as "this device") 1 that implements an interactive intention information extraction program. When a user inputs a Japanese sentence into this device 1, (2) this input sentence and each factor in an externally uploaded factor list are read. If an expression similar to a factor is found in the input sentence, this factor is extracted as intention information. Since this system is characterized by asking multiple questions to understand the user's intention, (3) the system controls the dialogue to determine what kind of probing question to ask or to terminate the dialogue, and asks a new question to the user or sends an end message. The dialogue history, including the questions to the user, their answers, and the extracted factors, is stored in a log file. Then, (4) the system refers to the log file to aggregate and analyze the data and output it visually in an easy-to-understand format.

[0015] An overview of this system from the perspective of data transition is shown in Figure 2. In other words, an input sentence, which is an answer to one question, is associated with a layer. The factor list is divided into several groups based on a specific perspective, and each group is called a layer. In the example of Figure 3, the factor list consists of three layers: Layer A, which collects factors from the perspective of "attributes," Layer B, which collects factors from the perspective of "function," and Layer C, which collects factors from the perspective of "emotion." (Layer Z will be discussed later.) Each layer has one or more pieces of intent information, i.e., expressions that express the user's intent. For example, layer A (attribute) has intent information such as "product packaging" and "product logo," layer B (function) has intent information such as "good" and "delicious," and layer C (emotion) has intent information such as "cheerful" and "fun," which are all "factors" in this system. In other words, a layer is a group that is associated with the user's intention information. The layer is determined by the factors included in the user's input sentence. For example, when the factor "delicious" is extracted, the user's answer is determined to belong to layer B (function), and the next layer and question to be asked are determined by the probing question determination rules. Usually, layer C or layer A, which are adjacent to layer B to which the extracted factor "delicious" belongs, are the layers to be asked. This is called laddering up or laddering down, which will be described later.

[0016] The above is an overview of the processing and data transition of this system. Next, the configuration and operation of this system will be described in detail with reference to the drawings.

[0017] 2. Functional block configuration of this system FIG. 4 is a functional block diagram of the device 1. The device 1 includes a text input unit 11, an intention information extraction unit 12, a dialogue control unit 13, and a log analysis unit 14. These functional components can be realized by executing a computer program installed in the device 1.

[0018] The text input unit 11 inputs text in Japanese sent by the user U. The user U may be a mobile terminal or a personal computer used by the user. The input format may be character input from a keyboard or touch panel, or voice input. In the case of voice input, the device 1 must be equipped with a voice recognition unit (not shown). Furthermore, questions to the user U are usually given by voice, so a voice synthesis unit (not shown) is also required. The input text can be one or more sentences, a single word, or a string of two or more words. The sentences can be simple, compound, complex, or overlapping.

[0019] The intention information extraction unit 12 comprises a factor extraction unit 121 that externally references a factor list 21, which is a comparative expression database, to extract factors that result in a match from the input text; an extracted factor file 122 that saves each factor as it is extracted; and an extracted factor filtering unit 123 that appropriately selects from the extracted factors. Note that "matching occurs" does not mean a word-for-word match, but rather a near match. For example, if the input text is "I'm going to buy beer for my friend" and the factor is "I'm going to buy beer for my friend," a match occurs. Similarly, if the input text is "It feels like beer" and the factor is "It feels like beer," a match occurs. The selected extraction factors, layers, and other information are input to the dialogue control unit 13.

[0020] The dialogue control unit 13 comprises a laddering processing unit 131 that acquires layers from the extracted factors, identifies the next layer to transition to based on the acquired information and the dialogue history information up to now in the history file 22, and determines the end of laddering, etc.; a probing question identification unit 132 that identifies the next question to ask the user U by referring to the information from the laddering processing unit 131 and the dialogue scenario DB 23; and a user output unit 133 that sends the next question or message to the user U.

[0021] The log analysis unit 14 collects and analyzes the data stored in the history file 22, and outputs it to the screen in the form of a table or diagram, prints it, or transmits it to an external computer via a communication line.

[0022] Next, the factor list 21 and the dialogue scenario DB 23 that store data necessary for executing this system will be described.

[0023] The factor list 21 needs to be updated depending on the nature of the information that the system wishes to obtain from user U, so it is uploaded from an external device each time processing by the system begins. This factor list 21 is created by an external system separate from the system, so an explanation of how it is created will be omitted. As already mentioned, the factor list 21 is a collection of layers consisting of one or more factors. Here, the data structure of one layer will be explained using Figure 5.

[0024] As shown in Figure 5, similar factors are grouped under the same factor ID. Each factor ID is also assigned a label. A higher-level collection of these is a group (=layer). The index in the figure is used to distinguish between factors with the same factor ID. Each factor is expressed as text in Japanese. The factor label "CM" contains the factors "CM" and "CM is interesting." In this case, "CM is interesting" contains two independent words, "CM" and "interesting," and "CM" consists of one independent word. Since one factor list 21 can include multiple layers, the layers A, B, and C illustrated in FIG. As shown in Figure 3, it is a good idea to also register Layer Z, which is a collection of factors that represent the intention of "indifference," in Factor List 21.

[0025] The dialogue scenario DB 23 stores rules necessary for dialogue control. One of them is the probing question determination rule shown in Figure 6(1). The rules for determining in-depth questions include the current layer, the direction of laddering from the previous layer, the case distinctions for question selection, and specific example questions. In Figure 6(1), the line marked with (※1) indicates that the layer has not yet been determined and the first question (primary question) is prepared: "What image do you have of ▲▲▲?" The lines marked with (※2) indicate that the transition to Layer A (attribute) has been made via a ladder-down process, and that the first question to obtain factors related to Layer A (attribute) is "What made you feel that way?" Lines marked with (※3) indicate that the player has progressed down the ladder to Layer A (attribute), that they were hoping to obtain a factor related to Layer A (attribute), but instead obtained a factor for another layer (B or C), and that the same question (although worded differently) is available to obtain a factor related to Layer A (attribute). The line marked with (※4) indicates that the element has slid from Leia A (attribute) to Leia A (attribute), and that the first question to obtain the factor related to Leia A (attribute) is "Is there anything else you know?"

[0026] As shown in Figure 6(2), the probing question determination rules also include questions to ask when an unintended answer to a question is received, that is, when no corresponding factor is found in the factor list 21, or when an answer indicating indifference is received. As with the factor list 21, the probing question determination rules also need to be changed in consideration of the nature of the information that the system wishes to obtain from user U, and so are uploaded from an external device each time the system is started.

[0027] Laddering, i.e., the direction of a series of dialogues to extract factors from each layer, is important information for identifying probing questions. In order to obtain user intent information appropriately and efficiently, this system assumes desirable transitions between layers. The order of "Layer A (Attributes)" → "Layer B (Function)" → "Layer C (Emotion)" shown in Figure 7 is desirable, and a transition in this direction is called "Ladder Up." The reverse order of transitions, "Layer C (Emotion)" → "Layer B (Function)" → "Layer A (Attributes)," is called "Ladder Down." However, when one laddering session ends and the next one begins, transitioning to the same layer is permitted, and this is called "Lader Slide" (sideways movement). These ladder up, ladder down, and ladder slide are extremely important concepts in the dialogue control process described later, so the main points will be explained here.

[0028] Laddering is performed according to the following rules, which should also be stored in the dialogue scenario DB 23. ★Rules for the first laddering session★ Laddering starts with an up (down) movement from the start. -The layer based on the factor obtained in the first response (hereafter referred to as the "first response factor" to distinguish it from the "factor" obtained during laddering) will go up (down) until there is no more. When there are no more layers to go up (down), it will search for an available layer and go down (up). ★Rules for the second ladder★ If there are multiple primary response factors, the layer based on the primary response factors Continue going up (down) until there is no more. In other words, it is the same as the first ladder. If there is only one primary answer factor, return to the first layer (the layer where you first obtained the factor) and slide. In other words, the questions are designed to extract factors of the same layer. ★Laddering afterwards★ - Attempts to ladder as much as possible, but fails when laddering is completed twice If so, the slide will not be performed and the presentation will end.

[0029] A specific example of operation according to this laddering rule is shown in FIG. When laddering is possible only from ladder-up, laddering begins by obtaining a primary question from the probing question determination rule and obtaining a primary answer from the user, as shown in the left column of Figure 8. For example, suppose the primary question is "What image do you have of XX Beer Co., Ltd.'s 'Beer Weather'?" (see line (*1) in Figure 6), and the user responds with a primary answer of "I always buy XX Beer products. It's my favorite brand." From this primary response, the factor "Favorite manufacturer" (see (※a) in Figure 3), which belongs to Layer A (attribute), can be obtained as the primary response factor. Since Layer B is the layer that ladders up from Layer A, the first question associated with Layer B (function) is obtained, "What good does that do you get from that?" (see line (※5) in Figure 6). Let's assume that this question is asked to the user and the response is "It's a product of XX Beer, so it should be delicious." From this response, the factor "Delicious" (see (※b) in Figure 3) for Layer B (function) can be obtained, and then the first question associated with Layer C (emotion), which ladders up from Layer B, is obtained, "How does that make you feel?" (see line (※6) in Figure 6). Let's assume that this question is asked to the user and the response is "It makes me feel positive and wants to do my best tomorrow." From this response, we were able to obtain the Layer C (emotion) factor of "feeling positive" (see (※c) in Figure 3), so this one round of laddering was a success. In this way, laddering usually starts from the beginning. Factors are obtained in the order of layer A, layer B, and layer C. When there are no layers beyond layer C, there are no layers for which factors have not yet been extracted, so one laddering run ends.

[0030] The center column in Figure 8 shows a case where a single laddering session contains a mixture of laddering up and laddering down. Once the primary response factor for layer B (function) is obtained from the answer to the primary question, the first question associated with the next layer, layer C (emotion), is obtained. This question is asked to the user, and once a factor for layer C (emotion) has been obtained, there is no further layer to go up to. When there are no more layers to go up to, the system searches for a layer from which a factor has not yet been extracted and goes down. Since a factor has not been extracted from layer A (attribute), a question for layer A (attribute) is obtained. This is a laddering down from layer C to layer A. If a question for layer A (attribute) is asked to the user and a factor for layer A (attribute) is obtained, one laddering session is successful.

[0031] When laddering is achieved only from ladder down, as shown in the right column of Figure 8, the concept is the same as ladder up (left column of Figure 8), except the order of the layers is reversed.

[0032] The block configuration of the device 1, the factor list 21, and the dialogue scenario DB 23 have been described above. Next, regarding the operation of this device 1, The explanation will focus on three points: (1) factor extraction, (2) dialogue control including identification of probing questions, and (3) analysis of log data.

[0033] 3. Operation of this system 3.1 Factor Extraction The processing by the factor extraction unit 121 will be described with reference to FIG. The user's input sentence is run through a syntactic parser, and the results of the analysis are converted to make matching easier. For example, if the original input text is a compound or complex sentence, it is split into multiple simple sentences. In this case, the subject "he" is added to the second sentence, such as "He went to the convenience store and bought a can of beer.", which becomes "He went to the convenience store." and "He bought a can of beer." This is an example of completing a structure and adding redundancy.

[0034] The factors belonging to each layer of the factor list 21 are also processed by the parser, and for example, the factor "I drank a lot of the newly released beer on my day off" is decomposed as follows: drink (predicate) <------- Many 《predicate modification》 <--Holiday {complement} <--beer {complement} <-- New Release 《Modifying a Complement》 In this way, auxiliary words (particles, auxiliary verbs) are excluded. Conjunctions may also be excluded. The reason why factors are simplified compared to the user's input sentence, which may become redundant as a result of syntactic analysis, is to ensure accurate matching.

[0035] The basis of factor extraction processing is Japanese language processing, which can process any Japanese sentence input using structural comparison of the sentence. Existing technology exists for this Japanese language processing, so it is sufficient to use this. For example, the technology disclosed in Patent Application No. 2021-79401, filed on May 8th of this year, is one such example.

[0036] The extracted factors may overlap in their intended meaning or may be included as part of other factors, so the extracted factor filtering unit 123 performs the following processing. First, when multiple different factors containing the same intention information are detected, we want to eliminate the duplication. There are various ways to do this, for example, selecting a factor with strong constraints. Specifically, it is advisable to count the number of independent words contained in one factor and select the factor with the largest number. Alternatively, it is also possible to select the factor with the largest number of characters. Whatever the method, the key is to select the appropriate one from multiple factors. If multiple factors with the same information are detected, the first one found will be selected. For example, if the factor "Showa" is detected three times in one input sentence, as "Showa," "Showa," and "Showa," only the first "Showa" will be retained.

[0037] The factors selected by filtering are output to the laddering processing unit 131 and also stored in the history file 22. If the answer is to the first question in the interview, the history file 22 registers the user's identification information and, if necessary, personal information. If the input is voice, the user's gender and approximate age may be automatically determined from their voice. For extracted factors, the serial number of the answered question, the layer to which the factor belongs, the factor ID, etc. are recorded. These are essential information for identifying the next question.

[0038] 3.2 Dialogue Control First, we will provide an overview of the interaction process with the user from start to finish, as shown in Figure 10. The entire interactive process is divided into two parts: Phase 1 and Phase 2. Phase 1 corresponds to the interview, which is the beginning of the laddering survey, and the interview begins with the display of a chat screen. The first question displayed on this chat screen is called the primary question, and in Phase 1, primary response factors are extracted from the user's answers to the primary question (Step F1). In other words, the primary question is a question posed to the user with the aim of obtaining primary response factors for laddering. The subsequent Phase 2 is a laddering process that starts from the primary response factors extracted here, and Phase 1 is distinguished from the laddering of Phase 2. The questions posed to the user in Phase 2 are called "deep dive questions" and are distinguished from the "primary questions" mentioned above.

[0039] Phase 1 will be described with reference to the processing flow diagram of FIG. At the start of the interview, the answer to the primary question posed to the user is called the primary answer. Primary answer factors are extracted from this primary answer (step S11). There may be multiple extracted primary answer factors. If an extracted factor is found (Yes in step S12), it is stored as a primary answer factor (step S13). Let us assume that two primary answer factors, Factor A (attribute) and Factor B (function), are extracted. In this case, laddering will begin in the order of Factor A and Factor B. The layer to which the extracted primary response factor belongs is obtained (step S14). Here, since laddering starts from primary response factor A, layer A (attribute) is obtained. The next transition destination of layer A (attribute) is determined to be layer B (function) by laddering up, and questions that can extract factors of layer B (function) are obtained from the probing question determination rules and displayed on the chat screen of the user terminal (step S15). Phase 2 starts with asking this probing question to the user (see step F2 in Figure 10).

[0040] If the number of primary answer factors is 0 in step S12 (No in step S12), the primary question is re-asked, but if the number of primary questions already asked is three or more (Yes in step S16), the interview ends (step S17). The multi-stage questioning ends without transitioning to phase 2, i.e., without performing laddering even once. If the number of times the primary question has been asked is less than three (No in step S16), the primary question is acquired again and displayed on the chat screen (step S18). Note that, based on the results of step S14, the first laddering is performed assuming that the factor of layer A (attribute) has been obtained, and when this is completed, the laddering (second laddering) of factor B (function) follows. If three or more primary answer factors are extracted in step S12, laddering is performed the same number of times as the number of primary answer factors. If multiple primary answer factors are obtained in phase 1 of Figure 10, each will be laddered separately in phase 2. These are called the first laddering, second laddering (, third laddering, etc.).

[0041] In Phase 2 of Figure 10, the first laddering is performed using the primary response factor. Phase 2 laddering is repeated the number of times there are primary response factors, but if only one primary response factor is obtained, lateral movement is performed using this primary response factor. In other words, when only one primary response factor is obtained, the second laddering is performed using lateral movement to reuse the same primary response factor as the first time.

[0042] Returning to Figure 10, Phase 2 will now be described. The primary response factors of Layer A (attribute) are assumed to be extracted from the primary responses given at the start of the interview. The dialogue scenario DB 23 is referenced to obtain probing questions (step F2), and if a factor belonging to layer B (function) is extracted from the answers returned by the user (step F3) or a factor belonging to layer C (emotion) is extracted (step F4), the ladder-up questions are deemed successful. On the other hand, if neither a factor from layer B (function) nor a factor from layer C (emotion) can be extracted, that is, if "no factor" is found, it is determined to have failed, and the ladder-up questions are posed again. This is repeated up to three times (step F5), and if no factor can be extracted, laddering for this round is terminated. "No factor" refers to the case where, for example, a question is asked in the hope of getting an answer about a layer (attribute), but the factor extracted is in a different layer (function or emotion). In this case, the question is asked three times until an answer about the layer (attribute) is obtained. However, this number of times can be changed.

[0043] If a factor of Layer B (function) can be extracted in step F3, a ladder-up question is obtained according to the probing question determination rules (step F6), and if a factor belonging to Layer C (emotion) is extracted from the answer returned by the user (step F7), the ladder-up question is deemed successful. On the other hand, if there is "no factor," the question is repeated up to three times (step F8), and if a factor cannot be extracted, laddering for this round is terminated. If a factor of Layer C (emotion) is extracted in step F4, a ladder-down question is obtained according to the probing question determination rules (step F9), and if a factor belonging to Layer B (function) is extracted from the answer returned by the user (step F10), the ladder-down question is deemed successful. On the other hand, if there is "no factor," the question is repeated up to three times (step F11), and if a factor cannot be extracted, laddering for this round is terminated.

[0044] When step F7 is reached, primary response factors have been extracted from layer A (attributes), and factors have been extracted from both layer B (functions) and layer C (emotions) during laddering. In this way, factors have been extracted from all layers, and this laddering session ends (step F12). When step F10 is reached, factors have already been extracted from all of layer A (attributes), layer C (emotions), and layer B (functions), and the laddering for this round ends (step F12). Here, when the number of primary response factors acquired in the primary response (F1) is two or more and the laddering corresponding to each primary response factor is completed (step F13), the series of dialogue processing is terminated.

[0045] On the other hand, if the number of laddering runs is less than two (step F14) and the number of primary response factors obtained in Phase 1 (step F1) is less than two (step F15), a ladder slide question is obtained from the probing question determination rule (step F16) and the second laddering starts (step F17).Laddering is performed at least twice, but since the second primary response factor was not obtained in Phase 1, laddering starts from the slide of the first primary response factor. If the number of laddering runs is less than two (step F14) and two or more primary response factors are obtained in phase 1 (step F1), a second laddering run is started (step F18). If a factor belonging to Layer Z (indifference) is extracted (step F19), the series of dialogue processes is terminated. If no factor is extracted or if indifference responses continue, laddering ends as a failure. In this way, laddering can end without ever being completed.

[0046] Next, Phase 2 in FIG. 10 will be described in detail in accordance with the processing flow in FIG. Factors are extracted from the user's answers to the probing questions (step S101). If an extracted factor is found (Yes in step S102), it is confirmed whether the factor belongs to the expected layer (step S103). If the extracted factor is found in the expected layer (Yes in step S104), the next transition destination layer is obtained (step S105). If factors are found in all layers, there are no more transition destination layers, but if there are still transition destinations remaining (Yes in step S106), the next probing question is identified and displayed on the chat screen (step S107). In step S107, the next probing question is determined by referring to the probing question determination rules. For example, if the expected layer is (attribute) and the next transition destination is (function), then (function) is extracted from the layer sequence and (up) is extracted from the laddering sequence in Figure 6(1), and the corresponding probing question, "What good comes from that?" is identified.

[0047] If no extracted factor is found in step S102, or if the expected layer is not found in step S104, the process proceeds to step S115, and if the number of probing questions is three or more (Yes in step S115), the interview ends (step S116). The total number of probing questions in one layer is three, and if a factor cannot be obtained even after this number has been exceeded, the interview ends. If the number of questions is less than three (No in step S115), probing questions are re-obtained (step S117). In the case of "no extraction factor" (No in step S102), probing questions include "Could you please be more specific?" as shown in FIG. 6(2). This "no extraction factor" occurs when the questioner's intent and the user's answer are at odds with each other, so the question is rephrased and probing questions are repeated up to three times. If the questioner's intended answer is still not obtained, the interview ends (step S116).

[0048] If there are no transition destination layers remaining in step S106, that is, if factors have been entered into all layers, the laddering for this round is terminated (step S108). If there are any primary answer factors that have not yet been laddered, they are acquired (step S109). If there are any unexecuted primary response factors (Yes in step S110), a new laddering run, i.e., Phase 2 in FIG. 10, is started from the beginning (step S111). If there are no unexecuted primary response factors (No in step S110), the history file 23 is referenced to obtain the number of laddering runs that have been executed (step S112). If the number of laddering runs completed is less than two (Yes in step S113), slide questions are obtained (step S114). If the number of laddering runs completed is two or more (No in step S113), the interview ends (step S116).

[0049] Incidentally, when the transition destination layer is acquired in step S104 of FIG. 12, it is advisable to restrict the factors of the planned transition destination layer by the factors extracted from the current layer. For example, the (attribute) factor, (function) factor, and (emotion) factor enclosed in ellipses in Figure 13 are related in meaning. In other words, if a different factor is selected at the laddering destination, this is treated as a failure, and more accurate laddering can be expected. From the viewpoint of the structure of the factor list 21, it is advisable to tag the factor ID of the factor extracted in the current layer, and also tag the factor ID of the factor in the planned transition destination layer, and link these tags. Whether the tags are linked is determined in step S104 of FIG. To explain this more specifically with reference to the factor list 21 in Figure 13, if the user intends the factor "Can be drunk in a can. Can be drunk in a bottle" in Layer A (attribute), and then intends the factor "easy to drink" in Layer B (function), a more natural response can be achieved. However, if the factor "delicious" is intended in Layer B (function), this cannot be said to be a function of this attribute, and therefore cannot be said to be a natural response. In this case, when the desired answer is not obtained, the dialogue control unit 13 repeats the question in a different form within a set number of times, which makes it easier to obtain a natural answer.

[0050] The dialogue control unit 13 controls the overall dialogue with the user. The number of laddering rounds and the number of re-questions within the same ladder can be set as parameters in a memory unit (not shown) of the device 1. The primary questions, probing questions, dialogue end messages, etc. displayed on the user's chat screen are sent to the user U via the user output unit 133. The dialogue control unit 13 stores a series of dialogue contents, such as questions to the user, answers from the user, and layers and factors extracted from the answers, in a history file 22. In conventional questionnaire surveys, questionnaire forms are collected and then entered into the system using keyboard input or OCR. However, in this system, as the dialogue with the user progresses, the contents are automatically entered into the system.

[0051] 3.3 Log Analysis The log analysis unit 14 visualizes the chain of factors by referring to the history file 22. An example of the display is shown in FIG. The mechanism for visualizing the chain of factors (elements) is as follows. (In Figure 14, a new layer D (life) has been added for the convenience of explaining the visualization.) In laddering, which has layers of attributes, functions, emotions, and lifestyles, the factors detected in the conversations with 10 users, User 1 to User 10, are denoted as An, Fn, En, and Ln, and the size of the circle represents the number of factors detected. In Figure 14, the factors are arranged from top to bottom in order of the most common in each layer. If we visualize the linkages between each factor as in Figure 14, we can immediately see the following. For example, we can see that the most important factor in layer D (lifestyle) is L2, followed by L3. We can also see that users who prioritize these factors also prioritize factors E8 and E10 in layer C (emotions). In a similar way, we can see the linkages between factors in layer B (functions) and layer A (attributes).

[0052] In Figure 14, the number of extracted factors is shown in descending order by the size of the circle, but the appearance ranking can also be shown using a table. Either method of expression allows for analysis of frequently occurring factors. Furthermore, by combining not only ladder hierarchy but also demographics (demographic attributes), it is possible to compile data from various perspectives.

[0053] From the perspective of factor chains, it is also possible to visualize the context around a factor, as shown in Figure 15. For example, the factor "The taste of alcohol is strong" can be broken down into the words "alcohol," "taste," and "strong" through morphological analysis, and within this factor, these words are called co-occurrences. In Figure 15, factors that share at least one of the co-occurring words are assumed to be related, and are mapped.

[0054] By performing a cluster analysis of the factors that appear, the factors can be grouped based on the similarity between them and visualized. The dendrogram in Figure 16(1) and the mapping diagram in Figure 16(2) are examples of visualization.

[0055] By analyzing the extracted factors along two axes, such as the rate of connection to a specific ladder and the number of occurrences of the factor itself, it is possible to derive suggestions for marketing and promotion. FIG. 17 shows (1) a visualization of the analysis results and (2) an example of a marketing interpretation.

[0056] Figures 14 to 17 above visualize factors obtained for the same product. It is also useful for marketing purposes to aggregate and visualize factors obtained from surveys of multiple products (different brands of the same company, brands of competing companies, etc.). Figure 18(1) is a graph comparing the number of factors that appeared between different brands. The number of factors can be the total of all extracted factors or the aggregate number for each specific factor. It is also possible to compare the diversity indexes of factors and graph them, as shown in Figure 18(2). For example, if factors related to the likability of a commercial, such as "the commercial has impact" or "the commercial is interesting," are extracted in particularly large numbers, and other factors are extracted in small numbers, the diversity index will take a low value. On the other hand, if various factors such as the commercial, appearance, and taste are extracted evenly, the diversity index will take a high value.

[0057] Figure 19 shows an example of normalizing and visualizing transitions for a specific user from the history stored in the log. If User6 transitions from A9 to F8 to L5 to E8 and laddering is successful, the order of transitions is normalized by laddering up, so the chain is represented by the dashed circle and arrow in the figure. The order in which factors are obtained varies depending on whether you are laddering up or down. However, the order in which they are obtained for each laddering is not an issue in the aggregation process, and a series of values ​​is used as the response of the user. This is because if the aggregation and analysis were to take into account the order in which factors were extracted for each individual user, it would become complicated and make it difficult to grasp the overall trend.

[0058] The aggregated results of dialogue history can be analyzed and visualized from various perspectives in addition to those mentioned above.

[0059] This concludes the description of one embodiment of the present invention. However, the above-described embodiment is merely an example, and therefore the structure of the database such as the factor list, the data contents stored, the processing flow, the method of displaying the tabulated results, etc. are merely examples. For example, in the above embodiment, laddering is standardized by starting from the attribute layer and digging deeper up to the emotion layer, but in actual investigations, the number of layers varies. Also, although the order of layer attribute → layer function → layer emotion is called laddering up, this order is an artificial arrangement and may be changed flexibly. The names of the layers (attributes, functions, etc.) are not limited either. Both the input sentences and factors were assumed to be written in Japanese, but this does not mean that they cannot be applied to other languages.

[0060] The key is that AI can automatically engage in dialogue with users without human intervention and extract information from the dialogue that is relevant to the listener's area of ​​interest. Furthermore, surveys and market research typically require a great deal of effort to compile and analyze the collected data, but it is also important that AI can take over this task. The present invention having such characteristics can be applied not only to laddering as in the above embodiment but also to a wide range of other applications. Some examples of applications are listed below: Map out the customer journey · Collect word-of-mouth information by digging deep into reviews Mental care (resolving anxiety by digging deeper into it) Brain dump tool Coaching Chatbot Auto-navigation Interactive recommendation service · Content development for diagnosis, fortune telling, matching, etc. In both application examples, it is possible to automatically extract information that is deemed to be meaningful from freely input sentences, thereby realizing services that are satisfactory for both the implementing body and the user.

[0061] FIG. 20 is a diagram illustrating the application of the present invention to mental care. In specific fields such as healthcare and psychology, the present invention can be used as a tool for analyzing and diagnosing the answers of individuals in whom specific factors have been detected, in collaboration with experts in the field. In the example shown in Figure 20, factors with negative tendencies such as "not interested" and "it doesn't matter" were detected from some respondents. By looking at the history files of these respondents, it is found that they responded with comments such as "the alcohol has a mild taste, so I don't feel like I've drunk it," "the name is too long to remember," and "the smell is strong, so it gives me a headache." By having industrial physicians and clinical psychologists analyze and diagnose such respondents from a professional perspective, it becomes possible to take an appropriate approach to the respondents. In this case, analysis and diagnosis are performed by human experts, but the preliminary work of extracting respondents who are likely to have problems is handled by this invention, i.e., AI.

[0062] Regarding the convenience of being able to automatically process free-form input sentences, I would like to use the example of a multiple-choice questionnaire (including two-choice questions). Suppose there is a question that asks, "Regarding today's class, please select one of the following: 1. Very good, 2. Good, 3. Average, 4. Bad, 5. Very bad." If there were both good and bad points about "today's class," it would be difficult to choose. Furthermore, the choice between "1. Very good" and "2. Good" is subjective and there is no clear standard for it. Furthermore, if there are no appropriate options, there is a possibility of no response. The present invention overcomes these limitations of multiple choice questionnaires. [Industrial Applicability]

[0063] As a system that can process Japanese dialogue accurately and with minimal effort, this invention has the potential to be used in a wide range of applications, including consumer preference surveys, questionnaires after various events and lectures, and even for checking proficiency levels at cram schools. [Explanation of symbols]

[0064] 1: Interactive intention information extraction device 12: Intention information extraction unit 13: Dialogue control section 14: Log analysis section 21: Comparative Expression Database (Factor List) 22: History file 23: Dialogue scenario database (laddering rules, probing question decision rules, etc.)

Claims

1. An interactive intention information extraction program that extracts, through multi-stage questioning of a user, information contained in the user's utterance that a party listening to the user's utterance wants to extract for a certain purpose, as user intention information, Computer, A text input means for capturing free-form responses by users in an interactive format; an intention information extraction means for extracting user intention information included in the user's reply sentence using a comparative expression database; a dialogue control means for controlling the progress of a dialogue with a user by referring to a dialogue scenario database in which rules used for dialogue control are stored, based on the extracted user intention information; This is an interactive intent information extraction program that operates as a

2. The comparative expression database has listener areas of interest classified by dialogue purpose, and comparison target sentences are registered in each area of ​​interest, When a comparison sentence belonging to a region of interest on the listener side is extracted as user intention information, the dialogue control means proceeds with a dialogue with the user, The comparative expression database has a user's interest area to which a comparison target sentence expressing user indifference belongs, 2. The program for extracting intention information in an interactive manner according to claim 1, wherein said dialogue control means terminates the dialogue with the user when a comparison sentence belonging to a region of interest on the user's side is extracted as the intention information of the user.

3. The dialogue control means starts with a first response to a first question posed to the user, The intention information extraction means extracts user intention information from the primary response, the dialogue control means executes multi-stage questions the number of which corresponds to the number of pieces of user intention information extracted from the primary response; 2. The interactive intention information extraction program according to claim 1, wherein the intention information extraction means extracts the user's intention information from answers to probing questions that constitute a multi-stage question.

4. The comparative expression database has one or more listener interest areas classified into any number of areas for each purpose of dialogue, and a comparison target sentence is registered in each interest area, 2. The interactive intention information extraction program according to claim 1, wherein when a match is established between all or part of a user's response sentence and any of the comparison sentences, the intention information extraction means extracts the comparison sentence as the user's intention information.

5. The dialogue scenario database stores rules that define the progress of the dialogue and the individual probing questions that constitute the multi-stage questions to the user; 3. The interactive intention information extraction program according to claim 2, wherein when the interaction control means proceeds with the interaction with the user, the next area of ​​interest to be asked and the content of the question are determined according to the rules that define the probing question.

6. An interactive intention information extraction device that extracts user intention information, which is information included in the user's utterance that a listener wants to extract for a certain purpose, through multi-stage questioning of the user, a text input means for interactively capturing free responses from users; an intention information extraction means for extracting user intention information included in the user's reply sentence using a comparative expression database; a dialogue control means for controlling the progress of a dialogue with a user by referring to a dialogue scenario database in which rules used for dialogue control are stored, based on the extracted user intention information; An interactive intention information extraction device comprising:

7. An interactive intention information extraction method for extracting user intention information, which is information included in a user's utterance that a listener of the user's utterance wants to extract for a certain purpose, through multi-stage questioning of the user, comprising: The computer capturing free-form user responses in an interactive format; extracting user intention information included in the user's answer sentence using a comparative expression database; a step of controlling the progress of a dialogue with the user based on the extracted user intention information, by referring to a dialogue scenario database in which rules used for dialogue control are stored; This method extracts interactive intent information.

Citation Information

Patent Citations

  • Information extraction device, method and program

    JP2009193533A

  • Data processing apparatus and data processing method

    JP2011022874A

  • Dialogue control system, method, and computer-readable recording medium; multidimensional ontology processing system, method, and computer-readable recording medium

    WO2010125709A1