Dialogue assistance program, dialogue assistance system, and dialogue assistance method

The dialogue assistance system addresses low speech recognition accuracy by verifying answers through second questions, ensuring accurate dialogue assistance and reducing staff burden.

JP2026056222APending Publication Date: 2026-04-01NEC CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing dialogue assistance systems face challenges in accurately assisting dialogue due to low speech recognition accuracy, particularly with proper nouns, leading to increased burden on staff in correcting misrecognized content.

Method used

A dialogue assistance system that includes a first answer acquisition unit, a second question generation unit, and a determination unit to verify the appropriateness of answers through a second question and answer comparison, reducing staff burden by confirming proper nouns and correcting misrecognitions.

Benefits of technology

The system effectively assists dialogue by accurately confirming proper nouns and reducing staff workload through automated verification and correction of speech recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026056222000001_ABST
    Figure 2026056222000001_ABST
Patent Text Reader

Abstract

To provide a dialogue assistance program that can appropriately support dialogue. [Solution] The dialogue assistance program according to this disclosure causes a computer to perform the following steps: a first answer acquisition step of obtaining a first answer to a first question; a second question generation step of generating a second question based on the first answer; a second answer acquisition step of obtaining a second answer to the second question; and a determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an interaction assistance program, an interaction assistance system, and an interaction assistance method.

Background Art

[0002] Techniques for assisting interactions among multiple people are known in the responses at the windows of various facilities and in the responses for reporting to a predetermined institution.

[0003] As a related technique, Patent Document 1 discloses a command system used for receiving reports, giving commands to the site, and supporting rescue activities in emergency activities such as fires. The command system includes an acoustic analysis device and an OA terminal used by an operator (receiver). The acoustic analysis device analyzes an acoustic signal input to the command system to predict an acoustic scene at the incident site and adjusts questions to the reporter according to the prediction result. For the analysis of acoustic signals, voice recognition techniques such as pattern matching and language models (for example, recurrent neural network language models) are used.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By using technologies such as Large Language Models (LLMs) in conjunction with speech recognition technology, real-time assistance for dialogue can be provided. For example, LLMs generate appropriate questions to facilitate efficient dialogue and present these questions to the speaker. LLMs also acquire the content of the dialogue through speech recognition and transcribe it into text in real time. However, if the accuracy of speech recognition is low, LLMs may have difficulty adequately assisting the dialogue.

[0006] The purpose of this disclosure is to provide a dialogue assistance program, a dialogue assistance system, and a dialogue assistance method that can appropriately assist dialogue, in light of the issues described above. [Means for solving the problem]

[0007] The dialogue support program related to this disclosure is: The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, The computer is instructed to perform a determination step that determines whether the first answer is appropriate based on the relationship between the first answer and the second answer.

[0008] The dialogue assistance system relating to this disclosure is A first answer acquisition unit that obtains a first answer to a first question, A second question generation unit that generates a second question based on the first answer, A second answer acquisition unit that obtains a second answer to the second question, The system includes a determination unit that determines whether the first answer is appropriate based on the relationship between the first answer and the second answer.

[0009] The method for assisting dialogue related to this disclosure is: The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, The method includes a determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer. [Effects of the Invention]

[0010] The dialogue assistance program, dialogue assistance system, and dialogue assistance method relating to this disclosure can appropriately assist dialogue. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a block diagram showing the configuration of the dialogue assistance system. [Figure 2] Figure 2 is a flowchart showing the processes performed by the dialogue assistance system. [Figure 3] Figure 3 is a block diagram showing the configuration of the dialogue assistance system. [Figure 4] Figure 4 is a flowchart showing the processes performed by the dialogue assistance system. [Figure 5] Figure 5 shows a specific example of dialogue in a dialogue assistance system. [Figure 6] Figure 6 is a block diagram illustrating the hardware configuration of a computer that implements a dialogue assistance system. [Modes for carrying out the invention]

[0012] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each drawing, the same or corresponding elements are denoted by the same reference numerals. For clarity of explanation, redundant explanations will be omitted where necessary.

[0013] <Embodiment 1> (Dialogue assistance system 100) FIG. 1 is a block diagram showing the configuration of a dialogue assistance system 100 according to the present disclosure. The dialogue assistance system 100 includes a first answer acquisition unit 101, a second question generation unit 102, a second answer acquisition unit 103, and a determination unit 104.

[0014] The first answer acquisition unit 101 acquires a first answer to a first question. The second question generation unit 102 generates a second question based on the first answer acquired by the first answer acquisition unit 101. The second answer acquisition unit 103 acquires a second answer to the second question generated by the second question generation unit 102. The determination unit 104 determines whether the first answer is appropriate based on the relevance between the first answer acquired by the first answer acquisition unit 101 and the second answer acquired by the second answer acquisition unit 103.

[0015] Note that the dialogue assistance system 100 includes a processor, a memory, and a storage device as components not shown in the figure. A computer program implementing the processing according to the present disclosure is stored in the storage device. The processor can read the computer program from the storage device into the memory and execute the computer program. Thereby, the processor realizes the functions of the first answer acquisition unit 101, the second question generation unit 102, the second answer acquisition unit 103, and the determination unit 104.

[0016] Alternatively, the first answer acquisition unit 101, the second question generation unit 102, the second answer acquisition unit 103, and the determination unit 104 may each be realized by dedicated hardware. Also, some or all of the components may be realized by general-purpose or dedicated circuitry, a processor, etc., or a combination thereof. These may be constituted by a single chip or by a plurality of chips connected via a bus. Some or all of the components may be realized by a combination of the above-described circuitry, etc., and a program.

[0017] (Processing of the dialogue assistance system 100) Figure 2 is a flowchart showing the processes performed by the dialogue assistance system 100. First, the first answer acquisition unit 101 acquires the first answer to the first question (S1).

[0018] Next, the second question generation unit 102 generates a second question based on the first answer (S2). Subsequently, the second answer acquisition unit 103 acquires the second answer to the second question (S3). Then, the determination unit 104 determines whether the first answer is appropriate based on the relationship between the first answer and the second answer (S4).

[0019] The dialogue assistance system 100 described herein generates a second question based on a first answer to a first question and obtains a second answer to the second question. The dialogue assistance system 100 can determine whether the first answer is appropriate based on the relationship between the first answer and the second answer. In this way, the dialogue assistance system 100 can appropriately assist the dialogue.

[0020] <Embodiment 2> Next, Embodiment 2 will be described. Embodiment 2 is a specific example of Embodiment 1 described above.

[0021] (Dialogue support system 1) Figure 3 is a block diagram showing the configuration of the dialogue assistance system 1 according to this disclosure. The dialogue assistance system 1 is an example of the dialogue assistance system 100 described above. The dialogue assistance system 1 comprises a speech recognition unit 11, an answer acquisition unit 12, a question generation unit 13, a determination unit 14, a text generation unit 15, a recognition result correction unit 16, a display unit 17, and a speech input unit 18.

[0022] Any speech recognition technology may be used in the speech recognition unit 11. Furthermore, technologies such as LLM may be used in the answer acquisition unit 12, question generation unit 13, judgment unit 14, text generation unit 15, and recognition result correction unit 16. The dialogue assistance system 1 can be implemented using, for example, a PC (Personal Computer), a tablet device, or a smartphone.

[0023] First, let me explain the overview of Dialogue Assistance System 1. Dialogue Assistance System 1 is a system that assists in dialogue between multiple people. Dialogue Assistance System 1 can be used in any environment where dialogue takes place. For example, Dialogue Assistance System 1 may be used in government, local authorities, public institutions, financial institutions, or medical institutions. However, it is not limited to these; Dialogue Assistance System 1 may also be used at public transportation ticket counters, reception desks in stores providing various services, and so on.

[0024] The dialogue assistance system 1 may be used in environments that do not have a designated contact point. For example, the dialogue assistance system 1 may be used in a reporting system that receives reports to designated agencies such as the police or fire department. The number of people who communicate using the dialogue assistance system 1 is arbitrary. In the following example, a conversation between two people is used, but there may be three or more people communicating.

[0025] The following explanation uses an example where Dialogue Assistance System 1 is used when a counter staff member receives a consultation from a person seeking advice at a government office counter. The counter staff member listens to the details of the consultation by engaging in a dialogue with the person seeking advice. The counter staff member is, for example, a government office employee. The person seeking advice is, for example, a resident living in the area. The counter staff member is an example of a questioner who asks questions generated by Dialogue Assistance System 1 to the other party. The person seeking advice is an example of a respondent who answers the questions received from the questioner. The people engaging in the dialogue are not limited to the counter staff member and the person seeking advice. Dialogue Assistance System 1 can also be used in places other than government offices, such as retail stores and banks.

[0026] The dialogue assistance system 1 acquires speech between the counter staff and the client using a voice input unit 18 located near the counter, and performs speech recognition on the speech. The dialogue assistance system 1 also generates a document recording the content of the consultation by converting the content of the conversation into text using artificial intelligence (AI) or the like. For example, the dialogue assistance system 1 uses LLM to create a report to report on the content of the consultation.

[0027] In the dialogue assistance system 1, if the speech recognition result of the utterance is incorrect, it is difficult for the LLM to create a report based on the correct dialogue content. Therefore, the counter staff needs to correct the report created by the LLM as needed. If there are many misrecognitions, the burden on the counter staff increases.

[0028] The dialogue assistance system 1 appropriately assists in the dialogue to confirm the content of utterances that may have been misrecognized in speech recognition. For example, the dialogue assistance system 1 generates questions to confirm proper nouns included in the caller's utterances and presents the generated questions to the counter staff. By having the counter staff ask questions to the caller according to the presented questions, the dialogue assistance system 1 determines whether the proper nouns included in the utterances were correctly recognized by speech recognition. The counter staff can then proceed with the dialogue by asking further questions based on the determination result. In this way, the dialogue assistance system 1 reduces the burden on the counter staff in correcting reports. The following describes each functional part of the dialogue assistance system 1 in detail.

[0029] The speech recognition unit 11 performs speech recognition on a dialogue between multiple people. The dialogue between multiple people includes questions and answers from each of the multiple people. The speech recognition unit 11 performs speech recognition on the utterances input from the speech input unit 18. The speech recognition unit 11 performs speech recognition on the answer to the first question and outputs the speech recognition result as the first answer to the answer acquisition unit 12. The speech recognition unit 11 also performs speech recognition on the answer to the second question and outputs the speech recognition result as the second answer to the answer acquisition unit 12.

[0030] The answer acquisition unit 12 is an example of the first answer acquisition unit 101 and the second answer acquisition unit 103 described above. The answer acquisition unit 12 (first answer acquisition unit) acquires the first answer to the first question. The answer acquisition unit 12 (second answer acquisition unit) acquires the second answer to the second question. The answer acquisition unit 12 outputs the acquired first answer to the question generation unit 13. The answer acquisition unit 12 also outputs the acquired second answer to the determination unit 14.

[0031] Here, we describe an example in which the answer acquisition unit 12 includes the functions of a first answer acquisition unit 101 and a second question generation unit 102, but it is not limited to this. The functions of the first answer acquisition unit 101 and the functions of the second question generation unit 102 may be provided separately.

[0032] The first question includes a question that requires the answer to provide a proper noun in the first answer. A proper noun is information that identifies the subject indicated in the first answer (hereinafter simply referred to as the "subject"). Examples of proper nouns include place names, personal names, facility names, building names, organization names, company names, or product names. However, proper nouns are not limited to these; any proper noun that identifies the subject indicated in the first answer is acceptable. For example, a proper noun could be a common name for a region, or the name of an area around a station represented by a station name.

[0033] The subject can be anything indicated by a proper noun. The subject may be a specific place or space, a specific item, a specific person or organization, etc. For example, the subject may be a place, facility, building, organization, company, or product corresponding to the place name, person's name, facility name, building name, organization name, company name, or product name mentioned above.

[0034] The first question might be something like, "Where is the location?" By asking such a question, the person seeking advice can provide a first answer that includes a proper noun. The first answer might be something like, "It's a park in Tamachi." Here, "Tamachi" refers to the area around "Tamachi Station" in Minato Ward, Tokyo.

[0035] The question generation unit 13 is an example of the second question generation unit 102 described above. The question generation unit 13 generates a second question based on the first answer.

[0036] As mentioned above, the first question includes a question to prompt the user to provide a proper noun in the first answer; therefore, in the following explanation, the first answer will be described as including a proper noun. Proper nouns are relatively likely to be misrecognized in speech recognition. For this reason, the question generation unit 13 generates a second question that includes a question to confirm the object indicated by the proper noun included in the first answer. The second question is to confirm that the proper noun included in the first answer was correctly recognized by speech recognition. The second question is different from the first question.

[0037] For example, suppose "Tamachi" is the location of Company X's headquarters. The question generation unit 13 generates a second question, such as "Is this the area where Company X's headquarters are located?", in response to the first answer, "It's a park in Tamachi."

[0038] For example, the dialogue assistance system 1 stores proper nouns and related terms for each target in advance as a proper noun database. The proper noun database identifies targets using target IDs and other identifiers.

[0039] For example, in the above example, the proper noun database stores the correspondence between "Tamachi" and "X Company's headquarters." The question generation unit 13 identifies the proper noun included in the first answer, refers to the proper noun database, and retrieves the term associated with that proper noun. The question generation unit 13 generates the second question using the retrieved term. In this way, the question generation unit 13 can incorporate "X Company's headquarters" into the second question.

[0040] The question generation unit 13 may generate a second question that includes a question for identifying the object indicated by the proper noun included in the first answer from among multiple target candidates. The proper noun DB identifies and manages each of multiple different objects that have the same proper noun by their target ID. The proper noun DB may treat multiple proper nouns as the same proper noun if they have the same pronunciation. The proper noun DB may store not only identical place names or area names, but also homophones that could be misrecognized.

[0041] For example, suppose the conversation is taking place in Tokyo, and the first response includes the proper noun "Chuo Ward." In this case, the question generation unit 13 extracts "Chuo Ward, Tokyo," "Chuo Ward, Chiba City," and "Chuo Ward, Saitama City," all located within the area that could be the subject of the conversation, as target candidates. The question generation unit 13 may include a question in the second question to specify which area the target is. The question generation unit 13 may extract multiple target candidates based on the location of the government office, the address of the person seeking advice, or the location information of the person seeking advice.

[0042] The question generation unit 13 may generate a second question that corresponds to the attributes of the person seeking advice (respondent). Attributes include, for example, the person seeking advice's age, gender, whether they are an adult or a child, and whether they are a traveler or not. Attributes may be input into the dialogue assistance system 1 by the counter staff via an input unit (not shown), or they may be input into the dialogue assistance system 1 through image recognition using an image of the person seeking advice. The question generation unit 13 may also add attributes as input to the LLM using techniques such as Retrieval Augmented Generation (RAG).

[0043] For example, if the person seeking advice is a child or a tourist, they may not know the correct answer to the question, "Is this the area where Company X's headquarters are located?" In such cases, the question generation unit 13 may generate a second question, such as, "Is this the station next to Station Y?" Note that the second question does not have to be a question that can be answered with "yes" or "no." For example, the question generation unit 13 may generate a second question such as, "Are there any company headquarters in Tamachi?" which would elicit the answer, "There is Company X's headquarters here."

[0044] The question generation unit 13 may repeatedly generate a second question in response to the repetition of questions and answers in the dialogue. The LLM learns from the question-answer patterns and can continuously improve the accuracy of the second question through this process.

[0045] The question generation unit 13 outputs the generated second question in a manner that is perceptible to the counter staff. For example, the question generation unit 13 displays the second question on the display unit 17. The question generation unit 13 may also transmit the second question to a terminal device used by the counter staff via a communication unit (not shown) and display the second question on the terminal device. The communication unit is a communication interface for communication by wire or wireless.

[0046] The question generation unit 13 may present the second question to the counter staff using any method. The question generation unit 13 may also convey the second question to the counter staff by voice via an intercom or the voice output unit of a tablet terminal. The person seeking advice will respond to the second question, "Is this the area where Company X's headquarters are located?" with a second answer, for example, "Yes, it is."

[0047] The determination unit 14 is an example of the determination unit 104 described above. The determination unit 14 determines whether the first answer is appropriate based on the relationship between the first answer and the second answer. For example, the determination unit 14 determines that the first answer is appropriate if the first answer and the second answer are related, and determines that the first answer is inappropriate if they are not related. "The first answer is appropriate" indicates that the first answer was correctly recognized by speech recognition.

[0048] The relationship between the first answer and the second answer may be determined by whether or not the second answer is positive. For example, if the second answer contains positive words such as "yes" or "that's right," the determination unit 14 determines that there is a relationship between the first answer and the second answer. Conversely, if the second answer contains negative words such as "no" or "that's wrong," the determination unit 14 determines that there is no relationship between the first answer and the second answer.

[0049] If the text generation unit 15 determines that the first response is appropriate, it generates text related to the dialogue based on the first response. For example, the text generation unit 15 generates text to be displayed on the display unit 17. The text generation unit 15 also creates a report recording the content of the dialogue.

[0050] The text generation unit 15 may generate text including supplementary information to supplement the content regarding the subject indicated by the proper noun included in the first response. For example, if the proper noun is "Chuo Ward" in Tokyo, the text generation unit 15 generates text such as "Chuo Ward (Tokyo)," associating the supplementary information "(Tokyo)" with the subject text "Chuo Ward." The text generation unit 15 generates text in a manner that allows the counter staff to identify the subject identified from multiple candidate subjects. The text generation unit 15 may indicate the supplementary information in parentheses as described above, or it may be highlighted with underlining or bolding.

[0051] The recognition result correction unit 16 corrects the misrecognition in the first answer based on the second answer. For example, in the above example, the consultant's answer to the first question was "Tamachi Park," but due to a misrecognition by the speech recognition unit 11, "Tamaki Park" was obtained as the first answer. The recognition result correction unit 16 may determine whether the first answer is appropriate or not based on the proper noun database or the location of the government office, and if it determines that it is inappropriate, it may correct the first answer.

[0052] For example, the recognition result correction unit 16 may have the question generation unit 13 generate a question to confirm whether "Tamaki" is "Tamachi". Based on the answer to the question, the recognition result correction unit 16 determines that "Tamaki" is correctly "Tamachi" and corrects the first answer to "It's a park in Tamachi".

[0053] The display unit 17 displays information related to the dialogue assistance system 1. The display unit 17 is a display device such as a liquid crystal display. The display unit 17 may also have the function of an input unit that receives input from the person engaging in dialogue. For example, the display unit 17 may be a touch panel on a tablet terminal. The display unit 17 may also be a display on a wearable device such as smart glasses.

[0054] The display unit 17 displays the questions generated by the question generation unit 13. The display unit 17 may also display the speech recognition results from the speech recognition unit 11. If the speech recognition results are corrected, the display unit 17 may also display the corrected content. The display unit 17 may also display the text generated by the text generation unit 15. This allows the display unit 17 to display the dialogue in chronological order. The display unit 17 may also include supplementary information in the text. This allows the counter staff to accurately understand the content of the dialogue.

[0055] The voice input unit 18 accepts voice input of a conversation between multiple people. The voice input unit 18 is, for example, a microphone. The voice input unit 18 is installed in a position where it can collect the conversation. For example, the voice input unit 18 is installed near the multiple people having the conversation. The voice input unit 18 may be installed within a predetermined distance from the multiple people.

[0056] The configuration of the dialogue assistance system 1 has been described above. Note that the configuration of the dialogue assistance system 1 described above is merely an example and can be modified as appropriate. For example, if some or all of the components of the dialogue assistance system 1 are implemented by multiple information processing devices or circuits, these devices may be centrally located or distributed. For example, the information processing devices or circuits may be implemented in a form where each is connected via a communication network, such as a client-server system or a cloud computing system. Furthermore, some of the functions of the dialogue assistance system 1 may be provided in SaaS (Software as a Service) format.

[0057] Furthermore, although the above explanation used an example where the dialogue takes place at a counter, the dialogue may also take place using various communication devices such as telephones. For example, the person seeking advice may make a phone call to the counter staff at the government office from their smartphone, and the dialogue between the person seeking advice and the counter staff may take place. Alternatively, a virtual assistant or chatbot may interact with the person seeking advice, either on behalf of or together with the counter staff. When using a virtual assistant or chatbot, their utterances may be output from at least one of the display unit 17 and the voice output unit, and may be translated into a foreign language as appropriate.

[0058] (Processing by Dialogue Assistance System 1) Next, we will explain the processes performed by the dialogue assistance system 1 with reference to Figures 4 and 5. Figure 4 is a flowchart showing the processes performed by the dialogue assistance system 1. Figure 5 is a diagram showing a specific example of dialogue in the dialogue assistance system 1. In the following explanation, we will use an example in which a person visiting a government office counter asks the counter staff for advice regarding playground equipment in a park.

[0059] As shown in Figure 5, a display unit 17 is provided near the counter staff member. A voice input unit 18 is also provided near the counter staff member and the person seeking advice, but it is not shown in Figure 5. The voice input unit 18 constantly accepts input of the conversation between the counter staff member and the person seeking advice and outputs the conversation to the voice recognition unit 11. The voice recognition unit 11 performs voice recognition on the conversation acquired from the voice input unit 18 and outputs the voice recognition result to the answer acquisition unit 12.

[0060] First, as shown in Figure 5, let's assume that the person seeking advice consults with the counter staff, saying, "The playground equipment in the park is aging, so I would like to discuss what to do about it in the future." The counter staff then asks the person seeking advice the first question, Q1, "Which park is it?" The first question, Q1, includes a question to elicit a proper noun in the first answer, A1.

[0061] Next, the person seeking advice gives the first answer A1, "It's the park in Tamachi," to the counter staff. As shown in Figure 4, the answer acquisition unit 12 acquires the first answer A1 to the first question Q1 (S11). The answer acquisition unit 12 outputs the first answer A1 to the question generation unit 13.

[0062] The question generation unit 13 generates a second question Q2 based on the first answer A1 (S12). Here, the question generation unit 13 generates a second question Q2, "Is this the area where Company X's headquarters are located?", which includes a question to confirm the subject indicated by the proper noun "Tamachi" included in the first answer A1.

[0063] The question generation unit 13 may generate a second question Q2 that includes a question to identify the target indicated by the proper noun "Tamachi" in the first answer A1 from among multiple target candidates. For example, if there are multiple "Tamachi" areas that are target candidates, the question generation unit 13 may generate a second question Q2 such as "Is it Tamachi in Minato Ward, Tokyo?". The question generation unit 13 may also generate a second question Q2 that is tailored to the respondent's attributes. For example, the question generation unit 13 may generate a second question Q2 that is tailored to the respondent's age. The question generation unit 13 outputs the generated second question Q2 to the display unit 17.

[0064] Next, the display unit 17 displays the second question Q2 (S13). The counter staff member looks at the display unit 17 and asks the second question Q2, "Is this the area where Company X's headquarters are located?" The person seeking advice gives the second answer A2, "Yes, it is." The answer acquisition unit 12 acquires the second answer A2 to the second question Q2 (S14). The answer acquisition unit 12 outputs the second answer A2 to the determination unit 14.

[0065] Furthermore, if in the first response A1 "It is Tamachi Park" is misrecognized as "It is Tamaki Park," the recognition result correction unit 16 may correct the misrecognition in the first response A1 based on the second response A2.

[0066] Next, the determination unit 14 determines whether the first answer A1 is appropriate based on the relationship between the first answer A1 and the second answer A2 (S15). For example, the determination unit 14 determines the relationship by determining whether the second answer A2 contains positive content. If the determination unit 14 determines that the second answer A2 contains positive content, it determines that the first answer A1 and the second answer A2 are related. If the determination unit 14 determines that the second answer A2 contains negative content, it determines that the first answer A1 and the second answer A2 are not related. If the determination unit 14 determines that the first answer A1 and the second answer A2 are related, it determines that the first answer A1 is appropriate.

[0067] Next, the determination unit 14 determines whether the first answer A1 is appropriate (S16). If the first answer A1 is determined to be inappropriate (NO in S16), the process returns to step S12 and the subsequent processing is repeated. If the first answer A1 is determined to be appropriate (YES in S16), the text generation unit 15 generates text based on the first answer A1 (S17). The text generation unit 15 may generate text that includes supplementary information about the area indicated by the proper noun "Tamachi" included in the first answer A1. For example, the text generation unit 15 generates text that is written as "Tamachi (Tokyo)". The text generation unit 15 may also display the generated text on the display unit 17.

[0068] As described above, the dialogue assistance system 1 in this disclosure obtains a first answer to a first question and generates a second question based on the first answer. The dialogue assistance system 1 obtains a second answer to the second question and determines whether the first answer is appropriate based on the relationship between the first and second answers. If the first answer is determined to be appropriate, the system generates and outputs a text related to the dialogue based on the first answer. With this configuration, the dialogue assistance system 1 can appropriately assist in dialogue.

[0069] <Example Hardware Configuration> Each functional component of the dialogue assistance system 1 may be implemented by hardware (e.g., hardwired electronic circuits) or by a combination of hardware and software (e.g., a combination of an electronic circuit and a program to control it). The following will further explain the case where each functional component of the dialogue assistance system 1 is implemented by a combination of hardware and software.

[0070] Figure 6 is a block diagram illustrating the hardware configuration of computer 900, which implements the dialogue assistance system 1. Computer 900 may be a dedicated computer designed to implement the dialogue assistance system 1, or it may be a general-purpose computer. Computer 900 may also be a portable computer such as a smartphone or tablet device.

[0071] For example, by installing a predetermined application on computer 900, each function of the dialogue assistance system 1 is realized on computer 900. The above application consists of a program for realizing the functional components of the dialogue assistance system 1.

[0072] Computer 900 includes a bus 902, a processor 904, memory 906, a storage device 908, an input / output interface 910, and a network interface 912. The bus 902 is a data transmission path for the processor 904, memory 906, storage device 908, input / output interface 910, and network interface 912 to send and receive data to and from each other. However, the method of connecting the processor 904 and other components to each other is not limited to bus connection.

[0073] The processor 904 is a variety of processors such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), or quantum processor (quantum computer control chip). The memory 906 is the main memory, implemented using RAM (Random Access Memory), etc. The storage device 908 is the auxiliary storage, implemented using a hard disk, SSD (Solid State Drive), memory card, or ROM (Read Only Memory), etc.

[0074] The input / output interface 910 is an interface for connecting the computer 900 with input / output devices. For example, input devices such as keyboards and output devices such as display devices are connected to the input / output interface 910. In addition, the audio input unit 18 is connected to the input / output interface 910.

[0075] The network interface 912 is an interface for connecting computer 900 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0076] The storage device 908 stores programs that implement each functional component of the dialogue assistance system 1 (programs that implement the aforementioned applications). The processor 904 reads these programs into memory 906 and executes them to implement each functional component of the dialogue assistance system 1.

[0077] Each processor executes one or more programs containing a set of instructions for causing the computer to perform the algorithms described with reference to the drawings. These programs, when loaded into the computer, contain a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The programs may be stored in various types of non-transitory computer-readable medium or tangible storage medium. Examples, but not limited to, include non-transitory computer-readable medium or tangible storage medium, such as RAM, ROM, flash memory, SSD or other memory technologies, CD-ROM, DVD (Digital Versatile Disc), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The programs may also be transmitted over various types of transient computer-readable medium or communication medium. Examples, but not limited to, include transient computer-readable medium or communication medium, such as electrical, optical, acoustic or other forms of propagating signals.

[0078] <Specific example> The program (interaction assistance program) relating to this disclosure may be implemented in the following manner, for example.

[0079] The program is installed, for example, on a mobile device. Alternatively, the program is installed on an external storage device (such as an on-premise server or cloud) that can exchange information with an information processing device installed on the mobile device. A mobile device includes, but is not limited to, automobiles, aircraft, or ships. Furthermore, the mobile device is not necessarily limited to those capable of carrying a person. For example, the program can be installed on unmanned aerial vehicles such as drones.

[0080] In this example, the car can, for instance, determine whether the driver's speech is appropriately recognized. Furthermore, the car can control its behavior based on the determination result. Controlling the car's behavior means performing some kind of control over the various devices that make up the car.

[0081] For example, when controlling the drive of an automobile, controlling the behavior of the automobile includes changing the output of the electric motor and controlling the gear ratio of the transmission. Also, when controlling electronic equipment inside the vehicle, controlling the behavior of the automobile includes, for example, controlling the operation of the car navigation system and controlling the output to air conditioning equipment, AV (Audio Visual) equipment, and seats. Furthermore, when applying this disclosure to an unmanned aircraft, the program may be installed on a controller for operating the unmanned aircraft.

[0082] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure are possible, as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0083] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments rather than with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps shown in any of the drawings may be changed as appropriate.

[0084] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, The computer is made to perform a determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer. Dialogue assistance program. (Note 2) If the first response is determined to be appropriate, the computer is instructed to perform a text generation step that generates text related to the dialogue based on the first response. The dialogue assistance program described in Appendix 1. (Note 3) The first question above includes a question to prompt the respondent to provide a proper noun in their answer to the first question above. Dialogue assistance programs as described in Appendix 1 or 2. (Note 4) The second question generation step generates the second question, which includes a question to confirm the subject indicated by the proper noun included in the first answer. A dialogue support program described in any one of the following appendices 1 to 3. (Note 5) In the second question generation step, the second question is generated, which includes a question for identifying the object indicated by the proper noun included in the first answer from among several candidate objects. A dialogue assistance program described in any one of the following appendices 1 to 4. (Note 6) In the second question generation step, the second question is generated according to the attributes of the respondent. A dialogue assistance program described in any one of the appendices 1 to 5. (Note 7) In the text generation step, the text is generated including supplementary information to supplement the content concerning the subject indicated by the proper noun included in the first answer. The dialogue assistance program described in Appendix 2. (Note 8) Based on the second answer, the computer is further instructed to perform a recognition result correction step to correct the misrecognition in the first answer. A dialogue support program described in any one of the appendices 1 to 7. (Note 9) A first answer acquisition unit that obtains a first answer to a first question, A second question generation unit that generates a second question based on the first answer, A second answer acquisition unit that obtains a second answer to the second question, The system includes a determination unit that determines whether the first answer is appropriate based on the relationship between the first answer and the second answer. Dialogue assistance system. (Note 10) The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, A determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer, is included. Methods for assisting dialogue.

[0085] Some or all of the elements (e.g., configuration and function) described in Appendices 2 to 8 that are dependent on Appendice 1 may also be dependent on Appendices 9 and 10 in the same manner as those described in Appendices 2 to 8. Some or all of the elements described in any appendice may be applicable to various hardware, software, recording means, systems, and methods for recording software. [Explanation of Symbols]

[0086] 1. Dialogue Assistance System 11. Voice Recognition Unit 12 Answer acquisition part 13 Question generation part 14 Judgment section 15 Sentence generation section 16 Recognition result correction section 17 Display section 18. Voice input section 100 Dialogue Assistance Systems 101 First Answer Acquisition Section 102 Second Question Generation Unit 103 Second response acquisition section 104 Judgment section 900 Computers 902 Bus 904 Processor 906 memory 908 Storage Devices 910 Input / Output Interface 912 Network Interface A1 First answer A2 Second answer Q1 First question Q2 Second question

Claims

1. The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, The computer is made to perform a determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer. Dialogue assistance program.

2. If the first response is determined to be appropriate, the computer is instructed to perform a text generation step that generates text related to the dialogue based on the first response. The dialogue assistance program according to claim 1.

3. The first question includes a question to prompt the respondent to provide a proper noun in their answer to the first question. The dialogue assistance program according to claim 1 or 2.

4. The second question generation step generates the second question, which includes a question to confirm the subject indicated by the proper noun included in the first answer. The dialogue assistance program according to claim 1 or 2.

5. In the second question generation step, the second question is generated, which includes a question for identifying the object indicated by the proper noun included in the first answer from among several candidate objects. The dialogue assistance program according to claim 1 or 2.

6. In the second question generation step, the second question is generated according to the attributes of the respondent. The dialogue assistance program according to claim 1 or 2.

7. In the text generation step, the text is generated including supplementary information to supplement the content concerning the subject indicated by the proper noun included in the first answer. The dialogue assistance program according to claim 2.

8. Based on the second answer, the computer is further instructed to perform a recognition result correction step to correct the misrecognition in the first answer. The dialogue assistance program according to claim 1 or 2.

9. A first answer acquisition unit that obtains a first answer to a first question, A second question generation unit that generates a second question based on the first answer, A second answer acquisition unit that obtains a second answer to the second question, The system includes a determination unit that determines whether the first answer is appropriate based on the relationship between the first answer and the second answer. Dialogue assistance system.

10. The first answer acquisition step involves obtaining the first answer to the first question, A second question generation step, which generates a second question based on the answer to the first question, A second answer acquisition step for obtaining a second answer to the second question, The process includes a determination step of determining whether the first answer is appropriate based on the relationship between the first answer and the second answer. Methods for assisting dialogue.

Citation Information

Patent Citations

  • Sound analysis device, sound analysis method, and recording medium

    WO2023112668A1