Dialogue support system, method and program

The dialogue support system addresses the challenge of multiple speakers by identifying the main speaker and maintaining context, ensuring accurate and contextually relevant responses in environments with multiple sound sources.

JP2026123565APending Publication Date: 2026-07-30CLASSIX CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CLASSIX CO LTD
Filing Date
2025-01-17
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Conventional voice response systems struggle with accuracy in environments with multiple sound sources due to their reliance on responding to a single speaker, leading to inaccurate responses when multiple speakers speak alternately.

Method used

A dialogue support system that identifies the main speaker, determines context, and prioritizes responses based on the main speaker's speech, while managing noise and secondary speakers' utterances through voice analysis and real-time context maintenance.

Benefits of technology

The system effectively detects dynamic speaker switching, eliminates noise, and provides contextually appropriate responses, enhancing accuracy and user experience in multi-speaker environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026123565000001_ABST
    Figure 2026123565000001_ABST
Patent Text Reader

Abstract

This system provides a dialogue support system that detects dynamic speaker switching, eliminates background noise, and enables responses that are in line with the context of the main speaker. [Solution] A dialogue support system used in a dialogue situation, comprising: a first step of determining whether a new speaker has spoken by voice analysis; a second step of determining whether the new speaker's speech is in line with the context of the dialogue; a third step of registering the new speaker if it is in line with the context; a fourth step of reflecting the content of the registered speaker's speech in the context; and a fifth step of responding in accordance with the speech of the registered speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dialogue support system, method, and program for supporting dialogue among multiple speakers.

Background Art

[0002] Conventionally, voice response systems have focused on extracting information from a single voice signal, and there is a problem that the accuracy significantly decreases in an environment where there are multiple sound sources. That is, since conventional voice response systems are premised on responding to a single speaker, responses tend to be inaccurate in situations where multiple speakers speak alternately.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to solve this problem, there is a need for a technology that identifies the main speaker from noise and the voices of multiple speakers and appropriately responds based on the speech content of the main speaker.

[0005] The present invention has been made in view of such circumstances, and its object is to provide a dialogue support system, method, and program that can detect dynamic speaker switching, eliminate noise, etc., and respond in accordance with the context of the main speaker.

Means for Solving the Problems

[0006] The present invention is a dialogue support system used in a situation in which multiple registered speakers are conversing, and the system performs the following steps: a first step of determining whether a new speaker has spoken by voice analysis; a second step of determining whether the new speaker's speech is in line with the context of the dialogue; a third step of registering the new speaker if it is in line with the context; a fourth step of reflecting the content of the registered speaker's speech in the context; and a fifth step of responding in accordance with the speech of the registered speaker.

[0007] Preferably, the process involves identifying the main speaker based on predetermined conditions, and determining the context based on the content of the main speaker's speech.

[0008] Preferably, the first speaker is identified as the main speaker.

[0009] Preferably, the context is determined by prioritizing the content of the utterance of the main speaker.

[0010] Preferably, responses to the utterance of the main speaker are given priority over responses to utterances of speakers other than the main speaker.

[0011] Preferably, the system maintains the content of the main speaker's utterance as context, while analyzing the content of other speakers' utterances and generating an appropriate response in real time.

[0012] Preferably, it is determined whether the utterances of speakers other than the main utterancer maintain the context, and if it is determined that the context is maintained, it is used in determining the context.

[0013] Preferably, the characteristics of the speaker's voice signal, volume, timing of speech, or voiceprint analysis are used to determine whether or not the new speaker has spoken.

[0014] Preferably, the dialogue takes place in real time.

[0015] It is particularly suitable for a wide range of multi-speaker applications, including conference systems, customer service, home AI assistants, educational systems, and meeting minute-taking systems.

[0016] The present invention relates to a dialogue support system used in a situation in which multiple registered speakers are conversing, and is a dialogue support method in which a computer performs the following steps: a first step of determining whether a new speaker has spoken by voice analysis; a second step of determining whether the new speaker's speech is in line with the context of the dialogue; a third step of registering the new speaker if it is in line with the context; a fourth step of reflecting the content of the registered speaker's speech in the context; and a fifth step of responding in accordance with the speech of the registered speaker.

[0017] The present invention relates to a dialogue support system used in a situation in which multiple registered speakers are conversing, and is a program that causes a computer to perform the following steps: a first step of determining whether a new speaker has spoken by voice analysis; a second step of determining whether the new speaker's speech is in line with the context of the dialogue; a third step of registering the new speaker if it is in line with the context; a fourth step of reflecting the content of the registered speaker's speech in the context; and a fifth step of responding in accordance with the speech of the registered speaker. [Effects of the Invention]

[0018] According to the present invention, it is possible to provide a dialogue support system, method, and program that can detect dynamic speaker switching, eliminate noise and other unwanted sounds, and provide responses that are in line with the context of the main speaker. [Brief explanation of the drawing]

[0019] [Figure 1] Figure 1 is a diagram showing the configuration of a communication environment in which a dialogue support system 1 according to an embodiment of the present invention is employed. [Figure 2] Figure 2 is a functional block diagram of the dialogue support system 1 shown in Figure 1. [Figure 3]FIG. 3 is a flowchart for explaining the dialogue support using the settlement system 1 according to the embodiment of the invention. [Figure 4] FIG. 4 is a flowchart for explaining the process of reflecting the utterance of the speaker in the context in the dialogue support system 1 shown in FIG. 1. [Figure 5] FIG. 5 is a functional block diagram of the speaker terminal device 4 shown in FIG. 1. [Figure 6] FIG. 6 is a functional block diagram of the dialogue support system 1 shown in FIG. 1.

Embodiments for Carrying Out the Invention

[0020] Hereinafter, a dialogue support system according to an embodiment of the present invention will be described. FIG. 1 is a configuration diagram of a communication environment in which the dialogue support system 1 according to the embodiment of the present invention is employed. As shown in FIG. 1, the dialogue support system 1 communicates with a plurality of speaker terminal devices 4 via a network 9. The dialogue support system 1 is applied in a wide range of scenes involving multiple speakers, such as a conference system, customer service, a home AI assistant, and an education system.

[0021] FIG. 2 is a functional block diagram of the dialogue support system 1 shown in FIG. 1. As shown in FIG. 2, the dialogue support system 1 includes a speaker determination function, a speaker registration function, a function of generating a context based on the utterance content, and a response function to the utterance.

[0022] The speaker determination function is a function of analyzing the voice of the speaker received from the speaker terminal device 4 to determine the speaker. For example, by analyzing the voice signal of the speaker, a plurality of speakers are distinguished. For identification, for example, voiceprint feature analysis, acoustic characteristics (pitch, volume, etc.) of the utterance content, patterns of utterance start timing and voice intervals, etc. are used.

[0023] The speaker registration function is a function of registering the speaker determined by the speaker determination function under predetermined conditions. The speaker switching detection function detects the timing when the primary speaker stops speaking and another speaker begins speaking, based on the speaker's voice data. This is achieved through the identification of continuous speech with short intervals between utterances, spatial localization based on the source of the sound (multiple microphone placement), and real-time matching with voiceprint data for each specific speaker. At this time, the speaker's characteristics are identified using the characteristics of the speaker's voice signal, volume, timing of speech, or voiceprint analysis.

[0024] The function that generates context based on utterances analyzes the utterances of registered speakers and reflects them in the context of conversations involving multiple speakers. The speech response function is a feature that responds to speech based on the content of the registered speaker's speech. For example, it can answer questions asked by the speaker. Furthermore, the decision of whether or not to answer questions from speakers other than the main speaker is also made, taking into consideration the context that primarily determined the content of the main speaker's utterance. Furthermore, as an AI response generation function, the dialogue support system 1 generates appropriate responses using an AI model, taking into account speaker switching. This includes analyzing the context, extracting highly relevant information, and processing statements from non-primary speakers independently if they cover new topics.

[0025] The dialogue in dialogue support system 1 is, for example, in real time. Dialogue support system 1 determines the context by prioritizing the content of the main speaker's utterance. Furthermore, the dialogue support system 1 prioritizes responses to the primary speaker's utterances over responses to utterances by speakers other than the primary speaker. Furthermore, the dialogue support system 1 retains the content of the main speaker's utterance as context, and when other speakers speak, it analyzes their utterances and generates appropriate responses in real time.

[0026] The following will explain dialogue support using dialogue support system 1. Figure 3 is a flowchart illustrating dialogue support using the payment system 1 according to an embodiment of the invention. Let's explain each step.

[0027] Step ST11: When the dialogue support system 1 receives a speech signal from the speaker terminal device 4, it proceeds to step ST12.

[0028] Step ST12: The dialogue support system 1 determines whether the speaker is already registered or not. If it determines that the speaker is not registered, it proceeds to step ST13; if it determines that the speaker is registered, it proceeds to step ST14.

[0029] Step ST13: The dialogue support system 1 registers the speaker as the primary speaker. At this time, the dialogue support system 1 stores the voice analysis data of the speaker.

[0030] Step ST14: Dialogue support system 1 determines, based on context, whether the speaker is a new speaker or not. In other words, dialogue support system 1 analyzes the speaker's voice and determines whether the content is consistent with the context associated with previous main speakers.

[0031] If the dialogue support system 1 determines that a new speaker is present, it registers them as a secondary speaker. Step ST15: The dialogue support system 1 determines whether the speaker is new or not based on the speech signal received in step ST11. Specifically, it analyzes the audio of the speech signal and makes a determination based on whether it matches an already registered speaker, etc. If the dialogue support system 1 determines that the speaker is a new speaker, it proceeds to step ST16.

[0032] Step ST16: The dialogue support system 1 registers the speaker as a secondary speaker. At this time, the dialogue support system 1 stores the speaker's voice analysis data.

[0033] The following explains how the utterance of a speaker is reflected in the context in Dialogue Support System 1. Figure 4 is a flowchart illustrating the process of reflecting the speaker's utterance in context in the dialogue support system 1 shown in Figure 1. Let's explain each step.

[0034] Step ST21: When the dialogue support system 1 receives a speech signal from the speaker terminal device 4, it proceeds to step ST22.

[0035] Step ST22: The dialogue support system 1 determines whether the utterance received in step ST21 is the utterance of the main speaker. This determination is made using voice analysis, etc. If the dialogue support system 1 determines that the utterance is from the main speaker, it proceeds to step ST23; otherwise, it proceeds to step ST24.

[0036] Step ST23: Dialogue support system 1 reflects the speaker's utterance in the context of the ongoing conversation.

[0037] Step ST24: The dialogue support system 1 reverses whether to reflect the utterance of the subordinate speaker in context. Specifically, it determines whether the utterance is in context or not, and if it is, proceeds to step ST23; otherwise, terminates the process.

[0038] The configuration of the speaker terminal device 4 shown in Figure 1 will be described below. Figure 5 is a functional block diagram of the speaker terminal device 4 shown in Figure 1. As shown in Figure 5, the speaker terminal device 4 includes, for example, a display 51, a camera 52, an operation unit 53, a communication unit 55, a microphone 57, a memory 59, and a processing unit 61.

[0039] The display 51 displays an image based on the signal from the processing unit 61. For example, it displays images of the speakers. Camera 52 captures an image of the object to be photographed. For example, it can capture the speaker's face. The operation unit 53 is an operating means such as a touch panel, keyboard, or mouse. The communication unit 55 communicates with other speaker terminal devices 4 and the dialogue support system 1 using voice and images. Mike 57 inputs the speaker's voice. Memory 59 stores the program that the processing unit 61 will execute. The processing unit 61 executes the program PRG1 stored in the memory 59 to perform the processing of the speaker terminal device 4 as defined in this embodiment.

[0040] Figure 6 is a functional block diagram of the dialogue support system 1 shown in Figure 1. As shown in Figure 6, the dialogue support system 1 includes, for example, a communication unit 75, an input unit 57, a memory 59, and a processing unit 61.

[0041] The communication unit 75 communicates with the speaker terminal device 4. The input section 77 is a terminal or similar device for inputting data from an external source. Memory 79 stores the program that the processing unit 81 will execute. The processing unit 81 executes the program PRG2 stored in the memory 79 to perform the processing of the dialogue support system 1 as defined in this embodiment.

[0042] Dialogue support system 1 can be applied to cases such as the following: Case 1: System demonstration in a conference room (1) The main speaker makes a statement about the topic. (2) If the subordinate speaker asks a related question, the system will respond appropriately based on the context of the topic, ensuring that the response is not out of context.

[0043] Case 2: Use as a home assistant This corresponds to a situation in the home where the main speaker checks the schedule and the secondary speaker asks for more details. (1) Subordinate speaker: "What's the weather like today?" (2) Main speaker: "What is the probability of rain?" (3) System response: "Today's weather is cloudy. There is a 50% chance of rain." Prioritizing responses to the main speaker.

[0044] As explained above, the dialogue support system 1 can detect dynamic speaker switching with high accuracy, eliminate noise and other unwanted elements, and enable responses that are in line with the context of the main speaker. When there are multiple speakers, it can accurately maintain the context of the main speaker while generating responses that correspond to speaker switching, contributing to improved work efficiency and user experience.

[0045] The present invention is not limited to the embodiments described above. In other words, those skilled in the art may make various modifications, combinations, subcombinations, and substitutions with respect to the components of the embodiments described above, within the technical scope of the present invention or its equivalents.

[0046] Furthermore, the following techniques may be used to detect changes in the speaker. Speech analysis: This involves analyzing the characteristics of speech to identify the patterns and features of the speaker's voice. This includes the frequency spectrum, pitch, and rhythm of the speech. Machine Learning: A machine learning model is trained using a large amount of audio data to detect changes in the speaker. This allows the model to learn the speaker's individuality and identify speech in new audio data. Dictionary Method: This method uses a pre-created speaker dictionary to compare audio data and identify speakers. It matches new audio data with audio samples of speakers included in the dictionary. Cross-coupled training: This method uses audio data from multiple speakers to perform cross-coupled training and detect speaker changes. This technique trains a model using audio data from different speakers to identify speaker characteristics. [Industrial applicability]

[0047] This invention is applicable to systems in which multiple speakers interact via a network. [Explanation of symbols]

[0048] 1…Dialogue support system 4…Speaker terminal device

Claims

1. A dialogue support system used in situations where multiple registered speakers are having a conversation, The first step involves determining whether a new speaker has spoken through voice analysis, A second step of determining whether the utterance of the new speaker is in accordance with the context of the dialogue, If the context is appropriate, the third step is to register the new speaker, A fourth step involves reflecting the content of the registered speaker's speech in the context, A fifth step involves responding to the speech of the registered speaker, A dialogue support system that performs the following.

2. The process involves identifying the primary speaker based on predetermined conditions, The process of determining the context based on the content of the utterance of the main speaker. The dialogue support system according to claim 1, which performs the following:

3. Identify the first utterance as the primary utterance. The dialogue support system according to claim 2.

4. The context is determined by prioritizing the content of the utterance of the main speaker. The dialogue support system according to claim 2.

5. Prioritize responses to the utterances of the primary speaker over responses to utterances of other speakers. The dialogue support system according to claim 1.

6. While retaining the content of the main speaker's utterance as context, the system analyzes the content of other speakers' utterances and generates appropriate responses in real time. The dialogue support system according to claim 1.

7. Determine whether the utterances of speakers other than the main utterancer maintain the context. On the condition that the aforementioned context is maintained, the following is used to determine the aforementioned context. The dialogue support system according to claim 1.

8. The characteristics of the speaker's voice signal, volume, timing of speech, or voiceprint analysis are used to determine whether or not the new speaker has spoken. The dialogue support system according to claim 7.

9. The aforementioned dialogue takes place in real time. The dialogue support system according to claim 1.

10. It can be applied to a wide range of scenarios involving multiple speakers, such as conference systems, customer service, home AI assistants, educational systems, and meeting minute creation systems. The dialogue support system according to claim 1.

11. A dialogue support system used in situations where multiple registered speakers are having a conversation, The first step involves determining whether a new speaker has spoken through voice analysis, A second step is to determine whether the utterance of the new speaker is in accordance with the context of the dialogue, If the context is appropriate, the third step is to register the new speaker, A fourth step involves reflecting the content of the registered speaker's speech in the context, A fifth step involves responding to the speech of the registered speaker, A method of providing interactive support that a computer can perform.

12. A dialogue support system used in situations where multiple registered speakers are having a conversation, The first step involves determining whether a new speaker has spoken through voice analysis, A second step is to determine whether the utterance of the new speaker is in accordance with the context of the dialogue, If the context is appropriate, the third step is to register the new speaker, A fourth step involves reflecting the content of the registered speaker's speech in the context, A fifth step involves responding to the speech of the registered speaker, A program that causes a computer to execute something.