Terminal device and information output method

The terminal device addresses the challenge of assisting conversations by generating and outputting explanations during silent periods or detecting non-listener statements, improving conversation flow and user engagement.

JP2026054173APending Publication Date: 2026-03-26JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing technologies fail to appropriately assist conversations by providing explanations for unknown words during discussions, potentially disrupting the flow and making the other party uncomfortable, and may cause the user to miss the content of the conversation.

Method used

A terminal device equipped with audio data acquisition, determination, and output control units that generate and output audio information when silent periods exceed a threshold or detect filler sounds or statements from others, assisting the user in understanding the conversation.

Benefits of technology

The device effectively supports conversations by providing timely explanations, enhancing user engagement and ensuring the user does not miss important dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026054173000001_ABST
    Figure 2026054173000001_ABST
Patent Text Reader

Abstract

To appropriately assist in conversation. [Solution] The terminal device comprises: an audio data acquisition unit that acquires audio data indicating voice; an audio determination unit that determines whether the length of the silent period during which the audio data acquisition unit does not acquire audio data is longer than a predetermined length; an information generation unit that generates output information if it is determined that the length of the silent period is longer than a predetermined length; and an output control unit that causes the audio output information generated by the information generation unit to be output from the output unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a terminal device and an information output method.

Background Art

[0002] Techniques for providing information to users are known. For example, Patent Document 1 discloses a technique for providing recommended items to a user based on keywords included in a conversation.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For example, an unknown word may come up during a conversation with another user. In such a case, it is possible to search for information about the unknown topic using the search function of a smartphone, but searching during a conversation may make the other party feel uncomfortable. Also, even when a function for assisting the conversation content is available, if one focuses too much on the assisting function, one may miss the content of the other party's conversation. There is a need to assist the conversation by outputting an explanation of an unknown word at an appropriate timing.

[0005] An object of the present disclosure is to provide a terminal device and an information output method that can appropriately assist a conversation.

Means for Solving the Problems

[0006] The terminal device of this disclosure includes: an audio data acquisition unit that acquires audio data indicating voice; an audio determination unit that determines whether the length of the silent period during which the audio data acquisition unit does not acquire the audio data is longer than a predetermined length; an information generation unit that generates audio output information if it is determined that the length of the silent period is longer than a predetermined length; and an output control unit that causes the audio output information generated by the information generation unit to be output from an output unit.

[0007] The information output method disclosed herein includes the steps of: acquiring audio data indicating sounds emitted around the user; determining whether the length of the silent period during which the audio data is not acquired is longer than a predetermined period; if it is determined that the period during which the audio data is not acquired is longer than a predetermined period, generating audio output information; and causing the generated audio output information to be output from the output unit. [Effects of the Invention]

[0008] According to this disclosure, it is possible to appropriately assist in conversation. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a schematic diagram showing an example of a terminal device according to the first embodiment. [Figure 2] Figure 2 is a block diagram showing an example configuration of a terminal device according to the first embodiment. [Figure 3] Figure 3 is a flowchart showing the flow of audio output processing according to the first example of the first embodiment. [Figure 4] Figure 4 is a flowchart showing the flow of audio output processing according to a second example of the first embodiment. [Figure 5] Figure 5 is a flowchart showing the flow of audio output processing according to a third example of the first embodiment. [Figure 6] Figure 6 is a schematic diagram showing an example of a terminal device according to the second embodiment. [Figure 7] Figure 7 is a block diagram showing an example configuration of a terminal device according to the second embodiment. [Figure 8] Figure 8 is a diagram illustrating the averaging method according to the second embodiment. [Figure 9] Figure 9 shows the detection results of the event-related potential P300 according to the second embodiment. [Figure 10] Figure 10 is a flowchart showing the flow of audio output processing according to the first example of the second embodiment. [Figure 11] Figure 11 is a flowchart showing the flow of audio output processing according to a second example of the second embodiment. [Figure 12] Figure 12 is a flowchart showing the flow of audio output processing according to a third example of the second embodiment. [Figure 13] Figure 13 is a flowchart showing the flow of audio output processing according to the fourth example of the second embodiment. [Modes for carrying out the invention]

[0010] Embodiments relating to this disclosure will be described in detail below with reference to the attached drawings. However, this embodiment does not limit this disclosure, and in the following embodiments, the same parts are denoted by the same reference numerals to avoid redundant explanations.

[0011] [First Embodiment] The terminal device according to the first embodiment will be described using Figure 1. Figure 1 is a schematic diagram showing an example of the terminal device according to the first embodiment.

[0012] As shown in Figure 1, the terminal device 10 is a headphone-type device worn by the user on their head. The user, for example, converses with another user while wearing the terminal device 10. The terminal device 10 outputs audio data to the user wearing the terminal device 10 to assist in the conversation, based on the content of the conversation between the user and the other user. In the following description, the terminal device 10 is described as a headphone-type device, but this disclosure is not limited to this. The terminal device 10 may be, for example, an earphone-type device or a head-mounted display-type device.

[0013] (Terminal device) Using FIG. 2, a configuration example of the terminal device according to the first embodiment will be described. FIG. 2 is a block diagram showing a configuration example of the terminal device according to the first embodiment.

[0014] As shown in FIG. 2, the terminal device 10 includes an input unit 12, a microphone 14, a speaker 16, a communication unit 18, a storage unit 20, and a control unit 22.

[0015] The input unit 12 receives various input operations to the terminal device 10. The input unit 12 is composed of various input devices such as buttons and switches, for example.

[0016] The microphone 14 detects sounds around the terminal device 10. The microphone 14 detects, for example, sounds emitted around the user of the terminal device 10. The microphone 14 detects, for example, the sounds of conversations taking place around the terminal device 10. The microphone 14 converts the detected sound into audio data and outputs it to the control unit 22, for example.

[0017] The speaker 16 outputs various sounds to the user of the terminal device 10. The speaker 16 is worn on the ear of the user of the terminal device 10.

[0018] The communication unit 18 is a communication interface that executes communication between the terminal device 10 and an external device. The communication unit 18 executes communication, for example, between the terminal device 10 and another device such as a smartphone held by the user of the terminal device 10. The communication unit 18 may execute communication between the terminal device 10 and a server device storing various information, for example.

[0019] The memory unit 20 stores various types of information. The memory unit 20 stores the calculation contents of the control unit 22 and information such as programs. The memory unit 20 includes, for example, at least one of the following: RAM (Random Access Memory), main memory such as ROM (Read Only Memory), and external memory such as HDD (Hard Disk Drive).

[0020] The control unit 22 controls each part of the terminal device 10. The control unit 22 includes, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as RAM or ROM. The control unit 22 executes a program that controls the operation of the terminal device 10 according to this disclosure. The control unit 22 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 22 may be implemented by a combination of hardware and software.

[0021] The control unit 22 comprises an audio data acquisition unit 30, an audio determination unit 32, an information generation unit 34, and an output control unit 36.

[0022] The voice data acquisition unit 30 acquires voice data indicating sounds emitted around the user of the terminal device 10. For example, the voice data acquisition unit 30 acquires sounds emitted by other users when the user of the terminal device 10 is having a conversation with another user. For example, the voice data acquisition unit 30 controls the microphone 14 to detect sounds emitted around the user of the terminal device 10 and acquires voice data indicating the detected sounds from the microphone 14. For example, the voice data acquisition unit 30 continuously acquires voice data.

[0023] The voice determination unit 32 performs various determination processes on the voice data acquired by the voice data acquisition unit 30. For example, the voice determination unit 32 determines whether the length of the silent period during which the voice data acquisition unit 30 has not acquired voice data is longer than a predetermined length. In other words, the voice determination unit 32 determines whether the conversation has been paused for a predetermined period of time or longer. Here, the silent period may be defined as the period during which the volume level is determined to be below a certain level. Alternatively, the silent period may be defined as the period during which the speech was not recognized as speech by the speech recognition process. The length at which the length of the silent period is determined to be longer than a predetermined length may be set arbitrarily.

[0024] The voice determination unit 32 may, for example, determine whether the voice data acquired by the voice data acquisition unit 30 contains filler sounds such as "um" or "well." Here, in order to determine whether or not filler sounds are included, it is preferable to store information about filler sounds in the storage unit 20 and use it to determine whether or not it is a filler sound. This information about filler sounds may be at least one of the following: voice data of the filler sound, frequency information or feature quantities of frequency information, or text information of the filler sound. The voice determination unit 32 may, for example, determine whether or not sounds indicating the end of a statement, such as "That's all," are included. The voice determination unit 32 may, for example, determine whether or not sounds soliciting opinions, such as "Do you have any comments?" In this case, the voice determination unit 32 may exclude from the determination any voice data that indicates questions posed to the user of the terminal device 10 or statements made by the user of the terminal device 10. The voice determination unit 32, for example, performs voice recognition processing on the voice data acquired by the voice data acquisition unit 30 to convert it into text information, and uses the information about filler voices stored in the storage unit 20 to determine whether or not filler voices are included.

[0025] The voice determination unit 32 may, for example, determine whether the voice data acquired by the voice data acquisition unit 30 is voice data of a person who spoke. The voice determination unit 32 may, for example, determine whether the voice data acquired by the voice data acquisition unit 30 is voice data of a person who spoke to the user of the terminal device 10. The voice determination unit 32 may, for example, perform sound source localization on the voice data acquired by the voice data acquisition unit 30 to identify the direction from which the sound was emitted, and determine whether the person located in the identified direction is the listener. The voice determination unit 32 may, for example, store voice data such as the voiceprint of the listener in the memory unit 20, and by comparing the voice data acquired by the voice data acquisition unit 30 with the voice data stored in the memory unit 20, determine whether the voice data acquired by the voice data acquisition unit 30 is voice data of a person who spoke. For example, electroencephalogram (EEG) data showing the user's brainwaves when audio data is emitted can be acquired, and the user's level of attentiveness can be determined based on the EEG data for a predetermined interval that includes the interval in which the event-related potentials described later occurred. Specifically, if the P300 event-related potential described later is detected, it can be determined that the level of attentiveness is above a predetermined level. Audio data that has been determined to have a level of attentiveness above a predetermined level can then be determined to be audio data emitted by a listener that the user is listening to attentively.

[0026] The information generation unit 34 generates output information to assist the user conversation of the terminal device 10. For example, the information generation unit 34 generates voice output information when the voice determination unit 32 determines that the length of the silent period is longer than a predetermined length. For example, the information generation unit 34 generates voice output information when the voice determination unit 32 determines that the voice data acquired by the voice data acquisition unit 30 contains filler voice and that the length of the silent period is longer than a predetermined length. For example, the information generation unit 34 generates voice output information when the voice determination unit 32 determines that the voice data acquired by the voice data acquisition unit 30 is voice data of speech spoken by someone other than the listener. For example, the information generation unit 34 generates explanatory information as voice output information that explains the content of the voice data acquired by the voice data acquisition unit 30. For example, the information generation unit 34 generates explanatory information based on the content of the conversation immediately preceding the determination that the silent period was longer than a predetermined length. The information generation unit 34 generates, for example, a summary of explanatory information that explains the content of the audio data acquired by the audio data acquisition unit 30, as audio output information.

[0027] The output information generated by the information generation unit 34 is not limited to audio output information. The information generation unit 34 may also generate text information and image information as output information, for example.

[0028] The output control unit 36 ​​causes the output information generated by the information generation unit 34 to output. For example, the output control unit 36 ​​causes the audio output information generated by the information generation unit 34 to output from the speaker 16. For example, the output control unit 36 ​​causes the audio output information generated by the information generation unit 34 to output from the speaker 16 when it is determined that the conversation has been interrupted. For example, the output control unit 36 ​​causes the audio output information generated by the information generation unit 34 to output from the speaker 16 when filler speech is detected and it is determined that the conversation has been interrupted. For example, the output control unit 36 ​​causes the audio output information generated by the information generation unit 34 to output from the speaker 16 when it is detected by the terminal device 10 that the utterance is from someone other than the listener whose speech the terminal device 10 is listening to.

[0029] The output control unit 36 ​​may, for example, transmit the text information or image information generated by the information generation unit 34 to a smartphone or other device held by the user of the terminal device 10 via the communication unit 18.

[0030] (Audio output processing) (Example 1) The audio output processing according to the first example of the first embodiment will be explained using Figure 3. Figure 3 is a flowchart showing the flow of the audio output processing according to the first example of the first embodiment.

[0031] Figure 3 shows, for example, the flow of processing performed by terminal device 10 after a conversation has started around the user of terminal device 10.

[0032] The audio data acquisition unit 30 acquires audio data indicating the sound detected by the microphone 14 (step S10). Specifically, the audio data acquisition unit 30 continuously acquires audio data indicating the sound detected by the microphone 14.

[0033] The sound determination unit 32 determines whether the length of the silent period is longer than or equal to a predetermined length (step S12). If it is determined that the length of the silent period is longer than or equal to a predetermined length (step S12; Yes), the process proceeds to step S14. If it is determined that the length of the silent period is not longer than or equal to a predetermined length (step S12; No), the process proceeds to step S18.

[0034] If the result in step S12 is determined to be Yes, the information generation unit 34 generates audio output information to be output to the user of the terminal device 10 (step S14).

[0035] The output control unit 36 ​​causes the audio output information generated by the information generation unit 34 to be output from the speaker 16 (step S16).

[0036] The control unit 22 determines whether or not to terminate the process (step S18). The control unit 22 determines to terminate the process, for example, when the user conversation on the terminal device 10 ends or when it receives an operation to turn off the power of the terminal device 10. If it is determined that the process should be terminated (step S18; Yes), the process shown in Figure 3 is terminated. If it is not determined that the process should be terminated (step S18; No), the process proceeds to step S10.

[0037] As described above, in the first example of the first embodiment, when the silent period in the conversation exceeds a predetermined length, audio output information is output to the user. This allows the first example of the first embodiment to appropriately support the conversation.

[0038] (Second example) Using Figure 4, the audio output processing according to the second example of the first embodiment will be explained. Figure 4 is a flowchart showing the flow of the audio output processing according to the second example of the first embodiment.

[0039] Figure 4 shows, for example, the flow of processing performed by terminal device 10 after a conversation has started around the user of terminal device 10.

[0040] The process in step S30 is the same as the process in step S10 shown in Figure 3, so the explanation will be omitted.

[0041] The voice determination unit 32 determines whether or not filler voice is included in the voice data acquired by the voice data acquisition unit 30 (step S32). If it is determined that filler voice is included in the voice data acquired by the voice data acquisition unit 30 (step S32; Yes), the process proceeds to step S34.

[0042] The processes from steps S34 to S40 are the same as the processes from steps S12 to S18 shown in Figure 3, so their explanation is omitted. Note that the process in step S34 in which the sound determination unit 32 determines whether the length of the silent period is longer than a predetermined value may be omitted.

[0043] As described above, in the second example of the first embodiment, when filler voice is detected and the silence period in the conversation is longer than a predetermined amount, voice output information is output to the user. This allows the second example of the first embodiment to appropriately support the conversation.

[0044] (Third example) The audio output processing according to the third example of the first embodiment will be explained using Figure 5. Figure 5 is a flowchart showing the flow of the audio output processing according to the third example of the first embodiment.

[0045] Figure 5 shows, for example, the flow of processing performed by terminal device 10 after a conversation has started around the user of terminal device 10.

[0046] The process in step S50 is the same as the process in step S10 shown in Figure 3, so the explanation will be omitted.

[0047] The voice determination unit 32 determines whether the voice data acquired by the voice data acquisition unit 30 is voice data spoken by someone other than the listener (step S52). If the voice data acquired by the voice data acquisition unit 30 is determined to be voice data spoken by someone other than the listener (step S52; Yes), the process proceeds to step S54. If the voice data acquired by the voice data acquisition unit 30 is not determined to be voice data spoken by someone other than the listener (step S52; No), the process proceeds to step S58.

[0048] The processes from step S54 to step S58 are the same as the processes from step S14 to step S18 shown in Figure 3, so their explanation will be omitted.

[0049] As described above, in the third example of the first embodiment, when a statement by someone other than the listener is detected, audio output information is output to the user. This allows the third example of the first embodiment to appropriately support the conversation.

[0050] [Second Embodiment] The terminal device according to the second embodiment will be described using Figure 6. Figure 6 is a schematic diagram showing an example of the terminal device according to the second embodiment.

[0051] As shown in Figure 6, terminal device 10A differs from terminal device 10 shown in Figure 1 in that it is equipped with multiple sensors 24. Multiple sensors 24 are provided so as to contact the user's head from the top to the sides when the user wears terminal device 10A on their head. Sensors 24 are electroencephalogram (EEG) sensors that detect the user's brain waves. Sensors 24 may be, for example, EEG sensors formed in the shape of a pad.

[0052] In the following description, sensor 24 is assumed to be an electroencephalogram (EEG) sensor, but this disclosure is not limited thereto. Sensor 24 may include various biosensors, such as a body temperature sensor, a pulse sensor, a blood pressure sensor, and a respiration sensor. Sensor 24 may also include, for example, an accelerometer for detecting a predetermined action of the user of terminal device 10A.

[0053] (Terminal device) An example of the configuration of a terminal device according to the second embodiment will be explained using Figure 7. Figure 7 is a block diagram showing an example of the configuration of a terminal device according to the second embodiment.

[0054] As shown in Figure 7, terminal device 10A differs from terminal device 10 shown in Figure 2 in that it is equipped with a sensor 24 and the control unit 22A is equipped with an electroencephalogram data acquisition unit 38, a listening level determination unit 40, and a comprehension level determination unit 42.

[0055] The electroencephalogram (EEG) data acquisition unit 38 acquires EEG data showing the brainwaves of the user wearing the terminal device 10A. The EEG data acquisition unit 38, for example, controls the sensor 24 to detect the brainwaves of the user wearing the terminal device 10, and acquires EEG data showing the detected brainwaves from the sensor 24. The EEG data acquisition unit 38 continuously acquires EEG data, for example.

[0056] The listening level determination unit 40 determines the listening level of the voices emitted around the user of the terminal device 10A. The listening level determination unit 40 determines the listening level of conversations taking place around the user of the terminal device 10A. The listening level determination unit 40 determines the listening level of the user of the terminal device 10A based, for example, on the electroencephalogram data acquired by the electroencephalogram data acquisition unit 38.

[0057] The listening level determination unit 40 determines the listening level of the user of the terminal device 10A based on electroencephalogram (EEG) data for a predetermined interval that includes the interval in which an event-related potential (ERP) occurred. An event-related potential is a transient, minute change in electrical potential that occurs when a specific event (for example, a statement made by another user in a conversation) occurs. Specifically, the listening level determination unit 40 determines the listening level of the user of the terminal device 10A based on the P300 of the event-related potential. P300 is a positive electrical potential that appears approximately 300 milliseconds after a specific event occurs. The listening level determination unit 40 determines that the listening level is above a predetermined level when P300 is detected. The listening level determination unit 40 may also determine the listening level in steps according to the electrical potential level of the detected P300.

[0058] The listening level determination unit 40 detects P300 by removing noise components contained in the electroencephalogram (EEG) data acquired by the EEG data acquisition unit 38. The listening level determination unit 40 detects P300 by removing noise components contained in the EEG data, for example, by performing an averaging method on multiple EEG data sets.

[0059] Figure 8 is a diagram illustrating the averaging method according to the second embodiment. The averaging method is a method of storing multiple unique EEG data of a user and calculating the average of the multiple EEG data. Figure 8 shows EEG data D1, EEG data D2, EEG data D3, ... as EEG data of the user of terminal device 10A. The listening level determination unit 40 detects P300 by adding up each EEG data and dividing the sum by the number of data. In this case, the multiple EEG data used are each acquired at the same time interval by the EEG data acquisition unit 38.

[0060] The listening level determination unit 40 can determine the listening level of the user of the terminal device 10A by detecting P300 based on electroencephalogram (EEG) data including the interval 300 milliseconds after a specific event occurs. Here, the conditions for P300 include that it is a positive potential, that the appearance time is between 250 milliseconds and 700 milliseconds after the specific event occurs, and that the potential is larger than the potential approximately 100 milliseconds before the specific event occurs. For this reason, it is preferable for the EEG data acquisition unit 38 to acquire EEG data for the interval from approximately 100 milliseconds before the specific event occurs to approximately 700 milliseconds after the specific event occurs.

[0061] Figure 9 shows the detection result of event-related potential P300 according to the second embodiment. In Figure 9, the horizontal axis represents time and the vertical axis represents potential. Waveform 101 represents the result calculated by performing an averaging method on multiple electroencephalogram (EEG) data by the listening level determination unit 40, and peak 102 represents P300. As shown in waveform 101, P300 can be appropriately detected by appropriately setting the interval in which the EEG data acquisition unit 38 acquires EEG data.

[0062] The comprehension determination unit 42 determines the user of the terminal device 10A's comprehension of the audio data acquired by the audio data acquisition unit 30. For example, the comprehension determination unit 42 determines whether the audio data acquired by the audio data acquisition unit 30 contains any unfamiliar words or phrases that the user of the terminal device 10A is determined not to understand.

[0063] The comprehension determination unit 42 determines, for example, that if the electroencephalogram (EEG) data shows a predetermined change when sound is emitted around the user of the terminal device 10A, the audio data acquired during a predetermined period before and after the predetermined change in the EEG data contains an unknown word or phrase. The comprehension determination unit 42 determines the comprehension level of the user of the terminal device 10A based on the EEG data for a predetermined section that includes the section in which an event-related potential of the EEG occurred. Specifically, the comprehension determination unit 42 determines the comprehension level of the user of the terminal device 10A based on the N400 of the event-related potential. N400 refers to a negative potential that appears approximately 400 milliseconds after a specific event occurs.

[0064] The comprehension level determination unit 42 determines, for example, that the level of comprehension is above a predetermined level when N400 is detected. The comprehension level determination unit 42 determines, for example, that when N400 is detected, the voice data acquired by the voice data acquisition unit 30 does not contain any unfamiliar words used by the user of the terminal device 10A. The comprehension level determination unit 42 may also determine the level of comprehension in stages, for example, according to the potential level of the detected N400.

[0065] The comprehension determination unit 42 may, for example, determine whether the audio data acquired by the audio data acquisition unit 30 contains any unfamiliar words of the user of the terminal device 10A, based on information previously stored in the storage unit 20. In this case, the storage unit 20 may store, for example, information about genres that the user of the terminal device 10A is not interested in or is not good at. Such information may be stored, for example, in a server device that can communicate with the terminal device 10A via the communication unit 18. The comprehension determination unit 42 may, for example, determine that a keyword related to a topic previously stored in the storage unit 20 is an unfamiliar word of the user of the terminal device 10A if it contains that keyword.

[0066] Specifically, the comprehension determination unit 42 performs speech recognition processing on the audio data acquired by the audio data acquisition unit 30 to generate text data. The comprehension determination unit 42 may determine that the generated text data contains keywords related to a topic that have been previously stored in the storage unit 20, and that these keywords are unfamiliar words to the user of the terminal device 10A.

[0067] The memory unit 20 may store, for example, a vocabulary database corresponding to the attributes of the user using the terminal device 10A. Specifically, the vocabulary database may be a database that associates, for example, the generation of the user using the terminal device 10A and the fields in which each generation is expected to be weak. Such a vocabulary database may be generated, for example, based on the results of a survey conducted on multiple users. Alternatively, the vocabulary database may be stored in a server device that can communicate with the terminal device 10A via the communication unit 18. In this case, the comprehension determination unit 42 may refer to the vocabulary database corresponding to the attributes of the user using the terminal device 10A and determine whether or not the voice data acquired by the voice data acquisition unit 30 contains any unfamiliar words of the user of the terminal device 10A.

[0068] The comprehension determination unit 42 may, for example, determine whether the audio data acquired by the audio data acquisition unit 30 contains any unfamiliar words used by the user of the terminal device 10A, based on the user's actions on the terminal device 10A. In this case, the comprehension determination unit 42 may, for example, determine that the audio data acquired by the audio data acquisition unit 30 contains any unfamiliar words used by the user of the terminal device 10A if the audio data acquisition unit 30 detects actions such as tapping the device body with a finger or tilting its head when acquiring audio data. In this case, the terminal device 10A is equipped with an acceleration sensor (not shown), and the comprehension determination unit 42 may determine that the user of the terminal device 10A has performed a predetermined action when the acceleration sensor detects a predetermined acceleration.

[0069] The information generation unit 34A generates summary information of the conversation indicated by the audio data acquired by the audio data acquisition unit 30 as output information. The information generation unit 34A also generates explanatory information for the user's unfamiliar words on the terminal device 10A as output information. The information generation unit 34A generates, for example, multiple explanatory pieces of information with different levels of detail. For example, the information generation unit 34A generates the shortest and simplest explanation as the first explanatory piece. For example, it may omit explanations of terms included in the explanation and related explanations other than those that directly explain the user's unfamiliar words, and generate the shortest and simplest explanation as the first explanatory piece. Alternatively, for example, it may generate an explanation that fits within a predetermined first number of characters or words as the first explanatory piece. The information generation unit 34A also generates, for example, a more detailed explanation than the first explanatory piece as the second explanatory piece. For example, it may generate the second explanatory piece including explanations that were omitted in the first explanatory piece. Furthermore, for example, a second explanatory information may be generated that is set to fit within a predetermined second number of characters or words, which is greater than a predetermined first number of characters or words. The information generation unit 34A may, for example, sequentially generate explanatory information that provides a more detailed explanation than the second explanatory information.

[0070] The information generation unit 34A may, for example, change the level of detail of the explanation in the first explanation information according to the result of the comprehension level determination unit 42's determination of the level of comprehension of the unknown word. For example, if the information generation unit 34A determines that the level of comprehension of the unknown word is low, it may increase the level of detail of the first explanation information to a higher level than usual.

[0071] The information generation unit 34A may change the level of detail of the first explanatory information based, for example, on information stored in the storage unit 20 or a server device that can communicate with the terminal device 10A. For example, if information about genres that the user of the terminal device 10A is not familiar with is stored in the storage unit 20 or a server device that can communicate with the terminal device 10A, the information generation unit 34A may increase the level of detail of the first explanatory information for keywords in those genres more than usual.

[0072] The information generation unit 34A may, for example, use a generating AI (Artificial Intelligence) to generate explanatory information for unfamiliar terms used by the user of the terminal device 10A. The information generation unit 34A may, for example, generate prompts for unfamiliar terms used by the user of the terminal device 10A. A prompt refers to a request or instruction to be input to the generating AI. The information generation unit 34A may, for example, input requests for multiple explanatory information of different levels of detail to the prompt, causing the generating AI to generate multiple explanatory information of different levels of detail. The information generation unit 34A may, for example, input a request for a summary of multiple explanatory information of different levels of detail to the prompt, causing the generating AI to generate a summary of multiple explanatory information of different levels of detail.

[0073] The output control unit 36A causes the information generation unit 34A to output explanatory information from the speaker 16. For example, the output control unit 36A causes the user of terminal device 10A to output explanatory information about unfamiliar words from the speaker 16. For example, the output control unit 36A causes the user of terminal device 10A to output explanatory information about unfamiliar words from the speaker 16 when it is determined that the user of terminal device 10A is not listening attentively to the speaker. In other words, the output control unit 36A causes the user of terminal device 10A to output explanatory information about unfamiliar words from the speaker 16 when it is determined that the user of terminal device 10A is not listening attentively to the speaker.

[0074] (Audio output processing) (Example 1) The audio output processing according to the first example of the second embodiment will be explained using Figure 10. Figure 10 is a flowchart showing the flow of the audio output processing according to the first example of the second embodiment.

[0075] Figure 10 shows the flow of processing performed by terminal device 10A when, for example, a user of terminal device 10A engages in conversation.

[0076] The process in step S70 is the same as the process in step S10 shown in Figure 3, so the explanation will be omitted.

[0077] The comprehension determination unit 42 determines whether or not an unknown word used by the user of terminal device 10A has been detected from the audio data acquired by the audio data acquisition unit 30 (step S72). If it is determined that an unknown word used by the user of terminal device 10A has been detected (step S72; Yes), the process proceeds to step S74. If it is determined that an unknown word used by the user of terminal device 10A has not been detected (step S72; No), the process proceeds to step S86.

[0078] If the answer in step S72 is determined to be Yes, the information generation unit 34A generates the simplest possible explanatory information about the detected unknown phrase (step S74).

[0079] The listening level determination unit 40 measures the listening level and determines whether the user of the terminal device 10A is listening to the speaker's statement at or below a predetermined level (step S76). If it is determined that the user of the terminal device 10A is listening to the speaker's statement at or below a predetermined level (step S76; Yes), the process proceeds to step S78. If it is determined that the user of the terminal device 10A is listening to the speaker's statement at or below a predetermined level (step S76; No), the process proceeds to step S86.

[0080] If the answer in step S76 is determined to be Yes, the output control unit 36A will output the simplest explanatory information about the unfamiliar word used by the user of the terminal device 10A from the speaker 16 (step S78).

[0081] The listening level determination unit 40 measures the listening level and determines whether the user's listening level to the explanatory information on the terminal device 10A is equal to or greater than their listening level to the speaker's remarks (step S80). Specifically, the listening level determination unit 40 identifies the timing when the explanatory information was output, the timing when the speaker made a remark, and the timing when P300 was detected from the electroencephalogram data. The listening level determination unit 40 then determines whether P300 was caused by the explanatory information or the speaker's remarks. For example, since P300 occurs between 250 milliseconds and 700 milliseconds after a specific event occurs, the listening level determination unit 40 determines, based on this time interval, whether the specific event was the output of the explanatory information or the speaker's remarks. For example, if the listening level determination unit 40 determines that P300 was caused by the explanatory information, it determines that the user's listening level to the explanatory information on the terminal device 10A is equal to or greater than their listening level to the speaker's remarks. If it is determined that the user's level of attentiveness to the explanatory information of terminal device 10A is equal to or greater than the user's level of attentiveness to the speaker's remarks (Step S80; Yes), the process proceeds to Step S82. If it is determined that the user's level of attentiveness to the explanatory information of terminal device 10A is not equal to or greater than the user's level of attentiveness to the speaker's remarks (Step S80; No), the process proceeds to Step S86. In Step S80, the attentiveness determination unit 40 may determine, based on the flow of the conversation, whether or not the user's level of attentiveness to the explanatory information of terminal device 10A is equal to or greater than the user's level of attentiveness to the speaker's remarks.

[0082] If the answer in step S80 is determined to be Yes, the information generation unit 34A generates more detailed explanatory information than the output explanatory information (step S82).

[0083] The output control unit 36A outputs the detailed explanatory information generated by the information generation unit 34A in step S82 through the speaker 16 (step S84). Then, the process proceeds to step S80.

[0084] In other words, as long as the terminal device 10A determines that the level of attention given to the explanatory information is greater than or equal to the level of attention given to the speaker's remarks, it generates and outputs more detailed explanatory information to the user one after another. As a result, the terminal device 10A can help the user properly understand the meaning of unfamiliar words and thus appropriately support the conversation.

[0085] The process in step S86 is the same as the process in step S18 shown in Figure 3, so the explanation will be omitted.

[0086] As described above, in the first example of the second embodiment, if it is determined that the user's level of attentiveness to explanatory information about unfamiliar words is greater than or equal to the user's level of attentiveness to the speaker's statements, more detailed explanatory information is output to the user. This allows the first example of the second embodiment to appropriately support the conversation.

[0087] (Second example) The audio output processing according to the second example of the second embodiment will be explained using Figure 11. Figure 11 is a flowchart showing the flow of the audio output processing according to the second example of the second embodiment.

[0088] Figure 11 shows the flow of processing performed by terminal device 10A when, for example, a user of terminal device 10A engages in conversation.

[0089] The processes from step S90 to step S100 are the same as the processes from step S70 to step S80 shown in Figure 10, so their explanation will be omitted.

[0090] If the result in step S100 is "Yes", the audio data acquisition unit 30 acquires ambient sound data indicating the ambient sound detected by the microphone 14 (step S102).

[0091] The output control unit 36A changes the output voice of the explanatory information output from the speaker 16 based on the ambient sound data acquired by the audio data acquisition unit 30 (step S104). Specifically, the output control unit 36A changes the volume and sound quality of the output voice so that the user can hear the explanatory information more easily on the terminal device 10A. For example, if the ambient sound data contains many male (female) voices, the output control unit 36A sets the output voice to a female (male) voice. For example, the output control unit 36A performs a frequency characteristic analysis on the ambient sound data to identify the frequency of the ambient sound and sets the frequency of the output voice to a value that does not overlap with the frequency of the ambient sound. In this case, the output control unit 36A may select the appropriate output voice from multiple frequency band output voices stored in the storage unit 20 in advance, or it may perform an equalization process on the currently set output voice to adjust the frequency. For example, if the ambient sound level is above a predetermined level, the output control unit 36A increases the volume of the output voice. The output control unit 36A may reduce the effective volume of ambient noise audible to the user by performing noise cancellation or sound masking processing if the ambient noise level exceeds a predetermined level. The output control unit 36A may also adjust the playback speed of the explanatory text to the terminal device 10A so that the user can hear the explanatory information more easily.

[0092] The processes from step S106 to step S110 are the same as the processes from step S82 to step S86 shown in Figure 10, so their explanation will be omitted.

[0093] As described above, in the second example of the second embodiment, if it is determined that the user's level of attentiveness to explanatory information about unfamiliar words is greater than or equal to the user's level of attentiveness to the speaker's statements, the output voice is modified based on ambient sounds, and then more detailed explanatory information is output to the user. This allows the second example of the second embodiment to appropriately support conversation.

[0094] (Third example) The audio output processing according to the third example of the second embodiment will be explained using Figure 12. Figure 12 is a flowchart showing the flow of the audio output processing according to the third example of the second embodiment.

[0095] Figure 12 shows the flow of processing performed by terminal device 10A when, for example, a user of terminal device 10A engages in conversation.

[0096] The processes from step S120 to step S130 are the same as the processes from step S70 to step S80 shown in Figure 10, so their explanation will be omitted.

[0097] If the result in step S30 is Yes, the output control unit 36A determines whether or not there has been silence for a predetermined period (step S132). Specifically, the output control unit 36A determines that there has been silence for a predetermined period if, for example, the voice data acquisition unit 30 does not acquire voice data for 3 seconds or more. The predetermined period can be set arbitrarily. If it is determined that there has been silence for a predetermined period (step S132; Yes), the process proceeds to step S134. If it is not determined that there has been silence for a predetermined period (step S132; No), the process in step S132 is repeated. In other words, detailed explanatory information is not output until it is determined that there has been silence for a predetermined period.

[0098] The processes from steps S134 to S138 are the same as the processes from steps S82 to S86 shown in Figure 10, so their explanation will be omitted. That is, in the third example of the second embodiment, the simplest explanatory information is output when it is first determined that the user's level of listening to the voice is below a predetermined level. Subsequently, if it is determined that the user's level of listening to the voice is below a predetermined level and that the user has been silent for a predetermined period of time, the output control unit 36A outputs more detailed explanatory information.

[0099] As described above, in the third example of the second embodiment, if it is determined that the user's level of attentiveness to explanatory information about unfamiliar words is greater than or equal to the user's level of attentiveness to the speaker's statements, more detailed explanatory information is output to the user at the time it is determined that there has been silence for a predetermined period of time. This prevents the timing of the speaker's statements from overlapping with the timing of the output of explanatory information in the third example of the second embodiment, thereby enabling more appropriate support for conversation.

[0100] (Fourth example) The audio output processing according to the fourth example of the second embodiment will be explained using Figure 13. Figure 13 is a flowchart showing the flow of the audio output processing according to the fourth example of the second embodiment.

[0101] Figure 13 shows the flow of processing performed by terminal device 10A when, for example, a user of terminal device 10A engages in conversation.

[0102] The processes from step S150 to step S160 are the same as the processes from step S70 to step S80 shown in Figure 10, so their explanation will be omitted.

[0103] If the result in step S160 is Yes, the listening level determination unit 40 determines whether the user of the terminal device 10A is listening to the speaker's statement at a predetermined level or lower (step S162). If it is determined that the user of the terminal device 10A is listening to the speaker's statement at a predetermined level or lower (step S162; Yes), the process proceeds to step S164. If it is not determined that the user of the terminal device 10A is listening to the speaker's statement at a predetermined level or lower (step S162; No), the process in step S162 is repeated.

[0104] The processes from steps S164 to S168 are the same as the processes from steps S82 to S86 shown in Figure 10, so their explanation will be omitted. That is, in the fourth example of the second embodiment, if it is first determined that the user's level of listening to the voice is below a predetermined level, the simplest explanatory information is output. Then, if it is determined that the level of listening to the explanatory information is greater than or equal to the level of listening to the speech, and the user's level of listening to the voice is below a predetermined level, then detailed explanatory information is output.

[0105] As described above, in the fourth example of the second embodiment, if it is determined that the user's level of attentiveness to explanatory information about unfamiliar words is greater than or equal to the user's level of attentiveness to the speaker's statement, then more detailed explanatory information is output to the user at the time it is determined that the user's level of attentiveness to the speaker's statement is below a predetermined level. As a result, in the fourth example of the second embodiment, after the explanatory information is output, more detailed explanatory information is output at the time when it is assumed that the user's level of attentiveness to the speaker's statement is below a predetermined level and the unfamiliar words have not yet been understood, thus enabling more appropriate support for conversation.

[0106] [Modified version of the second embodiment] A modified example of the second embodiment will now be described.

[0107] (First variation) In the second embodiment, the explanatory information was described as being output as audio, but this disclosure is not limited thereto. For example, the information generation unit 34A may generate explanatory information for unknown words as text or images. In this case, the output control unit 36A may transmit the text or images generated by the information generation unit 34A to the user's smartphone or the like via the communication unit 18. As a result, the first modification of the second embodiment can output explanatory information for unknown words as text or images to the user of the terminal device 10A, thereby appropriately supporting the conversation.

[0108] (Second variation) In the second embodiment, it was described that the user of the terminal device 10A detects an unknown word during a conversation, but the disclosure is not limited thereto. For example, the voice data acquisition unit 30 may acquire voice data of the voices of meeting participants detected by the microphone 14 in a meeting in which the user of the terminal device 10A is participating. In this case, if an unknown word is detected during the meeting, the terminal device 10A outputs explanatory information about the unknown word to the user of the terminal device 10A. This allows the second modification of the second embodiment to appropriately support the meeting.

[0109] (Third variation) In the second embodiment, the voice data acquisition unit 30 was described as acquiring voice data from the microphone 14, but this disclosure is not limited thereto. For example, the voice data acquisition unit 30 may acquire voice data indicating the voice spoken by a speaker via the communication unit 18 during online conversations or online meetings. As a result, the third modification of the second embodiment can output explanatory information about unfamiliar words to the user of the terminal device 10A online, thereby appropriately supporting online conversations and online meetings.

[0110] (Fourth variation) In the second embodiment, the terminal device 10A was described as generating explanatory information, but the disclosure is not limited thereto. For example, the explanatory information may be generated by a server device that can communicate with the terminal device 10A. In this case, the terminal device 10A transmits the audio data acquired by the audio data acquisition unit 30 and the electroencephalogram data acquired by the electroencephalogram data acquisition unit 38 to the server device via the communication unit 18. The server device then generates explanatory information based on the audio data and electroencephalogram data acquired from the terminal device 10A and transmits it to the terminal device 10A. In this case, the terminal device 10A outputs the explanatory information generated by the server device to the user of the terminal device 10A. Thus, the fourth modification of the second embodiment can appropriately support the meeting.

[0111] Each component of the illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. Furthermore, this distribution and integration configuration may be performed dynamically.

[0112] While embodiments of the present disclosure have been described above, the present disclosure is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that are readily conceivable to those skilled in the art, those that are substantially identical, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the embodiments described above. [Explanation of Symbols]

[0113] 10,10A Terminal device 12 Input section 14 Mike 16 speakers 18 Communications Department 20 Memory section 22,22A Control Unit 24 sensors 30 Audio data acquisition unit 32 Voice Judgment Unit 34,34A Information generation section 36,36A Output Control Unit 38. Electroencephalogram (EEG) data acquisition unit 40 Listening level determination unit 42 Comprehension level judgment part

Claims

1. A voice data acquisition unit that acquires voice data representing speech, A voice determination unit that determines whether the length of the silent period during which the voice data acquisition unit does not acquire voice data is longer than a predetermined length, If it is determined that the length of the silent period is greater than or equal to a predetermined length, an information generation unit generates audio output information. An output control unit that causes the audio output information generated by the information generation unit to be output from the output unit, A terminal device equipped with the following features.

2. The voice determination unit determines whether or not the voice data contains a predetermined phrase. The information generation unit generates the audio output information when it is determined that the audio data contains the predetermined phrase and that the length of the silent period is greater than or equal to a predetermined length. The terminal device according to claim 1.

3. The information generation unit generates explanatory information for the audio data as audio output information. The terminal device according to claim 1 or 2.

4. The steps include acquiring audio data that represents sounds emitted around the user, A step of determining whether the length of the silent period during which the aforementioned audio data is not acquired is greater than or equal to a predetermined length, If it is determined that there has been a period of time longer than a predetermined period during which the aforementioned audio data has not been acquired, the step of generating audio output information is performed. The steps include: outputting the generated audio output information from the output unit; Information output methods, including those mentioned above.

Citation Information

Patent Citations

  • Information recommendation system, information search device, information recommendation method, and program

    WO2021250833A1