Information output device, information output method, and program for information output device

The information output device in vehicles addresses the issue of conversation interruptions by using voice data analysis to determine optimal response times, ensuring timely and context-aware interactions.

JP2026016688APending Publication Date: 2026-02-03PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025184335
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-03

Smart Images

  • Figure 2026016688000001_ABST
    Figure 2026016688000001_ABST
Patent Text Reader

Abstract

To provide an information output device, an information output method and a program for the information output device for preventing conversation from being disturbed as much as possible.SOLUTION: In this method, a voice acquiring means acquires the voice of a conversation, a keyword extracting means extracts a keyword from the voice (S1, S2), a voiceless time counting means counts a voiceless time (S5, S6), and an outputting means outputs information related to the keyword (S7YES) when the voiceless time reaches a prescribed time or longer (S8).SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application belongs to the technical field of an information output device, an information output method, and a program for an information output device. [Background technology]

[0002] Using a smart speaker installed in a vehicle, voice instructions are given, music is played, and responses to questions are provided. Patent Document 1 below discloses an information processing device that is provided so that multiple voice assistants can be used, and that includes a master control unit that generates voice instructions for each voice assistant based on the content of a user's utterance and transmits the instructions to a server device for the voice assistant. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-4950 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the technology of Patent Document 1, the voice assistant system simply outputs a response to a voice instruction regardless of the conversation situation, which can sometimes interfere with the conversation.

[0005] The present invention has been made in view of the above problems, and one example of the object of the present invention is to provide an information output device or the like that prevents interruptions to conversation as much as possible. [Means for solving the problem]

[0006] In order to solve the above problem, the invention described in claim 1 comprises a voice acquisition means for acquiring voice data of a conversation, a keyword extraction means for extracting keywords from the voice data, a silent time counting means for counting silent time from the voice data, an output means for outputting information related to the keyword when the silent time reaches a predetermined time or more, and a conversation state analysis means for analyzing the state of the conversation related to the tempo or excitement of the conversation from the voice data, and is characterized in that when the conversation state analysis means analyzes that the tempo of the conversation is fast or that the conversation is exciting, the predetermined time is shortened.

[0007] The invention described in claim 2 comprises a voice acquisition means for acquiring voice data of a conversation, a keyword extraction means for extracting keywords from the voice data, a silent time counting means for counting silent time from the voice data, an output means for outputting information related to the keyword when the silent time reaches a predetermined time or more, and a degree determination means for determining the degree of physiological needs of the speaker in the conversation based on the voice data, wherein the predetermined time is set according to the degree of physiological needs of the speaker.

[0008] The invention described in claim 3 comprises a voice acquisition means for acquiring voice data of a conversation, a keyword extraction means for extracting keywords from the voice data, a silent time counting means for counting silent time from the voice data, an output means for outputting information related to the keyword when the silent time reaches a predetermined time or more, and a driving condition detection means for detecting the driving condition of a mobile body in which the speaker of the conversation is riding, and is characterized in that when the driving condition detection means analyzes that the driving condition is congested, the predetermined time is shortened.

[0009] The invention described in claim 4 is characterized by comprising an audio acquisition means for acquiring audio data of a conversation, a desire determination means for determining whether the conversation includes the possibility of the speaker's wish based on the audio data, a no-answer time counting means for counting the time during which no answer to the wish is included in the conversation after the desire determination when the desire determination means determines that the wish is included, and an output means for outputting information corresponding to the wish when the no-answer time exceeds a predetermined time.

[0010] The invention described in claim 5 includes a voice acquisition step in which a voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts keywords from the voice data; a voiceless time counting step in which a voiceless time counting means counts voiceless times from the voice data; an output step in which an output means outputs information related to the keyword when the voiceless time reaches a predetermined time or more; and a conversation state analysis step in which a conversation state analysis means analyzes the state of the conversation related to the tempo or excitement of the conversation from the voice data, and is characterized in that when the conversation state analysis step analyzes that the tempo of the conversation is fast or that the conversation is exciting, the predetermined time is shortened.

[0011] The invention described in claim 6 includes a voice acquisition step in which a voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts a keyword from the voice data; a voiceless time counting step in which a voiceless time counting means counts a voiceless time from the voice data; an output step in which an output means outputs information related to the keyword when the voiceless time reaches a predetermined time or more; and a degree determination step in which a degree determination means determines a degree of a physiological need of a speaker in the conversation based on the voice data, and the predetermined time is set in accordance with the degree of the physiological need of the speaker.

[0012] The invention described in claim 7 includes a voice acquisition step in which a voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts a keyword from the voice data; a voiceless time counting step in which a voiceless time counting means counts a voiceless time from the voice data; an output step in which an output means outputs information related to the keyword when the voiceless time reaches a predetermined time or more; and a driving condition detection step in which a driving condition detection means detects the driving condition of a mobile body in which a speaker of the conversation is riding, and is characterized in that when the driving condition detection step analyzes that the driving condition is congested, the predetermined time is shortened.

[0013] The invention described in claim 8 is characterized in that it includes a voice acquisition step in which a voice acquisition means acquires voice data of the conversation; a desire determination step in which a desire determination means determines whether the conversation includes the possibility of the speaker's desire based on the voice data; a no-response time counting step in which a no-response time counting means counts the time during which no response to the desire is included in the conversation after the determination by the desire determination means, if the possibility of the desire is included; and an output step in which an output means outputs information corresponding to the desire when the no-response time reaches a predetermined time or more.

[0014] The invention described in claim 9 is characterized in that a computer is made to function as the information output device described in any one of claims 1 to 4. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a block diagram illustrating an example of a configuration of an information output device according to an embodiment. [Figure 2] 1 is a schematic diagram illustrating an example of an information output system using an information output device according to an embodiment. [Figure 3] 1 is a block diagram showing an example of a schematic configuration of an information output device according to an embodiment; [Figure 4]FIG. 2 is a schematic diagram illustrating an example of a database of an information output device. [Figure 5] FIG. 2 is a schematic diagram illustrating an example of a database of an information output device. [Figure 6] 10 is a flowchart illustrating an example of an operation of the information output device according to the embodiment. [Figure 7] 10 is a flowchart illustrating a modified example of the operation of the information output device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] An embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of an information output device according to an embodiment.

[0017] As shown in FIG. 1, the information output device 1 includes a voice acquisition unit 1a, a keyword extraction unit 1b, a silent time counting unit 1c, and an output unit 1d.

[0018] In this configuration, the voice acquisition means 1a acquires voice data of a conversation.

[0019] The keyword extraction means 1b extracts keywords from the voice data.

[0020] The silent time counting means 1c counts the silent time.

[0021] When the silent time reaches or exceeds a predetermined time, the output means 1d outputs information related to the keyword.

[0022] As described above, the information output device 1 according to the embodiment acquires the audio data of the conversation, extracts keywords from the audio data, counts the silent time, and outputs information related to the keyword when the silent time reaches a predetermined time or more. As a result, the information output device 1 does not output the information related to the keyword immediately after the conversation stops, but outputs it after an interval of at least a predetermined time, thereby making it possible to prevent interruptions to the conversation as much as possible. [Example]

[0023] [1. Information output system]

[0024] Next, specific examples corresponding to the above-described embodiments will be described with reference to Figures 2 to 6. Note that the examples described below are examples in which the present application is applied to an information output device 10.

[0025] (1.1 Configuration and Overview of Information Output System) The configuration and overview of the information output system will be explained using FIG.

[0026] FIG. 2 is a schematic diagram illustrating an example of an information output system using an information output device according to an embodiment.

[0027] As shown in Figure 2, the information output system S includes an information output device 10 (an example of an information output device 1) mounted on a vehicle Vh, a camera Cm that photographs a passenger Ps inside the vehicle Vh, a microphone Mc that collects the voice of the passenger Ps, and a speaker Sp that outputs synthesized voice, etc.

[0028] The information output device 10 is, for example, a computer having a navigation function, an audio function, an AI assistant function, and the like.

[0029] The camera Cm has an imaging element such as a CMOS image sensor. The camera Cm captures video and still images of the passengers Ps. The camera Cm is installed in a position where all passengers Ps inside the vehicle Vh can be easily captured, for example, near the rearview mirror of the vehicle Vh. Note that multiple cameras Cm may be installed inside the vehicle Vh. The information output system S may also have a camera that captures the outside of the vehicle Vh.

[0030] The microphone Mc is, for example, an electret condenser microphone or a MEMS (Micro-Electro-Mechanical System) microphone. The microphone Mc converts sounds inside the vehicle Vh into electrical signals. The microphone Mc is installed in a position where it is easy to pick up the sounds of the passengers Ps inside the vehicle Vh, for example, near the rearview mirror of the vehicle Vh.

[0031] The speaker Sp is, for example, an audio speaker in the vehicle Vh. The speaker Sp outputs music, synthesized voice guidance for navigation, etc. The speaker Sp may be a speaker in an additional device installed in the vehicle Vh, a speaker in the audio system of the vehicle Vh, or a speaker in a handheld device that the user has brought into the car and that is set to work with this system.

[0032] In addition to the vehicle Vh, examples of the moving body include a train, a ship, an airplane, and the like.

[0033] (1.2 Configuration and Function of Information Output Device 10) Next, the configuration and functions of the information output device 10 will be described with reference to FIGS.

[0034] Fig. 3 is a block diagram showing an example of a schematic configuration of an information output device according to an embodiment. Fig. 4 and Fig. 5 are schematic diagrams showing an example of a database of the information output device.

[0035] As shown in FIG. 3, the information output device 10 includes a communication unit 11, a storage unit 12, a display unit 13, an operation unit 14, an interface unit 15, a sensor unit 16, and a control unit 17.

[0036] The communication unit 11 is connected to a network such as a wireless communication network and controls the state of communication with an external server device. The information output device 10 and the external server device can transmit and receive data to and from each other via the network using, for example, TCP / IP as a communication protocol. The network is constructed, for example, by the Internet, a dedicated communication line (for example, a CATV (Community Antenna Television) line), a mobile communication network (including base stations, etc.), a gateway, etc. The external server device is, for example, a search server device, a traffic information providing server device, etc.

[0037] The storage unit 12 is configured by, for example, a hard disk drive, a solid state drive, or the like.

[0038] The storage unit 12 stores various programs for controlling the information output device 10. The various programs include an operating system, application software for navigation and music playback, and audio programs. Examples of audio programs include a speech recognition program that converts audio data into text data using acoustic analysis, an acoustic model, and the like, a natural language processing program such as morphological analysis, syntactic analysis, and semantic analysis, and a program that generates synthetic speech from text data. The various programs may be obtained, for example, via a network, or may be recorded on a recording medium such as a CD or DVD and read via a drive device.

[0039] A database for information output is also constructed in the storage unit 12. For example, as shown in Fig. 4, a keyword database is constructed in the storage unit 12, in which keywords are stored in association with the degree of possibility of the speaker's desire.

[0040] Here, the degree of possibility of the speaker's desire is set so that the likelihood of desire increases in the order of, for example, "I'm so tired," "I'm so tired," and "I'm so tired," or "I'm so hungry," "I'm so hungry," and "I'm so hungry." If a person says "I'm so tired," it is considered that the possibility of the desire or intention to take a break is still low (degree of possibility of desire: 1). If a person says "I'm so tired," it is considered that fatigue is building up and the possibility of the desire or intention to take a break is increasing (degree of possibility of desire: 2). If a person says "I'm so tired," it is considered that the feeling of fatigue is strong and the possibility of the desire or intention to take a break is high (degree of possibility of desire: 3). If a person says "I'm so hungry," it is considered that the possibility of the desire or intention to eat is still low (degree of possibility of desire: 1). If a person says "I'm so hungry," it is considered that the feeling of hunger is getting a little stronger and the possibility of the desire or intention to eat is increasing (degree of possibility of desire: 2). When someone says "I'm hungry," it is highly likely that they are feeling very hungry and have a desire or intention to eat (degree of possibility of desire: 3). This database may be a database that classifies endings such as "naa," "runaa," "tanaa," and "taa" of the above-mentioned phrases such as "I'm tired," "tired," and "tired." The memory unit 12 also stores psychologically predictable keywords such as "I've been running for an hour already" and "What should I do for lunch?" Furthermore, the degree of desire may be set based on the time period of use and usual behavioral tendencies.

[0041] Furthermore, in the storage unit 12, as a database for information output, a predetermined time database is constructed, as shown in Fig. 5, in which predetermined times are stored in association with the degree of possibility of the wish. The predetermined times are set, for example, to 5 seconds, 4 seconds, 3.5 seconds, and 3 seconds, respectively, according to the degree of possibility of the wish, from "1" indicating a low possibility to "4" indicating a high possibility. When the degree of possibility of the wish is low, the information output system S does not need to respond immediately when it detects the speaker's wish, and the predetermined time until it responds to the wish is set long. When the degree of possibility of the wish is high, the predetermined time until it responds to the wish is set short.

[0042] The storage unit 12 may have a database that realizes an AI assistant function.

[0043] The display unit 13 is a monitor display configured with a liquid crystal display element, an EL element, or the like, which is used when operating the information output device 10. The display unit 13 displays route guidance information and the like.

[0044] The operation unit 14 includes, for example, various buttons such as a mechanical power button and a volume button, and the display unit 14 is a touch switch type display panel such as a touch panel.

[0045] The interface unit 15 connects the information output device 10 with the camera Cm, microphone Mc, and speaker Sp.

[0046] The sensor unit 16 is, for example, a variety of sensors such as a speed sensor, an acceleration sensor, a gyro sensor, a GPS sensor, a direction sensor, a steering angle sensor, a seat sensor, a temperature sensor, etc. The sensor unit 16 may also have a timer function for counting time and a clock function for measuring time.

[0047] The speed sensor detects the speed of the vehicle Vh. The acceleration sensor detects the acceleration of the vehicle Vh. The GPS sensor acquires latitude and longitude information as the current position of the vehicle Vh. The gyro sensor detects the angular acceleration of the body of the vehicle Vh. The orientation sensor detects the orientation of the vehicle Vh. The steering angle sensor detects the steering angle. The seat sensor detects whether a passenger in the vehicle Vh is sitting in a seat. The temperature sensor detects the temperature inside the vehicle Vh, the outside temperature, etc.

[0048] The control unit 17 includes, for example, a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory). The control unit 17 reads and executes various programs stored in the ROM, RAM, and storage unit 12. The control unit 17 may also have a chip for AI functions.

[0049] The control unit 17 performs route calculation, speech recognition, realizes an AI assistant function, and controls the information output device 10.

[0050] 2. Operation of Information Output Device 10 Next, the operation of the information output device 10 according to the embodiment will be described with reference to FIG.

[0051] FIG. 6 is a flowchart illustrating an example of the operation of the information output device according to the embodiment.

[0052] As shown in FIG. 6, the information output device 10 acquires and analyzes audio data (step S1). Specifically, the control unit 17 receives an electric signal of the sound in the vehicle Vh from the microphone Mc, and converts it into a digital signal of the sound by A / D conversion. The included voice data is converted into text data using a voice recognition program. Next, the control unit 17 extracts words from the text data using a natural language processing program. Note that this operation is initiated after the user gives permission to acquire the voice data.

[0053] The control unit 17 analyzes the voice data of the voice uttered by the passenger Ps to determine whether or not there is voice. For example, if the volume of the voice data is equal to or less than a predetermined value, the control unit 17 determines that there is no voice. If there is sound but it cannot be recognized as voice by the voice recognition program, the control unit 17 determines that there is no voice.

[0054] The control unit 17 also uses a voice recognition program to identify each person from the voice data and estimate the number of passengers Ps. The control unit 17 may also count the number of people in the vehicle Vh by extracting people from images captured by the camera Cm using an image recognition program. The control unit 17 may also count the number of passengers Ps from information from the seat sensors of the sensor unit 16 of the vehicle Vh. The control unit 17 may also identify the age, gender, and individual of the passengers Ps using an image recognition program. The control unit 17 may also identify the individual and acquire user information about the passengers Ps.

[0055] The control unit 17 may analyze the conversational state of the passengers Ps, such as the tempo of the conversation and the level of excitement of the conversation, from the voice data. For example, the tempo of the conversation is calculated from the number of words uttered per unit time. The more words uttered per unit time, the faster the tempo of the conversation. The level of excitement of the conversation is calculated, for example, from the volume, the number of people participating in the conversation, and frequency analysis of the voice data. The higher the volume [dB], the more exciting the conversation. The more people participating in the conversation, the more exciting the conversation. Frequency analysis indicates that the stronger the power spectrum of high-frequency components, the more exciting the conversation. When determining the level of excitement based on the volume, the volume threshold for determining that the conversation is exciting may be set higher when external noise is loud, such as when driving on a highway, on a road with poor road surface, or on a rainy day. The control unit 17 may also determine the conversational state, such as the level of excitement of the conversation, by analyzing each person's facial expression from images captured by the camera Cm.

[0056] Furthermore, the control unit 17 may take statistics on the extracted words and set frequently occurring nouns, verbs, and other words as trending keywords. For example, if the word "America" ​​is frequently used, the control unit 17 may extract "America" ​​as a trending keyword. In the case of a trending keyword, the degree of possibility of the desire may be set lower than "1" as exemplified previously.

[0057] Furthermore, the control unit 17 may detect the traveling state of the vehicle Vh based on traffic information acquired via the communication unit 11, speed information from the sensor unit 16, position information, and the like.

[0058] Next, the information output device 10 determines whether or not the extracted word is a keyword (step S2). Specifically, the control unit 17 determines whether or not the extracted word is a keyword by referring to the keyword database in the storage unit 12. For example, the control unit 17 determines whether or not the converted text contains keywords such as "I'm so tired," "I'm so tired," or "I'm so tired."

[0059] If the extracted word is not a keyword (step S2; NO), the information output device 10 returns to the processing of step S1.

[0060] If the extracted word is a keyword (step S2; YES), the information output device 10 sets a predetermined time (step S3). Specifically, the control unit 17 references the keyword database and the predetermined time database in the storage unit 12 and sets the predetermined time based on the degree of possibility that the keyword is a desire. For example, if the keyword is "I'm so tired" or "I'm so hungry," the degree of possibility that the desire is 2, and the control unit 17 sets the predetermined time to 4 seconds. If the keyword is "I'm so tired" or "I'm so hungry," the degree of possibility that the desire is 3, and the control unit 17 sets the predetermined time to 3.5 seconds. If the keyword is "I'm so tired" or "I'm so hungry," the degree of possibility that the desire is 4, and the control unit 17 sets the predetermined time to 3 seconds.

[0061] Next, the information output device 10 acquires and analyzes the voice data (step S4). Specifically, the control unit 17 acquires and analyzes the voice data as in step S1, and analyzes whether or not there is voice.

[0062] Next, the information output device 10 determines whether or not there is no voice (step S5). Specifically, the control unit 17 determines whether or not there is voice based on the analysis result of the voice data in step S4.

[0063] If there is no silent voice (step S5; NO), the information output device 10 returns to the processing of step S1.

[0064] If there is no voice (step S5; YES), the information output device 10 counts the time (step S6). Specifically, the control unit 17 starts measuring the time during which there is no voice. Note that the control unit 17 may use a clock function to compare the time at which it was determined in step S5 that there is voice with the current time, and use this to determine the counted time.

[0065] Next, the information output device 10 determines whether or not the counted time is equal to or greater than a predetermined time (step S7). Specifically, the control unit 17 determines whether or not the counted time is equal to or greater than a set predetermined time.

[0066] If the time is not equal to or longer than the predetermined time (step S7; NO), the information output device 10 returns to the processing of step S4.

[0067] If the predetermined time is equal to or longer than the predetermined time (step S7; YES), the information output device 10 outputs information related to the keyword (step S8). Specifically, the control unit 17 generates a response sentence in the form of text data using the AI ​​assistant function based on the keyword determined in step S2 or a sentence uttered by the speaker that includes the keyword, synthesizes the voice, and outputs the voice from the speaker Sp. If the response content is a song, the control unit 17 acquires music data, plays the song, and outputs the music from the speaker Sp.

[0068] For example, if the uttered words are "I'm hungry," "I'm hungry," or "I'm so hungry," restaurants along the route set by the navigation function are searched for, and the control unit 17 generates a sentence recommending a meal at a specific restaurant based on the search results, and makes the suggestion audibly from the speaker Sp. If the uttered words are "I'm tired," "I'm so tired," or "I'm so tired," parking areas, coffee shops, etc. along the route set by the navigation function are searched for, and the control unit 17 generates a sentence recommending a break at a specific parking area, coffee shop, etc. based on the search results, and makes the suggestion audibly from the speaker Sp. In this case, the control unit 17 may make a suggestion audibly from the speaker Sp, or may automatically play music such as soothing music, nostalgic music, uplifting music, or music to sing along to. The control unit 17 may also suggest relaxing by talking to distant family or friends over the phone. If the topic keyword is the name of a specific place, a search is performed by the place name, and location information is provided as a search result.

[0069] Furthermore, the information output device 10 may generate and output an additional question in addition to the suggestion as information related to the keyword, thereby further increasing the accuracy of the content of the answer to the keyword uttered by the speaker.

[0070] The control unit 17 may change the search content depending on the degree of possibility of the desire. In the case of "I'm hungry," suggestions for a cafe are made, and in the case of "I'm hungry," suggestions for a restaurant where you can have a proper meal are made. Also, instead of suggestions, the control unit 17 may output voice responses such as "I see," "Hmm," etc.

[0071] The control unit 17 may output information related to the keyword to the display unit 13. For example, the control unit 17 may display the information in a text data format as a response, or may play a related video or the like.

[0072] Additionally, an external server device may use a search or AI assistant function to find information related to the keyword.

[0073] The information output device 10 may set the predetermined time period according to the state of the conversation of the passenger Ps. For example, if it is analyzed in step S1 that the tempo of the conversation is fast or that the conversation is lively, the control unit 17 may shorten the predetermined time period in step S3 to match the rhythm of the conversation. Also, if it is analyzed in step S1 that the number of passengers is one, the control unit 17 may shorten the predetermined time period in step S3 because the conversation partner is not inside the vehicle Vh.

[0074] Furthermore, the information output device 10 may set the predetermined time period according to the driving state of the vehicle Vh. For example, if it is analyzed in step S1 that the driving state of the vehicle Vh is in a traffic jam, the control unit 17 may shorten the predetermined time period in step S3. If the driving state of the vehicle Vh is not in a traffic jam, the control unit 17 may lengthen the predetermined time period. Because it is easy to become irritated when in a traffic jam, the predetermined time period is shortened to speed up the response from the information output device 10.

[0075] As described above, according to the operation of the embodiment, voice data of a conversation is acquired, keywords are extracted from the voice data, and silent periods are counted. When the silent periods exceed a predetermined time, information related to the keyword is output. This prevents the information related to the keyword from being output immediately after the conversation stops, but rather after an interval of at least a predetermined time, thereby preventing interruptions to the conversation as much as possible. By preventing interruptions to the conversation as much as possible, it is possible to achieve the effect of reading the mood of the situation.

[0076] Furthermore, since the answer is based on the possibility of the speaker's wish before the speaker gives a clear instruction, it is possible to build a so-called clever system that can respond in advance of the speaker.

[0077] Furthermore, when the state of conversation is analyzed from the voice data and the silent time exceeds a predetermined time set according to the state of conversation, information related to the keyword is output, as the predetermined time is set to a length according to the state of conversation, it is possible to prevent interruptions to the conversation as much as possible.

[0078] Furthermore, the degree of possibility of the speaker's wish in the conversation is judged based on the voice data, and when the silent time exceeds a predetermined time set according to the degree of possibility of the speaker's wish, information related to the keyword is output, since the predetermined time is set to a length according to the degree of possibility of the speaker's wish, it is possible to prevent the conversation from being disturbed as much as possible.

[0079] In addition, when the traveling state of a moving body such as a vehicle Vh in which the speaker is riding is detected and the silent period exceeds a predetermined period set according to the traveling state, information related to the keyword is output, and the predetermined period is set to a length according to the traveling state, thereby making it possible to prevent interruptions to the conversation as much as possible.

[0080] (Variation) Next, a modified example of the operation of the information output device 10 will be described with reference to Fig. 7. Note that the same reference numerals will be used for the same or corresponding parts as in the above embodiment, and only different configurations and operations will be described.

[0081] FIG. 7 is a flowchart illustrating a modified example of the operation of the information output device according to the embodiment.

[0082] As shown in Fig. 7, the information output device 10 acquires and analyzes voice data (step S11). Specifically, as in step S1, the control unit 17 receives an electric signal of the sound inside the vehicle Vh from the microphone Mc, and converts it into a digital signal of the sound by A / D conversion. However, the voice data contained in the digital sound signal is converted into text data, and the meaning is interpreted using AI functions such as natural language processing to analyze whether the conversation may contain the speaker's wishes.

[0083] The control unit 17 may analyze the intonation of the voice of the passenger Ps from the digital sound signal. The control unit 17 may also interpret the meaning using an AI function such as natural language processing and calculate the degree of possibility of the wish of the speaker, the passenger Ps.

[0084] Furthermore, the control unit 17 may use an AI function to identify each passenger Ps, estimate the number of passengers Ps, and identify the age, gender, and individuality of the passengers Ps. The control unit 17 may use an AI function to analyze the state of the conversation, such as the tempo and excitement of the conversation, from the voice data. The control unit 17 may also identify the topic of the conversation by interpreting the meaning using the AI ​​function.

[0085] Next, the information output device 10 determines whether or not it is a wish (step S12). Specifically, the control unit 17 interprets the meaning using an AI function such as natural language processing, and determines whether or not the conversation or the sentence of the conversation contains the possibility of the speaker's wish based on the voice data.

[0086] If it is not a wish (step S12; NO), the information output device 10 returns to the processing of step S11.

[0087] If it is a wish (step S12; YES), the information output device 10 sets a predetermined time (step S13). For example, the control unit 17 sets the predetermined time as in step S3, depending on the degree of possibility of the speaker's wish calculated in step S11. Alternatively, the information output device 10 may set the predetermined time depending on the state of the conversation. The information output device 10 may set the predetermined time depending on the running state of the vehicle Vh.

[0088] Next, the information output device 10 acquires and analyzes the voice data (step S14). Specifically, the control unit 17 acquires the voice data as in step S11, interprets the meaning using an AI function such as natural language processing, and analyzes whether the conversation includes a response to the possibility of the speaker's desire.

[0089] Next, the information output device 10 determines whether or not the conversation contains an answer (step S15). Specifically, the control unit 17 determines whether or not the conversation contains an answer to the possibility of the speaker's wish, based on the analysis result of the voice data in step S14.

[0090] If an answer is included (step S15; YES), the information output device 10 returns to the processing of step S11.

[0091] If no answer is included (step S15; NO), the information output device 10 counts time (step S16). Specifically, the control unit 17 starts measuring the time (no answer time) during which no answer to the wish is included in the conversation after the determination in step S15. Note that the control unit 17 may use a clock function to compare the time at which it was determined in step S15 that there was a possibility of a wish with the current time, and use this to count the time.

[0092] Next, the information output device 10 determines whether the time during which no answer is included is equal to or longer than a predetermined time (step S17). Specifically, the control unit 17 determines whether the counted time is equal to or longer than a set predetermined time.

[0093] If the time is not equal to or longer than the predetermined time (step S17; NO), the information output device 10 returns to the process of step S14.

[0094] If the predetermined time has elapsed (step S17; YES), the information output device 10 outputs information corresponding to the speaker's desire, as in step S8 (step S18). Specifically, the control unit 17 uses the AI ​​assistant function to generate a sentence in text data format that responds as information corresponding to the speaker's desire, based on the sentence uttered by the speaker including the desire in step S12, synthesizes the voice, and outputs the voice from the speaker Sp. If the content of the response is a song, the control unit 17 acquires music data, plays the song, and outputs the music from the speaker Sp.

[0095] For example, if a speaker utters "I'm hungry," "I'm hungry," or "I'm so hungry," which is determined to be a possible desire, restaurants near the route set by the navigation function are searched for, and the control unit 17 generates a sentence recommending a meal at a specific restaurant based on the search results, and makes the suggestion aloud from the speaker Sp. If a speaker utters "I'm tired," "I'm so tired," or "I'm so tired," which is determined to be a possible desire of the speaker, the navigation function searches for parking areas, coffee shops, etc. near the route set by the navigation function, or music to play, and the control unit 17 generates a sentence recommending a break at a specific parking area, coffee shop, etc. based on the search results, or a sentence recommending music to play, and makes the suggestion aloud from the speaker Sp. If the topic is the name of a place, a search is performed by the place name, and location information is provided as a search result.

[0096] The control unit 17 may change the information corresponding to the desire according to the degree of possibility of the speaker's desire calculated in step S11. For example, in the case of "I'm hungry," a suggestion of a cafe is made, and in the case of "I'm hungry," a suggestion of a restaurant where you can have a proper meal is made.

[0097] As explained above, according to the operation of the modified example, audio data of the conversation is acquired, and it is determined based on the audio data whether the conversation includes the possibility of the speaker's wish. If it is determined that the wish is included, the time during which the conversation does not include an answer to the wish is counted, and if the time without answer exceeds a predetermined time, information corresponding to the wish is output. This means that information corresponding to the wish is not output immediately after the speaker's wish is detected, but is output after an interval of a predetermined time or more during which there is no answer from the speakers in the conversation, thereby preventing interruptions to the conversation as much as possible. [Explanation of symbols]

[0098] 1, 10... Information output device 1a. Audio acquisition means 1b Keyword extraction method 1c... Silent time counting means 1d... Output means

Claims

1. A voice acquisition means for acquiring voice data of a conversation; a keyword extraction means for extracting keywords from the voice data; a silent time counting means for counting silent time from the voice data; an output means for outputting information related to the keyword when the silent time period reaches or exceeds a predetermined time period; a conversation state analysis means for analyzing a conversation state related to the tempo or excitement of the conversation from the voice data; Equipped with An information output device characterized in that, when the conversation state analysis means analyzes that the tempo of the conversation is fast or that the conversation is getting lively, the predetermined time is shortened.

2. A voice acquisition means for acquiring voice data of a conversation; a keyword extraction means for extracting keywords from the voice data; a silent time counting means for counting silent time from the voice data; an output means for outputting information related to the keyword when the silent time period reaches or exceeds a predetermined time period; a degree determining means for determining a degree of physiological needs of the speaker in the conversation based on the voice data; Equipped with An information output device, wherein the predetermined time is set in accordance with the degree of physiological needs of the speaker.

3. A voice acquisition means for acquiring voice data of a conversation; a keyword extraction means for extracting keywords from the voice data; a silent time counting means for counting silent time from the voice data; an output means for outputting information related to the keyword when the silent time period reaches or exceeds a predetermined time period; a running state detection means for detecting a running state of a vehicle in which the speaker of the conversation is riding; Equipped with An information output device characterized in that, when the driving condition detection means analyzes that the driving condition is a traffic jam, the predetermined time is shortened.

4. A voice acquisition means for acquiring voice data of a conversation; a desire determination means for determining whether the conversation includes a possibility of the speaker's desire based on the voice data; a no-answer time counting means for counting a time during which no answer to the desire is included in the conversation after the desire determination means determines that the desire is possibly included; and an output means for outputting information corresponding to the desire when the no-response time period reaches or exceeds a predetermined time period; An information output device comprising:

5. a voice acquisition step in which voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts keywords from the voice data; a silent time counting step in which silent time counting means counts silent time from the audio data; an output step in which, when the silent time period reaches or exceeds a predetermined time period, an output means outputs information related to the keyword; a conversation state analysis step in which a conversation state analysis means analyzes a conversation state related to the tempo or excitement of the conversation from the voice data; Including, An information output method characterized in that, when the conversation state analysis step determines that the tempo of the conversation is fast or that the conversation is getting lively, the predetermined time is shortened.

6. a voice acquisition step in which voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts keywords from the voice data; a silent time counting step in which silent time counting means counts silent time from the audio data; an output step in which, when the silent time period reaches or exceeds a predetermined time period, an output means outputs information related to the keyword; a degree determining step in which a degree determining means determines a degree of a physiological need of the speaker in the conversation based on the voice data; Including, An information output method, characterized in that the predetermined time is set according to the degree of physiological needs of the speaker.

7. a voice acquisition step in which voice acquisition means acquires voice data of a conversation; a keyword extraction step in which a keyword extraction means extracts keywords from the voice data; a silent time counting step in which silent time counting means counts silent time from the audio data; an output step in which, when the silent time period reaches or exceeds a predetermined time period, an output means outputs information related to the keyword; a running state detection step in which a running state detection means detects a running state of a mobile body on which a speaker of the conversation is riding; Including, The information output method is characterized in that, when the driving state is analyzed to be a traffic jam in the driving state detection step, the predetermined time is shortened.

8. a voice acquisition step in which voice acquisition means acquires voice data of a conversation; a desire determination step in which a desire determination means determines whether or not the conversation includes a possibility of the speaker's desire based on the voice data; a no-answer time counting step in which, when the desire determination means determines that the possibility of the desire is included, a no-answer time counting means counts a time during which no response to the desire is included in the conversation after the determination; an output step in which an output means outputs information corresponding to the desire when the no-answer time is equal to or longer than a predetermined time; An information output method comprising:

9. 5. A program for an information output device, which causes a computer to function as the information output device according to claim 1.

Citation Information

Patent Citations

  • Information processing device, information processing system, and information processing method

    JP2021004950A