Voice output device

The voice output device addresses the challenge of elderly users struggling with fast or complex voice outputs by adjusting its output mode based on user response times, resulting in improved understanding and usability for elderly users.

JP2025084011APending Publication Date: 2025-06-02TOYOTA JIDOSHA KK
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2023197744
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-06-02

AI Technical Summary

Technical Problem

Some speakers, such as the elderly, may struggle to understand voice outputs from agent devices that are too fast or complex.

Method used

A voice output device that adjusts its output mode based on the response time of the user, switching to a dedicated mode for users with longer response times, such as the elderly, by slowing, shortening, or amplifying the voice output, or simplifying the language used.

Benefits of technology

The solution makes voice outputs easier for elderly users to understand by adapting the output mode to accommodate their response times, thereby improving comprehension and usability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084011000001_ABST
    Figure 2025084011000001_ABST
Patent Text Reader

Abstract

To make voice output easier to understand for occupants of a vehicle, such as elderly people.SOLUTION: A voice output device 20 comprises a control unit that controls a specific function when voice outputting a question to occupants of the vehicle 12 and receiving a voice input from the occupants to the question as a response, that acquires time data indicating a time taken for a response from the occupants to the question voice output at a first time point, and switches a mode of voice output to the occupants at a second time point later than the first time point, based on the time indicated in the acquired time data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a voice output device.

Background Art

[0002] Patent Document 1 discloses an agent device that provides various types of information based on the requests of vehicle occupants and controls in-vehicle devices while interacting with the vehicle occupants. The agent device estimates the speaker based on the content of the utterance made to the agent device. For example, when the content of the utterance is related to the driving or operation of the vehicle, the agent device estimates that the speaker is the driver of the vehicle.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Some speakers, such as the elderly, may not be able to understand that the voice output from the agent device is too fast or too difficult.

[0005] An object of the present disclosure is to make voice output easier for vehicle occupants such as the elderly to understand.

Means for Solving the Problems

[0006] The voice output device according to the present disclosure is A control unit that outputs a question as voice to an occupant of a vehicle and receives a voice input from the occupant in response to the question, the control unit acquiring time data indicating the time required for the response from the occupant to the question output as voice at a first time point, and switching the mode of the voice output to the occupant at a second time point after the first time point based on the time indicated by the acquired time data.

Advantages of the Invention

[0007] According to the present disclosure, voice output becomes easier for vehicle occupants such as the elderly to understand.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Modes for Carrying Out the Invention

[0009] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings.

[0010] In each figure, the same or corresponding parts are denoted by the same reference numerals. In the description of this embodiment, the description of the same or corresponding parts will be omitted or simplified as appropriate.

[0011] With reference to FIG. 1, the configuration of the system 10 according to this embodiment will be described.

[0012] The system 10 according to this embodiment includes a voice output device 20 and a server device 30. The voice output device 20 can communicate with the server device 30 via a network 40.

[0013] The voice output device 20 is a computer with a voice recognition function mounted on the vehicle 12. The voice output device 20 is used by the user 11. The user 11 is a passenger in the vehicle 12.

[0014] The server device 30 is a computer belonging to a cloud computing system or other computing system installed in a facility such as a data center. The server device 30 is operated by an operator providing services such as web services.

[0015] The vehicle 12 is, for example, an automobile of any type such as a gasoline vehicle, a diesel vehicle, a hydrogen vehicle, an HEV, a PHEV, a BEV, or an FCEV. "HEV" is an abbreviation for hybrid electric vehicle. "PHEV" is an abbreviation for plug-in hybrid electric vehicle. "BEV" is an abbreviation for battery electric vehicle. "FCEV" is an abbreviation for fuel cell electric vehicle. The vehicle 12 may be driven by the user 11 or may have its driving automated at any level. The level of automation is, for example, any one of levels 1 to 5 in the SAE level classification. "SAE" is the abbreviation of Society of Automotive Engineers. The vehicle 12 may also be a vehicle dedicated to MaaS. "MaaS" is an abbreviation for Mobility as a Service.

[0016] Network 40 includes the Internet, at least one WAN, at least one MAN, or any combination thereof. "WAN" is an abbreviation for wide area network. "MAN" is an abbreviation for metropolitan area network. Network 40 may include at least one wireless network, at least one optical network, or any combination thereof. The wireless network may be, for example, an ad hoc network, a cellular network, a wireless LAN, a satellite communication network, or a terrestrial microwave network. "LAN" is an abbreviation for local area network.

[0017] Referring to FIG. 1, the outline of this embodiment will be described.

[0018] When the voice output device 20 outputs a question to the user 11 as voice and receives the voice input from the user 11 in response to the question as an answer, it controls a specific function Fp. The voice output device 20 acquires time data indicating the time required for the answer from the user 11 to the question output as voice at the first time point. Based on the time indicated by the acquired time data, the voice output device 20 switches the mode of voice output to the user 11 at the second time point after the first time point.

[0019] According to this embodiment, for a passenger with a large response time lag, such as an elderly person, the mode of voice output can be switched to a dedicated mode. Therefore, the voice output becomes easier for such a passenger to understand. For example, when the passenger is an elderly person, the voice output can be made easier for the passenger to understand by making the voice output slower, shorter, or louder. Alternatively, it is also conceivable to replace the words or expressions included in the voice output with those that are easier for the passenger to understand.

[0020] Referring to FIG. 2, the configuration of the voice output device 20 according to this embodiment will be described.

[0021] The voice output device 20 includes a control unit 21, a storage unit 22, a communication unit 23, an input unit 24, and an output unit 25.

[0022] The control unit 21 includes at least one processor, at least one programmable circuit, at least one dedicated circuit, or any combination thereof. The processor is a general-purpose processor such as a CPU or GPU, or a dedicated processor specialized for specific processing. "CPU" is an abbreviation for central processing unit. "GPU" is an abbreviation for graphics processing unit. The programmable circuit is, for example, an FPGA. "FPGA" is an abbreviation for field-programmable gate array. The dedicated circuit is, for example, an ASIC. "ASIC" is an abbreviation for application specific integrated circuit. The control unit 21 executes processes related to the operation of the voice output device 20 while controlling each part of the voice output device 20.

[0023] The storage unit 22 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or any combination thereof. The semiconductor memory is, for example, RAM, ROM, or flash memory. "RAM" is an abbreviation for random access memory. "ROM" is an abbreviation for read only memory. RAM is, for example, SRAM or DRAM. "SRAM" is an abbreviation for static random access memory. "DRAM" is an abbreviation for dynamic random access memory. ROM is, for example, EEPROM. "EEPROM" is an abbreviation for electrically erasable programmable read only memory. The flash memory is, for example, SSD. "SSD" is an abbreviation for solid-state drive. The magnetic memory is, for example, HDD. "HDD" is an abbreviation for hard disk drive. The storage unit 22 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 22 stores data used for the operation of the voice output device 20 and data obtained by the operation of the voice output device 20.

[0024] The communication unit 23 includes at least one communication module. The communication module is a module compatible with mobile communication standards such as LTE, 4G standard, or 5G standard, or wireless LAN communication standards such as IEEE802.11. "LTE" is an abbreviation for Long Term Evolution. "4G" is an abbreviation for 4th generation. "5G" is an abbreviation for 5th generation. "IEEE" is the abbreviation of Institute of Electrical and Electronics Engineers. The communication unit 23 communicates with the server device 30. The communication unit 23 receives data used for the operation of the voice output device 20 and transmits data obtained by the operation of the voice output device 20.

[0025] The input unit 24 includes at least one input device. The input device is, for example, a physical key, a capacitive key, a pointing device, a touch screen provided integrally with a display, a visible light camera, a depth camera, LiDAR, or a microphone. "LiDAR" is an abbreviation for light detection and ranging. The input unit 24 receives an operation of inputting data used for the operation of the voice output device 20. Instead of being provided in the voice output device 20, the input unit 24 may be connected to the voice output device 20 as an external input device. As the connection interface, an interface corresponding to a standard such as USB, HDMI (registered trademark), or Bluetooth (registered trademark) can be used. "USB" is an abbreviation for Universal Serial Bus. "HDMI (registered trademark)" is an abbreviation for High-Definition Multimedia Interface.

[0026] The output unit 25 includes at least one output device. The output device is, for example, a display or a speaker. The display is, for example, an LCD or an organic EL display. "LCD" is an abbreviation for liquid crystal display. "EL" is an abbreviation for electro luminescent. The output unit 25 outputs the data obtained by the operation of the voice output device 20. Instead of being provided in the voice output device 20, the output unit 25 may be connected to the voice output device 20 as an external output device such as a display audio. As the connection interface, an interface corresponding to a standard such as USB, HDMI (registered trademark), or Bluetooth (registered trademark) can be used.

[0027] The functions of the voice output device 20 are realized by executing the program according to this embodiment on a processor as the control unit 21. That is, the functions of the voice output device 20 are realized by software. The program causes a computer to execute the operations of the voice output device 20, thereby causing the computer to function as the voice output device 20. That is, the computer functions as the voice output device 20 by executing the operations of the voice output device 20 according to the program.

[0028] The program can be stored in a non-transitory computer-readable medium. The non-transitory computer-readable medium is, for example, a flash memory, a magnetic recording device, an optical disc, a magneto-optical recording medium, or a ROM. The distribution of the program is performed, for example, by selling, transferring, or lending a portable medium such as an SD card, a DVD, or a CD-ROM storing the program. "SD" is an abbreviation for Secure Digital. "DVD" is an abbreviation for digital versatile disc. "CD-ROM" is an abbreviation for compact disc read only memory. The program may be stored in the storage of a server and transferred from the server to other computers to distribute the program. The program may be provided as a program product.

[0029] The computer temporarily stores, for example, a program stored in a portable medium or a program transferred from a server in the main memory device. Then, the computer reads the program stored in the main memory device with the processor and executes processing according to the read program with the processor. The computer may directly read the program from the portable medium and execute processing according to the program. The computer may sequentially execute processing according to the received program each time a program is transferred from the server to the computer. The processing may be executed by a so-called ASP type service that realizes functions only by execution instructions and result acquisition without transferring the program from the server to the computer. "ASP" is an abbreviation for application service provider. The program includes information for use in processing by an electronic computer and things conforming to the program. For example, data that is not a direct instruction to the computer but has the property of defining the processing of the computer corresponds to "things conforming to the program".

[0030] Some or all of the functions of the voice output device 20 may be realized by a programmable circuit or a dedicated circuit as the control unit 21. That is, some or all of the functions of the voice output device 20 may be realized by hardware.

[0031] Referring to FIG. 3, the operation of the voice output device 20 according to the present embodiment will be described. The operations described below correspond to the control method according to the present embodiment. That is, the control method according to the present embodiment includes steps S1 to S7 shown in FIG. 3.

[0032] When the user 11 issues an activation command such as "Hey, car!" or presses an activation button displayed on the screen or physically arranged, the step of S1 is started.

[0033] In S1, the control unit 21 outputs a question as voice to the user 11. Specifically, the control unit 21 outputs, as voice from the speaker as the output unit 25, a question asking about the desire of the user 11.

[0034] In S2, the control unit 21 receives, as a response, the voice input from the user 11 to the question output by voice in S1. Specifically, the control unit 21 receives, via the microphone as the input unit 24, a response that conveys the request of the user 11.

[0035] In S3, the control unit 21 controls a specific function Fp. Specifically, the control unit 21 analyzes the response received in S2 to recognize the request of the user 11. As a method for voice analysis of the response, a known method can be used. Machine learning such as deep learning may be used. The control unit 21 controls, as the specific function Fp, a function corresponding to the recognized request. The specific function Fp is a function of the device mounted on the vehicle 12. For example, the specific function Fp is a function of an information terminal device such as presentation of information such as weather or news, a function of an audio device such as music playback, a function of a navigation device such as search for a destination, or a function of an air conditioning device such as adjustment of temperature or air volume. Alternatively, the specific function Fp may be a function of the vehicle 12 itself, such as opening and closing of a window or sunroof equipped on the vehicle 12, or presentation of information regarding the vehicle 12 such as fuel consumption.

[0036] In S4, the control unit 21 records, as time data, the time required for the response from the user 11 to the question output by voice in this S1. Assuming that the time when the question was output by voice in this S1 is the first time point, the time data is data indicating the time required for the response from the user 11 to the question output by voice at the first time point. Specifically, the control unit 21 counts the time elapsed from the time when the question was output by voice in S1 to the time when the response was received in S2. The control unit 21 stores, in the storage unit 22 as time data, the data indicating the counted time.

[0037] Not only the time when the question was audibly output in the current S1, but also the time when the question was audibly output in S1 before the previous time may be regarded as the first time point respectively. For example, after identifying the user 11, the control unit 21 may obtain the cumulative value of the time counted so far for the identified user 11 by adding the time counted this time to the cumulative value of the time counted up to the previous time. Then, the control unit 21 may store the obtained cumulative value in the storage unit 22, and may also store in the storage unit 22, as time data, data indicating the average time obtained by dividing the cumulative value by the number of counts.

[0038] In S5 to S7, the control unit 21 acquires the time data. The control unit 21 sets the mode of the audible output to the user 11 in the next S1 to be different modes according to the time indicated by the acquired time data. Assuming that the time when the question is audibly output in the next S1 is the second time point after the first time point, the control unit 21 switches the mode of the audible output to the user 11 at the second time point based on the time required for the user 11 to reply to the question audibly output at the first time point. Specifically, in S5, the control unit 21 acquires the time data from the storage unit 22. The control unit 21 compares the time indicated by the acquired time data with a threshold value Th. The threshold value Th is a fixed value preset in this embodiment, but may be adjusted as appropriate. When the time indicated by the time data is not longer than the threshold value Th, in S6, the control unit 21 sets the mode of the audible output to the user 11 to the normal mode. After S6, the steps after S1 are executed again. On the other hand, when the time indicated by the time data is longer than the threshold value Th, in S7, the control unit 21 sets the mode of the audible output to the user 11 to the elderly mode. After S7, the steps after S1 are executed again.

[0039] As examples of the elderly mode, the first example to the sixth example will be described.

[0040] In the first example, the senior mode is a mode in which the voice output is slower than in the normal mode. For example, assuming that the normal mode was applied at startup, in S1, the control unit 21 outputs the general question "What can I do for you?" in the normal speed as voice. On the other hand, assuming that the senior mode was applied at startup, in S1, the control unit 21 outputs the same question "What can I do for you?" in a slower voice than normal.

[0041] In the second example, the senior mode is a mode in which the voice output is shorter than in the normal mode. For example, assuming that the normal mode was applied when the general request of user 11, "Tell me the weather today," was transmitted as voice, in S1, the control unit 21 outputs the specific question "Which area's weather should I inform you of?" as voice. On the other hand, assuming that the senior mode was applied when the general request of user 11, "Tell me the weather today," was transmitted as voice, in S1, the control unit 21 outputs the shorter question than normal, "Which area?" as voice.

[0042] In the third example, the senior mode is a mode in which the voice output is louder than in the normal mode. For example, assuming that the normal mode was applied when the general request of user 11, "Open the window," was transmitted as voice, in S1, the control unit 21 outputs the specific question "Which seat's window should I open?" in the normal volume as voice. On the other hand, assuming that the senior mode was applied when the general request of user 11, "Open the window," was transmitted as voice, in S1, the control unit 21 outputs the same question "Which seat's window should I open?" in a louder volume than normal as voice.

[0043] In the fourth example, the elderly mode is a mode in which there are fewer technical terms in the questions output as voice than in the normal mode. For example, assuming that when a general request of user 11 such as "Play music" is conveyed by voice, the normal mode is applied, in S1, the control unit 21 outputs as voice a specific question "Shall we switch to Bluetooth (registered trademark) audio?" On the other hand, assuming that when a general request of user 11 such as "Play music" is conveyed by voice, the elderly mode is applied, in S1, the control unit 21 outputs as voice a question "Shall we make it possible to listen to the music in your mobile phone?" which has fewer technical terms than normal.

[0044] In the fifth example, the elderly mode is a mode in which the words included in the questions output as voice in the normal mode are replaced with other words having the same meaning within the utterance of user 11. At least while the elderly mode is applied, the control unit 21 monitors the utterance of user 11 via the microphone as the input unit 24. For example, when the elderly mode is applied, in S1, before outputting the question as voice, the control unit 21 replaces the words usually included in the question with other words having the same meaning used within the utterance of user 11 as much as possible.

[0045] In the sixth example, the elderly mode is a mode in which the expressions included in the questions output as voice in the normal mode are replaced with dialects having the same meaning within the utterance of user 11. At least while the elderly mode is applied, the control unit 21 monitors the utterance of user 11 via the microphone as the input unit 24. For example, when the elderly mode is applied, in S1, before outputting the question as voice, the control unit 21 replaces the expressions usually included in the question with dialects having the same meaning used within the utterance of user 11 as much as possible.

[0046] Among the first to sixth examples, any combination of two or more may be applied simultaneously. The mode of voice output directed to the user 11 at the time of factory shipment of the voice output device 20 or the vehicle 12 is assumed to be set to the normal mode in the present embodiment, but it may be set to the elderly mode.

[0047] As described above, the control unit 21 switches the mode of voice output directed to the user 11 at the second time point after the first time point between the normal mode and the elderly mode depending on whether the time required for the response from the user 11 to the question voice - output at the first time point is longer than the threshold Th. That is, the control unit 21 estimates the age of the user 11 to some extent based on whether the time lag of the past response is large, and sets the mode of voice output to a mode suitable for the estimated age. Therefore, according to the present embodiment, the voice output becomes easier for the user 11 to understand.

[0048] Instead of the threshold Th, two or more thresholds may be set. For example, assume that a first threshold and a second threshold larger than the first threshold are set. In such a modification, when the time indicated by the time data is not longer than the first threshold, the control unit 21 sets the mode of voice output directed to the user 11 to the normal mode. When the time indicated by the time data is longer than the first threshold but not longer than the second threshold, the control unit 21 sets the mode of voice output directed to the user 11 to the first elderly mode. When the time indicated by the time data is longer than the second threshold, the control unit 21 sets the mode of voice output directed to the user 11 to the second elderly mode. The second elderly mode is a mode in which the voice output is slower, shorter, or louder than the first elderly mode. Alternatively, the second elderly mode may be a mode in which there are fewer technical terms included in the questions for which the voice output is made than in the first elderly mode.

[0049] The present disclosure is not limited to the above-described embodiments. For example, two or more blocks described in the block diagram may be integrated, or one block may be divided. Instead of executing two or more steps described in the flowchart in time series according to the description, they may be executed in parallel or in a different order according to the processing capabilities of the device that executes each step, or as necessary. In addition, changes can be made without departing from the spirit of the present disclosure.

Explanation of Signs

[0050] 10 System 11 User 12 Vehicle 20 Voice Output Device 21 Control Unit 22 Storage Unit 23 Communication Unit 24 Input Unit 25 Output Unit 30 Server Device 40 Network

Claims

1. A control unit that outputs a question as voice to a vehicle occupant and receives a voice input from the occupant in response to the question to control a specific function. The control unit obtains time data indicating the time required for the response from the occupant to the question output as voice at a first time point, and based on the time indicated by the obtained time data, switches the mode of voice output to the occupant at a second time point after the first time point. A voice output device comprising the control unit.

2. The control unit switches the mode between a normal mode and a senior mode in which voice output is slower, shorter, or louder than the normal mode depending on whether the time indicated by the time data is longer than a threshold value. The voice output device according to Claim 1.

3. The control unit switches the mode between a normal mode and a senior mode in which the technical terms included in the question for which voice output is performed are fewer than in the normal mode depending on whether the time indicated by the time data is longer than a threshold value. The voice output device according to Claim 1.

4. The control unit monitors the speech of the occupant and switches the mode between a normal mode and a senior mode in which words included in the question for which voice output is performed in the normal mode are replaced with other words having the same meaning in the speech of the occupant depending on whether the time indicated by the time data is longer than a threshold value. The voice output device according to Claim 1.

5. The control unit monitors the speech of the occupant and switches the mode between a normal mode and a senior mode in which expressions included in the question for which voice output is performed in the normal mode are replaced with dialects having the same meaning in the speech of the occupant depending on whether the time indicated by the time data is longer than a threshold value. The voice output device according to Claim 1.

Citation Information

Patent Citations

  • Voice recognition responder

    JP1985247697A

  • Voice converting device

    JP2000112488A

  • Interactive system

    JP2002123385A

  • Heating device

    JP2006153420A

  • Equipment controller of voice recognition type and vehicle

    JP2006208460A