Voice recognition device
The voice recognition device addresses the inefficiencies in predicting passenger intentions by automatically moving second-row seats based on speech recognition, ensuring accurate detection and smooth access to third-row seats.
Patent Information
- Application Number
- JP2024002860
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
AI Technical Summary
Conventional devices have low accuracy in predicting the boarding and alighting intentions of passengers for third-row seats, and require manual operation methods that are often unknown to crew members, leading to inefficiencies.
A voice recognition device that detects boarding or alighting intentions through speech recognition, controlling the movement of second-row seats in vehicles to facilitate easy access to third-row seats based on the presence or absence of occupants in the second row.
Accurately detects intentions to get off or board the third-row seat, enabling automatic seat movement or providing operation guidance, reducing the need for manual intervention and enhancing accessibility.
Smart Images

Figure 2025109129000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a voice recognition device.
Background Art
[0002] Patent Document 1 discloses a seat control device for a vehicle that controls a seat drive mechanism to increase the distance between the second-row seat and the third-row seat when it is predicted that a passenger scheduled to sit in the third-row seat has the intention to board the vehicle.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional devices have problems such as low accuracy in predicting the boarding intention of a passenger scheduled to sit in the third-row seat, inability to predict the alighting intention of a passenger already sitting in the third-row seat, and the possibility that the operation method of the second-row seat may not be known when it becomes necessary to manually increase the distance between the second-row seat and the third-row seat.
[0005] An object of the present disclosure is to accurately detect either the intention to get off from the third-row seat of a vehicle or the intention to board the third-row seat, and when such an intention is detected, to make the second-row seat of the vehicle move automatically or to make the method of moving the second-row seat known.
Means for Solving the Problems
[0006] The voice recognition device according to the present disclosure is a voice recognition device that controls equipment mounted on a vehicle having second-row and third-row seats by voice recognition, When there is no occupant in the seats in the second row, based on the speech content of the speaker, if any intention to get off from the seats in the third row or to board the seats in the third row is detected, a control unit is provided to control the movement of the seats in the second row or to guide how to move the seats in the second row.
Advantages of the Invention
[0007] According to the present disclosure, either the intention to get off from the seats in the third row of the vehicle or the intention to board the seats in the third row is detected with high accuracy. And when such an intention is detected, the seats in the second row of the vehicle will move automatically or the way to move the seats in the second row will be known.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Modes for Carrying Out the Invention
[0009] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings.
[0010] In each figure, the same or corresponding parts are denoted by the same reference numerals. In the description of this embodiment, the description of the same or corresponding parts will be omitted or simplified as appropriate.
[0011] Referring to FIG. 1, the configuration of the vehicle 1 according to this embodiment will be described.
[0012] Vehicle 1 is any type of automobile such as, for example, a gasoline vehicle, a diesel vehicle, a hydrogen vehicle, an HEV, a PHEV, a BEV, or an FCEV. "HEV" is an abbreviation for hybrid electric vehicle. "PHEV" is an abbreviation for plug-in hybrid electric vehicle. "BEV" is an abbreviation for battery electric vehicle. "FCEV" is an abbreviation for fuel cell electric vehicle. Vehicle 1 may be driven by a driver or the driving may be automated at any level. The level of automation is, for example, any one of levels 1 to 5 in the SAE level classification. "SAE" is the abbreviation for Society of Automotive Engineers. Vehicle 1 may also be a vehicle dedicated to MaaS. "MaaS" is an abbreviation for Mobility as a Service.
[0013] In addition to the seats in the first row such as the driver's seat, Vehicle 1 has at least the seats in the second and third rows. In the present embodiment, by automatically or manually moving the seats in the second row, it is assumed that a person sitting in the third row can get out of Vehicle 1 or can more easily get out through the side door of the seats in the second row. Similarly, by automatically or manually moving the seats in the second row, it is assumed that a person outside Vehicle 1 can get into the third row or can more easily get in through the side door of the seats in the second row. Moving the seats in the second row may include reclining the backrests of the seats in the second row forward, sliding the seats in the second row forward, or both. The side door of the seats in the second row is, for example, a sliding door.
[0014] Vehicle 1 is equipped with a microphone 10 attached to each seat, an occupant detection device 20 attached to each seat, a voice recognition device 30, a seat 40 to be operated, a display device 50 attached to each seat, and a speaker 60 attached to each seat. The occupant detection device 20 includes a seating sensor 21, a seat belt wearing sensor 22, a weight sensor 23, or any combination thereof. The voice recognition device 30 is a computer with a voice recognition function. The voice recognition device 30 can communicate with the microphone 10, the occupant detection device 20, the seat 40, the display device 50, and the speaker 60 via a network such as a cable or CAN, or wirelessly. "CAN" is an abbreviation for Controller Area Network. The seat 40 is the seat in the second row in this embodiment. The display device 50 is, for example, an LCD or an organic EL display. "LCD" is an abbreviation for liquid crystal display. "EL" is an abbreviation for electro luminescent.
[0015] Referring to FIG. 1, the outline of this embodiment will be described.
[0016] The voice recognition device 30 is a device that controls the equipment mounted on the vehicle 1 by voice recognition. When the voice recognition device 30 detects an intention to get on or off the seat in the third row based on the speech content of the speaker, it executes the movement of the seat 40. Therefore, according to this embodiment, the seat 40 can be moved by voice, and getting on and off the seat in the third row can be performed smoothly. There is no need to manually move the second-row seat to get into the third-row seat, and the annoyance can be reduced.
[0017] In this embodiment, the microphone 10 can determine from which seat the occupant spoke. The occupant detection device 20 can determine whether an occupant is sitting on the seat 40 to be operated. The voice recognition device 30 determines the feasibility of the electric operation of the seat 40 and the necessity of operation guidance, and can execute what is determined to be necessary.
[0018] For example, if a person who wants to get off at the seat in the third column performs a voice operation saying "want to get off", after determining that there is no person sitting in the seat in the second column using the seating sensor 21, the seat belt wearing sensor 22, or the weight sensor 23, the seat 40 can be reclined electrically. If a person who wants to board the seat in the third column performs a voice operation saying "want to board the third column", after determining that there is no person sitting in the seat in the second column with the seating sensor 21, the seat belt wearing sensor 22, or the weight sensor 23, the seat 40 can be reclined electrically. Therefore, when there is a person who wants to board the seat in the third column or a person who wants to get off from the seat in the third column, the crew member does not need to know the operation method of the seat in the second column that needs to be reclined for boarding and alighting. That is, the crew member does not need to have sufficient knowledge to directly operate the seat in the second column. The crew member does not need to take time-consuming and uncertain measures such as asking an experienced person about the operation method, searching for the operation method by trial and error, or searching the vehicle 1's instruction manual to know the operation method. It is not necessary for an experienced person to be nearby, nor is it necessary for the experienced person to convey the operation method verbally.
[0019] For example, when the seat 40 cannot be moved electrically, a manual guide on how to recline the seat 40 can be displayed on a display device 50 such as a rear seat display. Therefore, when there is a person who wants to board the seat in the third column or a person who wants to get off from the seat in the third column, the crew member can know the operation method of the seat in the second column that needs to be reclined for boarding and alighting.
[0020] Referring to FIG. 1, the configuration of the voice recognition device 30 according to the present embodiment will be described.
[0021] The voice recognition device 30 includes a control unit 31, a storage unit 32, and a communication unit 33.
[0022] The control unit 31 includes at least one processor, at least one programmable circuit, at least one dedicated circuit, or any combination thereof. The processor is a general-purpose processor such as a CPU or GPU, or a dedicated processor specialized for specific processing. "CPU" is an abbreviation for central processing unit. "GPU" is an abbreviation for graphics processing unit. The programmable circuit is, for example, an FPGA. "FPGA" is an abbreviation for field-programmable gate array. The dedicated circuit is, for example, an ASIC. "ASIC" is an abbreviation for application specific integrated circuit. The control unit 31 executes processes related to the operation of the voice recognition device 30 while controlling each part of the voice recognition device 30.
[0023] The storage unit 32 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or any combination thereof. The semiconductor memory is, for example, a RAM, ROM, or flash memory. "RAM" is an abbreviation for random access memory. "ROM" is an abbreviation for read only memory. The RAM is, for example, an SRAM or DRAM. "SRAM" is an abbreviation for static random access memory. "DRAM" is an abbreviation for dynamic random access memory. The ROM is, for example, an EEPROM. "EEPROM" is an abbreviation for electrically erasable programmable read only memory. The flash memory is, for example, an SSD. "SSD" is an abbreviation for solid-state drive. The magnetic memory is, for example, an HDD. "HDD" is an abbreviation for hard disk drive. The storage unit 32 functions as, for example, a main memory device, an auxiliary storage device, or a cache memory. The storage unit 32 stores data used for the operation of the voice recognition device 30 and data obtained by the operation of the voice recognition device 30.
[0024] The communication unit 33 includes at least one communication module. The communication module is a module corresponding to an in-vehicle communication standard such as CAN, a wired LAN communication standard such as Ethernet (registered trademark), a wireless LAN communication standard such as IEEE 802.11, a mobile communication standard such as LTE, 4G standard, or 5G standard, or a short-range wireless communication standard such as Bluetooth (registered trademark). "LAN" is an abbreviation for local area network. "IEEE" is an abbreviation for Institute of Electrical and Electronics Engineers. "LTE" is an abbreviation for Long Term Evolution. "4G" is an abbreviation for 4th generation. "5G" is an abbreviation for 5th generation. The communication unit 33 receives data used for the operation of the voice recognition device 30 and transmits data obtained by the operation of the voice recognition device 30.
[0025] The functions of the voice recognition device 30 are realized by executing the program according to the present embodiment on a processor as the control unit 31. That is, the functions of the voice recognition device 30 are realized by software. The program causes a computer to execute the operations of the voice recognition device 30, thereby enabling the computer to function as the voice recognition device 30. That is, the computer functions as the voice recognition device 30 by executing the operations of the voice recognition device 30 according to the program.
[0026] The program can be stored in a non-transitory computer-readable medium. The non-transitory computer-readable medium is, for example, a flash memory, a magnetic recording device, an optical disc, a magneto-optical recording medium, or a ROM. The distribution of the program is performed, for example, by selling, transferring, or lending a portable medium such as an SD card, a DVD, or a CD-ROM that stores the program. "SD" is an abbreviation for Secure Digital. "DVD" is an abbreviation for digital versatile disc. "CD-ROM" is an abbreviation for compact disc read only memory. The program may be stored in the server's storage and distributed by transferring the program from the server to other computers. The program may be provided as a program product.
[0027] The computer stores, for example, a program stored in a portable medium or a program transferred from a server, once, in the main memory device. Then, the computer reads the program stored in the main memory device with the processor and executes the processing according to the read program with the processor. The computer may directly read the program from the portable medium and execute the processing according to the program. The computer may sequentially execute the processing according to the received program each time a program is transferred from the server to the computer. The processing may be executed by a so-called ASP-type service that realizes the function only by execution instructions and result acquisition without transferring the program from the server to the computer. "ASP" is an abbreviation for application service provider. The program is information for use in processing by an electronic computer and includes those conforming to the program. For example, data that is not a direct instruction to the computer but has the property of defining the processing of the computer corresponds to "those conforming to the program".
[0028] Some or all of the functions of the voice recognition device 30 may be realized by a programmable circuit or a dedicated circuit as the control unit 31. That is, some or all of the functions of the voice recognition device 30 may be realized by hardware.
[0029] Referring to FIG. 2, the operation of the voice recognition device 30 according to the present embodiment will be described. The operations described below correspond to the control method according to the present embodiment. That is, the control method according to the present embodiment includes the steps of S101 to S110 shown in FIG. 2.
[0030] When the user utters a wake-up phrase such as "Hey, car!" or presses a start button that is displayed on the screen or physically arranged, the voice recognition device 30 is activated and the step of S101 is executed.
[0031] In S101, the control unit 31 determines the seat of the speaker from the position of the microphone 10. For example, the control unit 31 receives voice data from the microphone 10 that has picked up sound via the communication unit 33. When the control unit 31 receives voice data from the microphone 10 of the seat in the first row such as the driver's seat, it determines that the seat of the speaker is the seat in the first row. When the control unit 31 receives voice data from the microphone 10 of the seat in the second row, it determines that the seat of the speaker is the seat in the second row. When the control unit 31 receives voice data from the microphone 10 of the seat in the third row, it determines that the seat of the speaker is the seat in the third row.
[0032] In S102, the control unit 31 recognizes the voice of the speaker. For example, the control unit 31 recognizes utterances such as "want to go out" or "want to sit in the third row" included in the voice data received in S101.
[0033] In S103, the control unit 31 estimates the speaker's intention. For example, the control unit 31 analyzes the content of the utterance recognized in S102 to determine whether the speaker intends to get off the seat in the third row or intends to board the seat in the third row. That is, the control unit 31 detects either an utterance intending to get off or an utterance intending to board the seat in the third row. "Getting off" may be reinterpreted as moving to a seat other than the third row. "Boarding" may be reinterpreted as moving from a seat other than the third row.
[0034] If an utterance intending to get off is detected in S103 and it is determined in S101 that the speaker's seat is the seat in the third row, the step of S104 is executed. Also, if an utterance intending to board the seat in the third row is detected in S103, the step of S104 is executed. If an utterance intending to get off is detected in S103 but it is determined in S101 that the speaker's seat is a seat other than the third row, the flow shown in FIG. 2 ends. If an utterance intending to get off is not detected in S103 and an utterance intending to board the seat in the third row is not detected, the flow shown in FIG. 2 also ends.
[0035] In S104, the control unit 31 determines whether there is an occupant in the seat 40 to be operated. For example, the control unit 31 receives sensor data from the occupant detection device 20 of the seat in the second row via the communication unit 33. The control unit 31 refers to the received sensor data to determine whether there is an occupant in the seat in the second row. If there is an occupant in the seat in the second row, the condition for not being able to execute the function is satisfied.
[0036] In S105, the control unit 31 determines whether the electric control of the seat 40 is possible. For example, the control unit 31 receives status data from the seat in the second row via the communication unit 33. The control unit 31 refers to the received status data to determine whether the seat in the second row can be electrically reclined. Alternatively, the control unit 31 may refer to the specification data pre-stored in the storage unit 32 to determine whether the seat in the second row can be electrically reclined. If the seat in the second row cannot be electrically reclined, the condition for not being able to execute the function is satisfied.
[0037] If the conditions for disabling the execution of the functions of both S104 and S105 are not met, the steps after S106 are executed. If the conditions for disabling the execution of the functions of either or both of S104 and S105 are met, the steps after S108 are executed.
[0038] In S106, the control unit 31 outputs a response. For example, the control unit 31 causes a response such as "Please do not touch as the seat in the second row will be reclined" to be displayed on the display device 50 of the speaker's seat or output as audio through the speaker 60 of the speaker's seat via the communication unit 33.
[0039] In S107, the control unit 31 executes the function of moving the seat 40. For example, the control unit 31 electrically controls and reclines the seat in the second row via the communication unit 33. When the step of S107 is completed, the flow shown in FIG. 2 ends.
[0040] In S108, the control unit 31 selects a response. For example, if the condition for disabling the execution of the function in S104 is met, the control unit 31 selects a response such as "Please speak again after the passenger in the second row gets off the train". If the condition for disabling the execution of the function in S105 is met, the control unit 31 selects a response such as "I will guide you on how to recline the seat in the second row".
[0041] In S109, the control unit 31 outputs a response. For example, the control unit 31 causes the response selected in S108 to be displayed on the display device 50 of the speaker's seat or output as audio through the speaker 60 of the speaker's seat via the communication unit 33.
[0042] In S110, the control unit 31 displays an operation guidance regarding how to move the seat 40. For example, the control unit 31 causes the operation manual of the seat in the second row to be displayed on the display device 50 of the speaker's seat via the communication unit 33. When the step of S110 is completed, the flow shown in FIG. 2 ends.
[0043] As a modification, even when a person other than those who wish to get on or off speaks, if the speech clearly indicates an intention to recline the seat 40, the steps after S104 may be executed. For example, even when the sound is collected at the driver's seat or the front passenger seat, the function execution of S107 or the operation guidance display of S110 may be applicable. In this modification, the operation guidance may be displayed on the display device 50 closest to the speaker, or on the display device 50 in the rear seat, or on the display devices 50 for all seats. Alternatively, the speaker may be able to select as an option on which seat's display device 50 to display the operation guidance.
[0044] As described above, in the present embodiment, when there is no occupant in the second row seat of the vehicle 1, the control unit 31, based on the speech content of the speaker, detects either the intention to get off from the third row seat of the vehicle 1 or the intention to get on the third row seat, and then controls and moves the second row seat or guides how to move the second row seat. Therefore, according to the present embodiment, either the intention to get off from the third row seat or the intention to get on the third row seat can be detected with high accuracy. And when such an intention is detected, the second row seat automatically moves or how to move the second row seat becomes known.
[0045] The present disclosure is not limited to the above-described embodiments. For example, two or more blocks described in the block diagram may be integrated, or one block may be divided. Instead of executing two or more steps described in the flowchart in chronological order according to the description, each step may be executed in parallel or in a different order according to the processing ability of the device that executes each step or as necessary. In addition, changes can be made without departing from the spirit of the present disclosure.
Description of Reference Numerals
[0046] 1 Vehicle 10 Microphone 20 Occupant Detection Device 21 Seating Sensor 22 Seat Belt Wearing Sensor 23 Weight Sensor 30 Voice Recognition Device 31 Control Unit 32 Memory Unit 33 Communication Unit 40 Seat 50 Display Device 60 Speaker
Claims
【Claim 1】 A voice recognition device for controlling equipment mounted on a vehicle having seats in the second and third rows by voice recognition, When there is no occupant in the seat in the second row, based on the speaker's utterance content, if it detects either an intention to get off the seat in the third row or an intention to board the seat in the third row, it is provided with a control unit that controls and moves the seat in the second row or guides how to move the seat in the second row. A voice recognition device.
Citation Information
Patent Citations
Getting on / Off auxiliary device for vehicle
JP2003212011A
Information presentation apparatus
JP2022175096A
Sheet control device for vehicle
JP2004082832A