Speech recognition method and speech recognition device
The speech recognition method addresses the limitation of existing systems by estimating and guiding users on in-vehicle device operation through speech and input signal analysis, improving user interaction.
Patent Information
- Application Number
- JP2024545443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-05
- Filing Date
- 2023-06-05
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2043-06-05
AI Technical Summary
Existing speech recognition systems for in-vehicle devices fail to inform occupants about the name or purpose of operation input devices, limiting user interaction and guidance.
A speech recognition method that acquires speech content and input operation signals to estimate a target component, providing guidance on the operation input device based on the speech content and input operation signal.
Enables the notification of information related to operation input devices, enhancing user interaction by providing clear guidance on device operation.
Smart Images

Figure 0007764971000001 
Figure 0007764971000002 
Figure 0007764971000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech recognition method and a speech recognition device. [Background technology]
[0002] Patent Document 1 proposes a technique for operating an in-vehicle device and highlighting an operating section of the in-vehicle device when a voice command to the in-vehicle device is received from a vehicle occupant. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-097378 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology described in Patent Document 1 can inform the occupant of the location of the operation input device that accepts the occupant's operation input to the in-vehicle equipment, but cannot inform the occupant of the name or purpose of the operation input device. The present invention aims to notify an occupant of information relating to an operation input device that accepts an operation input from the occupant to an in-vehicle device. [Means for solving the problem]
[0005] In one aspect of the present invention, a speech recognition method acquires the speech content of a vehicle occupant, acquires an input operation signal generated by the occupant operating an operation input device of the vehicle, and, based on the speech content and the input operation signal, estimates a target component, which is a component mentioned in the speech content among multiple components that make up the vehicle, and outputs information about the target component. For example, in order to start the operation of a vehicle's driving assistance function, it is necessary to press a first switch of the steering switch group provided on the steering wheel that switches the driving assistance function on and off, and then press a second switch that starts the operation of the driving assistance function.Based on the input operation signal generated when the occupant presses the first switch and the occupant's speech content, ``What should I do next?'', it can be assumed that the component mentioned in the speech content is the steering switch group, and an explanatory message on how to use the steering switch group, ``Please press the second switch,'' can be output as information about the steering switch group. [Effects of the Invention]
[0006] According to the present invention, it is possible to notify the occupant of information relating to an operation input device that accepts an operation input from the occupant to an in-vehicle device. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a schematic diagram illustrating an example of a vehicle equipped with a voice recognition device according to an embodiment. [Figure 2] 2 is a block diagram showing an example of a functional configuration of a controller shown in FIG. 1. FIG. [Figure 3] 3 is a flowchart illustrating an example of a speech recognition method according to the first embodiment. [Figure 4] 10 is a flowchart of a speech recognition method according to a first modified example. [Figure 5] 10 is a flowchart of a speech recognition method according to a second modified example. [Figure 6] 10 is a flowchart illustrating an example of a speech recognition method according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the drawings are schematic and may differ from the actual product. Furthermore, the embodiments of the present invention shown below are examples of devices and methods for embodying the technical concept of the present invention, and the technical concept of the present invention does not limit the structure, arrangement, etc. of component parts to those described below. The technical concept of the present invention can be modified in various ways within the technical scope defined by the claims.
[0009] (First embodiment) (composition) 1 is a schematic diagram of an example of a vehicle equipped with a voice recognition device according to an embodiment. The vehicle 1 is equipped with an in-vehicle device 2, a plurality of operation input devices 3, a voice recognition device 4, a push-to-talk (PTT) switch 5, a speaker 6, and a display device 7. The in-vehicle devices 2 are various devices mounted on the vehicle 1. The in-vehicle devices 2 may be, for example, an air conditioning system, an audio device, an interior light, a glove box, a console lamp, an in-vehicle infotainment (IVI) system, or a navigation device.
[0010] The operation input device 3 is a device that accepts operation inputs from a passenger to the in-vehicle device 2. The operation input device 3 may be, for example, a push switch, a click switch, a toggle switch, a rocker switch, a magnetic non-contact switch, a capacitive non-contact switch, a jog dial, a jog lever, a knob, a slide bar, a dial controller, or a touch panel. The push switch may be, for example, an alternate type push switch that maintains the contact state even after the button is released after being pressed, or a momentary type push switch that returns to the state before the button was pressed when the button is released.
[0011] A jog dial is an operation input device that accepts a selection operation or an adjustment operation by rotating an operation unit such as a dial or wheel, and also accepts an operation by pressing the operation unit. The jog lever is an operation input device that accepts a selection operation by tilting the lever, and also accepts an operation by pushing the lever. A dial controller is an operation input device that accepts selection or adjustment operations by rotating a dial, selection operations by tilting a dial, operations by pressing a dial, and operations on the touchpad on the top surface of the dial (e.g., character input, etc.).
[0012] The voice recognition device 4 recognizes the content of speech of the occupant of the vehicle 1 and outputs a guidance message that answers the occupant's questions regarding the operation input device 3. The voice recognition device 4 includes a microphone 8 and a controller 9. The microphone 8 is a voice input device that acquires voice input from the occupant. The controller 9 is an electronic control unit (ECU) that executes voice recognition processing to recognize the content of the occupant's speech. The controller 9 includes a processor 9a and peripheral components such as a storage device 9b. The processor 9a may be, for example, a central processing unit (CPU) or a micro-processing unit (MPU). The storage device 9b may include a semiconductor storage device, a magnetic storage device, an optical storage device, or the like. The storage device 9b may include memories such as a read-only memory (ROM) and a random access memory (RAM), a register, and a cache memory. The functions of the controller 9 described below are realized, for example, by the processor 9a executing a computer program stored in the storage device 9b.
[0013] The PTT switch 5 is an operation input device that allows the occupant to instruct the start of voice recognition processing by the voice recognition device 4. As will be described later, if the start of voice recognition processing is instructed by a wake-up word, a dedicated voice command, or by operating the operation input device 3 other than the PTT switch 5, the PTT switch 5 may be omitted. The speaker 6 is an information presentation device that outputs the voice message generated by the voice recognition device 4. The display device 7 is an information presentation device that displays the text message generated by the voice recognition device 4, as well as images, symbols, and figures.
[0014] Fig. 2 is a block diagram showing an example of the functional configuration of the controller 9 in Fig. 1. The controller 9 includes a voice recognition unit 10, an input operation signal acquisition unit 11, an action determination unit 12, a response generation unit 13, and a device control unit 14. When the voice recognition device 4 is activated, the voice recognition unit 10 maintains the first standby mode until a predetermined voice recognition start event occurs. The voice recognition start event may be a voice input of a common wake-up word (e.g., "Hello XX") for starting the voice recognition process, or may be an input of a dedicated voice command (e.g., "I want to ask about the switch?") for accepting a voice question regarding the operation input device 3. Alternatively, the voice recognition start event may be an operation of the PTT switch 5.
[0015] When a voice recognition start event occurs, the voice recognition unit 10 starts voice recognition processing. The voice recognition unit 10 recognizes voice input from the occupant acquired by the microphone 8 and converts it into linguistic information such as text. The voice recognition unit 10 analyzes the linguistic information using natural language processing to acquire the content of the user's utterance. For example, the voice recognition unit 10 extracts keywords (such as "switch," "lever," and "dial") that refer to the operation input device 3 from the utterance content.
[0016] Furthermore, the voice recognition unit 10 may extract, from the utterance content, the type of question regarding the operation input device 3. For example, when the utterance content is "What is this switch?", the voice recognition unit 10 may determine that the type of question from the occupant is a "question regarding the name" of the operation input device 3. For example, if the utterance content is "Which switch does XX?" or "Where is the switch for XX?", the voice recognition unit 10 may determine that the type of question from the occupant is a "question regarding the use and location" of the operation input device 3.
[0017] Furthermore, for example, when the utterance content is "Is this switch called XX, correct?", the voice recognition unit 10 may determine that the type of question from the occupant is "confirmation of the name of the operation input device 3." For example, if the utterance content is "I want to do XX, is this switch okay?", the voice recognition unit 10 may determine that the type of question from the occupant is "to confirm the purpose and position" of the operation input device 3. The voice recognition unit 10 outputs the acquired utterance content to the action determination unit 12.
[0018] The input operation signal acquisition unit 11 acquires, for each of the multiple operation input devices 3, an input operation signal generated by an occupant operating the operation input device 3. The input operation signal acquisition unit 11 determines whether or not the input operation signal satisfies a predetermined operation determination condition for each operation input device 3. When an operation input device 3 that satisfies the operation determination condition is found, the input operation signal acquisition unit 11 generates an operation detection signal that identifies the operation input device 3 that satisfies the operation determination condition. The operation detection signal may include identification information of the operation input device 3 that satisfies the operation determination condition.
[0019] For example, the input operation signal acquiring unit 11 determines that the operation determination condition is satisfied in the following cases. (1) When you press a push switch, click switch, jog dial or dial controller dial, or jog lever. (2) When a toggle switch, rocker switch, jog lever, or dial controller dial is pushed to a position where it is in one of the operating states.
[0020] (3) When the magnet is removed from the magnetic non-contact switch (4) When a change in capacitance is detected by placing a hand over the capacitive non-contact switch or placing an object on it. (5) When the jog dial or dial on the dial controller is rotated (6) When the knob is rotated (7) When you slide the bar (8) When a change in capacitance of the touchpad on the dial of the dial controller is detected (9) When the state of the Graphical User Interface (GUI) on the touch panel screen is changed or a selection operation is performed by touching the surface of the touch panel or sliding a finger on the surface.
[0021] There are operation input devices 3 that can accept multiple types of operations with a single operation unit. For example, a jog dial can accept a selection operation or adjustment operation by rotating an operation unit such as a dial or wheel, and an operation of pressing the operation unit. A jog lever can accept a selection operation by tilting the lever, and an operation of pressing the lever. A dial controller can accept a selection operation or adjustment operation by rotating the dial, a selection operation by tilting the dial, an operation of pressing the dial, and an operation on the touchpad on the top of the dial (for example, character input, etc.). In the case of such an operation input device 3, different operation detection signals may be generated for different types of operations. For example, the operation detection signal may include identification information for identifying the type of operation. The input operation signal acquisition unit 11 outputs the input operation signal and the operation detection signal acquired from the operation input device 3 to the action determination unit 12 .
[0022] The operation determination unit 12 switches the operation of the voice recognition device 4 according to the acquired result of the speech content of the occupant and the acquired result of the input operation signal of the operation input device 3. That is, when the input operation signal acquisition unit 11 acquires an input operation signal from the operation input device 3 and the voice recognition unit 10 acquires a speech content including a question about the operation input device 3, the operation determination unit 12 outputs a response generation command to the response generation unit 13 to generate a guidance message that answers the speech content, and causes the response generation unit 13 to output a guidance message that answers the estimated question about the operation input device 3. On the other hand, even if the input operation signal acquisition unit 11 acquires an input operation signal from the operation input device 3, if the voice recognition unit 10 does not acquire a speech content including a question regarding the operation input device 3, the action determination unit 12 outputs the input operation signal acquired from the operation input device 3 to the equipment control unit 14. The equipment control unit 14 controls the in-vehicle equipment 2 in accordance with the input operation signal.
[0023] Specifically, when the input operation signal acquisition unit 11 acquires an input operation signal from the operation input device 3 and the voice recognition unit 10 acquires a speech content including a question regarding the operation input device 3, the operation determination unit 12 estimates the operation input device 3 mentioned in the speech content of the occupant among the multiple operation input devices 3 that constitute the vehicle 1, based on the speech content acquired by the voice recognition unit 10 and the input operation signal acquired by the input operation signal acquisition unit 11. For example, the operation input device 3 mentioned in the utterance content is estimated based on the utterance content acquired by the voice recognition unit 10 and the operation detection signal output by the input operation signal acquisition unit 11. The operation input device 3 is an example of the "component ... constituting a vehicle" described in the claims.
[0024] In the first embodiment, when an utterance content including a question about an operation input device 3 is acquired after an input operation signal output from a certain operation input device 3 is acquired, the action determination unit 12 estimates that the operation input device 3 that output the input operation signal is the operation input device 3 mentioned in the utterance content of the occupant. For example, when an utterance content including a question about the operation input device 3 is acquired before a predetermined time has elapsed after the acquisition of the input operation signal, it may be estimated that the operation input device 3 that output the input operation signal is the operation input device 3 mentioned in the utterance content of the occupant. Furthermore, the operation determination unit 12 may determine that an input operation signal has been acquired when it receives an operation detection signal from the input operation signal acquisition unit 11, for example.
[0025] When the operation determination unit 12 estimates the operation input device 3 mentioned in the utterance content, it outputs a response generation command to the response generation unit 13 to generate a guidance message that responds to the utterance content. For example, the response generation command may include identification information of the estimated operation input device 3 and identification information of the type of question included in the occupant's utterance content (for example, "question regarding name," "question regarding use and location," "confirmation of name," or "confirmation of use and location"). Based on the response generation command received from the operation determination unit 12, the response generation unit 13 outputs a guidance message including audio or an image representing information about the estimated operation input device 3 from the speaker 6 or the display device 7 as a response to the question included in the occupant's speech content. In this case, the controller 9 may stop outputting the input operation signal acquired from the operation input device 3 to the device control unit 14 until the response generation unit 13 outputs the guidance message. That is, even if the input operation signal is acquired by operating the operation input device 3, the controller 9 may stop controlling the in-vehicle device 2.
[0026] For example, the response generation unit 13 may generate a voice guidance message representing information related to the inferred operation input device 3 and output it from the speaker 6. Furthermore, for example, the response generation unit 13 may generate a textual guidance message, image, symbol, or graphic representing information related to the inferred operation input device 3 and output it from the display device 7. A specific example of a message generated by the response generation unit 13 will be described below. (Example 1) If the operation input device 3 is a volume control switch for an audio device, an operation detection signal is output when the switch is pressed. In this case, the operation input device 3 may be, for example, a push switch, a click switch, a jog lever (when pressed), or a dial controller (when pressed). When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, such as "This is a volume control switch. Press + to increase the volume and - to decrease the volume."
[0027] When the utterance is "Which switch controls the volume?", the speech recognition unit 10 determines that the type of question is "question about use and location." The response generation unit 13 outputs a guidance message including information about the use, location, and usage method, such as "You can control the volume with the switch marked + and - on the left side of the steering wheel. Press + to increase the volume and - to decrease the volume." When the utterance is "Is this switch a volume control switch?", the speech recognition unit 10 determines that the type of question is "confirm the name." The response generation unit 13 outputs a guidance message saying "Yes, that's right. You can turn up the volume with + and turn down the volume with -." When the utterance is "I want to adjust the volume. Is this button okay?", the speech recognition unit 10 determines that the question type is "Confirm use and location." The response generation unit 13 outputs the guidance message "Yes, that's right. You can increase the volume with + and decrease the volume with -."
[0028] (Example 2) When the operation input device 3 is an item selection switch of a navigation device, an operation detection signal is output when the lever is pushed to a position where one of the operation states is achieved. In this case, the operation input device 3 may be, for example, a toggle switch, a rocker switch, a jog lever (when the lever is pushed), or a dial controller (when the dial is pushed). When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, saying, "This is an item selection switch. You can focus on the item you want to select by tilting it up and down / left and right / pressing it." When the utterance is "Which is the cursor movement / item selection switch?", the speech recognition unit 10 determines that the type of question is "question about use and location." For example, the response generation unit 13 outputs a guidance message including information about use, location, and usage, such as "Item selection can be performed using the round knob-shaped dial on the console. You can focus on the item you want to select by tilting / pressing it up / down / left / right, or by rotating the dial left / right."
[0029] When the utterance is "Is this switch a cursor move / item selection switch?", the speech recognition unit 10 determines that the type of question is "confirm name." For example, the response generation unit 13 outputs a guidance message such as "Yes, that's right. You can focus on the item you want to select by tilting / pressing up / down / left / right, or by rotating the dial left / right." When the utterance content is "I want to select an item, is this button okay?", the speech recognition unit 10 determines that the type of question is "Confirm use and location." For example, the response generation unit 13 outputs a guidance message such as "Yes, that's right. You can focus on the item you want to select by tilting / pressing up / down / left / right, or by rotating the dial left / right."
[0030] (Example 3) When the operation input device 3 is an open / close interlocking switch for a glove box, an operation detection signal is output when the magnet is separated from the magnetic non-contact switch that is the open / close interlocking switch. When the utterance is "What is the switch to turn on the light in the storage compartment in front of the passenger seat?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, saying, "This is the glove box open / close interlock switch. When you open the box, the light will turn on, and when you close it, the light will turn off." When the utterance is "Which switch turns on the glove box light?", the speech recognition unit 10 determines that the type of question is a "question about use and location." The response generation unit 13 outputs a guidance message including information about use, location, and how to use it, saying, "The glove box is a drawer in front of the passenger seat. It can be operated by opening and closing the lid of the glove box. When you open the box, the light comes on, and when you close it, the light goes out."
[0031] When the utterance is "Is this the switch to turn on the glove box light?", the speech recognition unit 10 determines that the type of question is "confirm the name." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. The light will turn on when you open the box, and turn off when you close it." When the utterance is "I want to turn on the glove box light, where is it?", the speech recognition unit 10 determines that the type of question is "confirmation of purpose and location." The response generation unit 13 outputs a guidance message saying, "The glove box is a drawer in front of the passenger seat. It can be operated by opening and closing the glove box lid. When you open the box, the light will turn on, and when you close it, the light will turn off."
[0032] (Example 4) If the operation input device 3 is a capacitance-type non-contact switch that turns a console lamp on and off, an operation detection signal is output when a change in capacitance is detected by holding a hand over the capacitance-type non-contact switch or placing an object on it. When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, such as "This is a switch for the console lamp inside the car. You can turn it on and off by waving your hand over it." When the utterance is "Which is the switch for the console lamp inside the car?", the speech recognition unit 10 determines that the type of question is "question about use and location." The response generation unit 13 outputs a guidance message including information about the use, location, and usage method, saying, "The console lamp can be operated with a switch on the center console. You can turn it on and off by waving your hand over it." When the utterance is "Is this switch a console lamp switch?", the speech recognition unit 10 determines that the type of question is "confirm the name." The response generation unit 13 outputs a guidance message saying "Yes, that's right. You can turn it on and off by waving your hand." When the utterance is "I want to turn on the console lamp, is this button okay?", the speech recognition unit 10 determines that the question type is "Confirm use and location." The response generation unit 13 outputs the guidance message "Yes, that's right. You can turn it on and off by holding your hand over it."
[0033] (Example 5) When the operation input device 3 is a volume control dial of an audio device, an operation detection signal is output when the passenger rotates the dial. In this case, the operation input device 3 may be, for example, a jog dial (when rotated) or a dial controller (when rotated). When the utterance is "What is this dial?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, such as "This is a volume control dial. Turn the dial to the left to decrease the volume, and turn it to the right to increase the volume." When the utterance is "Which dial controls the volume?", the speech recognition unit 10 determines that the type of question is a "question about use and location." The response generation unit 13 outputs a guidance message including information about the use, location, and usage method, such as "You can adjust the volume using the round knob-shaped dial on the bottom left of the IVI screen. Turn the dial to the left to decrease the volume, and turn it to the right to increase the volume."
[0034] When the utterance content is "Is this dial the volume control dial?", the speech recognition unit 10 determines that the type of question is "confirm the name." The response generation unit 13 outputs a guidance message saying "Yes, that's right. Turn the dial left to decrease the volume, and turn it right to increase the volume." When the utterance is "I want to adjust the volume. Is this button okay?", the speech recognition unit 10 determines that the question type is "Confirm use and location." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. Turn the dial left to decrease the volume, and turn it right to increase the volume."
[0035] (Example 6) When the operation input device 3 is an air volume adjustment knob for an air conditioner, an operation detection signal is output when the passenger turns the knob. When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, saying, "This is an airflow control switch. Turn it left to decrease the airflow, and turn it right to increase the airflow." When the utterance is "Which switch controls the airflow?", the speech recognition unit 10 determines that the question type is "Question about use and location." The response generation unit 13 outputs a guidance message including information about the use, location, and usage method, such as "You can adjust the airflow using the knob on the left side below the IVI. Turn it left to decrease the airflow, and turn it right to increase the airflow."
[0036] When the utterance is "Is this switch the airflow control switch?", the speech recognition unit 10 determines that the type of question is "Confirm the name." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. Turn it left to decrease the airflow, and turn it right to increase the airflow." When the utterance is "I want to adjust the air volume. Is this button okay?", the voice recognition unit 10 determines that the question type is "Confirm use and location." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. Turn it left to decrease the air volume, and turn it right to increase the air volume."
[0037] (Example 7) When the operation input device 3 is a slide bar used as a vehicle interior light switch, an operation detection signal is output when the passenger slides the bar. When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, saying, "This is the interior light switch. You can operate it by sliding it to the left for off, to the center for door interlocking, and to the right for on." When the utterance is "Which switch is used for the interior lights?", the speech recognition unit 10 determines that the type of question is a "question about use and location." The response generation unit 13 outputs a guidance message including information about the use, location, and usage method, saying, "The interior lights can be operated with the slide switch near the ceiling rearview mirror. Slide to the left to turn them off, to the center to link them to the door, and to the right to turn them on." When the utterance is "Is this switch the interior light switch?", the speech recognition unit 10 determines that the type of question is "confirm the name." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can operate them by sliding left for off, center for door interlocking, and right for on." When the utterance is "I want to use the interior lights. Is this button okay?", the speech recognition unit 10 determines that the question type is "Confirm use and position." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can operate them by sliding left for off, center for door interlocking, and right for on."
[0038] (Example 8) When the operation input device 3 is a dial controller used for inputting operations into a navigation device or operating an audio device, an operation detection signal is output when a change in capacitance of the touch pad on the top of the dial is detected. When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name and how to use it, saying, "This is a dial controller. You can manually input characters on the dial surface. You can also select items and adjust the volume by rotating the knob left and right, and tilting / pressing it forward, backward, left and right." When the utterance is "Which switch allows you to manually input characters?", the speech recognition unit 10 determines that the type of question is "Question about use and location." The response generation unit 13 outputs a guidance message including information about use, location, and usage, saying, "You can manually input characters on the dial surface. You can also select items and adjust the volume by rotating the knob left and right, or tilting / pressing it forward, backward, left, or right."
[0039] When the spoken content is "Is this the switch that can input characters?", the voice recognition unit 10 determines that the type of question is "Confirm the name." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can manually input characters on the surface of the dial. You can also select items and adjust the volume by rotating the knob left and right, and tilting / pressing it forward, backward, left and right." When the spoken content is "I want to input characters manually, is this button okay?", the voice recognition unit 10 determines that the type of question is "Confirm use and location." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can input characters manually on the dial surface. You can also select items and adjust the volume by rotating the knob left and right, and tilting / pressing it forward, backward, left and right."
[0040] (Example 9) If the operation input device 3 is a touch panel on the IVI screen, an operation detection signal is output when the occupant touches the surface of the touch panel or slides their finger, causing the GUI state of the touch panel to change or a selection operation to be performed. When the utterance is "What is this switch?", the speech recognition unit 10 determines that the type of question is a "question about the name." The response generation unit 13 outputs a guidance message including information about the name, such as "This is the IVI setting icon. You can configure settings related to language setting, navigation, telephone, etc." When the utterance is "Which switch is used to set up the IVI?", the speech recognition unit 10 determines that the type of question is "question about use and location." The response generation unit 13 outputs a guidance message including information about use, location, and usage, such as "You can set up the IVI using the gear icon in the upper right / left corner of the IVI screen. You can set up language, navigation, telephone, etc."
[0041] When the utterance is "Is this switch the correct IVI setting switch?", the voice recognition unit 10 determines that the type of question is "Confirm name." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can configure settings related to language setting, navigation, telephone, etc." When the utterance is "I want to set up the IVI. Is this button OK?", the voice recognition unit 10 determines that the question type is "Confirm use and location." The response generation unit 13 outputs a guidance message saying, "Yes, that's right. You can set up language settings, navigation, telephone settings, etc."
[0042] It should be noted that there are cases where a single operation input device 3 can accept multiple types of operations, such as a jog dial, a jog lever, or a dial controller. When different names or uses are assigned to different types of operations of such an operation input device 3, the response generation unit 13 may generate a guidance message including information on the different names and uses for a single operation input device 3. For example, when the dial controller is pressed as shown above (Example 1), if an utterance including a question from the occupant to the operation input device 3 is acquired, a guidance message may be generated informing the driver that the name and purpose of the dial controller are "volume control switch" and "volume control," respectively. On the other hand, when the dial of the dial controller is pushed to a position that results in one of the operation states (when operating the lever) as shown above (Example 2), if an utterance including a question from the occupant to the operation input device 3 is acquired, a guidance message may be generated informing the driver that the name and purpose of the dial controller are the "item selection switch" and "focus for the item to be selected," respectively.
[0043] Furthermore, when a single operation input device 3 can accept multiple types of operations, a unique use may be assigned to a combination or sequence of a series of different types of operations. For example, a first use may be assigned when the dial controller is tilted while being rotated, and a second use may be assigned when the dial controller is pressed down while being tilted. In this case, when a series of different types of operations are performed on the operation input device 3, if an utterance including a question from the occupant to the operation input device 3 is acquired, a guidance message may be generated to inform the user of the purpose assigned to the combination and sequence of these operations.
[0044] If the occupant wishes to interrupt the guidance message while the response generation unit 13 is outputting the guidance message (i.e., after the output of the guidance message has started but before the output is completed), the occupant can perform a predetermined interrupt instruction operation. For example, the occupant may perform the interrupt instruction operation by again operating the operation input device 3 mentioned in the utterance, or by operating an operation input device other than the operation input device mentioned in the utterance among the multiple operation input devices 3, or by pressing and holding the PTT switch 5, or by uttering a specific keyword (e.g., "interrupt the guidance"). When the interrupt instruction operation is received, the response generation unit 13 interrupts the output of the guidance message. Furthermore, the operation determination unit 12 outputs the input operation signal acquired from the operation input device 3 to the equipment control unit 14. The equipment control unit 14 controls the in-vehicle equipment 2 in accordance with the input operation signal.
[0045] Even if an input operation signal is acquired, if the occupant's speech content is not acquired within a predetermined period (for example, 3 seconds), the voice recognition unit 10 terminates the voice recognition process. In this case, the operation determination unit 12 does not estimate the operation input device 3 based on the occupant's speech, and the response generation unit 13 does not output a guidance message including information about the operation input device 3, but outputs an end guide message "Speech recognition is ending" to inform the occupant that the voice recognition process is ending. Furthermore, the operation determination unit 12 outputs the input operation signal acquired from the operation input device 3 to the equipment control unit 14. The equipment control unit 14 controls the in-vehicle equipment 2 in response to the input operation signal.
[0046] In addition, even if the voice recognition unit 10 detects a voice recognition start event, if the input operation signal acquisition unit 11 does not acquire an input operation signal within a predetermined period, the operation determination unit 12 does not estimate the operation input device 3 based on the occupant's speech. The response generation unit 13 does not output a guidance message including information about the operation input device 3, but outputs an end guidance message. Furthermore, if the voice recognition unit 10 acquires an input operation signal before detecting a voice recognition start event (i.e., before the voice recognition process is started), the action determination unit 12 does not estimate the operation input device 3 based on the occupant's speech, but outputs the input operation signal acquired from the operation input device 3 to the equipment control unit 14. As a result, the response generation unit 13 does not output a guidance message including information about the operation input device 3, and the equipment control unit 14 controls the in-vehicle equipment 2 in accordance with the input operation signal.
[0047] (operation) FIG. 3 is a flowchart of an example of the voice recognition method of the first embodiment. In step S1, the voice recognition unit 10 determines whether or not a voice recognition start event has occurred. If a voice recognition start event has occurred (step S1: Y), the process proceeds to step S4. If a voice recognition start event has not occurred (step S1: N), the process proceeds to step S2. In step S2, the input operation signal acquisition unit 11 determines whether or not an input operation signal has been acquired. If an input operation signal has been acquired (step S2: Y), the process proceeds to step S3. If an input operation signal has not been acquired (step S2: N), the process proceeds to step S12. In step S3, the operation determination unit 12 outputs the input operation signal to the device control unit 14. The device control unit 14 controls the in-vehicle device 2 in accordance with the input operation signal. Thereafter, the process proceeds to step S12.
[0048] In step S4, the input operation signal acquisition unit 11 determines whether or not an input operation signal has been acquired. If an input operation signal has been acquired (step S4: Y), the process proceeds to step S6. If an input operation signal has not been acquired (step S4: N), the process proceeds to step S5. In step S5, the response generation unit 13 outputs an end guide message. Thereafter, the process proceeds to step S12. In step S6, the operation determination unit 12 determines whether or not the occupant's utterance content has been acquired. If the utterance content has been acquired (step S6: Y), the process proceeds to step S7. If the utterance content has not been acquired (step S6: N), the process proceeds to step S9.
[0049] In step S7, the operation determination unit 12 estimates, based on the utterance content and the input operation signal, the operation input device 3 mentioned in the occupant's utterance content among the multiple operation input devices 3 constituting the vehicle 1. The response generation unit 13 outputs a guidance message including information about the estimated operation input device 3. In step S8, the operation determination unit 12 determines whether or not the occupant has performed an interrupt instruction operation. If an interrupt instruction operation has been performed (step S8: Y), the process proceeds to step S9. If an interrupt instruction operation has not been performed (step S8: N), the process proceeds to step S11. In step S9, the response generation unit 13 outputs an end guide message. In step S10, the operation determination unit 12 outputs an input operation signal to the device control unit 14. The device control unit 14 controls the in-vehicle device 2 in accordance with the input operation signal. Thereafter, the process proceeds to step S12.
[0050] In step S11, the response generator 13 determines whether the output of the guidance message has been completed. If the output of the guidance message has been completed (step S11: Y), the process proceeds to step S12. If the output of the guidance message has not been completed (step S11: N), the process returns to step S7. In step S12, the controller 9 determines whether the ignition (IGN) switch of the vehicle is turned off. If the IGN switch is not turned off (step S12: N), the process returns to step S1. If the IGN switch is turned off (step S12: Y), the process ends.
[0051] (First Modification) In the first modification, when the operation input device 3 (i.e., an operation input device other than the PTT switch 5) is operated, it is determined that a voice recognition start event has occurred, and the voice recognition process is started. That is, an input operation signal is acquired before starting the voice recognition process. For example, the voice recognition unit 10 may determine that a voice recognition start event has occurred when it receives an operation detection signal from the input operation signal acquisition unit 11, and start the voice recognition process. If the occupant wishes to terminate the voice recognition process even after operating the operation input device 3 (for example, if the occupant does not need the guidance message of the operation input device 3 and wishes to immediately operate the in-vehicle device 2), the occupant can perform a predetermined interrupt instruction operation. Also, for example, when a predetermined operation (long press of a button, repeated taps, turning a dial left and right, etc.) is received, the operation of the operation device may be interrupted for a certain period of time (without issuing an input operation signal) and only wait for voice.
[0052] For example, the occupant may perform the interrupt instruction operation by again operating the operation input device 3 mentioned in the utterance, or may perform the interrupt instruction operation by operating one of the multiple operation input devices 3 other than the operation input device mentioned in the utterance. When the interrupt instruction operation is received, the response generation unit 13 interrupts the voice recognition. Furthermore, the action determination unit 12 outputs the input operation signal acquired from the operation input device 3 to the equipment control unit 14. The equipment control unit 14 controls the in-vehicle equipment 2 in accordance with the input operation signal.
[0053] FIG. 4 is a flowchart of the voice recognition method of the first modified example. In step S20, the input operation signal acquisition unit 11 determines whether or not an input operation signal has been acquired. If an input operation signal has been acquired (step S20: Y), the process proceeds to step S21. If an input operation signal has not been acquired (step S20: N), the process proceeds to step S28. In step S21, the operation determination unit 12 determines whether or not the occupant has performed an interrupt instruction operation. If an interrupt instruction operation has been performed (step S21: Y), the process proceeds to step S22. If an interrupt instruction operation has not been performed (step S21: N), the process proceeds to step S24. In step S22, the response generation unit 13 outputs an end guide message. In step S23, the operation determination unit 12 outputs the input operation signal to the device control unit 14. The device control unit 14 controls the in-vehicle device 2 in accordance with the input operation signal. Thereafter, the process proceeds to step S28. The processes in steps S24 to S28 are similar to the processes in steps S6 to S8, S11, and S12 in FIG. 1, respectively.
[0054] (Second Modification) In the second modification, similarly to the first modification, when the operation input device 3 (i.e., an operation input device other than the PTT switch 5) is operated, it is determined that a voice recognition start event has occurred, and the voice recognition process is started. That is, an input operation signal is acquired before the voice recognition process is started. In the second variant, when the input operation signal is acquired and then the occupant's speech content is acquired, control of the in-vehicle equipment 2 is performed according to the input operation signal, and a guidance message regarding the operation input device 3 mentioned in the occupant's speech content is output.
[0055] 5 is a flowchart of the voice recognition method of the second modified example. In step S30, the input operation signal acquisition unit 11 determines whether or not an input operation signal has been acquired. If an input operation signal has been acquired (step S30: Y), the process proceeds to step S31. If an input operation signal has not been acquired (step S30: N), the process proceeds to step S37. In step S31, the operation determination unit 12 outputs the input operation signal to the device control unit 14. The device control unit 14 controls the in-vehicle device 2 in accordance with the input operation signal. In step S32, the operation determination unit 12 determines whether or not the occupant's utterance content has been acquired. If the utterance content has been acquired (step S32: Y), the process proceeds to step S34. If the utterance content has not been acquired (step S6: N), the process proceeds to step S33. In step S33, the response generator 13 outputs an end guide message, after which the process proceeds to step S37. The processes in steps S34 to S37 are similar to the processes in steps S7, S8, S11 and S12 in FIG. 1, respectively.
[0056] (Third Modification) The voice recognition unit 10 may determine whether the type of question is a "question about how to use" the operation input device 3. For example, when the utterance content is "How do I use this switch?", the voice recognition unit 10 may determine that the type of question from the occupant is a "question about how to use" the operation input device 3. When the type of question is a "question about how to use" the operation input device 3, the response generation unit 13 may output a guidance message including information about how to use the operation input device 3.
[0057] For example, if the occupant's utterance after receiving an operation detection signal from the input operation signal acquisition unit 11 is a question, the voice recognition unit 10 may determine that the type of question is a ``question about how to use'' the operation input device 3. For example, suppose that in order to start the operation of the driving assistance function of vehicle 1, it is necessary to press a first switch among the steering switches provided on the steering wheel, which switches the driving assistance function on and off, and then press a second switch, which starts the operation of the driving assistance function. In this case, if the occupant's utterance after operating the first operation input device is a question such as "What should I do next?", it may be determined that the type of question from the occupant is a "question about how to use" the operation input device 3. Then, an explanatory message about how to use the steering switch group, "Please press the second switch," may be output.
[0058] (Fourth Modification) The voice recognition unit 10 may extract, from the utterance content, an operation instruction for the in-vehicle device 2. For example, if the utterance content is "move this" or "set this to XX", the voice recognition unit 10 may determine that the utterance content from the occupant is an operation instruction for the in-vehicle device 2. When the operation determination unit 12 receives an operation detection signal from the input operation signal acquisition unit 11 and acquires the speech content including an operation instruction for the in-vehicle device 2, the operation determination unit 12 may estimate that the operation input device 3 mentioned in the speech content of the occupant (i.e., the operation input device 3 used to operate the in-vehicle device 2 to be operated) is the operation input device 3 that output the input operation signal. Then, the operation determination unit 12 outputs a control signal for operating the in-vehicle device 2 in accordance with the operation instruction in the speech content to the device control unit 14. The device control unit 14 controls the in-vehicle device 2 in accordance with the control signal from the operation determination unit 12.
[0059] For example, when the operation determination unit 12 acquires an utterance content including an operation instruction for an in-vehicle device 2 after outputting a guidance message related to the operation input device 3 that output an input operation signal as described above, the operation determination unit 12 may estimate that the operation input device 3 mentioned in the utterance content including the operation instruction is the operation input device 3 of the guidance message. Then, the operation determination unit 12 may operate the in-vehicle device operated by this operation input device 3 according to the operation instruction in the utterance content. For example, assume that the in-vehicle device 2 is an interior light, the operation input device 3 is an interior light switch, and the occupant operates the interior light switch. In response to the utterance content "What is this switch?", the guidance message "This is the interior light switch. You can operate them by sliding to the left for off, to the center for door interlocking, and to the right for on" is output as described above, and then if the occupant utters "Set this to on," the operation determination unit 12 may infer that the operation input device 3 mentioned in the utterance content including the operation instruction is the interior light switch and the in-vehicle device 2 to be operated is the interior light, and may control the interior light to be on.
[0060] (Second embodiment) In the first embodiment, after an input operation signal is acquired, if a speech content including a question regarding the operation input device 3 is acquired, the operation input device 3 mentioned in the speech content of the occupant is estimated, and a guidance message regarding the estimated operation input device 3 is output. In contrast, in the second embodiment, when an input operation signal is acquired after acquiring a speech content including a question about the operation input device 3, the operation input device 3 mentioned in the speech content of the occupant is estimated, and a guidance message about the estimated operation input device 3 is output. In the voice recognition unit 10 of the second embodiment, voice input of a wake-up word, input of a dedicated voice command for accepting a question (e.g., "I'd like to ask about the switch?"), and operation of the PTT switch 5 may also be detected as voice recognition start events.
[0061] Alternatively, the voice recognition unit 10 of the second embodiment may constantly recognize voice input from the occupant acquired by the microphone 8, analyze the content of the speech using natural language processing, and determine whether a question regarding the operation input device 3 (for example, "What is this switch?", "Which switch does XX?", "Where is the switch for XX?", "Is this switch XX correct?", "I want to do XX, is this switch okay?") has been input. When a question regarding the operation input device 3 is input, the operation determination unit 12 transitions to a standby mode in which it monitors whether the input operation signal acquisition unit 11 acquires an input operation signal. When the input operation signal is acquired in the standby mode, the operation determination unit 12 estimates the operation input device 3 mentioned in the utterance content of the occupant. The response generation unit 13 outputs a guidance message regarding the estimated operation input device 3.
[0062] Fig. 6 is a flowchart of an example of a voice recognition method according to the second embodiment. The processes of steps S40 to S42 are the same as the processes of steps S1 to S3 in Fig. 1. If a voice recognition start event occurs (step S40: Y), the process proceeds to step S43. In step S43, the operation determination unit 12 determines whether or not the occupant's utterance content has been acquired. If the occupant's utterance content has been acquired (step S43: Y), the process proceeds to step S44. If the occupant's utterance content has not been acquired (step S43: N), the process proceeds to step S45.
[0063] In step S44, the input operation signal acquisition unit 11 determines whether or not an input operation signal has been acquired. If an input operation signal has been acquired (step S44: Y), the process proceeds to step S46. If an input operation signal has not been acquired (step S44: N), the process proceeds to step S45. In step S45, the response generation unit 13 outputs an end guide message. Thereafter, the process proceeds to step S51. The processing in steps S46 to S51 is the same as that in steps S7 to S12 in FIG.
[0064] (Effects of the embodiment) (1) In the voice recognition method, the speech content of the occupant of the vehicle 1 is acquired, an input operation signal generated by the occupant operating the operation input device 3 of the vehicle 1 is acquired, and based on the speech content and the input operation signal, a target component, which is the component mentioned in the speech content among multiple components that make up the vehicle 1, is estimated, and information about the target component is output. This allows the occupant to be notified of information about the operation input device 3 that accepts the occupant's operation input to the in-vehicle device 2.
[0065] (2) For example, the utterance content may be acquired after the input operation signal is acquired, thereby making it possible to estimate the operation input device 3 that generated the input operation signal as the target component. (3) For example, if the speech content is not acquired even after a predetermined time has elapsed since the input operation signal was acquired, the in-vehicle device 2 may be controlled in accordance with the input operation signal. As a result, when the speech content is not acquired, the in-vehicle device 2 can be controlled in the same way as when the operation input device 3 is operated normally.
[0066] (4) For example, it may be possible to determine whether a voice recognition process for acquiring the contents of an occupant’s speech has started, and if an input operation signal is acquired before the voice recognition process has started, to control an in-vehicle device according to the input operation signal without outputting information about the target component. As a result, when the voice recognition process has not started, the in-vehicle device 2 can be controlled in the same way as when the operation input device 3 is operated normally.
[0067] (5) For example, it may be possible to determine whether a voice recognition process for acquiring the contents of an occupant's speech has started, and if an input operation signal is acquired before the voice recognition process is started, to control an in-vehicle device according to the input operation signal and output information about the target component. As a result, even in a configuration in which voice recognition processing for a question regarding the operation input device 3 is started based on the operation of the operation input device 3, control of the in-vehicle device 2 and voice recognition processing can be performed at the same time.
[0068] (6) For example, the input operation signal may be acquired after acquiring the speech content, thereby making it possible to estimate the operation input device 3 that generated the input operation signal as the target component. (7) For example, if an occupant speaks or operates an input device while information about the target component is being output, the output of information about the target component may be interrupted and control of the in-vehicle equipment may be performed in accordance with the input operation signal. This allows the control of the in-vehicle device 2 to be started immediately when the information on the target component becomes unnecessary.
[0069] (8) For example, the target component may be the operation input device 3. For example, the target component may be a switch, a lever, a dial, a knob, a slide bar, or a touch panel. This allows information about the operation input device 3 to be communicated to the occupant. (9) For example, it may be possible to determine whether the content of the utterance is a question about the name, the method of use, or the purpose, and if it is determined that the content of the utterance is a question about the name, the method of use, or the purpose, output the name, the method of use, or the purpose of the target component as information about the target component. This allows the occupant to be informed of the name, the method of use, or the purpose of the operation input device 3. (10) For example, a sound or an image representing information about the target component may be output, thereby notifying the occupant of information about the operation input device 3. [Explanation of symbols]
[0070] 1...vehicle, 2...in-vehicle equipment, 3...operation input device, 4...voice recognition device, 5...push-to-talk switch, 6...speaker, 7...display device, 8...microphone, 9...controller, 9a...processor, 9b...storage device, 10...voice recognition unit, 11...input operation signal acquisition unit, 12...action determination unit, 13...response generation unit, 14...equipment control unit
Claims
1. Acquire the speech content of the vehicle occupant, acquiring an input operation signal generated by the occupant operating an operation input device of the vehicle; estimating a target component that is a component mentioned in the utterance content among a plurality of components that configure the vehicle based on the utterance content and the input operation signal; outputting information about the target component; A voice recognition method characterized by the fact that, when an utterance by the occupant or an operation of the operation input device is detected while information about the target component is being output, the output of information about the target component is interrupted and control of in-vehicle equipment is performed in accordance with the input operation signal obtained from the operation input device.
2. 2. The speech recognition method according to claim 1, wherein the speech content is acquired after the input operation signal is acquired.
3. 3. The speech recognition method according to claim 2, wherein, if the speech content is not acquired even after a predetermined time has elapsed since the input operation signal was acquired, control of an in-vehicle device is executed in accordance with the input operation signal.
4. determining whether a voice recognition process for acquiring the speech of the occupant has been started; If the input operation signal is acquired before the voice recognition process is started, control of an in-vehicle device according to the input operation signal is executed without outputting information about the target component.
3. The speech recognition method according to claim 2.
5. determining whether a voice recognition process for acquiring the speech of the occupant has been started; When the input operation signal is acquired before the voice recognition process is started, the control of an in-vehicle device is executed in accordance with the input operation signal, and information about the target component is output.
3. The speech recognition method according to claim 2.
6. 2. The speech recognition method according to claim 1, wherein the input operation signal is acquired after the speech content is acquired.
7. 2. The speech recognition method according to claim 1, wherein the target component is the operation input device.
8. 2. The speech recognition method according to claim 1, wherein the target component is a switch, a lever, a dial, a knob, a slide bar, or a touch panel.
9. determining whether the utterance is a question about a name, a method of use, or an application; When it is determined that the content of the utterance is a question about a name, a method of use, or a purpose, the name, a method of use, or a purpose of the target component is output as information about the target component.
2. The speech recognition method according to claim 1.
10. 2. The speech recognition method according to claim 1, wherein a voice or an image representing information about the target component is output.
11. A process of acquiring the speech content of a vehicle occupant; a process of acquiring an input operation signal generated by the occupant operating an operation input device of the vehicle; a process of estimating a target component, which is a component mentioned in the utterance content, among a plurality of components constituting the vehicle, based on the utterance content and the input operation signal; outputting information about the target component; a process of interrupting the output of the information about the target component and controlling an in-vehicle device in accordance with an input operation signal acquired from the operation input device when an utterance by the occupant or an operation of the operation input device is detected during the output of the information about the target component; A speech recognition device comprising a controller that executes the above.
Citation Information
Patent Citations
Speech recognition device
JP2007017839A
Control system for in-vehicle apparatus
JP2010247799A
Voice recognition control system
JP2017090613A
Voice recognition control system
JP2017090614A
On-vehicle device operation system
JP2020097378A