Speech recognition method and speech recognition device
The speech recognition method enhances object identification in user speech by integrating vehicle device and sensor inputs, improving accuracy and responsiveness in dynamic environments.
Patent Information
- Application Number
- JP2023576247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing speech recognition systems struggle to accurately identify objects mentioned in user speech, especially in situations like driving a vehicle, where it's difficult for users to articulate specific keywords.
A speech recognition method that utilizes input signals from vehicle devices and sensors to detect and estimate objects mentioned in user speech by recognizing states or positions, enhancing accuracy through natural language understanding and analysis of control signals and sensor outputs.
Improves the accuracy of estimating objects in user speech by correlating keyword information with detected states or positions, enabling precise object identification and responsive actions.
Smart Images

Figure 0007722475000001 
Figure 0007722475000002 
Figure 0007722475000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech recognition method and a speech recognition device. [Background technology]
[0002] The following Patent Document 1 describes an in-vehicle system that, when a warning light on a meter panel comes on, displays on a display device an explanation of the warning related to the lighted warning light and how to deal with the warning. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-193138 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, input systems that use speech recognition to respond to user questions and operate devices have been proposed. In such systems, the instructions that the user intends to input to the system are inferred from the content of the user's speech. In this case, in order for the input system to accurately identify the instruction, it is necessary for the user to accurately speak several keywords. However, it is difficult for the user to accurately speak the instruction in all situations. For example, when a user uses a voice input system while performing other tasks, such as driving a vehicle, it is difficult to imagine the keywords required to execute the instruction. The present invention aims to improve the accuracy of estimating objects mentioned in a speech in a speech recognition system that acquires the speech of a vehicle user and estimates the objects mentioned in the speech. [Means for solving the problem]
[0005] According to one aspect of the present invention, there is provided a speech recognition method for acquiring an utterance from a vehicle user and estimating an object mentioned in the utterance, which acquires at least one of a control signal of a device mounted on the vehicle and an output signal of a sensor mounted on the vehicle as an input signal, recognizes an expression representing a state or position from the utterance, detects a state or position of the object candidate based on the input signal, and estimates the object candidate that matches the state or position recognized from the utterance as the object mentioned in the utterance. [Effects of the Invention]
[0006] According to the present invention, in a speech recognition system that acquires the speech content of a vehicle user and estimates an object mentioned in the speech content, it is possible to improve the accuracy of estimating an object mentioned in the speech content. The objects and advantages of the invention will be realized and attained by means of the elements and combinations set forth in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a schematic diagram illustrating an example of a vehicle equipped with a voice recognition device according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of a voice recognition device. [Figure 3] FIG. 10 is a schematic diagram of an example of a command list. [Figure 4] FIG. 10 is a schematic diagram of an example of a response list. [Figure 5] 1 is a flowchart illustrating an example of a speech recognition method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] (composition) 1 is a schematic diagram of an example of a vehicle equipped with a voice recognition device according to an embodiment. The vehicle 1 includes an in-vehicle device 2, an in-vehicle device controller 3, an in-vehicle sensor 4, an external sensor 5, a human-machine interface (hereinafter referred to as "HMI") 6, and a voice recognition device 7. The in-vehicle device 2 is various devices mounted on the vehicle 1. The in-vehicle device 2 may be, for example, a warning light disposed on an instrument panel in front of the driver's seat or near an A-pillar of the vehicle 1. The warning light is an example of a visual information presentation device that is provided inside the vehicle 1 and presents visual information to a user.
[0009] Furthermore, for example, the in-vehicle device 2 may be an alarm device that outputs an alarm sound to the user of the vehicle 1. The alarm device is an example of an auditory information presentation device that is provided inside the vehicle and presents auditory information to the user. Furthermore, for example, the in-vehicle device 2 may be a window provided in a door of the vehicle 1, an engine of the vehicle 1, or a braking device.
[0010] The in-vehicle device controller 3 is an electronic control unit (ECU) that controls the operation of the in-vehicle device 2 and generates control signals for controlling the in-vehicle device 2. The in-vehicle device controller 3 includes, for example, a processor and peripheral components such as a storage device. The processor may be, for example, a CPU (Central Processing Unit) or an MPU (Micro-Processing Unit). The storage device may include a semiconductor storage device, a magnetic storage device, an optical storage device, etc. The storage device may include a register, a cache memory, a memory such as a ROM (Read Only Memory) and a RAM (Random Access Memory) used as a main memory device.
[0011] The in-vehicle device controller 3 may be formed of dedicated hardware for executing each of the information processes described below. For example, the in-vehicle device controller 3 may include a functional logic circuit configured in a general-purpose semiconductor integrated circuit. For example, the in-vehicle device controller 3 may include a programmable logic device (PLD) such as a field-programmable gate array (FPGA).
[0012] The interior sensor 4 is a sensor that detects the state inside the vehicle 1. For example, the interior sensor 4 may be an interior camera that takes pictures of the interior of the vehicle, a pressure sensor or a seat belt sensor that is provided in a seat to determine whether an occupant is seated, a biosensor that detects biometric information of an occupant, or a microphone that detects sounds generated from the vehicle 1. The external sensor 5 is a sensor that detects objects present around the vehicle 1. For example, External Sensor 5 may be, for example, an external camera that captures the surrounding environment of the vehicle 1, or may be a distance measurement sensor such as a laser range finder (LRF), radar, or a laser radar of a LiDAR (Light Detection and Ranging).
[0013] The HMI 6 is an interface device that exchanges information between the voice recognition device 7 and the user. The HMI 6 includes a display device that can be seen by the user of the vehicle 1 (for example, a display screen of a navigation system), as well as a speaker and a buzzer for outputting warning sounds, notification sounds, and voice information. The HMI 6 also includes a voice input device (for example, a microphone) that acquires voice input from the user.
[0014] The voice recognition device 7 is an electronic control unit (ECU) that operates as a controller that executes voice recognition to recognize the contents of speech uttered by the user of the vehicle 1. The voice recognition device 7 estimates an object mentioned in the contents of speech uttered by the user and outputs information related to the object from the HMI 6 to provide it to the user. Alternatively, the voice recognition device 7 operates the object mentioned in the contents of speech uttered by the user.
[0015] The speech recognition device 7 includes a processor 8 and peripheral components such as a storage device 9. The processor 8 may be, for example, a CPU or an MPU. The storage device 9 may include a semiconductor storage device, a magnetic storage device, an optical storage device, etc. The storage device 9 may include memories such as a register, a cache memory, and a ROM and a RAM used as a main storage device. The functions of the speech recognition device 7 described below are realized by, for example, the processor 8 executing a computer program stored in the storage device 9. The speech recognition device 7 may be formed of dedicated hardware for executing each of the information processes described below. For example, the speech recognition device 7 may include a functional logic circuit configured in a general-purpose semiconductor integrated circuit. For example, the speech recognition device 7 may include a programmable logic device such as a field programmable gate array.
[0016] 2 is a block diagram showing an example of the functional configuration of the speech recognition device 7. The speech recognition device 7 operates as a speech recognition unit 10, a natural language understanding unit 11, an input signal acquisition unit 12, an analysis unit 13, and a response generation unit 14. The speech recognition unit 10 recognizes a speech input from a user acquired by the HMI 6 and converts it into linguistic information such as text. The speech recognition unit 10 converts the linguistic information generated by converting the speech input into Natural Language Understanding Department 11 Output to.
[0017] The natural language understanding unit 11 analyzes the linguistic information output from the speech recognition unit 10 by natural language processing, and extracts the user's speech intention and keywords related to the speech intention. For example, the natural language understanding unit 11 extracts keywords indicating the state or position of an object mentioned in the utterance. The natural language understanding unit 11 may also supplementarily extract keywords indicating the aspect (shape, color, position) of the object.
[0018] For example, keywords and their synonyms may be defined in advance, and synonyms contained in the user's utterance may be converted into keywords. For example, if a user utters "What is that red car light that just came on?" to ask about the meaning of a warning light, the natural language understanding unit 11 extracts "inquiry of meaning" as the intention of the utterance, and extracts "red," "lit," and "car" as keywords.
[0019] In this case, for example, synonyms for the keyword "red" may be predefined as "red," "red," "red," "vermilion," etc., synonyms for the keyword "car" may be predefined as "car," "Car," "automobile," "passenger car," etc., and synonyms for the keyword "lights" may be predefined as "turned on," "turned on now," "lit," "on," etc.
[0020] In addition, the user's speech intentions extracted by the natural language understanding unit 11 include various speech intentions such as "inquiry about status" to inquire about the status of the in-vehicle device 2, operation instructions to operate the in-vehicle device 2 (for example, "open the window"), "inquiry about the cause of abnormal sound" to inquire about the cause of abnormal sound generated from the vehicle 1, and "inquiry about surrounding conditions" to inquire about the conditions around the vehicle 1. The natural language understanding unit 11 outputs the extracted information on the intention of the utterance and the extracted information on the keywords to the analysis unit 13 .
[0021] The input signal acquisition unit 12 acquires, as an input signal, a control signal for the in-vehicle device 2 generated by the in-vehicle device controller 3. For example, the control signal may be an on / off signal for a warning light. Also, for example, the control signal may be a signal instructing an alarm device to output or stop an alarm sound. Furthermore, the input signal acquisition unit 12 acquires output signals from the in-vehicle sensor 4 and the external sensor 5 as input signals. The input signal acquisition unit 12 converts the acquired control signals of the in-vehicle devices 2 and the output signals of the in-vehicle sensors 4 and the external sensors 5 into a specific data format determined in advance to represent the detected situation.
[0022] For example, the input signal acquiring unit 12 may convert the control signal into flag information and set the value of the flag according to the control state of the in-vehicle device 2. For example, if an EV (Electric Vehicle) system warning light is on, the value of the flag F1 may be set to "True," and if not, the value of the flag F1 may be set to "False." Also, for example, when the alarm device operates and outputs an alarm sound, the value of flag F3 may be set to "True," and when the alarm sound is not output, the value of flag F3 may be set to "False."
[0023] For example, the input signal acquisition unit 12 may convert the output signals of the in-vehicle sensor 4 and the external sensor 5 into flag information and set the value of the flag according to the state and position of the object detected by the in-vehicle sensor 4 and the external sensor 5. For example, a flag may be set according to the position of the user inside the vehicle detected based on the output signal of the in-vehicle sensor 4, such as an in-vehicle camera, a pressure sensor, a seat belt sensor, a biological sensor, etc. For example, the value of flag F4 may be set to "True" when the user is sitting in the driver's seat, and "False" when the user is sitting in the passenger seat.
[0024] Furthermore, for example, the input signal acquisition unit 12 may set a flag according to the position of an object around the vehicle 1 detected based on an output signal from an external sensor 5 such as an external camera or a distance measurement sensor. For example, the value of flag F6 may be set to "True" when another vehicle is approaching to the right rear of the vehicle 1, and the value of flag F6 may be set to "False" when no other vehicle is approaching. Furthermore, for example, the value of flag F6 may be set to "True" when the speed of the other vehicle traveling to the right rear of the vehicle 1 exceeds a threshold Vth, and may be set to "False" when the speed does not exceed the threshold Vth. For example, the input signal acquisition unit 12 may analyze sound information output by the microphone of the interior sensor 4 and estimate the in-vehicle device 2 that is the source of an abnormal sound generated by the vehicle 1 and the cause of the abnormal sound based on the characteristics of the sound information. The input signal acquisition unit 12 may set a flag based on the in-vehicle device 2 that is the sound source and the cause of the abnormal sound. For example, if it is estimated that the source of the abnormal sound is the engine of the vehicle 1 and that the cause of the abnormal sound is a lack of engine oil, the input signal acquisition unit 12 may set the value of flag F5 to "True," and if no abnormal sound is detected, the input signal acquisition unit 12 may set the value of flag F5 to "False." A similar flag may be set for abnormal sounds generated by the braking system. Furthermore, separate flags may be set for multiple abnormal sounds generated by the same in-vehicle device 2 and caused by different reasons. Here, to estimate the cause of the abnormal sound, the input signal acquisition unit 12 may perform frequency analysis on the sound information acquired from the microphone of the interior sensor 4 and pre-stored sound information of the in-vehicle device in a normal state, and determine that an abnormality has occurred if a predetermined frequency pattern or a parameter pattern including the frequency pattern is detected. Furthermore, if the source of the abnormal sound is the engine, sound information for a state in which engine oil is low can be stored in advance, and frequency analysis can be performed between this and sound information acquired from a microphone.If different frequency characteristics beyond a certain range are obtained compared to the frequency pattern of a normal engine sound source, it can be determined that the cause is a lack of engine oil.
[0025] The input signal acquisition unit 12 may also convert the control signals of the in-vehicle devices 2 and the output signals of the in-vehicle sensors 4 and the external sensors 5 into numerical data, identification information, text data, etc., indicating the extracted information. For example, the input signal acquisition unit 12 may convert the control signals of the in-vehicle devices 2 and the output signals of the in-vehicle sensors 4 and the external sensors 5 into numerical data, such as distance information (e.g., "10 m") or speed information (e.g., "60 km / h") to other vehicles detected based on the output signals of the external sensors 5, or into identification information or text data indicating the vehicle type. The input signal acquisition unit 12 outputs the converted input signal (hereinafter simply referred to as the “input signal”) to the analysis unit 13.
[0026] The analysis unit 13 receives the input signal output from the input signal acquisition unit 12 and the information on the intention of the utterance and the information on keywords output from the natural language understanding unit 11 . The analysis unit 13 detects the state or position of a candidate object mentioned in the user's speech based on the input signal output from the input signal acquisition unit 12. For example, the analysis unit 13 detects the control state of the in-vehicle device 2 by the control signal as the state of the candidate object. For example, the analysis unit 13 may detect whether a warning light is on or off (i.e., the display state of visual information by the visual information presentation device).
[0027] When detecting the state or position of a candidate object, the analysis unit 13 refers to the command list 15 stored in the storage device 9. FIG. The command list 15 stores multiple rows of records. Each record records a command ID, information about a candidate object, a keyword related to the candidate object, and information specifying an input signal used to detect the state or position of the candidate object. That is, the command list 15 records a command ID, information about a candidate object, a keyword, and information specifying an input signal in association with each other. Note that keywords related to the candidate object are keywords indicating the state or position of the candidate object. Keywords indicating the state of the candidate object may also be recorded for each candidate object.
[0028] For example, the record on the first line specifies input signal flag F1 as input information indicating the state of an EV system warning light, which is an example of a warning light. Based on flag F1, analysis unit 13 detects whether the EV system warning light is on or off. Furthermore, for example, the analysis unit 13 detects whether the alarm device is in an alarm output state or in a stopped state (that is, whether the auditory information presentation device is in an audio information presentation state). For example, the record on the third line of the command list 15 specifies the input signal flag F3 as input information indicating the state of the alarm device. The analysis unit 13 detects whether the alarm device is in an output state or a stopped state based on the flag F3.
[0029] Furthermore, for example, the analysis unit 13 detects an in-vehicle device 2 placed at a specific position as a candidate for the object mentioned in the user's utterance. That is, the analysis unit 13 detects the position of the in-vehicle device 2 that is a candidate for the object. For example, the record on the fourth line of the command list 15 specifies the flag F4 of the input signal as information indicating whether the window that is a candidate for the object is the driver's window. The flag F4 is set to "True" when the user is sitting in the driver's seat, and is set to "False" when the user is not sitting in the driver's seat. The analysis unit 13 detects that the driver's window is a candidate for the object when the flag F4 is "True," and detects that the driver's window is not a candidate for the object when the flag F4 is "False."
[0030] Furthermore, for example, the analysis unit 13 may detect whether the source of an abnormal sound generated from the vehicle 1 is a specific in-vehicle device 2. In other words, whether an in-vehicle device 2 that is a candidate for the object is the source of the abnormal sound may be detected as the state of the candidate object. The analysis unit 13 may also estimate the cause of the abnormal sound. For example, the record on the fifth line of the command list 15 specifies the flag F5 of the input signal as information indicating whether the engine, which is a candidate object, is the source of the abnormal sound. If the flag F5 is "True," the analysis unit 13 estimates that the engine is the source of the abnormal sound and that the cause of the abnormal sound is a lack of engine oil. If the flag F5 is "False," the analysis unit 13 detects that the engine is not the source of the abnormal sound.
[0031] Furthermore, for example, the analysis unit 13 may detect the state or position of an object around the vehicle 1 as the state or position of the target object candidate. For example, the record on the sixth line of the command list 15 specifies the flag F6 of the input signal as information indicating whether or not another vehicle is approaching to the right rear of the vehicle 1. When the flag F6 is "True," the analysis unit 13 detects that another vehicle is approaching to the right rear, and when the flag F6 is "False," it determines that another vehicle is not approaching to the right rear. Furthermore, the distance to another vehicle traveling to the right rear of the vehicle 1 (i.e., the position of the other vehicle) may be detected based on distance information (e.g., "10 m") included in the input signal. Furthermore, the speed of another vehicle traveling to the right rear of the vehicle 1 (i.e., the speed of the other vehicle) may be detected based on speed information (e.g., "60 km / h") included in the input signal. The analysis unit 13 may store the received input signal in the storage device 9. The analysis unit 13 may detect the state or position of the candidate object based on the input signal stored in the storage device 9 in addition to or instead of the currently input input signal. Alternatively, for example, the analysis unit 13 may detect the state or position of the candidate object based on a time series of previously input input signals and currently input input signals. The state of the candidate object may be estimated by storing previously input signals and detecting the difference (the difference between True and False) between the previously input signals and the current input signal. Alternatively, for example, distance information to another vehicle on the right rear, which is included in a past input signal, may be stored, and if the current distance information is smaller than the past distance information, it may be estimated that another vehicle is approaching on the right rear.
[0032] The analysis unit 13 estimates that the candidate object that matches the state or position indicated by the keyword information output from the natural language understanding unit 11 (i.e., the state or position of the object mentioned in the user's utterance) is the object mentioned in the utterance. Specifically, if the state or position indicated by the keyword information output from the natural language understanding unit 11 matches the state or position of the candidate object detected from the input signal, the candidate object is presumed to be the object mentioned in the utterance content.
[0033] For example, suppose a user utters, "What is the red car light that just came on?" and the natural language understanding unit 11 extracts the keyword "lit" which indicates the state of the candidate object, and the keywords "red" and "car" which indicate the aspect of the object (shape, color, position). The analysis unit 13 refers to the command list 15 and selects the record in the first line (EV system warning light) and the record in the second line (water temperature warning light) that contain the same keyword as the keyword "on" extracted by the natural language understanding unit 11.
[0034] The analysis unit 13 determines whether the EV system warning light is on or not based on the flag F1 specified in the record on the first line. That is, the analysis unit 13 determines whether the state of the object candidate is the same as the keyword "on" that indicates the state of the object candidate included in the command list 15. If the state of the candidate object is the same as the keyword "on" included in command list 15, analysis unit 13 determines that the state of the object mentioned in the user's utterance matches the state of the EV system warning light, and estimates that the object mentioned in the utterance is the EV system warning light.
[0035] The analysis unit 13 outputs the command ID "id0001" of the record in the first row to the response generation unit 14. The command ID is associated with information about the candidate object, keywords related to the candidate object, and the input signal, so that the object mentioned in the user's utterance and the state and position of the object can be identified based on the command ID. The analysis unit 13 also outputs the information on the intention of the utterance output from the natural language understanding unit 11 to the response generation unit 14.
[0036] Assume that the water temperature warning light is also on in addition to the EV system warning light. In this case, the state of the water temperature warning light will be the same as the keyword "on" included in command list 15. Therefore, it is not possible to distinguish whether the object mentioned in the utterance is the EV system warning light or the water temperature warning light based solely on the keyword "on" that indicates the state of the candidate object. In this case, the analysis unit 13 may determine the object mentioned in the utterance content by supplementarily using keywords "red" and "car" that indicate the aspect of the object.
[0037] Next, assume that the user utters "What just beeped?" and the natural language understanding unit 11 extracts the keyword "beep" indicating the state of the candidate object. The analysis unit 13 refers to the command list 15 and selects the record (alarm device) in the third row that contains the same keyword as the keyword "beep" extracted by the natural language understanding unit 11.
[0038] The analysis unit 13 3 Based on the flag F3 specified in the record of the line 13, the analysis unit 13 determines whether the alarm device is in an output state. That is, the analysis unit 13 determines whether the state of the candidate object is the same state (operating state) as the keyword "beep" that indicates the state of the candidate object included in the command list 15. If the state of the candidate object is the same as the state of the keyword included in the command list 15, the analysis unit 13 determines that the state of the object mentioned in the user's speech matches the state of the alarm device, and estimates that the object mentioned in the speech is an alarm device. The analysis unit 13 outputs the command ID “id0003” of the record on the third line and the information on the utterance intention output from the natural language understanding unit 11 to the response generation unit 14.
[0039] Also, for example, it is assumed that the user utters "Open this window" and the natural language understanding unit 11 extracts the keyword "here" indicating the location of the candidate object. The analysis unit 13 refers to the command list 15 and selects the record (driver's window) in the fourth line that includes the same keyword as the keyword "here" extracted by the natural language understanding unit 11.
[0040] Based on the flag F4 specified in the record on the fourth line, the analysis unit 13 determines whether the position of the candidate object (driver's seat window) (i.e., near the driver's seat) is the same as the candidate object included in the command list 15. position If flag F4 is "True," the user is seated in the driver's seat, so it is determined that the position of the candidate object is the same as the keyword included in command list 15.
[0041] If the position of the candidate object is the same as the position of a keyword included in the command list 15, the analysis unit 13 determines that the position of the object mentioned in the user's utterance matches the position of the driver's window, and estimates that the object mentioned in the utterance is the driver's window. The analysis unit 13 4 The command ID “id0004” of the record in the th line and the information on the utterance intention output from the natural language understanding unit 11 are output to the response generation unit 14.
[0042] Also, for example, assume that the user utters "It's making a strange noise, but it's okay," and the natural language understanding unit 11 extracts the keyword "strange noise" indicating the state of the candidate object. The analysis unit 13 refers to the command list 15 and selects the record (engine) in the fifth line that includes the same keyword as the keyword "strange sound" extracted by the natural language understanding unit 11.
[0043] The analysis unit 13 determines whether the engine is the source of the abnormal sound based on the flag F5 specified in the record on the fifth line. That is, the analysis unit 13 determines whether the state of the candidate object (engine) is the same as the keyword "(making) a strange sound" that indicates the state of the candidate object included in the command list 15. If the state of the candidate object is the same as the state of the keyword included in the command list 15, the analysis unit 13 determines that the state of the object mentioned in the user's utterance matches the state of the engine, and presumes that the object mentioned in the utterance is the engine. It also presumes that the cause of the abnormal sound is a lack of engine oil. The analysis unit 13 5 The command ID “id0005” of the record in the th line and the information on the utterance intention output from the natural language understanding unit 11 are output to the response generation unit 14.
[0044] Also, for example, assume that the user utters "What is that approaching at a great speed?" and the natural language understanding unit 11 extracts the keyword "approaching" which indicates the state of the candidate object. The analysis unit 13 refers to the command list 15 and selects the record in the sixth line (right rear vehicle) that includes the same keyword as the keyword "approaching" extracted by the natural language understanding unit 11.
[0045] The analysis unit 13 determines whether the vehicle behind on the right is approaching the vehicle 1 based on the flag F6 specified in the record on the sixth line. That is, the analysis unit 13 determines whether the state of the object candidate (the vehicle behind on the right) is the same as the keyword "approaching" that indicates the state of the object candidate included in the command list 15. The analysis unit 13 may determine whether the vehicle behind on the right is approaching the vehicle 1 based on the position information and speed information specified in the record on the sixth line. If the state of the candidate object is the same as the state of the keyword included in the command list 15, the analysis unit 13 determines that the state of the object mentioned in the user's utterance matches the state of the vehicle behind on the right, and estimates that the object mentioned in the utterance is the vehicle behind on the right. The analysis unit 13 6 The command ID “id0006” of the record in the th line and the information on the utterance intention output from the natural language understanding unit 11 are output to the response generation unit 14.
[0046] See Fig. 2. The response generation unit 14 outputs a response message and a response command based on the information on the utterance intention extracted by the natural language understanding unit 11 and input via the analysis unit 13, and the command ID output from the analysis unit 13. The response message is a voice signal or text information of a message presented to the user in response to the user's utterance. The response command is a command signal that causes the HMI 6 to output a response message in response to the user's utterance or causes the in-vehicle device 2 to perform a predetermined operation.
[0047] When generating a response message and a response command, the response generation unit 14 refers to a response list 16 stored in the storage device 9. FIG. The response list 16 stores multiple rows of records. Each record stores information about an utterance intention, a command ID, a response message, and a response command. That is, the response list 16 stores information about an utterance intention, a command ID, a response message, and a response command in association with each other.
[0048] For example, if a user utters "What is the red car light that just came on?", the natural language understanding unit 11 extracts "inquiry about meaning" as the utterance intention as described above. The analysis unit 13 outputs the command ID "id0001". The response generation unit 14 extracts the record in the first line that matches the utterance intention "inquiry about meaning" and the command ID "id0001". The response generator 14 outputs the response command "command C001" to the HMI 6 to notify the user of the meaning of the warning light stored in the record on the first line, and causes the response message "This means that an abnormality has occurred in the EV system" to be output as audio or text information from the speaker of the HMI 6 or displayed on the display device. In this way, command C001 is a command signal that causes the HMI 6 to output a response message, and the same is true for commands C0002, C003, C005, and C006 shown in Figure 4.
[0049] Also, assume that when a user utters, "The red thermometer is on, what's going on?", the natural language understanding unit 11 extracts the utterance intention "inquire about status" and the analysis unit 13 outputs the command ID "id0002." The response generation unit 14 extracts the record in the second line that matches the utterance intention "inquire about status" and the command ID "id0002."
[0050] The response generating unit 14 outputs to the HMI 6 a response message "The temperature of the engine coolant is high" notifying the state of the radiator stored in the record on the second line, and a response command "Command C002". Note that a response message regarding the state of an object may be associated with the utterance intention "inquiry about meaning" and stored in the response list 16. In this case, the response generator 14 can output a response message regarding the state of an object in response to an utterance with the utterance intention "inquiry about meaning".
[0051] Similarly, a response message regarding a method of dealing with a target object's state may be stored in association with the utterance intention of "inquiring about meaning." For example, the record on the third line stores the utterance intention "query of meaning" and the command ID "id0003." For example, when a user utters "What is the red light on the thermometer that just came on?", the natural language understanding unit 11 extracts the utterance intention "query of meaning," and the analysis unit 13 extracts the command ID "id0003." 3 " is output. In this case, the response generation unit 14 selects the record on the third line and outputs the response message "Please park the car in a safe place" and the response command "Command C003" to the HMI 6, thereby notifying the user of the appropriate action to take when the engine coolant temperature is high.
[0052] Also, assume that when a user utters "Open this window," the natural language understanding unit 11 outputs the utterance intention "Open the window," and the analysis unit 13 outputs the command ID "id0004." The response generation unit 14 extracts the record in the fourth line that matches the utterance intention "Open the window" and the command ID "id0004." The response generation unit 14 outputs a response command "command C004" which is a command signal to open the driver's seat window to the in-vehicle device controller 3. The in-vehicle device controller 3 opens the driver's seat window, which is an example of the in-vehicle device 2, in accordance with the response command "command C004." Note that the response generation unit 14 may output a response command to close the driver's seat window to the in-vehicle device controller 3 when the user utters "close this window."
[0053] Also, assume that the user utters "There's a strange noise, but everything's OK," the natural language understanding unit 11 extracts the utterance intention "Inquire about the cause of the abnormal noise," and the analysis unit 13 outputs the command ID "id0005." The response generation unit 14 extracts the record in the fifth line that matches the utterance intention "Inquire about the cause of the abnormal noise" and the command ID "id0005." The response generator 14 outputs to the HMI 6 a response message "It appears that the engine oil is low" notifying the cause of the abnormal sound stored in the record on the fifth line, and a response command "Command C005".
[0054] Also, assume that the user utters "What is that approaching at a great speed?", the natural language understanding unit 11 extracts the utterance intention "inquire about surrounding circumstances", and the analysis unit 13 outputs the command ID "id0006". The response generation unit 14 extracts the record in the sixth line that matches the utterance intention "inquire about surrounding circumstances" and the command ID "id0006". The response generation unit 14 outputs to the HMI 6 the response message "A vehicle is approaching from the rear right" that notifies the surrounding situation stored in the record on the sixth line, and the response command "Command C006".
[0055] (operation) FIG. 5 is a flowchart of an example of a speech recognition method according to an embodiment. In step S1, the input signal acquisition unit 12 acquires the control signal for the in-vehicle device 2 generated by the in-vehicle device controller 3 and the output signals of the in-vehicle sensor 4 and the external sensor 5 as input signals. In step S2, the speech recognition unit 10 recognizes speech input from the user acquired by the HMI 6 and converts it into linguistic information such as text. The natural language understanding unit 11 analyzes the linguistic information output from the speech recognition unit 10 by natural language processing and extracts the user's speech intention. In step S3, the natural language understanding unit 11 extracts keywords related to the linguistic intention from the linguistic information output from the speech recognition unit 10.
[0056] In step S4, the analysis unit 13 detects the state or position of a candidate object mentioned in the user's speech based on the input signal acquired by the input signal acquisition unit 12. In step S5, the analysis unit 13 estimates, based on the information of the keywords extracted by the natural language understanding unit 11, that a candidate object that matches the state or position recognized from the utterance content is the object mentioned in the utterance content. In step S6, the response generation unit 14 outputs a response message or operates the in-vehicle device 2 in accordance with the utterance intention extracted by the natural language understanding unit 11 and the target object estimated by the analysis unit 13.
[0057] (Effects of the embodiment) (1) The voice recognition device 7 acquires the content of an utterance made by a vehicle user and estimates the object mentioned in the utterance. The voice recognition device 7 acquires at least one of a control signal of a device mounted on the vehicle 1 or an output signal of a sensor mounted on the vehicle 1 as an input signal, recognizes an expression representing a state or position from the utterance, detects the state or position of a candidate object based on the input signal, and estimates that the candidate object that matches the state or position recognized from the utterance is the object mentioned in the utterance. This makes it possible to improve the accuracy of estimating the object mentioned in the speech content in speech recognition that acquires the speech content of the vehicle user and estimates the object mentioned in the speech content.
[0058] (2) For example, the candidate object may be a device controlled by a control signal acquired as an input signal. The speech recognition device 7 may detect the control state of the control signal as the state of the candidate object. This allows the state of the candidate object to be determined based on the control signal that controls the device. (3) For example, the input signal may be a control signal for a visual information presentation device that is installed inside the vehicle 1 and presents visual information to the user, and the control state may be a display state of the visual information. This allows the state of the visual information presentation device to be determined as a candidate for the object. (4) For example, the visual information presentation device may be a warning light, and the control state may be the on or off state of the warning light. This allows the state of the warning light to be determined as a candidate object.
[0059] (5) For example, the input signal may be a control signal for an auditory information presentation device installed inside the vehicle 1 to present auditory information to the user, and the control state may be a notification state of the auditory information. This allows the state of the auditory information presentation device to be determined as a candidate for the object. (6) For example, the audio information presentation device may be an alarm device, and the control state may be an alarm output state or a stop state. This allows the state of the alarm device to be determined as a candidate for the object.
[0060] (7) The speech recognition device 7 may store the acquired input signal and detect the state or position of the candidate object based on the stored past input signal and the currently acquired input signal. This allows the object to be estimated based on the past state or position before the user speaks, even if the state or position of the object changes before the user speaks. (8) The speech recognition device 7 may output information about an object mentioned in the speech content, or may output information about the state of the object mentioned in the speech content. The speech recognition device 7 may store a countermeasure according to the state of a candidate object in a predetermined storage device, and output information about the countermeasure according to the state of the object mentioned in the speech content. This allows information about the object mentioned in the user's speech to be provided.
[0061] (9) The candidate object may be a device mounted on the vehicle 1. The voice recognition device 7 may acquire, as an input signal, an output signal from a sensor that detects the state inside the vehicle 1, and detect the state or position of the device based on the acquired output signal. This makes it possible to determine the state or position of the equipment mounted on the vehicle 1 based on the output signals of sensors that detect the state inside the vehicle 1. (10) The voice recognition device 7 receives as an input signal the output signal of a sensor that detects the seating position of an occupant of the vehicle 1, detects that the window that is a candidate for the object is a window near the seating position, recognizes an expression indicating the position of the window to be opened from the speech content that includes an opening instruction to open the window of the vehicle 1, and if the window position recognized from the speech content indicates the vicinity of the seating position, may estimate that the window near the seating position is the object. From the output signal of the sensor that detects the seating position of the occupant and the speech content that includes an instruction to open the window of vehicle 1, it can be estimated that the window to be opened is the window near the seating position of the user. (11) The voice recognition device 7 may receive as an input signal an output signal of a sensor that detects sound information of an abnormal sound from the vehicle 1, and may detect a state in which a device that is a candidate for the target is emitting an abnormal sound by estimating the device that is the source of the abnormal sound based on the sound information. This allows the state of the device installed in the vehicle 1 to be estimated based on the output signal of the sensor that detects sound information. (12) For example, the candidate object may be an object around the vehicle 1. The voice recognition device 7 may acquire, as an input signal, an output signal from a sensor that detects the surrounding object, and detect the state or position of the surrounding object based on the acquired output signal. For example, the voice recognition device 7 may acquire, as an input signal, a captured image generated by a camera capturing an image of the surroundings of the vehicle 1, and recognize an object approaching the vehicle 1 as a candidate object based on the captured image. This makes it possible to determine the state or position of objects around the vehicle 1 based on output signals from sensors that detect the objects around the vehicle 1. (13) For example, the sensors may include a pressure sensor, a seat belt sensor, a camera, a distance sensor, a microphone, or a biometric sensor, which can detect the state and position of various potential objects inside or outside the vehicle.
[0062] All examples and conditional terms described herein are intended for educational purposes to aid the reader in understanding the present invention and the concepts provided by the inventor for the advancement of technology, and should be construed without limitation to the specifically described examples and conditions above, and the configuration of examples herein for illustrating the advantages and disadvantages of the present invention. Although the embodiments of the present invention have been described in detail, it should be understood that various changes, substitutions, and alterations can be made thereto without departing from the spirit and scope of the present invention. [Explanation of symbols]
[0063] 1...vehicle, 2...in-vehicle equipment, 3...in-vehicle equipment controller, 4...in-vehicle sensor, 5...external sensor, 6...human-machine interface, 7...speech recognition device, 8...processor, 9...storage device, 10...speech recognition unit, 11...natural language understanding unit, 12...input signal acquisition unit, 13...analysis unit, 14...response generation unit, 15...command list, 16...response list
Claims
1. A speech recognition method for acquiring a speech content of a vehicle user and estimating an object mentioned in the speech content, comprising: acquiring, as an input signal, at least one of a control signal of a device mounted on the vehicle and an output signal of a sensor mounted on the vehicle; Recognizing an expression that indicates a state or a position from the utterance content; Detecting a candidate state or position of the object based on the input signal; determining whether or not the state or position recognized from the speech content matches the state or position of the candidate object detected based on the input signal; The candidate object that matches the state or position recognized from the speech content is estimated to be the object mentioned in the speech content. A speech recognition method comprising:
2. the candidate object is a device controlled by the control signal acquired as the input signal, 2. The speech recognition method according to claim 1, wherein a control state according to the control signal is detected as a candidate state of the object.
3. the input signal is a control signal for a visual information presentation device that is provided inside the vehicle and presents visual information to the user, the control state is a display state of the visual information; 3. The speech recognition method according to claim 2.
4. 4. The speech recognition method according to claim 3, wherein the visual information presentation device is a warning light, and the control state is a lighting state or an extinguishing state of the warning light.
5. the input signal is a control signal for an auditory information presentation device that is provided in the vehicle and presents auditory information to the user, The control state is a state in which the auditory information is notified.
3. The speech recognition method according to claim 2.
6. 6. The speech recognition method according to claim 5, wherein the auditory information presentation device is an alarm device, and the control state is an alarm output state or a stop state.
7. storing the acquired input signal; detecting a state or position of the candidate object based on the stored past input signal and the currently acquired input signal; 7. The speech recognition method according to claim 1, wherein the speech recognition method is a speech recognition method for speech recognition.
8. 8. The speech recognition method according to claim 1, further comprising outputting information about an object mentioned in the speech content.
9. 8. The speech recognition method according to claim 1, further comprising outputting information about a state of an object mentioned in the speech content.
10. storing a countermeasure according to the state of the candidate object in a predetermined storage device; The speech recognition method according to any one of claims 1 to 7, further comprising outputting information relating to the method of dealing with the state of the object mentioned in the speech content.
11. the candidate object is a device mounted on the vehicle, 2. The speech recognition method according to claim 1, wherein an output signal of a sensor that detects a state inside the vehicle is acquired as the input signal, and the state or position of the device is detected based on the acquired output signal.
12. As the input signal, an output signal of a sensor that detects the seating position of an occupant of the vehicle is acquired, Detecting that the window that is the candidate for the object is a window near the seating position; Recognizing an expression representing a position of a window to be opened or closed from the speech content including an opening / closing instruction for opening or closing a window of the vehicle; When the position of a window recognized from the content of the utterance indicates the vicinity of the seating position, the window in the vicinity of the seating position is estimated as the object.
12. The speech recognition method according to claim 11.
13. As the input signal, an output signal of a sensor that detects sound information of an abnormal sound from the vehicle is acquired, and detecting a state in which the device that is the candidate for the target object is generating the abnormal sound by estimating the device that is the source of the abnormal sound based on the sound information.
12. The speech recognition method according to claim 11.
14. the candidate object is an object around the vehicle; 2. The speech recognition method according to claim 1, wherein an output signal of a sensor that detects the surrounding object is acquired as the input signal, and the state or position of the surrounding object is detected based on the acquired output signal.
15. Acquiring a captured image generated by a camera capturing an image of the surroundings of the vehicle as the input signal; Recognizing an object approaching the vehicle as a candidate object based on the captured image; 15. The speech recognition method according to claim 14.
16. 16. The speech recognition method according to claim 1, wherein the sensor includes any one of a pressure sensor, a seat belt sensor, a camera, a distance measurement sensor, a microphone, and a biosensor.
17. A speech recognition device that acquires a speech content of a vehicle user and estimates an object mentioned in the speech content, A process of acquiring at least one of a control signal of a device mounted on the vehicle and an output signal of a sensor mounted on the vehicle as an input signal; A process of recognizing an expression representing a state or position from the content of the utterance; detecting a candidate state or position of the object based on the input signal; a process of determining whether or not the state or position recognized from the speech content matches the state or position of the candidate object detected based on the input signal; a process of estimating that the candidate object that matches the state or position recognized from the speech content is the object mentioned in the speech content; A speech recognition device comprising a controller that executes the above.
Citation Information
Patent Citations
On-vehicle system, warning light detailed information notifying system, and server system
JP2006193138A
Device and method for photographing around vehicle
JP2006343829A
Voice interactive device and voice interactive method
JP2010281855A
Vehicular voice recognition apparatus
JP2015089697A
On-vehicle device
JP2019127192A