Dialog method and dialog device

The dialogue system in vehicles automatically detects the number of occupants to activate the dialogue processor without a wake word when alone, improving convenience and reducing noise-related activation issues.

WO2025224837A1PCT designated stage Publication Date: 2025-10-30NISSAN MOTOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/015894
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing vehicle dialogue systems struggle to distinguish between a driver's speech and external noise, particularly when multiple occupants are present, necessitating the driver to repeatedly utter a wake word for activation, which is cumbersome and inconvenient.

Method used

A dialogue method and device that employs an occupant number detection system to activate the dialogue processor without a wake word when there is one occupant, and requires a wake word when there are two or more occupants, using a combination of microphones, cameras, and sensors to differentiate between internal and external speech.

Benefits of technology

Enhances convenience by allowing interaction without the need for a wake word when alone, preventing unintended activation due to noise, and reducing unnecessary vehicle control or response processing when multiple occupants are present.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024015894_30102025_PF_FP_ABST
    Figure JP2024015894_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a dialog method by a controller (20) which is mounted on a vehicle (1) and which is provided with a dialog processing unit for performing predetermined processing in response to an utterance of an occupant of the vehicle (1). The controller (20) performs: a number-of-occupants detection step for detecting or estimating the number of occupants of the vehicle (1); and a dialog start step for starting the dialog processing unit (225). In the dialog start step, if the number of occupants is one, the dialog processing unit (225) is started at a timing when it has been detected or estimated that the number of occupants is one.
Need to check novelty before this filing date? Find Prior Art

Description

Dialogue method and dialogue device

[0001] The present invention relates to a dialogue method and a dialogue device using a voice recognition device mounted on a vehicle.

[0002] Conventionally, an on-vehicle dialogue device is known that allows a vehicle occupant to input various commands by interacting with the vehicle occupant using a voice recognition device mounted on the vehicle (see, for example, Patent Document 1). In Patent Document 1, the dialogue device is activated when a driver or other occupant utters a predetermined wake word (trigger phrase). Patent Document 1 also allows the wake word to be customized. By activating the dialogue device using such a wake word, the driver does not need to operate a mechanical operation unit, such as a button, to activate the dialogue device, allowing the driver to concentrate on driving.

[0003] Special Publication No. 2015-520409

[0004] Incidentally, in an interactive device in a vehicle, the occupant may be a single driver or may be a passenger. When multiple occupants are in the vehicle, it is difficult to distinguish between external noise, such as conversations between occupants, and the driver's speech, making it difficult to omit the wake word. However, for example, when the vehicle is only occupied by the driver, there is a low possibility of external noise due to conversations between occupants being introduced, and there is a problem in that it is cumbersome and inconvenient for the driver to utter the wake word every time the interactive device is started.

[0005] An object of the present invention is to provide a more convenient dialogue method and dialogue device.

[0006] In the dialogue method of the present disclosure, a computer executes an occupant number detection step of detecting or estimating the number of occupants in a vehicle and a dialogue activation step of activating a dialogue processor. In the dialogue activation step, if the number of occupants is one, the dialogue processor is activated at the timing of detecting the number of occupants, and if the number of occupants is two or more, the previous speech processor is activated when a predetermined wake word is included in an utterance by an occupant.

[0007] As a result, when there is only one occupant, the dialogue processor is activated without the occupant needing to utter a predetermined wake word or perform any mechanical operation to activate the dialogue processor, thereby enabling the occupant to interact with the dialogue device without performing any specific operation, thereby improving convenience.

[0008] 1 is a schematic configuration diagram showing a vehicle of a first embodiment; 2 is a flowchart showing processing related to the startup of a dialogue processing unit according to a dialogue method of the first embodiment; 3 is a flowchart showing dialogue processing in single mode according to the dialogue method of the first embodiment; 4 is a flowchart showing dialogue processing in multi mode according to the dialogue method of the first embodiment; 5 is a flowchart showing dialogue stop processing in single mode according to the dialogue method of the first embodiment; 6 is a flowchart showing dialogue processing in single mode according to a dialogue method of a second embodiment; 7 is a flowchart showing processing related to the startup of a dialogue processing unit according to a dialogue method of a third embodiment;

[0009] [First Embodiment] A first embodiment of the present invention will be described below. Fig. 1 is a schematic configuration diagram showing a vehicle of this embodiment. As shown in Fig. 1, the vehicle 1 of this embodiment includes an ECU (Electronic Control Unit) 11, a microphone 12 (voice recognition unit), a speaker 13, an IVI (In-Vehicle Infotainment) controller 14, a switch interface 15, a camera 16, a display 17 (display unit), a group of sensors 18, and a controller 20.

[0010] The ECU 11 controls each component of the vehicle 1 based on commands from the controller 20. For example, the ECU 11 controls driving control for driving the vehicle 1 (drive control of a drive source such as an engine or a motor, reduction ratio control of a drive transmission gear that transmits driving force from the drive source to the drive wheels, steering control of a steering mechanism, braking control of a braking mechanism, etc.), air conditioning control of an in-vehicle air conditioner, audio control of an audio device, etc.

[0011] The microphone 12 detects sound and outputs the detected sound data (sound input data) to the controller 20. Only one microphone 12 may be provided inside the vehicle 1, or directional microphones may be provided for each seat inside the vehicle. The speaker 13 outputs sound based on the sound data (sound output data) input from the controller 20.

[0012] The IVI controller 14 provides various data in response to occupant operations, such as a navigation system that displays a map and provides route guidance based on the map. The IVI controller 14 includes hardware components, such as a storage device (e.g., memory) and a central processing unit (CPU), and reads and executes software programs stored in the storage device. This functions as an information provider that processes various data in response to occupant operations. The IVI controller 14 also functions as a call detector 141 that communicates with an external communication device, such as a mobile phone owned by the occupant, and detects calls made using the external communication device. For example, when a call application that makes a call via a communication line is executed on the external communication device, the IVI controller 14, which communicates with the external communication device, is notified that a call is in progress. This allows the call detector 141 to detect the occupant's call activity. Alternatively, when a call application is executed on the external communication device, the IVI controller 14 may connect the microphone 12 and speaker 13 to the external communication device and process the call via the IVI controller 14.

[0013] The switch interface 15 is a switch provided at a predetermined position inside the vehicle 1, and detects operation of the switch by an occupant. Examples of the switch include a switch for locking the doors and a switch for controlling the power windows, and the switch interface 15 outputs switch operation data relating to the operated switch to the controller 20. Therefore, the controller 20 can determine the position of the switch that has been operated from the input switch operation data, and can estimate that an occupant is present in a seat near the position of the switch.

[0014] The camera 16 is an observation unit of the present disclosure and captures images of the interior of the vehicle 1. The camera 16 has an imaging element such as a charge coupled device (CCD) and an optical system such as a lens that guides light to the imaging element, and captures images of the interior of the vehicle. The range captured by the camera 16 may be a range within the vehicle where the occupants and their speech can be confirmed. For example, the camera 16 may be provided on the ceiling at the front of the vehicle interior, or a camera 16 may be provided for each seat. Note that, although the present embodiment exemplifies the camera 16 as the observation unit, other devices such as millimeter wave radar and laser radar may also be used to observe the occupants within the vehicle.

[0015] The display 17 is a display unit that displays various types of information. The display 17 displays various types of information under the control of the controller 20.

[0016] The sensor group 18 includes various sensors provided in the vehicle 1, and inputs measurement data obtained by these sensors to the controller 20. In this embodiment, a door detection sensor 181 provided in a door and detecting whether the door is open or closed is exemplified as one of the sensors in the sensor group 18. The controller 20 can estimate whether an occupant is inside the vehicle based on a detection signal related to the door opening or closing from the door detection sensor 181. Note that other sensors included in the sensor group 18 provided in the vehicle 1 may also include a speed sensor, an acceleration sensor, an accelerator opening sensor, a torque sensor for detecting torque in the engine or drive motor, a room temperature sensor for measuring room temperature, etc.

[0017] The controller 20 is a computer that functions as the interactive device of the present disclosure. The controller 20 includes, for example, a storage unit 21 including a memory or the like, a processor 22 including a CPU (Central Processing Unit) or the like, and an input / output interface (not shown). The processor 22 loads and executes various programs stored in the storage unit 21, thereby functioning as a voice processing control unit 221, an occupant number detection unit 222, a speech determination processing unit 223, a startup unit 224, a dialogue processing unit 225, a control command generation unit 226, and a display control unit 227, as shown in FIG. 1 . While an example is shown in which the processor 22 executes the programs to realize the respective functional components of the voice processing control unit 221, the occupant number detection unit 222, the speech determination processing unit 223, the startup unit 224, the dialogue processing unit 225, the control command generation unit 226, and the display control unit 227, some or all of these components may be realized by individual hardware configurations. 1 shows a single controller 20, the controller 20 may be configured to include multiple controllers with different functions. For example, a voice controller that processes voice input data from the microphone 12, a camera controller that controls the camera 16 and processes images, an interactive system controller that interactively processes commands from the occupant, and the like may be provided as separate units.

[0018] The voice processing control unit 221 analyzes the content of speech uttered by the occupant based on the voice input data input from the microphone 12. Well-known techniques can be used to analyze the content of speech. For example, an acoustic model, dictionary data, and language model are stored in advance in the storage unit 21. The acoustic model outputs phonemes in response to voice input. The dictionary data records words corresponding to combinations of phonemes. The language model constructs sentences from word connections. The acoustic model and language model can be trained in advance by machine learning. Then, the voice processing control unit 221 extracts phonemes from the voice input data using the acoustic model, infers words from combinations of consecutive phonemes, and identifies the sentence from the inferred words before and after.

[0019] The occupant number detection unit 222 detects the number of occupants in the vehicle 1. In this embodiment, the occupant number detection unit 222 detects the number of occupants based on captured images captured by the camera 16. A known technique can be used to detect occupants using captured images. For example, objects are detected from the captured images using edge detection processing or the like, and people (occupants) are identified and the number of occupants is detected based on the similarity between the shapes of people learned by machine learning and the objects detected in the captured images. Note that, while an example of detecting the number of occupants based on captured images is shown here, this is not limiting. The occupant number detection unit 222 may detect or estimate the number of occupants based on other data. For example, the occupant number detection unit 222 may estimate the number of occupants by detecting the positions of occupants from switch operation data of each switch obtained by the switch interface 15. Alternatively, the number of occupants may be estimated from the number or positions of doors opened and closed by the door detection sensor 181. Alternatively, the voice processing control unit 221 may analyze the voice input data to identify each occupant who has spoken, and the occupant number detection unit 222 may estimate the number of occupants from the number of speakers identified by the voice processing control unit 221. The number of occupants may also be estimated based on multiple data such as the image captured by the camera 16, switch operation data from the switch interface 15, door opening and closing detected by the door detection sensor 181, and the number of speakers analyzed by the voice processing control unit 221.

[0020] The speech determination processing unit 223 determines whether the voice input data analyzed by the voice processing control unit 221 is a speech by the occupant or external noise. For example, the speech determination processing unit 223 monitors the mouth movement of the occupant based on the captured image from the camera 16. Then, it determines whether the occupant is speaking based on the mouth movement of the occupant when the voice input data is obtained. Alternatively, the speech determination processing unit 223 may perform machine learning in advance on sounds generated in response to human mouth movement, and compare the mouth movement based on the captured image with the speech content analyzed by the voice processing control unit 221 to determine whether the occupant is speaking.

[0021] Furthermore, when multiple occupants are in the vehicle, the utterance determination processing unit 223 may further determine which occupant is speaking. For example, the dialogue device of this embodiment may be configured to recognize control commands (command words) from utterances made only by the driver, and in this case, the utterance determination processing unit 223 may determine whether the utterance is directed at the driver.

[0022] The activation unit 224 activates and stops the dialogue processing unit 225, which will be described later. Specifically, when the number of occupants is two or more, the activation unit 224 activates the dialogue processing unit 225 when a predetermined wake word is included in the voice input data analyzed by the voice processing control unit 221. On the other hand, when the number of occupants is one, the activation unit 224 activates the dialogue processing unit 225 at the timing when it is detected that the number of occupants is one, regardless of whether the occupant has uttered the wake word.

[0023] Furthermore, as described above, when the utterance determination processing unit 223 can identify a specific occupant as the speaker, the activation unit 224 may activate the dialogue processing unit 225 in response to only the wake word of the specific occupant when there are two or more occupants. For example, the activation unit 224 may activate the dialogue processing unit 225 when the utterance of the wake word by the driver is detected.

[0024] The dialogue processing unit 225 performs processing on the occupant's utterance according to the content of the utterance. Examples of processing according to the content of the utterance include vehicle control processing related to the control of the vehicle 1 and response processing to utterances unrelated to the control of the vehicle 1. The vehicle control processing is performed when the occupant's utterance contains a predetermined command word. For example, command data recording control codes for the command words is stored in the storage unit 21 in advance. As a result, the dialogue processing unit 225 determines whether the occupant's utterance contains a command word contained in the command data, and if so, outputs a control code corresponding to the command word to the control command generation unit 226.

[0025] The response process is performed when the occupant's utterance does not include a command word. The response process is a conversation process in which a voice is returned to the occupant in response to the occupant's utterance. The response process can be performed, for example, by learning, in advance, by machine learning, a response sentence to be returned to the occupant in response to a sentence included in the utterance. For example, an automatic response model for an utterance may be stored in the storage unit 21 in advance, and the dialogue processing unit 225 outputs a response obtained by the automatic response model from the speaker 13. Alternatively, the automatic response model may be recorded in a predetermined dialogue processing server that is communicable via a communication line such as the Internet. In this case, the dialogue processing unit 225 transmits voice input data to the dialogue processing server via the Internet and outputs a response returned from the dialogue processing server from the speaker 13.

[0026] The dialogue processing unit 225 further outputs a command to the display control unit 227 to indicate whether or not the dialogue processing unit 225 is running on the display 17. As a result, the display control unit 227 causes the display 17 to display, using characters, an icon image, or the like, that the dialogue processing unit 225 is running.

[0027] When a control code is input from the dialogue processing unit 225, the control command generating unit 226 generates a control command corresponding to the control code and outputs it to the ECU 11. This causes the ECU 11 to carry out processing corresponding to the control command. For example, when a control code corresponding to the command word "audio playback" is input, the control command generating unit 226 generates a control command that prompts an audio device (not shown) installed in the vehicle 1 to perform playback and outputs it to the ECU 11.

[0028] The display control unit 227 controls the display on the display 17. Information displayed on the display 17 by the display control unit 227 may include, for example, text or an icon image, as described above, input from the dialogue processing unit 225, indicating that the dialogue processing unit 225 has been started. The display control unit 227 may also display information based on a command from another component such as the ECU 11. For example, the display control unit 227 may display map data input from the IVI controller 14.

[0029] [Dialogue Method] Next, a dialogue method for the vehicle 1 as described above will be described. Fig. 2 is a flowchart relating to the activation of the dialogue processing unit 225 in the dialogue method for the vehicle 1 of this embodiment. In the vehicle 1 of this embodiment, when the vehicle 1 is powered on, the occupant number detection unit 222 detects the number of occupants in the vehicle 1 (step S1: occupant number detection step). As a method for detecting the number of occupants, as described above, the occupant number detection unit 222 may perform detection using an image of the interior of the vehicle captured by the camera 16. Alternatively, the occupant number detection unit 222 may estimate the number of occupants from input signals from the switch interface 15 or the door detection sensor 181, the number of speakers analyzed by the voice processing control unit 221, etc.

[0030] Thereafter, the activation unit 224 determines whether the number of occupants is one (step S2). If the determination in step S2 is YES, the activation unit 224 activates the dialogue processing unit 225 (step S3). As a result, the dialogue processing unit 225 is activated in a single mode in which it waits for an occupant to speak. In this single mode, the occupant does not need to utter a wake word, and the dialogue processing unit 225 performs processing according to the content of the occupant's utterance. The processing of the dialogue processing unit 225 in this single mode will be described later. Also, in step S3, the dialogue processing unit 225 outputs a command to the display control unit 227 to display that the dialogue processing unit 225 is activated. As a result, the display control unit 227 displays, on the display 17, an icon image or character data indicating that dialogue processing is available.

[0031] On the other hand, if the determination in step S2 is NO, the activation unit 224 enters an activation standby mode in which it waits without activating the dialogue processing unit 225 until an activation trigger for the dialogue processing unit 225 is input (step S4). In the activation standby mode, when voice is input from the microphone 12 (step S5), the speech determination processing unit 223 detects the movement of the occupant's mouth at the time the voice input data is input, based on image data captured by the camera 16, and determines whether the voice input data is the occupant's speech (step S6). If the determination in step S6 is NO, the process returns to step S4.

[0032] If step S6 returns a "YES" result, the voice processing control unit 221 analyzes the voice input data to acquire the speech content (step S7). The activation unit 224 then determines whether the speech content analyzed in step S7 includes a wake word (step S8). If step S8 returns a "YES" result, the activation unit 224 activates the dialogue processing unit 225 (step S9). In step S9, the dialogue processing unit 225 is activated in a multi-mode in which it waits for the occupant to speak. In this multi-mode, the occupant utters a predetermined content, and the dialogue processing unit 225 performs dialogue processing. After that, the system returns to the activation standby mode of step S4. Details of this processing in the multi-mode will be described later. Also, in step S9, similar to step S3, the dialogue processing unit 225 outputs a signal indicating that the dialogue processing unit 225 is activated to the display control unit 227, and the display control unit 227 may display an icon image or text data indicating that dialogue processing is available on the display 17. The processes from step S2 to step S9 correspond to the dialogue starting step of the present disclosure.

[0033] Next, the operation of the vehicle 1 in the single mode will be described. Fig. 3 is a flowchart showing the dialogue processing in the single mode of the vehicle 1 of this embodiment. In the single mode, the dialogue processing unit 225 first determines whether or not voice has been input from the microphone 12 (step S11). If the determination in step S11 is YES (voice has been input from the microphone 12), the speech determination processing unit 223 determines whether or not the voice input data is the speech of an occupant (step S12), similarly to step S6. If the determination in step S11 is NO, the process returns to step S11 and waits until voice is input.

[0034] If the determination in step S12 is NO, the dialogue processor 225 determines that the input voice input data is not the speech of an occupant, ignores this, and returns to step S11, thereby suppressing dialogue processing for noise such as voice output from the speaker 13 by an audio device or external voice from outside the vehicle.

[0035] If the determination in step S12 is YES, similar to step S7, the voice processing control unit 221 analyzes the voice input data and acquires the speech content (step S13). Thereafter, the dialogue processing unit 225 performs various processes corresponding to the speech content acquired in step S13.

[0036] For example, the dialogue processor 225 determines whether the utterance includes a command word (step S14). If the determination in step S14 is YES, the dialogue processor 225 determines whether the utterance includes a stop word indicating that the dialogue process should be terminated (step S15). If the determination in step S15 is YES, the dialogue processor 225 terminates the dialogue process in single mode (step S16). In this case, after the single mode process is terminated, the device transitions to the startup standby mode in step S4 of FIG. 2.

[0037] If the determination in step S15 is NO, the dialogue processing unit 225 performs vehicle control processing (step S17). In step S17, as described above, the dialogue processing unit 225 reads a control code corresponding to the command word from the command data recorded in the storage unit 21 and outputs the control code to the control command generation unit 226. As a result, the control command generation unit 226 generates a control command corresponding to the control code and outputs it to the ECU 11, and electronic control is performed in a predetermined configuration of the vehicle 1. Thereafter, the process returns to step S11 and waits for voice input from the microphone 12.

[0038] If step S14 returns NO (the utterance does not include a command word), the dialogue processor 225 performs response processing (step S18). In step S18, the dialogue processor 225 uses the automatic response model to obtain a response to the utterance, and outputs the response as voice from the speaker 13. Thereafter, the process returns to step S11, and the process waits for voice input from the microphone 12.

[0039] Next, the operation of the vehicle 1 in the multi-mode will be described. Figure 4 is a flowchart showing the dialogue processing in the multi-mode of the vehicle 1 of this embodiment. In the multi-mode, the dialogue processing unit 225 counts the elapsed time from the activation of the dialogue processing unit 225 in step S9 (step S21). Then, it is determined whether the elapsed time is within a predetermined limit time and whether voice has been input from the microphone 12 (step S22). If the elapsed time exceeds the limit time without voice input from the microphone 12, the determination in step S22 is NO, and the vehicle returns to the activation standby mode in step S4.

[0040] On the other hand, if step S22 returns YES, the utterance determination processor 223 determines whether the voice input data is the speech of an occupant (step S23), similarly to steps S6 and S12. In multi-mode, in step S23, the utterance determination processor 223 may further determine whether the speech is from a specific occupant (for example, the driver only, or the driver or the passenger in the front passenger seat). If step S23 returns NO, the dialogue processor 225 determines that no voice input has been made, and returns to step S21, where it continues counting the elapsed time and waits for voice input.

[0041] If the determination in step S23 is YES, similar to step S13 in the single mode, the voice processing control unit 221 analyzes the voice input data and acquires the speech content (step S24). Furthermore, the dialogue processing unit 225 performs various processes corresponding to the speech content acquired in step S24.

[0042] That is, the dialogue processor 225 determines whether the utterance contains a command word (step S25). If the determination in step S25 is YES, the dialogue processor 225 performs vehicle control processing (step S26), similar to step S17. Note that in the multi-mode, after step S26, the process returns to the startup standby mode of step S4 in FIG. 2. That is, the dialogue processor 225 ends the dialogue processing and waits until the wake word is again uttered by the occupant and the activation unit 224 activates the dialogue processor 225.

[0043] If the determination in step S25 is NO, similar to step S18, the dialogue processor 225 performs response processing (step S27). Note that in the multi-mode, after step S27, the dialogue processor 225 returns to the startup standby mode of step S4. That is, the dialogue processor 225 ends the dialogue processing and waits until the wake word is again uttered by the occupant and the activation unit 224 activates the dialogue processor 225.

[0044] Next, the operation of stopping the dialogue processing in the single mode will be described. FIG. 5 is a flowchart showing the dialogue stop processing in the single mode of the vehicle 1 according to this embodiment. During the dialogue processing in the single mode shown in FIG. 3, when the dialogue processing unit 225 acquires a predetermined dialogue stop trigger (step S31), it stops the dialogue processing based on the stop trigger (step S32). After this, the system transitions to the startup standby mode in step S4. Steps S31 and S32 are stop steps in the present disclosure. Examples of stop triggers include an occupant answering a phone call, an occupant uttering a stop word, and an increase in the number of occupants. Note that, in this embodiment, when an occupant utters a stop word, the stop word is considered to be included in the command words, and the processing is performed as described in steps S15 and S16 above. Therefore, stop triggers other than an occupant uttering a stop word will be described here.

[0045] The call detection unit 141 of the IVI controller 14 detects whether the occupant is answering a phone call. For example, an external communication device carried by the occupant is connected to the IVI controller 14 wirelessly or via a wire. When the occupant makes a call using the external communication device, a signal indicating that the occupant is answering a phone call is input from the external communication device to the IVI controller 14. This allows the call detection unit 141 to detect the call made by the occupant. Furthermore, a call on the external communication device connected to the IVI controller 14 may be performed using the microphone 12 and the speaker 13. The call detection unit 141 may detect a call by connecting the external communication device to the microphone 12 and the speaker 13 in response to a call request from the external communication device. Furthermore, the call detection by the call detection unit 141 is not limited to the call detection. A call made by the occupant using the external communication device may also be detected from an image captured by the camera 16.

[0046] Regarding changes in the number of occupants, the occupant number detection unit 222 detects the number of occupants based on images captured by the camera 16. Alternatively, the occupant number detection unit 222 may estimate the number of occupants based on switch operation positions detected by the switch interface 15 and door opening and closing detected by the door detection sensor 181. The dialogue processing unit 225 maintains the single mode when the detected or estimated number of occupants is one, and determines that a stop trigger has been acquired when the number of occupants becomes two or more.

[0047] [Effects of the Present Embodiment] The vehicle 1 of the present embodiment includes a controller 20 configured by a computer. The processor 22 of the controller 20 functions as a voice processing control unit 221, an occupant number detection unit 222, a startup unit 224, a dialogue processing unit 225, and the like by reading and executing programs stored in the storage unit 21. The controller 20 then performs an occupant number detection step (step S1) and a dialogue startup step (steps S2 to S9). In the occupant number detection step, the occupant number detection unit 222 detects or estimates the number of occupants in the vehicle 1. In the dialogue startup step, if the startup unit 224 detects that there is one occupant, it starts the dialogue processing unit 225 at the timing when it detects that there is one occupant (the timing when it determines YES in step S2). As a result, when there is one occupant, the occupant does not need to utter a wake word every time he or she wants to interact with the dialogue device of the vehicle 1, thereby improving the convenience of the dialogue device.

[0048] In this embodiment, in the dialogue activation step, when the number of occupants is two or more, the activation unit 224 activates the dialogue processing unit 225 when the voice uttered by the occupant includes a wake word. In this case, when two or more occupants are aboard the vehicle 1, dialogue processing by the dialogue processing unit 225 is not performed unless the wake word is uttered. Therefore, conversation between occupants is not hindered by response processing by the dialogue processing unit 225. Furthermore, it is possible to prevent inconveniences such as unintended vehicle control processing or response processing by the dialogue processing unit 225 being performed due to conversation between occupants. Furthermore, when an occupant wishes to interact with the vehicle 1, the dialogue processing unit 225 is activated by the wake word, so the dialogue processing unit 225 can be easily activated without the occupant performing any special operation.

[0049] In this embodiment, in the dialogue activation step, if the number of occupants is two or more, the speech determination processing unit 223 determines the occupant's speech based on the occupant's speech actions (mouth movements, etc.) based on the image captured by the camera 16 at the time when the voice input data is input from the microphone 12. Then, when it is determined that the occupant is speaking, if the speech content includes a wake word, the dialogue processing unit 225 is activated. This prevents the dialogue processing unit 225 from being unintentionally activated due to noise, such as a voice from outside the vehicle or audio output from the speaker 13, and makes it possible to activate the dialogue processing unit 225 at a timing desired by the occupant.

[0050] In this embodiment, in the single mode, the dialogue processor 225 performs the stop step from step S31 to step S32, thereby stopping the active dialogue processor 225 when detecting at least one of a conversation by an occupant, an increase in the number of occupants, and a stop word by an occupant. This makes it possible to avoid unnecessary response processing and vehicle control processing by the dialogue processor 225, even in the single mode. For example, by detecting a conversation by an occupant and stopping the dialogue processing, it is possible to avoid the inconvenience of the dialogue processor 225 responding to a conversation by an occupant. Furthermore, by stopping the dialogue processing due to an increase in the number of occupants, it is possible to avoid the inconvenience of the dialogue processor 225 responding to a conversation between occupants. By stopping the dialogue processing due to a stop word by an occupant, it is possible to avoid unintended vehicle control processing and response processing by the dialogue processor 225 due to an occupant talking to themselves, for example.

[0051] In this embodiment, even if the dialogue processor 225 is stopped in the stopping step, the process returns to step S4 to enter the startup standby mode, and when the utterance of the wake word by the occupant is detected, the activation unit 224 activates the dialogue processor 225. This makes it possible to easily restart the dialogue processor 225 by uttering the wake word.

[0052] In this embodiment, while the dialogue processor 225 is activated, it displays a message that the dialogue processor 225 is activated on the display 17. This allows the occupant to easily determine, by checking the display 17, whether it is necessary to activate the dialogue processor 225 using a wake word or whether dialogue processing is possible without a wake word.

[0053] Second Embodiment Next, a second embodiment will be described. In the first embodiment, the activation unit 224 activates the dialogue processor 225 in single mode when there is one occupant. In this single mode, both vehicle control processing based on a command word and response processing responding to utterances other than command words can be performed without the occupant uttering a wake word. In contrast, the second embodiment differs from the first embodiment in that vehicle control processing that does not use a wake word is restricted in single mode. In the following description, items that have already been described will be designated by the same reference numerals, and their description will be omitted or simplified.

[0054] The vehicle 1 of this embodiment has the same configuration as the first embodiment, and therefore an overall configuration diagram is omitted. The vehicle 1 of this embodiment also performs the same startup process for the dialogue processor 225 as in FIG. 2 . That is, when there is one occupant, the dialogue processor 225 is started in single mode, and when there are two or more occupants, the dialogue processor 225 is started in multi mode. In this embodiment, the processing in single mode differs from that in the first embodiment.

[0055] 6 is a flowchart showing the dialogue processing in the single mode of the vehicle 1 according to this embodiment. In the single mode, the dialogue processing unit 225 performs the processes of steps S41 to S43, which are similar to steps S11 to S13. That is, the dialogue processing unit 225 determines whether or not voice has been input from the microphone 12 (step S41). If the determination in step S41 is NO, the process returns to step S41 and waits until voice is input. If the determination in step S41 is YES, the speech determination processing unit 223 determines whether or not the voice input data is the speech of the occupant, similar to step S12, etc. (step S42). If the determination in step S42 is NO, the process returns to step S41. If the determination in step S42 is YES, the voice processing control unit 221 analyzes the voice input data to acquire the content of the occupant's speech (step S43).

[0056] Next, in this embodiment, it is determined whether the utterance content is a wake word (step S44). If the determination in step S44 is YES, the dialogue processor 225 determines whether voice input data has been input within a predetermined time limit since the voice input in step S41 (step S45). If the time elapsed since step S41 exceeds the time limit in step S45, the process returns to step S41. If voice input data has been input within the time limit in step S45, the utterance determination processor 223 determines whether the voice input data is the result of an occupant's speech (step S46), similar to step S12, etc. If the determination in step S46 is NO, the dialogue processor 225 determines that there is no voice input data, and the process returns to step S45. If the determination in step S46 is YES, the processes in steps S13 to S18 of the first embodiment are performed. Furthermore, in this embodiment, the process returns to step S41 after steps S17 and S18.

[0057] On the other hand, if step S44 returns NO (no wake word is uttered), the dialogue processor 225 determines whether the utterance content analyzed in step S43 includes a command word (step S47). If step S47 returns YES, the dialogue processor 225 does not perform vehicle control processing and returns to step S41. In other words, vehicle control processing is not performed without the wake word.

[0058] Furthermore, if the determination in step S47 is NO, the response process is performed similarly to step S18 in the first embodiment. That is, in this embodiment, too, the dialogue processor 225 performs the response process regardless of whether a wake word is present. Note that in the example of FIG. 6 , if the determination in step S44 is NO, the dialogue processor 225 determines in step S47 whether the utterance content includes a command word, but this is not limiting. For example, if the determination in step S44 is NO, the dialogue processor 225 may perform the response process of step S18 regardless of whether a command word is present. For example, if the utterance content includes a command word, the dialogue processor 225 may perform the response process of outputting a sound from the speaker 13 indicating that a wake word is required.

[0059] In this embodiment, the process returns to step S41 after step S17 or step S18. That is, the response process by the dialogue processor 225 is performed regardless of whether the wake word is uttered, but the wake word must be uttered each time the vehicle control process is performed. Note that, once the wake word is input and step S44 returns YES, the process of steps S11 to S18 in the first embodiment may be performed instead. That is, once the wake word is input, the vehicle control process may be performed without the need for a wake word.

[0060] In the second embodiment described above, the state of step S41 in which the system waits for a voice input by the occupant without detecting the wake word corresponds to the first state of the present disclosure. Also, step S45 in which the system waits for a voice input by the occupant after detecting the wake word (after determining YES in step S44) corresponds to the second state of the present disclosure.

[0061] [Effects of the Present Embodiment] The present embodiment provides the same effects as the first embodiment, and further provides the following effects. In the present embodiment, when the number of occupants is one, the activation unit 224 activates the dialogue processing unit 225 in single mode at the timing of detecting the number of occupants, and enters a first state in which the dialogue processing unit 225 waits for an utterance from the occupant. If an utterance from the occupant is detected in step S41 and if the wake word is detected in the utterance (step S44: YES), the dialogue processing unit 225 transitions to a second state in which the dialogue processing unit 225 waits for an utterance from the occupant. In the first state, even if an utterance of a command word is detected, the dialogue processing unit 225 does not accept the command word and the vehicle control process is restricted by determining NO in step S44 and YES in step S47. On the other hand, in the second state, if an utterance of a command word is detected, the dialogue processing unit 225 performs the processes from step S14 to step S17, and thereby performs the vehicle control process. In either the first state or the second state, if an utterance of a word other than a command word is detected, the dialogue processor 225 performs a response process to output a voice response.

[0062] As a result, even in the single mode, by requiring a wake word for input of a command word related to vehicle control, it is possible to prevent the inconvenience of unintended vehicle control being performed due to an erroneous utterance by the occupant, etc. Furthermore, as with the first embodiment, the response process can be used regardless of the wake word, improving the convenience of dialogue processing.

[0063] [Third Embodiment] Next, a third embodiment will be described. In the first embodiment, when the number of occupants is one, the activation unit 224 activates the dialogue processing unit 225 in single mode at the timing when it detects or estimates that the number of occupants is one. In contrast, this embodiment differs from the first embodiment in that the timing when the dialogue processing unit 225 is activated in single mode is different.

[0064] The vehicle 1 of this embodiment has the same configuration as the first embodiment, and therefore an overall configuration diagram is omitted. Furthermore, the vehicle 1 of this embodiment activates the dialogue processor 225 in single mode when there is one occupant, and activates the dialogue processor 225 in multi mode when there are two or more occupants. However, in this embodiment, the activation timing of the single mode is the timing when an utterance by an occupant is detected. The activation process of the dialogue processor 225 of this embodiment will be described below.

[0065] 7 is a flowchart showing the processing up to activation of the dialogue processing unit 225 in the dialogue method of the third embodiment. In the vehicle 1 of this embodiment, as in the first embodiment, in step S1, the vehicle 1 is powered on, and the occupant number detection unit 222 detects the number of occupants in the vehicle 1. In addition, in step S2, the activation unit 224 determines whether the number of occupants is one. Here, if the determination in step S2 is NO, the processing of steps S4 to S9 is performed as in the first embodiment. In other words, in this embodiment, too, the method of activating the dialogue processing unit 225 in multi-mode is the same as in the first embodiment.

[0066] On the other hand, if step S2 returns YES, the activation unit 224 enters a single-activation standby mode, in which it waits until a activation trigger for activating the dialogue processor 225 is input (step S51). In the single-activation standby mode, when voice is input from the microphone 12 (step S52), the speech determination processor 223 determines whether the voice input data is the voice of an occupant (step S53). If step S53 returns NO, the process returns to step S51. On the other hand, if step S53 returns YES, the activation unit 224 activates the dialogue processor 225 in single mode (step S54). In other words, if there is one occupant, the activation unit 224 activates the dialogue processor 225 when the occupant utters any word. Therefore, in this embodiment, the occupant does not need to utter a specific wake word to activate the dialogue processor 225.

[0067] In this embodiment, the processing from step S13 onward may be performed for any word input in step S52, and processing by the dialogue processing unit 225 (vehicle control processing or response processing) may be performed.

[0068] [Effects of the Present Embodiment] In the vehicle 1 of the present embodiment, in the dialogue activation step, the activation unit 224 activates the dialogue processing unit 225 at the timing when it recognizes an utterance of any word by the occupant (at the timing when it is determined as YES in step S52 and as YES in step S53). As a result, as in the first embodiment, when there is only one occupant, the occupant does not need to utter the wake word every time they want to interact with the dialogue device of the vehicle 1, thereby improving the convenience of the dialogue device.

[0069] [Modifications] The present invention is not limited to the above-described embodiment, and includes the following modifications within the scope of achieving the object of the present invention.

[0070] [Variation 1] In the first to third embodiments described above, an example has been shown in which the activation unit 224 activates the dialogue processing unit 225 when detecting the utterance of a wake word by an occupant as a method of activating the dialogue processing unit 225 in multi-mode. However, this is not limiting. For example, the activation unit 224 may activate the dialogue processing unit 225 when a predetermined operation for activating the dialogue processing unit 225 is performed. Examples of the predetermined operation include pressing a switch for activating the dialogue processing unit 225 that is provided at a predetermined position inside the vehicle, operating a touch panel provided on the display 17, etc.

[0071] [Modification 2] In the first to third embodiments, after voice input data is obtained from the microphone 12, the speech determination processing unit 223 determines whether the voice input data is the speech of an occupant using the camera 16 or the like, which is an observation unit. However, this is not limiting. For example, the determination processing by the speech determination processing unit 223 may be omitted. Alternatively, the voice processing control unit 221 may analyze the voice input data and estimate the speaker from the voice feature quantities. In this case, the voice processing control unit 221 can filter out sounds that deviate from the voice feature quantities of the occupant as noise sounds.

[0072] [Variant 3] In the first embodiment, examples of stop triggers were described as detection of a conversation by an occupant, an increase in the number of occupants, and detection of an occupant uttering a stop word, but other triggers may also be included, such as an operational input to terminate the dialogue processing.

[0073] In addition, although an example has been described in which, when a stop trigger is detected in step S31, the dialogue processor 225 is stopped in step S32 and the vehicle transitions to the startup standby mode in step S4, the present invention is not limited to this. For example, if the stop trigger is the detection of a call by an occupant, the dialogue processor 225 may be temporarily stopped while the call is detected, and after the call ends, the process may return to step S11 to continue operation of the dialogue processor 225 in the single mode. In addition, in the first embodiment, an example has been described in which the vehicle transitions to the startup standby mode in response to the input of a stop trigger in the single mode. However, in the multi-mode or startup standby mode, if it is detected that the number of occupants has decreased to one, the process of transitioning to the single mode in step S11 may be performed.

[0074] 1...vehicle, 11...ECU, 12...microphone (voice recognition unit), 13...speaker, 14...IVI controller, 15...switch interface, 16...camera, 17...display, 20...controller, 21...memory unit, 22...processor, 141...call detection unit, 181...door detection sensor, 221...voice processing control unit, 222...occupant number detection unit, 223...speech determination processing unit, 224...startup unit, 225...dialogue processing unit, 226...control command generation unit, 227...display control unit.

Claims

1. A dialogue method using a computer that is mounted on a vehicle and has a dialogue processing unit that performs predetermined processing in response to utterances by occupants of the vehicle, wherein the computer carries out an occupant number detection step that detects or estimates the number of occupants in the vehicle, and a dialogue activation step that activates the dialogue processing unit, and in the dialogue activation step, if the number of occupants is one, the dialogue processing unit is activated at the timing when it is detected or estimated that the number of occupants is one, and if the number of occupants is two or more, the dialogue processing unit is activated when the utterance by the occupant includes a predetermined wake word.

2. The dialogue method according to claim 1, wherein, in the dialogue activation step, if the number of occupants is two or more, the speech behavior of the occupants is observed using an observation unit that observes the interior of the vehicle, and when the speech behavior of the occupant is detected at the timing when the speech by the occupant is detected and the occupant utters the wake word, the dialogue processing unit is activated.

3. The dialogue method described in claim 1, wherein the computer further performs a stopping step of stopping the dialogue processing unit that is running when, while the number of occupants is one and the dialogue processing unit is running, it detects at least one of a conversation by the occupant, an increase in the number of occupants, and the utterance of a predetermined stop word by the occupant.

4. The dialogue method according to claim 3, wherein the computer activates the dialogue processing unit when detecting the utterance of a predetermined wake word by the occupant while the dialogue processing unit is stopped by the stopping step.

5. The dialogue method according to claim 1, wherein in the dialogue activation step, if the number of occupants is one, the dialogue processing unit is activated at the timing when it is detected or estimated that the number of occupants is one, and a first state is entered in which a speech input by the occupant is awaited; and when a predetermined wake word is detected in the first state, a second state is entered in which a speech input by the occupant is awaited; in the first state, if a command word related to control of the vehicle is detected, the command word is not accepted, and if a word other than the command word is detected, a response process is performed to output a voice response; and in the second state, if the command word is detected, control of the vehicle corresponding to the command word is performed, and if a word other than the command word is detected, the response process is performed.

6. The dialogue method according to claim 1, wherein, when the dialogue processing unit is activated, the computer causes a display unit provided inside the vehicle to display that the dialogue processing unit is activated.

7. A dialogue method using a computer that is mounted on a vehicle and has a dialogue processing unit that performs predetermined processing in response to utterances by occupants of the vehicle, wherein the computer carries out an occupant number detection step that detects or estimates the number of occupants of the vehicle, and a dialogue activation step that activates the dialogue processing unit, and in the dialogue activation step, if the number of occupants is one, the dialogue processing unit is activated when an utterance of any word by the occupant is detected, and if the number of occupants is two or more, the dialogue processing unit is activated when the utterance by the occupant includes a predetermined wake word.

8. A dialogue device mounted on a vehicle, comprising: a voice recognition unit that recognizes the voices of occupants riding in the vehicle; an occupant number detection unit that detects or estimates the number of occupants; a dialogue processing unit that performs predetermined processing in response to the occupants' speech; and a startup unit that starts the dialogue processing unit, wherein the startup unit starts the dialogue processing unit when it detects or estimates that the number of occupants is one, if the number of occupants is one, and starts the dialogue processing unit when the occupants' speech includes a predetermined wake word, if the number of occupants is two or more.

9. A dialogue device mounted on a vehicle, comprising: a voice recognition unit that recognizes the voices of occupants riding in the vehicle; an occupant number detection unit that detects or estimates the number of occupants; a dialogue processing unit that performs predetermined processing in response to the occupant's speech; and a startup unit that starts the dialogue processing unit, wherein the startup unit starts the dialogue processing unit when it recognizes the utterance of any word by the occupant if the number of occupants is one, and starts the dialogue processing unit when the occupant's speech includes a predetermined wake word if the number of occupants is two or more.

Citation Information

Patent Citations

  • Enhanced noise reduction in voice activated device

    JP2023059845A

  • System and program etc.

    JP2023151512A

  • Voice recognition device, control method, program, and storage medium

    WO2023062817A1