Information processing apparatus, method, and vehicle
When switching from autonomous to manual driving, the information processing device provides partial prompts based on the driver's voice and preset correspondences, thus resolving the driver's discomfort and improving the driving experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2022-07-08
- Publication Date
- 2026-05-12
AI Technical Summary
When switching from autonomous driving to manual driving, drivers may experience discomfort, and current technology has failed to effectively mitigate this feeling.
During the autonomous driving control process, the information processing device prompts the driver with some reasons to guide the switch to manual driving. Based on the driver's voice and the pre-set correspondence, detailed explanations are provided step by step until the driver accepts the switch.
It reduces the driver's discomfort when switching from automatic to manual driving and improves the driving experience through reasonable prompts and information processing devices.
Smart Images

Figure CN122009241A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 202210805457.4, application date July 8, 2022, and invention title "Information Processing Apparatus, Method and Vehicle". Technical Field
[0002] This disclosure relates to an information processing apparatus, method, and vehicle. Background Technology
[0003] An autonomous driving assistance system is currently disclosed in which, when it is determined that autonomous driving control cannot be implemented, a reason for the inability to implement autonomous driving control is obtained, and the reason is also guided along with information indicating the meaning of the inability to implement autonomous driving control (e.g., Patent Document 1).
[0004] Prior art literature Patent documents Patent Document 1: Japanese Patent Application Publication No. 2016-28927 Summary of the Invention The problem that the invention aims to solve One of the challenges is to provide an information processing device, method, and vehicle that can reduce the unpleasant feelings experienced by the driver during the transition from automatic driving control to manual driving control.
[0005] Methods for solving problems One aspect of this disclosure is an information processing apparatus, wherein... The system includes a control unit that performs the following processing: During the autonomous driving control process of the first vehicle, the system detects situations requiring a switch to manual driving. When the switch is requested, the driver of the first vehicle is prompted with a portion of an explanation relating to the reason for the switch.
[0006] Another aspect of this disclosure is a method that includes, The information processing device performs the following processing: During the autonomous driving control process of the first vehicle, the system detects situations requiring a switch to manual driving. When the switch is requested, the driver of the first vehicle is prompted with a portion of an explanation relating to the reason for the switch.
[0007] Another aspect of this disclosure is a vehicle, wherein... The system includes a control unit that performs the following processing: During the autonomous driving control process of the first vehicle, the system detects situations requiring a switch to manual driving. When the switch is requested, a process is implemented to provide the driver of the first vehicle with a portion of an explanation related to the reason for the switch.
[0008] Invention Effects According to one aspect of this disclosure, the unpleasant sensations experienced by the driver can be reduced during the transition from automatic driving control to manual driving control. Attached Figure Description
[0009] Figure 1 This is a diagram illustrating an example of the system structure of the takeover guidance system and the system structure of the vehicle 10 according to the first embodiment.
[0010] Figure 2 This is an example of the hardware structure of a multimedia ECU.
[0011] Figure 3 A diagram illustrating an example of the functional structure of a vehicle and a central server.
[0012] Figure 4 An example of an intent numbering table.
[0013] Figure 5 A diagram illustrating an example of a dialogue script generated to explain the reasons for the takeover request.
[0014] Figure 6 An example of a dialogue level 1 correspondence table for situations where a takeover request arises because GNSS signal reception is difficult.
[0015] Figure 7 An example of a dialogue level 2 correspondence table for situations where a takeover request arises because GNSS signal reception is difficult.
[0016] Figure 8 An example of a dialogue level 3 correspondence table for situations where a takeover request arises because GNSS signal reception is difficult.
[0017] Figure 9 An example of a dialogue level 4 correspondence table for situations where a takeover request arises because GNSS signal reception is difficult.
[0018] Figure 10 This is an example of a flowchart illustrating the guidance process for taking over a vehicle according to the first embodiment.
[0019] Figure 11This is an example of a flowchart illustrating the bootstrapping process for the central server takeover as described in the first embodiment.
[0020] Figure 12 An example of a sequence diagram for downloading and processing dialogues for a corresponding table group.
[0021] Figure 13 An example of a sequence diagram for downloading and processing dialogues for a corresponding table group.
[0022] Figure 14 This diagram illustrates an example of the functional structure of the vehicle and central server involved in the second embodiment.
[0023] Figure 15 This is an example of a flowchart illustrating the guidance process for vehicle takeover according to the second embodiment.
[0024] Figure 16 An example of a flowchart of the bootstrapping process for the takeover of the central server according to the second embodiment. Detailed Implementation
[0025] In vehicles that switch between autonomous and manual driving modes, the transition from autonomous to manual mode is called takeover. The longer a driver has driven such a vehicle, the more experienced they become in recognizing when a takeover request will occur. A takeover request is a request from the driver to switch from autonomous to manual driving. Furthermore, the desire to know the reason for a takeover request sometimes depends on the driver's personality, mood, or driving state. For example, it's possible that a driver who has driven the same autonomous vehicle for many years and understands the conditions under which a takeover request occurs, or who is listening to music in the autonomous vehicle, or who is talking to other passengers, might become annoyed by the explanation of the reason for the takeover request. On the other hand, there are also situations where the driver wants to know more details about the reason for the takeover request.
[0026] One aspect of this disclosure is an information processing apparatus comprising a control unit that performs the following processing: when a request to switch to manual driving arises during the automatic driving control of a first vehicle, the control unit provides the driver of the first vehicle with a portion of the explanation related to the reason for the request. The information processing apparatus is, for example, an ECU (Electronic Control Unit) or an in-vehicle unit mounted on the first vehicle. However, it is not limited to this; the information processing apparatus may also be, for example, a server capable of communicating with the first vehicle. The control unit is, for example, a processor such as a CPU (Central Processing Unit). The method of providing the reason for the request to switch to manual driving can be, for example, voice output from a speaker in the first vehicle or message output to a display in the first vehicle. The request to switch to manual driving is also referred to as a takeover request.
[0027] According to one aspect of this disclosure, when a request to switch to manual driving arises during autonomous driving control, the unpleasant experience for the driver can be reduced by providing part, but not all, of an explanation related to the reason for the request.
[0028] In one embodiment of this disclosure, the control unit may acquire the voice of the driver of the first vehicle upon receiving the request. In this case, the control unit may also prompt the driver with a portion of the explanation related to the reason for the request, corresponding to the driver's voice. The driver's voice is acquired, for example, through speech recognition processing. The driver's voice reflects the degree of concern the driver has for the explanation related to the reason for the request to switch to manual driving. For example, if the driver desires to know the reason for the request, the driver's voice will express this intention, thus prompting a more detailed explanation. For example, if the driver is not considered to desire to know the reason for the request, since the driver's voice expresses this intention, a simple explanation will be prompted. Therefore, according to one embodiment of this disclosure, the explanation will be prompted based on the degree of concern the driver has for the explanation related to the reason for the request to switch to manual driving.
[0029] In one embodiment of this disclosure, the information processing device may include a storage unit that stores a correspondence between a first vocal content and a portion of the explanation related to the reason for the request. Alternatively, the control unit may prompt the driver with a portion of the explanation related to the reason for the request, which is associated with the first vocal content, when the driver's vocal content is at least similar to the first vocal content. "At least similar" includes cases where the driver's vocal content is similar to the first vocal content and cases where the driver's vocal content is consistent with the first vocal content. By maintaining a correspondence between a pre-conceived vocal content and a portion of the explanation related to the reason for the request, which serves as a response to the vocal content, the information processing device can further reduce the response delay to the driver's vocalizations.
[0030] In one embodiment of this disclosure, the information processing device may also be mounted on the first vehicle. That is, for example, the information processing device may be one of a plurality of ECUs mounted on the first vehicle, or an in-vehicle unit. Alternatively, the control unit may further perform the following processing: upon receiving the request, downloading from a predetermined device the correspondence between the driver's voice and a portion of the explanation related to the reason for the request, and storing this correspondence in a storage unit. Therefore, since it is only necessary to download the driver's voice and the correspondence between the explanation related to the reason for the request and store this correspondence in the storage unit when necessary, the resources of the storage unit can be utilized effectively.
[0031] In one embodiment of this disclosure, the control unit may also perform a process of obtaining the reason for the request. Alternatively, the control unit may download a mapping between the driver's voice and a portion of the explanation related to the reason for the request, corresponding to the reason for the request. Therefore, since it is unnecessary to download mappings other than those corresponding to the reason for the request to switch to manual driving, communication bandwidth and storage capacity of the storage unit can be saved.
[0032] Alternatively, the control unit can be configured to download from a predetermined device a mapping between the driver's verbal input and a portion of the explanation related to the reason for the request to switch to manual driving. This allows for improved response speed for each verbal input, even when the driver makes multiple inputs.
[0033] Alternatively, the correspondence between the driver's spoken content and a portion of the explanation related to the reason for the request may include a first correspondence and at least one second correspondence. The first correspondence is established by including multiple second spoken contents conceived when inquiring about the reason for the request to switch to manual driving, and the reason for the request, as part of the explanation related to the reason for the request. The second correspondence is established by including multiple third spoken contents conceived as accepting the reason for the request and further inquiring about it, and the response to the inquiry, as part of the explanation related to the reason for the request.
[0034] In this scenario, the control unit can also be configured such that, after generating the request, it refers to a first correspondence and provides the driver with an explanation of the reason for the request, which corresponds to a second voice message similar to the driver's voice. Furthermore, it can also be configured such that, after providing the driver with an explanation of the reason for the request, the control unit refers to at least one second correspondence and provides the driver with a response to an inquiry, which corresponds to a third voice message similar to the driver's voice. This allows for the phased provision of explanations related to the reasons for the request to switch to manual driving to the driver.
[0035] In one embodiment of this disclosure, the information processing device may be mounted on a first vehicle. Alternatively, the control unit may send the driver's voice to a predetermined device and receive from the predetermined device a portion of the explanation relating to the driver's voice and the reason for the request. Thus, the information processing device only needs to receive from the predetermined device a portion of the explanation relating to the driver's voice and the reason for the request to switch to manual driving, thereby reducing the use of storage space in the storage device.
[0036] In one embodiment of this disclosure, the control unit may repeatedly perform the following process until the driver begins manual driving: acquiring the driver's verbal input and providing a corresponding explanation related to the reason for the request. Whether the driver has begun manual driving can be detected, for example, by monitoring images captured by a camera installed in the first vehicle or by steering wheel operation. In another embodiment of this disclosure, the control unit may repeatedly perform the following process until the driver's verbal input indicates acceptance of the switch to manual driving: acquiring the driver's verbal input and providing a corresponding explanation related to the reason for the request. Thus, during the switch from automatic to manual driving, the process of providing a corresponding explanation related to the reason for the switch can be terminated.
[0037] Another aspect of this disclosure can be specifically defined as a method performed by the aforementioned information processing apparatus. This method includes processing performed by the information processing apparatus as follows: detecting a request to switch to manual driving during the automatic driving control of the first vehicle; and, if such a request is generated, providing the driver of the first vehicle with a portion of an explanation relating to the reason for the request. Furthermore, another aspect of this disclosure can be specifically defined as a program for causing a computer to execute the method, and a computer-readable, non-transitory recording medium for recording the program. Furthermore, another aspect of this disclosure can be specifically defined as a vehicle equipped with the aforementioned information processing apparatus.
[0038] The embodiments of this disclosure will now be described with reference to the accompanying drawings. The structures of the embodiments described below are illustrative, and this disclosure is not limited to the structures of the embodiments.
[0039] <First Implementation> Figure 1 This diagram illustrates an example of the system structure of the takeover guidance system 100 according to the first embodiment and the system structure of the vehicle 10. The takeover guidance system 100 is a system that guides the driver of the vehicle 10 to switch to manual driving when the vehicle 10 generates a request to do so. The switch of the vehicle 10 from automatic driving to manual driving is referred to as takeover.
[0040] The takeover guidance system 100 includes a vehicle 10 and a central server 50. The vehicle 10 is a connected vehicle equipped with a DCM (Data Communication Module) 1 capable of communication. The vehicle 10 is a vehicle that switches between and operates in autonomous driving and manual driving modes. The vehicle 10 can be either an engine-driven vehicle or an electric motor-driven vehicle. The vehicle 10 is an example of a "first vehicle".
[0041] The central server 50 provides support for the autonomous driving control of the vehicle 10 and provides pre-defined services to the vehicle 10 via communication. The vehicle 10 and the central server 50 can communicate via a network N1, such as the Internet. The vehicle 10's DCM1 connects to the wireless network via mobile wireless communication methods such as LTE (Long Term Evolution), 5G (5th Generation), and 6G (6th Generation), as well as wireless communication methods such as Wi-Fi or DSCR, and connects to the Internet through this wireless network.
[0042] Vehicle 10 is equipped with DCM1, multimedia ECU2, automatic driving control ECU3, microphone 4, speaker 5, sensor 6, and other ECU9. However, in Figure 1 In this document, as part of the system structure of vehicle 10, devices related to the processing involved in the first embodiment are extracted and shown. The system structure of vehicle 10 is not limited to... Figure 1 The structure shown.
[0043] DCM1, multimedia ECU2, automatic driving control ECU3, and other ECUs9 are connected together via, for example, CAN (Controller Area Network) or Ethernet (registered trademark). Other ECUs9 include various ECUs related to driving control and ECUs related to position management.
[0044] DCM1 is a device that includes an antenna, a transmitter and receiver, and a modulator and demodulator, and performs communication functions for vehicle 10. DCM1 accesses network N1 wirelessly and communicates with central server 50.
[0045] The multimedia ECU 2 is connected to, for example, the microphone 4 and the speaker 5, and controls them. The multimedia ECU 2 may include, for example, a car navigation system and an audio system. In the first embodiment, the multimedia ECU 2 accepts input of the driver's voice via the microphone 4. The multimedia ECU 2 outputs voice instructions for taking over control into the vehicle 10 via the speaker 5.
[0046] The autonomous driving control ECU 3 implements autonomous driving control of the vehicle 10. Various sensors 6 mounted on the vehicle 10 are connected to the autonomous driving control ECU 3, and the ECU 3 receives signal input from these sensors 6. These sensors 6 include, for example, cameras, LiDAR, Radar, GNSS (Global Navigation Satellite System) receivers, GNSS receiving antennas, acceleration sensors, yaw rate sensors, and rain sensors. HMI (Human Machine Interface) devices may also be included among the sensors 6. The autonomous driving control ECU 3 connects to the various sensors 6 via an in-vehicle network or directly to them.
[0047] The autonomous driving control ECU3 executes an autonomous driving control algorithm based on input signals from various sensors 6, and outputs control signals to actuators that drive the brakes, accelerator, steering wheel, headlights, turn indicators, brake lights, hazard lights, and drive circuits, thereby achieving autonomous driving. In addition to control signals, the autonomous driving control ECU3 also outputs information to HMI devices such as the instrument panel and displays.
[0048] In the first embodiment, the autonomous driving control ECU 3 determines, based on input signals from various sensors 6, whether autonomous driving will become technically difficult in the near future (e.g., a few seconds later) under certain driving conditions. If it is determined that autonomous driving will become technically difficult, the autonomous driving control ECU 3 generates a takeover request signal, along with this reason, to request the driver to switch to manual driving. The takeover request signal is input to the multimedia ECU 2.
[0049] When the multimedia ECU2 receives a takeover request signal, it outputs a voice prompt guiding the driver to switch to manual driving via speaker 5. Furthermore, in the first embodiment, the explanation related to the reason for the takeover request is presented in a dialogue format. The multimedia ECU2 downloads a mapping table from the central server 50 containing the driver's voice prompts intended to request an explanation related to the reason for the takeover request, and a response containing a portion of the explanation. The multimedia ECU2 then monitors the driver's voice prompts, retrieves the response from the mapping table, creates voice data based on the retrieved response, and outputs the voice data via speaker 5.
[0050] In the first embodiment, a portion of the explanation related to the reason for the takeover request is displayed in relation to the driver's verbal input related to the request. This explanation is prepared in stages in a mapping table obtained from the central server 50. Therefore, if the driver feels that the displayed explanation is insufficient, they may make a further verbal request seeking a more in-depth explanation, to which further explanation will be displayed. Conversely, if the driver feels that the displayed explanation is sufficient, they will accept the takeover request. Therefore, according to the first embodiment, a more appropriate explanation can be displayed to the driver regarding the reason for the takeover request, thereby reducing any unpleasantness to the driver.
[0051] Figure 2This is an example of the hardware structure of a multimedia ECU2. The multimedia ECU2, as a hardware structure, includes a CPU 201, a memory 202, an auxiliary storage device 203, an input interface 204, an output interface 205, and an interface 206. The memory 202 and the auxiliary storage device 203 are recording media that can be read by a computer.
[0052] The auxiliary storage device 203 stores various programs and the data used by the CPU 201 during the execution of each program. The auxiliary storage device 203 may be, for example, an EPROM (Electrically Erasable Programmable ROM) or flash memory. Among the programs stored in the auxiliary storage device 203 are, for example, a voice recognition program, a voice signal processing program, and a takeover and guidance control program. The voice signal processing program performs digital-to-analog conversion of voice signals and conversion between voice signals and data in a predetermined format. The takeover and guidance control program implements control for switching to manual driving.
[0053] Memory 202 is a storage area and working area for providing the CPU 201 with a storage device for downloading programs stored in auxiliary storage device 203, or a storage device used as a buffer. Memory 202 includes, for example, semiconductor memory such as ROM (Read Only Memory) or RAM (Random Access Memory).
[0054] CPU 201 performs various processes by loading the OS, which is stored in auxiliary storage device 203, and other various programs into memory 202 and executing them. CPU 201 is not limited to one; multiple CPUs may be present. CPU 201 includes a high-speed cache memory 201M.
[0055] Input interface 204 is for connecting microphone 4. Output interface 205 is for connecting speaker 5. Interface 206 is a circuit with a port for connecting to, for example, Ethernet (registered trademark), CAN, or other networks. Furthermore, the hardware structure of the multimedia ECU2 is not limited to... Figure 2 The structure shown.
[0056] Like the multimedia ECU2, the automatic driving control ECU3 also includes a CPU, memory, auxiliary storage device, and interface. The automatic driving control ECU3 stores various programs related to automatic driving control and takeover decision programs in its auxiliary storage device. The DCM1, like the multimedia ECU2, also includes a CPU, memory, auxiliary storage device, and interface. The DCM1 also includes a wireless communication unit. This wireless communication unit is a wireless communication circuit based on mobile communication methods such as 5G (5th Generation), 6G, 4G, and LTE (Long Term Evolution), as well as wireless communication methods such as WiMAX and WiFi. The wireless communication unit connects to the network N1 via wireless communication, thereby enabling communication with the central server 50.
[0057] Figure 3 This diagram illustrates an example of the functional structure of vehicle 10 and central server 50. Vehicle 10, as a functional structure, includes a communication unit 11, a control unit 21, a natural language processing unit 22, a mapping table storage unit 23, an automatic driving control unit 31, and a takeover decision unit 32. The communication unit 11 is a functional structural element equivalent to DCM1. The communication unit 11 serves as the interface for communication with external servers.
[0058] The automatic driving control unit 31 and the takeover judgment unit 32 are functional structural elements equivalent to the automatic driving control ECU 3. The processing of the automatic driving control unit 31 and the takeover judgment unit 32 is achieved by the CPU of the automatic driving control ECU 3 executing a predetermined program. The automatic driving control unit 31 implements automatic driving control of the vehicle 10. Automatic driving control includes, for example, engine or motor control, braking control, steering control, position management, and obstacle detection.
[0059] During the period when the vehicle 10 is driving in autonomous driving mode, the takeover judgment unit 32 determines whether autonomous driving can continue at predetermined intervals based on the detection values obtained by various sensors 6. For the vehicle 10 to continue autonomous driving, it needs to accurately identify the surrounding environment of the vehicle 10. For example, in adverse weather conditions, poor road conditions, or traffic congestion, the surrounding environment of the vehicle 10 may not be accurately identified by the sensors 6. In such cases, the takeover judgment unit 32 determines that it is difficult to continue autonomous driving. In addition, the conditions for the takeover judgment unit 32 to determine whether autonomous driving can continue depend on the structure of the autonomous driving control of the vehicle 10, and are not limited to specific conditions. Furthermore, the specific logic of the takeover judgment unit 32 in determining the cause of the takeover request is not limited to a specific method, and can be any of the following: a method according to predetermined rules, or logic using a machine learning model.
[0060] If the takeover determination unit 32 determines that continuing autonomous driving is difficult, it outputs a takeover request signal to the control unit 21. Furthermore, if the takeover determination unit 32 determines that continuing autonomous driving is difficult, it will send a takeover request generation notification and an intent number indicating the reason for the takeover request to the central server 50 via the communication unit 11. The intent number can be obtained, for example, by referring to the intent number table 32p described later. Alternatively, the takeover determination unit 32 may output the intent number along with the takeover request signal to the control unit 21.
[0061] The control unit 21, the natural language processing unit 22, and the correspondence table storage unit 23 are functional structural elements corresponding to the multimedia ECU 2. The control unit 21 implements control to guide the takeover. The control unit 21 receives a takeover request signal and an intent number from the takeover judgment unit 32. When the control unit 21 receives the takeover request signal, it outputs a voice prompting a switch to manual driving from the speaker 5. The voice data prompting the switch to manual driving is stored in the ultra-high-speed buffer memory 201M, for example, to reduce response latency. The processing of outputting the voice data prompting the switch to manual driving is also called the output of the takeover request.
[0062] Furthermore, when the control unit 21 receives a takeover request signal, it downloads a mapping table set corresponding to the intent number that caused the takeover request from the central server 50 via the communication unit 11, and stores the mapping table set in the mapping table storage unit 23. The mapping table set is a collection of mapping tables that establishes a correspondence between the driver's voice content envisioned in a dialogue explaining the reason for the takeover request and the response content to that driver's voice content. The mapping table has a number corresponding to the depth of the envisioned dialogue. The depth of the dialogue refers to the number of groups of voices and responses generated on a single topic, where voices and responses are grouped together. Details about the mapping table will be described later. The mapping table storage unit 23 corresponds to the ultra-high-speed cache memory 201M within the multimedia ECU 2.
[0063] After outputting a voice prompt to switch to manual driving, the control unit 21 begins dialogue processing, providing the driver with an explanation of the reasons for the takeover request in a conversational format. During this dialogue processing, the control unit 21 acquires the driver's spoken content, acquires response data that corresponds to the driver's spoken content, and outputs the response data via voice.
[0064] The driver's voice content can be obtained, for example, by recording the driver's voice through microphone 4, and then the control unit 21 performs speech recognition processing on the driver's voice data. The driver's voice content obtained as the speech recognition result of the driver's voice data can also be obtained, for example, as text data.
[0065] The response data representing the reply to the driver's spoken content is obtained, for example, by outputting the obtained spoken content data to the natural language processing unit 22 via the control unit 21 and receiving the response data representing the reply to the spoken content from the natural language processing unit 22. The response data to the driver's spoken content may also be obtained as text data, for example.
[0066] The control unit 21 generates speech data through speech synthesis based on the response data to the spoken content, and outputs the speech data to the speaker 5. Through the speaker 5, the speech data is output as speech.
[0067] The control unit 21 repeatedly performs dialogue processing until the driver's voice indicates acceptance of switching to manual driving, or until it detects that the driver has begun to engage manual driving. The driver's voice indicating acceptance of switching to manual driving may include phrases such as "Understood" or "Will drive." The correspondence table set described later contains a table showing the driver's expected voice when accepting a switch to manual driving, and the corresponding response to that voice. The control unit 21 detects that the driver's voice indicates acceptance of switching to manual driving, for example, by detecting that the driver's voice matches or is similar to the driver's voice in this correspondence table.
[0068] While initiating dialogue processing, the control unit 21 monitors the driver's actions, for example, using sensors that monitor the interior of the vehicle 10. Thus, the control unit 21 detects when the driver has begun manual driving. For example, it detects when the driver grips the steering wheel and when the driver's gaze is directed towards the front of the vehicle 10. Furthermore, the method for detecting when the driver has begun manual driving is not limited to a specific method; any known method may be used.
[0069] If a predetermined time has elapsed since the start of the dialogue process but no takeover has been implemented, the control unit 21 performs actions such as outputting a voice prompt to switch to manual driving, issuing a warning sound, or tightening the seat belt to request the driver to take over.
[0070] The natural language processing unit 22 retrieves the response data from the driver's voice input from the control unit 21 by searching the corresponding table group stored in the corresponding table storage unit 23, and outputs the response data to the control unit 21.
[0071] Next, the central server 50 includes a control unit 51 and a dialogue database 52 as part of its functional structure. These functional elements are achieved by the CPU of the central server 50 executing a predetermined program. The control unit 51 receives a notification of a takeover request from the vehicle 10. Along with the notification of the takeover request, it also receives an intent number. When the control unit 51 receives the notification of the takeover request, it specifies a correspondence table group corresponding to the received intent number and sends the correspondence table group to the vehicle 10. The dialogue database 52 is created, for example, in the storage area of the auxiliary storage device of the central server 50. The dialogue database 52 maintains a correspondence table group for each intent number.
[0072] In addition, in the first embodiment, the central server 50 pre-stores the corresponding table groups in the dialogue database 52. However, it is not limited to this. For example, the central server 50 may be equipped with a machine learning model instead of the dialogue database 52, and the machine learning model may be used to create the corresponding table groups. Specifically, when the control unit 51 receives a notification of the generation of a takeover request from the vehicle 10, it may use the machine learning model to create a corresponding table group corresponding to the received intent number, and send the corresponding table group to the vehicle 10.
[0073] Figure 4 This is an example of an intent number table 32p. The intent number table 32p is maintained in the auxiliary storage device of the automatic driving control ECU 3. The intent number table 32p maintains the assignment of intent numbers to the reasons for the takeover request.
[0074] exist Figure 4 In the examples shown, intent number 1 was assigned when the takeover request was triggered by difficulties in receiving GNSS signals. Intent number 2 was assigned when the takeover request was triggered by heavy rain. Intent number 3 was assigned when the takeover request was triggered by snowfall. Intent number 4 was assigned when the takeover request was triggered by a speed exceeding a threshold. Intent number 5 was assigned when the takeover request was triggered by difficulties in identifying the centerline. Additionally, Figure 4 The intent number assignment shown is an example; the intent number assignment for the reason for the takeover request can also be arbitrarily set by the administrator of the takeover guidance system 100.
[0075] <Regarding the correspondence table> Figure 5 This is a diagram illustrating an example of a dialogue script generated to explain the reasons for the takeover request. Figure 5 The example shown illustrates a dialogue script for a situation where the takeover request arises because GNSS signal reception is difficult.
[0076] When a takeover request is generated, the system first displays the voice prompt CV101, "Please switch to manual driving," indicating a message prompting a switch to manual control. It is envisioned that the system will also ask the driver the reason for the takeover request if the driver wishes to know why. Figure 5In the example shown, the question "Why?" is presented as an example of an inquiry into the reason for the takeover request. As a response to the inquiry, the voice CV102, stating the reason for the takeover request, is output: "Because the GNSS signal cannot be received successfully." In the first embodiment, a dialogue is constituted by the voice and the response to it. Furthermore, it is assumed that the dialogue depth increases for each group of dialogues. Hereinafter, the dialogue depth will be referred to as the dialogue level. Figure 5 In the example shown, dialogue level 1 is constituted by the driver’s utterance of “Why?” and the voice CV102 that responds to the utterance.
[0077] In cases where the driver requests more detailed information in response to the voice prompt CV102 stating "GNSS signal cannot be received successfully" explaining the reason for the takeover request, it is conceivable that derivative voice prompts may be generated to inquire about the GNSS signal or to inquire about the reason why the GNSS signal cannot be received.
[0078] exist Figure 5 In the example shown, a scenario is envisioned where, after the voice CV102 explaining the reason for the takeover request, the driver makes a verbal inquiry regarding the GNSS signal. Figure 5 In the example shown, the question "What is GNSS?" is displayed as an inquiry into GNSS signals. In response to this inquiry, a voice prompt CV103 explaining GNSS is output. Figure 5 In the example shown, the dialogue level is increased by 1 by the combination of a voice uttering an inquiry into the GNSS signal and a voice CV103 responding to that inquiry, thus becoming dialogue level 2.
[0079] exist Figure 5 In the example shown, a scenario is envisioned where, after the GNSS explanation in voice CV103, the driver makes a call to inquire why the GNSS signal cannot be received. Figure 5 In the example shown, the voice prompt asking why the GNSS signal cannot be received displays the question, "Why can't I receive it?". In response to this question, a voice prompt CV104 explaining the reason for the inability to receive the GNSS signal is output. Figure 5 In the example shown, the dialogue level is further increased by 1 by the voice asking why the GNSS signal cannot be received and the voice CV104 responding to the voice, thus becoming dialogue level 3.
[0080] exist Figure 5In the example shown, it is envisioned that after the voice CV104 explaining the reason for the inability to receive GNSS signals, the driver makes a voice indicating acceptance of switching to manual driving. Figure 5 In the example shown, the voice indicating acceptance of the switch to manual driving is "Understood." In response to this voice, the output includes a voice message CV105 confirming that the switch to manual driving has been accepted. Figure 5 In the example shown, the dialogue level is further increased by 1 by the voice CV105 indicating acceptance of the switch to manual driving and the response to this voice, thus becoming dialogue level 4.
[0081] exist Figure 5 In the example dialogue script shown, since dialogue levels range from 1 to 4, a mapping table corresponding to each dialogue level has been prepared. However, it is not necessary to follow this table. Figure 5 The dialogue will proceed in the order shown. Figure 5 The dialogue script shown also envisions a scenario where, after the output of voice CV101 indicating a switch to manual driving, voice CV102 responding to dialogue level 1, and voice CV103 responding to dialogue level 2, a voice indicating acceptance of the switch to manual driving is given, such as "Understood." If the voice indicating acceptance of the switch to manual driving is given, voice CV105 responding to dialogue level 4 is output.
[0082] Furthermore, for example, it is also possible to consider a scenario where, after the output of the voice CV102 in response to dialogue level 1, a voice prompt asking "Why can't I receive it?" is made to explain why the GNSS signal cannot be received. In this case, the voice CV103 in response to dialogue level 3 is output.
[0083] Figures 6 to 9 They are respectively, and Figure 5 The dialogue script shown is an example of an example of the correspondence tables contained in the correspondence table group when the takeover request is generated due to the difficulty in receiving GNSS signals. Figure 6 This is an example of a dialogue level 1 correspondence table for situations where the reason for a takeover request is the difficulty in receiving GNSS signals. If, after outputting a voice prompting a switch to manual driving, the driver questions the reason for the takeover request, it is first considered an inquiry into the reason for the takeover request. Therefore, in the first embodiment, regardless of the reason for the takeover request, the response in the dialogue level 1 correspondence table becomes the content indicating the reason for the takeover request.
[0084] exist Figure 6In the example shown, a correspondence is established between the imagined driver's voice when asked why a takeover request was made, and the response indicating that the reason for the takeover request was the difficulty in receiving GNSS signals. Figure 6 In the example shown, the driver's voice is imagined as a response to inquiring about the reason for the takeover request, and includes phrases such as "Why?", "Please tell me the reason", and "Why?". Figure 6 In the example shown, the message indicating that the takeover request was generated because GNSS signal reception was difficult is set to "because GNSS signal cannot be received successfully".
[0085] Figure 7 This is an example of a Level 2 correspondence table for a takeover request arising from difficulties in receiving GNSS signals. Correspondence tables for Level 2 and later correspond to inquiries derived from Level 1 responses and their corresponding answers.
[0086] exist Figure 7 The correspondence table for dialogue level 2, as shown, establishes a correspondence between the imagined driver's voice when questioned about GNSS signals and the corresponding explanation of the GNSS signals in response. Figure 7 In the example shown, the driver's voice is imagined to be responding to a question about GNSS signals, such as "What is GNSS?", "What is GNSS?", or "Please tell me what GNSS means." Figure 7 In the example shown, the message explaining the GNSS signal states, "GNSS refers to an artificial satellite system. It can accurately determine the latitude and longitude of your current location."
[0087] Figure 8 An example of a dialogue level 3 correspondence table for situations where a takeover request arises due to difficulties in receiving GNSS signals. Figure 8 The corresponding table for dialogue level 3 shows a correspondence between the imagined driver's voice when inquiring about the reason for not receiving GNSS signals and the message explaining the reason for not receiving GNSS signals in response. Figure 8 In the example shown, the driver's expected speech when inquiring about the reason for not receiving GNSS signals includes phrases such as "Why can't I receive it?", "Why can't I receive it?", and "What is the reason for not receiving it?". Figure 8In the example shown, the message explaining the reason for the inability to receive GNSS signals is set as "The received signal level is too weak. Your receiving device is working normally."
[0088] Figure 9 An example of a dialogue level 4 correspondence table for situations where a takeover request arises due to difficulties in receiving GNSS signals. Figure 9 The mapping table for dialogue level 4 shows a correspondence between the imagined driver's voice when accepting a switch to manual driving and the message confirming that the switch to manual driving has been accepted. Figure 9 In the example shown, the driver's voice is imagined to indicate a switch to manual driving, and phrases such as "Understood," "I'll drive," and "Got it" are provided. Figure 9 In the example shown, a message confirming the acceptance of the switch to manual driving is set as "Thank you. Please drive safely." If the takeover request is generated because GNSS signal reception is difficult, and if the driver's voice content is consistent with or similar to the voice content included in the correspondence table of dialogue level 4, and a response is given using the response content included in the correspondence table of dialogue level 4, then the control unit 21 determines that the dialogue processing has ended.
[0089] For example, the takeover request is generated because GNSS signal reception is difficult, and the corresponding table group contains... Figures 6 to 9 In the case of a mapping table, during dialogue processing, the response to the driver's spoken content is obtained in the following manner. The Natural Language Processing Unit 22 searches the mapping tables for at least each of dialogue levels 1 and 4 upon the start of dialogue processing and upon the initial input of the driver's spoken content. Furthermore, if a response is obtained to the input driver's spoken content, it is counted as one dialogue session. If no response is obtained to the input driver's spoken content, it is considered an error state and is not counted as one dialogue session.
[0090] When the driver inputs speech content a second or subsequent time, the natural language processing unit 22 can exclude the previously used response tables and refer to the remaining tables to obtain the response content for the input driver speech content. For example, if the driver's speech content initially matches the speech content included in the dialogue level 1 table and a response was given using the response content included in the dialogue level 1 table, then when the driver inputs speech content a second time, the tables for dialogue levels 2 to 4 will be referred to. If a response is given using the response content from the dialogue level 4 table, then the dialogue processing ends.
[0091] Furthermore, the driver's vocalizations in each correspondence table can be obtained either from past actual data or set by the administrator of the takeover guidance system 100. Moreover, it is not necessary for the actual driver's vocalizations to be completely identical to those in the correspondence table. Therefore, in the first embodiment, the natural language processing unit 22, in addition to obtaining the response content from the correspondence table under similar circumstances, will also use it as a response to the actual driver's vocalizations, except in cases where the actual driver's vocalizations are completely identical to those in the correspondence table.
[0092] Furthermore, the reason for the takeover request is that, in situations where GNSS signal reception is difficult, the correspondence table included in the correspondence table group is not necessarily limited to dialogue levels 1 to 4, but can be appropriately set according to the implementation method. Additionally, the reason for the takeover request is that, in situations where GNSS signal reception is difficult, the correspondence table for each dialogue level is not limited to... Figures 6 to 9 The corresponding table is shown below.
[0093] The corresponding tables are prepared in the dialogue database 52 of the central server 50 based on the reason for the takeover request. The maximum value of the dialogue level varies depending on the reason for each takeover request. However, for any given reason for a takeover request, the dialogue level 1 corresponding table represents the imagined voice of the driver when inquiring about the reason for the takeover request, and the corresponding message indicating the reason for the takeover request in response. Furthermore, for any given reason for a takeover request, the corresponding table at the maximum dialogue level represents the imagined voice of the driver when accepting manual driving, and the corresponding message confirming that manual driving has been accepted in response. The response content included in each corresponding table corresponds to "a portion of the explanation related to the reason for the request to switch to manual driving."
[0094] <Processing flow> Figure 10 This is an example of a flowchart illustrating the guidance process for taking over the vehicle 10 according to the first embodiment. Figure 10 The processing shown is repeatedly executed at predetermined cycles while vehicle 10 is driving in autonomous driving mode. Although Figure 10 The processing shown is performed by the automatic driving control ECU3, but for ease of understanding, the explanation will focus on its functional structural elements.
[0095] In OP101, the control unit 21 determines whether a takeover request has been generated. If a takeover request signal is received from the takeover determination unit 32, the control unit 21 detects that a takeover request has been generated. If a takeover request has been generated (OP101: Yes), the process proceeds to OP102. If no takeover request has been generated (OP101: No). Figure 10 The processing shown is now complete.
[0096] In OP102, the control unit 21 outputs a takeover request. Outputting a takeover request means outputting a message that prompts a switch to manual driving. In OP103, the control unit 21 begins downloading a mapping table from the central server 50 containing intent numbers corresponding to the reasons for the takeover request. The downloaded mapping table is stored in the mapping table storage unit 23.
[0097] The processing from OP104 to OP108 is equivalent to dialogue processing. In OP104, the control unit 21 determines whether the driver's voice is input from the microphone 4. If the driver's voice is input from the microphone 4 (OP104: Yes), the processing proceeds to OP105. If the driver's voice is not input (OP104: No), the processing proceeds to OP108.
[0098] In OP105, the control unit 21 performs speech recognition on the input spoken voice data to obtain the spoken content. In OP106, the control unit 21 outputs the driver's spoken content to the natural language processing unit 22 and obtains response data corresponding to the driver's spoken content from the natural language processing unit 22. The control unit 21 generates speech data through speech synthesis based on the response data and outputs the speech corresponding to the speech data from the speaker 5. The natural language processing unit 22 uses the driver's spoken content to search the correspondence table stored in the correspondence table storage unit 23 and outputs the response data contained in the correspondence table that matches or is similar to the driver's spoken content to the control unit 21.
[0099] In OP107, the control unit 21 determines whether the response output in OP106 is a response obtained from the table with the highest dialogue level in the mapping table group corresponding to the intent number corresponding to the cause of the takeover request. If the response output in OP106 is a response obtained from the table with the highest dialogue level in the mapping table group corresponding to the intent number corresponding to the cause of the takeover request (OP107: Yes), the dialogue processing ends, and the process proceeds to OP109. If the response output in OP106 is a response obtained from a mapping table other than the table with the highest dialogue level in the mapping table group corresponding to the intent number corresponding to the cause of the takeover request (OP107: No), the process proceeds to OP104.
[0100] In OP108, the control unit 21 determines whether the driver has begun manual driving. This is determined, for example, by detecting whether the driver is holding the steering wheel or looking directly ahead of the vehicle 10, based on images captured by a camera filming the interior of the vehicle 10. If manual driving is detected (OP108: Yes), the dialogue process ends, and the process proceeds to OP109. If manual driving is not detected (OP108: No), the process proceeds to OP104.
[0101] In OP109, the control unit 21 deletes the corresponding table group stored in the corresponding table storage unit 23. Afterwards, Figure 10 The processing shown is now complete. Furthermore, the processing of vehicle 10 is not limited to... Figure 10 The processing is shown. For example, in Figure 10 In this process, dialogue processing can be implemented until the driver's voice indicates acceptance of switching to manual driving (OP107) or until manual driving is detected to have started (OP108). However, it is not limited to these; either the driver's voice indicating acceptance of switching to manual driving or the start of manual driving can be set as the termination condition for dialogue processing.
[0102] Figure 11 An example of a flowchart of the boot process for the takeover of the central server 50 according to the first embodiment. Figure 11 The process shown is repeated at predetermined intervals. Although Figure 11 The processing shown is executed by the CPU of the central server 50, but for ease of understanding, the explanation will focus on the functional structural elements.
[0103] In OP201, the control unit 51 determines whether a takeover request generation notification has been received from vehicle 10. An intent number is also received along with the takeover request generation notification. If a takeover request generation notification has been received from vehicle 10 (OP201: Yes), the process proceeds to OP202. If no takeover request generation notification has been received from vehicle 10 (OP201: No). Figure 11 The processing shown is now complete.
[0104] In OP202, the control unit 51 reads the corresponding lookup table from the dialogue database 52, which corresponds to the intent number received from the vehicle 10, and sends the lookup table to the vehicle 10. Afterwards, Figure 11 The processing shown is now complete.
[0105] <Download the Correspondence Table> Figure 12 as well as Figure 13 An example of a sequence diagram for downloading and processing dialogues for a corresponding table group. Figure 12 as well as Figure 13 The reason for the takeover request was that GNSS signal reception was difficult and that data was downloaded from the central server 50. Figures 6 to 9 An example of the correspondence between dialogue levels 1 to 4 is shown.
[0106] exist Figure 12 In the example shown, the download of the correspondence table group from the central server 50 is implemented on a table-by-table basis. In S11, it is determined that continuing autonomous driving in vehicle 10 is difficult (takeover determination). In S12, vehicle 10 sends a notification of a takeover request to the central server 50, along with intent number 1 indicating that the reason for the takeover request is the difficulty in receiving GNSS signals (e.g., refer to...). Figure 4 In S13, in vehicle 10, a takeover request signal is output from the automatic driving control ECU3 to the multimedia ECU2. Figure 10 OP101). In S14, in vehicle 10, a voice message prompting a switch to manual driving is output. Figure 12 The middle part reads "Please switch to manual driving". Figure 10 OP102).
[0107] In S21, during the voice output process prompting a switch to manual driving, vehicle 10 downloads a mapping table from central server 50 for dialogue levels 1 and 4 corresponding to intent number 1. The mapping table for dialogue level 4 is the mapping table with the highest dialogue level in the mapping table group corresponding to intent number 1. In S22, the driver makes a sound inquiring about the reason for the takeover request. Figure 12The question mark in the middle is "Why?". Since the mapping table for dialogue levels 1 and 4 has been downloaded at time S22, in S23, vehicle 10 will output the response content contained in the mapping table for dialogue level 1 via voice (see...). Figure 6 , Figure 12 The middle part means "because the GNSS signal could not be received successfully".
[0108] In S31, after the voice prompting a switch to manual driving is output, and while waiting for or recognizing the driver's voice, vehicle 10 downloads a mapping table of dialogue level 2 corresponding to intent number 1 from central server 50. In S32, the driver makes a voice prompt to inquire about GNSS signals. Figure 12 The question in the middle is "What is GNSS?". Since the mapping table for dialogue level 2 has been downloaded at time S32, in S33, vehicle 10 outputs the response content contained in the mapping table for dialogue level 2 via voice (see...). Figure 7 , Figure 12 The text in the middle reads "GNSS refers to artificial satellite systems (hereinafter omitted)".
[0109] In S41, during the output of the voice-based response content included in the dialogue level 1 correspondence table in S23, vehicle 10 downloads the dialogue level 3 correspondence table corresponding to intent number 1 from central server 50. In S42, the driver makes a verbal inquiry regarding the reason for the inability to receive GNSS signals. Figure 12 The text in question is "Why can't I receive it?". Since the mapping table for dialogue level 3 was downloaded at time S42, in S43, vehicle 10 outputs the response content contained in the mapping table for dialogue level 3 via voice (see...). Figure 8 , Figure 12 The message in the middle is "The received signal level is too weak... (hereinafter omitted)". After that, vehicle 10 maintains the correspondence table group corresponding to intent number 1 and repeats the same process until the response content contained in the correspondence table of dialogue level 4 is output as a reply, or until it is detected that manual driving has begun.
[0110] exist Figure 12 In the example shown, by downloading a table of correspondences that are more likely to be spoken next, the delay in responding to the driver's speech can be shortened.
[0111] exist Figure 13 In the example shown, the download from the corresponding table group of the central server 50 is performed simultaneously for all corresponding tables contained in the corresponding table group corresponding to the reason for generating the takeover request. S11 to S14 and Figure 12 Same. Figure 13 In S51, during the voice output of the message prompting a switch to manual driving, vehicle 10 downloads all the correspondence tables contained in the correspondence table group corresponding to intent number 1 from central server 50.
[0112] Subsequently, in S52, S61, and S71, even though a questioning voice is issued, since vehicle 10 maintains all the correspondence tables in the correspondence table group corresponding to intention number 1, a response can be issued with a shorter delay, similar to S53, S62, and S72. Furthermore, the driver's voice content in S52, S61, and S71 is respectively... Figure 12 The responses to the driver's statements in S22, S32, and S42 are the same. Figure 12 The same applies to S23, S33, and S43. By downloading the corresponding table sets together, the impact of network latency can be further reduced, thereby enabling faster responses to the driver's voice.
[0113] Whether to download in units of a corresponding table or download all at once, and the corresponding corresponding table group for the reason for the takeover request, can be arbitrarily set by the administrator of the takeover boot system 100.
[0114] <Effects of the First Embodiment> In the first embodiment, when a takeover request is made, a portion of the explanation related to the reason for the takeover request is presented to the driver based on the content of the driver's speech. This allows for a more appropriate presentation of the explanation based on the driver's level of concern regarding the reason for the takeover request, thereby reducing any unpleasantness to the driver.
[0115] Furthermore, in the first embodiment, the vehicle 10 downloads a corresponding table set, along with the reason for the takeover request, from the central server 50 before the driver speaks, and stores this table set in the ultra-high-speed cache memory 201M. This allows for a more rapid response to the driver's voice.
[0116] <Second Implementation> In the first embodiment, vehicle 10 acquires response data to the driver's verbal input. Therefore, in the first embodiment, vehicle 10 downloads a corresponding table set from central server 50 before the driver's verbal input and stores the corresponding table set in ultra-high-speed cache memory 201M.
[0117] Instead of this approach, in the second embodiment, the central server acquires the response data that corresponds to the driver's verbal input. Therefore, in the second embodiment, the vehicle does not download the corresponding table set from the central server that corresponds to the reason for the takeover request. Furthermore, in the second embodiment, the descriptions common to the first embodiment are omitted.
[0118] Figure 14 This diagram illustrates an example of the functional structure of the vehicle 10B and the central server 50B according to the second embodiment. In the second embodiment, the system structure of the takeover guidance system 100, as well as the hardware structure of the vehicle 10B and the central server 50B, are the same as in the first embodiment. In the second embodiment, the vehicle 10B, as a functional structure, includes a communication unit 11, a control unit 21B, an automatic driving control unit 31, and a takeover determination unit 32. The communication unit 11, the automatic driving control unit 31, and the takeover determination unit 32 are the same as in the first embodiment.
[0119] The control unit 21B is a functional structural element equivalent to the multimedia ECU 2. When the control unit 21B receives a takeover request signal, it outputs a voice prompt from the speaker 5 to switch to manual driving and begins monitoring the voice input via the microphone 4. When the driver's voice is input from the microphone 4, the control unit 21B performs voice recognition processing on the voice data to obtain the driver's voice content. The control unit 21B sends the driver's voice content data to the central server 50B via the communication unit 11. Subsequently, when the control unit 21B receives response data from the central server 50B via the communication unit 11, it generates voice data through speech synthesis based on the response data to the driver's voice content and outputs this voice data to the speaker 5. Through the speaker 5, this voice data is output as speech. The driver's voice content data sent to the central server 50B is, for example, text data.
[0120] Furthermore, when the control unit 21B receives a takeover request signal, it begins monitoring the driver's actions, for example, using sensors that monitor the interior of the vehicle 10B. When the control unit 21B detects that the driver has begun manual driving, it sends a manual driving start notification to the central server 50B. The control unit 21B performs processing to acquire the driver's spoken content and monitors the driver's actions until it receives a dialogue end notification from the central server 50B. Aside from these points, the processing of the control unit 21B is the same as that of the control unit 21 in the first embodiment.
[0121] Furthermore, in the second embodiment, the central server 50B, as a functional structure, includes a control unit 51B, a dialogue database 52, and a natural language processing unit 53. When the control unit 51B receives data containing the driver's spoken content from the vehicle 10B, it outputs this data to the natural language processing unit 53 and obtains response data to the driver's spoken content from the natural language processing unit 53. The control unit 51B then sends the obtained response data back to the vehicle 10B. The response data can be, for example, text data or voice data in a predetermined format.
[0122] The natural language processing unit 53 retrieves the response data from the driver's speech input from the control unit 51B by searching a corresponding table stored in the natural language processing unit 53 for each intention number, thereby obtaining the response data and outputting it to the control unit 51B. Furthermore, the corresponding table is the same as in the first embodiment.
[0123] When the driver's voice indicates acceptance of switching to manual driving, that is, when the response to the driver's voice is obtained from the correspondence table with the maximum dialogue level in the correspondence table group, and when the manual driving start notification is received from the vehicle 10B, the control unit 51B sends a dialogue end notification to the vehicle 10B.
[0124] Figure 15 An example of a flowchart illustrating the guidance process for takeover of vehicle 10B according to the second embodiment. Figure 15 The processing shown is repeatedly executed at predetermined intervals during the operation of vehicle 10B in autonomous driving mode.
[0125] In OP301, control unit 21B determines whether a takeover request has been generated. If a takeover request has been generated (OP301: Yes), the process proceeds to OP302. If no takeover request has been generated (OP301: No). Figure 15 The processing shown is now complete.
[0126] In OP302, control unit 21B outputs a takeover request. In OP303, control unit 21B determines whether a driver's voice is input from microphone 4. If a driver's voice is input from microphone 4 (OP303: Yes), the process proceeds to OP304. If no driver's voice is input (OP303: No), the process proceeds to OP308.
[0127] In OP304, control unit 21B performs speech recognition on the input spoken voice data to obtain the spoken content data. In OP305, control unit 21B sends the spoken content data to central server 50B. In OP306, control unit 21B determines whether response data has been received from central server 50B. If response data has been received from central server 50B (OP306: Yes), processing proceeds to OP307. The period until response data is received from central server 50B (OP306: No) is a standby state; if no response data is received even after a predetermined time, an error state is entered.
[0128] In OP307, the control unit 21B generates speech data through speech synthesis based on the response data, and outputs the speech corresponding to the speech data from the speaker 5.
[0129] In OP308, control unit 21B determines whether the driver has started manual driving. If manual driving is detected (OP308: Yes), the process proceeds to OP309. In OP309, control unit 21B sends a manual driving start notification to central server 50B. If manual driving is not detected (OP308: No), the process proceeds to OP303.
[0130] In OP310, control unit 21B determines whether a conversation end notification has been received from central server 50B. If a conversation end notification has been received from central server 50B (OP310: Yes). Figure 15 The process shown has ended. If no session end notification is received from the central server 50B (OP310: No), proceed to OP303.
[0131] Figure 16 An example of a flowchart of the boot process for the takeover of the central server 50B according to the second embodiment. Figure 16 The process shown is repeated at predetermined intervals. Although Figure 16 The processing shown is executed by the CPU of the central server 50B, but for ease of understanding, the explanation will focus on the functional structural elements.
[0132] In OP401, control unit 51B determines whether a takeover request generation notification has been received from vehicle 10B. An intent number is also received along with the takeover request generation notification. If a takeover request generation notification has been received from vehicle 10B (OP401: Yes), the process proceeds to OP402. If no takeover request generation notification has been received from vehicle 10B (OP401: No). Figure 16 The processing shown is now complete.
[0133] In OP402, the control unit 51B determines whether it has received data on the driver's voice from the vehicle 10B. If the driver's voice data has been received from the vehicle 10B (OP402: Yes), the process proceeds to OP403. If the driver's voice data has not been received from the vehicle 10B (OP402: No), the process proceeds to OP406.
[0134] In OP403, the control unit 51B outputs data of the driver's spoken content to the natural language processing unit 53 and retrieves response data corresponding to the driver's spoken content from the natural language processing unit 53. The natural language processing unit 53 uses the driver's spoken content to search a lookup table stored in the natural language processing unit 53 that corresponds to the received intent number, and outputs the response data contained in the lookup table that matches or is similar to the driver's spoken content to the control unit 51B. In OP404, the control unit 51B sends the response data to the vehicle 10B.
[0135] In OP405, control unit 51B determines whether the content of the response data obtained in OP403 is obtained from the correspondence table with the maximum dialogue level in the correspondence table group of the received intent number. If the content of the response data is obtained from the correspondence table with the maximum dialogue level in the correspondence table group of the received intent number (OP405: Yes), the process proceeds to OP407. If the content of the response data is obtained from a correspondence table other than the correspondence table with the maximum dialogue level in the correspondence table group of the received intent number (OP405: No), the process proceeds to OP402.
[0136] In OP407, control unit 51B sends a dialogue termination notification to vehicle 10B. Afterwards, Figure 16 The processing shown is now complete. Additionally, Figure 15 as well as Figure 16 The processing shown is an example, and the processing of vehicle 10B and central server 50B involved in the second embodiment is not limited to this.
[0137] In the second embodiment, when a driver's voice is generated, the vehicle 10B sends the driver's voice content to the central server 50B and obtains response data to the driver's voice content from the central server 50B. Therefore, since the vehicle 10B does not need to download the corresponding table set from the central server 50B and store the corresponding table set in the ultra-high-speed cache memory 201M, the resources of the ultra-high-speed cache memory 201M can be saved.
[0138] <Other examples of changes> The above-described implementation is merely an example, and this disclosure may be implemented with appropriate modifications without departing from its spirit.
[0139] Although in the first and second embodiments, the dialogue processing is performed by the multimedia ECU2, it is also possible to replace this method by, for example, an in-vehicle unit such as a DCM1 or a car navigation system. In this case, the in-vehicle unit becomes an example of an "information processing device".
[0140] Although in the first and second embodiments, the explanation related to the reason for the takeover request is presented to the driver verbally through speaker 5, this is not a limitation. For example, the explanation related to the reason for the takeover request may also be presented to the driver in text form on a display screen inside the vehicle 10. The method of presenting the explanation related to the reason for the takeover request is not limited to a predetermined method.
[0141] The processing or methods described in this disclosure can be freely combined and implemented as long as they do not create technical contradictions.
[0142] Furthermore, processes described as being implemented by a single device can also be performed by multiple devices. Alternatively, processes described as being implemented by different devices can be performed by a single device. In a computer system, the hardware architecture (server architecture) used to implement each function can be flexibly changed.
[0143] This invention can also provide a computer with a computer program that has the functions described in the above embodiments, and enable one or more processors of the computer to read and execute the program. Such a computer program can be provided to the computer either through a non-transitory computer-readable storage medium that can be connected to the computer's system bus, or via a network. Non-transitory computer-readable storage media include, for example, any type of disk (Floopy, registered trademark), hard disk drive (HDD), optical disk (CD-ROM, DVD, Blu-ray disc, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, and any type of media suitable for storing electronic instructions.
[0144] Symbol Explanation 1…DCM; 2…Multimedia ECU; 3…Automatic driving control ECU; 4…microphone; 5…speakers; 10… vehicles; 11…Ministry of Communications; 21…Control Department; 22…Natural Language Processing Department; 23…corresponding table storage section; 31…Automatic driving control unit; 32…take over the judgment department; 50… central server; 51…Control Department; 52… Dialogue Database; 53…Natural Language Processing Department; 100…takes over the boot system; 201…CPU; 201M… Ultra-high-speed cache memory; 202…memory; 203… Auxiliary storage device.
Claims
1. An information processing device, wherein, The system includes a control unit that performs the following processing: If a request to switch to manual driving arises during the automatic driving control of the first vehicle, the driver of the first vehicle is provided with a portion of an explanation related to the reason for the request. When the request is generated, the control unit repeatedly executes a portion of the processing that provides a prompt and an explanation related to the reason for the request until it indicates to the driver that the driver accepts the switch to manual driving.
2. The information processing apparatus as described in claim 1, wherein, Upon receiving the request, the control unit further repeatedly executes the process of obtaining the driver's spoken content until the driver indicates acceptance of the switch to manual driving. The driver is prompted with a portion of an explanation relating to the reason for the request, which corresponds to the content of the driver's speech.
3. The information processing apparatus as described in claim 2, wherein, It also includes a storage unit that stores the correspondence between the first sound content and a portion of the explanation related to the reason for the request. The control unit performs the following processing: If the driver's voice is at least similar to the first voice, the driver is shown a portion of an explanation that corresponds to the first voice and relates to the reason for the request.
4. The information processing apparatus as described in claim 3, wherein, The information processing device is mounted on the first vehicle. The control unit further performs the following processing: When the request is made, the mapping between the driver's voice content and a portion of the explanation related to the reason for the request is downloaded from a predetermined device, and the mapping is stored in the storage unit.
5. The information processing apparatus as described in claim 4, wherein, The control unit performs the following processing: Further processing is performed to determine the cause of the aforementioned requirement. Download the corresponding relationship for the stated reason.
6. The information processing apparatus as described in claim 5, wherein, The control unit performs the following processing: The corresponding relationship for the cause is downloaded from the predetermined device.
7. The information processing apparatus according to any one of claims 3 to 6, wherein, The correspondence includes: The first correspondence establishes a correspondence between multiple second vocalizations conceived in the case of inquiring about the reason for the request and the reason for the request, as part of an explanation related to the reason for the request. At least one second correspondence, which will include multiple third statements, including those conceived as accepting the reasons that gave rise to the request and further generating inquiries, and responses to the inquiries, as part of a description related to the reasons that gave rise to the request, will establish a correspondence. The control unit performs the following processing: After the request is made, referring to the first correspondence, the driver is prompted with the reason for the request being generated, which corresponds to a second voice content similar to the driver's voice content. After informing the driver of the reason for the request, and referring to the at least one second correspondence, the driver is then prompted with an answer to the inquiry that corresponds to a third voice content similar to the driver's voice content.
8. The information processing apparatus as claimed in claim 2, wherein, The information processing device is mounted on the first vehicle. The control unit further performs the following processing: The driver's voice content is sent to a predetermined device. Receive from the predetermined device a portion of the explanation relating to the reason for the request, which is in relation to the content of the sound.
9. The information processing apparatus as claimed in claim 2, wherein, The control unit performs the following processing: When the request is made, after prompting the driver with the request, processing begins to obtain the driver's vocal content and prompt a portion of the explanation related to the reason for the request, corresponding to the vocal content.
10. An information processing method, comprising: The following processing is performed by the information processing device: If a request to switch to manual driving arises during the automatic driving control of the first vehicle, the driver of the first vehicle is provided with a portion of an explanation related to the reason for the request. When the request is generated, the information processing device repeatedly executes a portion of the explanation corresponding to the driver's vocalization and related to the reason for the request, until it indicates that the driver accepts the switch to manual driving.
11. The information processing method as described in claim 10, wherein, It also includes the following processing performed by the information processing device: If the aforementioned request is made, the process of obtaining the driver's spoken content is repeatedly executed until the driver indicates that they accept the switch to manual driving. The driver is prompted with a portion of an explanation relating to the reason for the request, which corresponds to the content of the driver's speech.
12. The information processing method as described in claim 11, wherein, The information processing device includes a storage unit that stores the correspondence between the first vocal content and a portion of the explanation related to the reason for the request. The information processing device performs the following processing: when the driver's voice content is at least similar to the first voice content, it prompts the driver with a portion of the explanation that corresponds to the first voice content and is related to the reason for the request.
13. The information processing method as described in claim 12, wherein, The information processing device is mounted on the first vehicle. The information processing device performs the following process: when the request is generated, it downloads from a predetermined device a correspondence between the driver's voice content and a portion of the explanation related to the reason for the request, and stores the correspondence in the storage unit.
14. The information processing method as described in claim 13, wherein, The information processing device performs the following processing: The reason for the aforementioned requirement was obtained. Download the corresponding relationship for the stated reason.
15. The information processing method as described in claim 14, wherein, The information processing device performs the following processing: The corresponding relationship for the cause is downloaded from the predetermined device.
16. The information processing method according to any one of claims 11 to 15, wherein, The correspondence includes: The first correspondence establishes a correspondence between multiple second vocalizations conceived in the case of inquiring about the reason for the request and the reason for the request, as part of an explanation related to the reason for the request. At least one second correspondence will include multiple third statements conceived as accepting the reasons that gave rise to the request and further generating inquiries, and responses to the inquiries, established as part of a description related to the reasons that gave rise to the request. The information processing device performs the following processing: After the request is generated, referring to the first correspondence, the driver is prompted with the reason for the generation of the request, which corresponds to a second voice content similar to the driver's voice content. After informing the driver of the reason for the request, and referring to the at least one second correspondence, the driver is then prompted with an answer to the inquiry that corresponds to a third voice content similar to the driver's voice content.
17. The information processing method as described in claim 11, wherein, The information processing device is mounted on the first vehicle. The information processing device performs the following processing: The driver's voice content is sent to a predetermined device. Receive from the predetermined device a portion of the explanation relating to the reason for the request, which is in relation to the content of the sound.
18. A vehicle, wherein, The information processing apparatus is provided with any one of claims 1 to 9.