Automated response apparatus, method and automated response computer program product
By integrating candidate identification and generation models into the vehicle and utilizing sensor information and occupant history, response processing can be quickly identified and executed, solving the problem of excessively long waiting time for LLM to generate vehicle occupant speech responses and achieving a faster response time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2025-12-24
- Publication Date
- 2026-06-26
AI Technical Summary
In existing technologies, there is a problem of excessively long waiting times when generating responses to vehicle occupants' speech based on large-scale language models (LLM).
By combining the candidate determination unit, identification unit, selection unit, and response processing unit, and utilizing the vehicle's sensor information and the passenger's speaking history, the system can quickly determine and execute response processing, including the use of candidate determination models and generation models, thereby shortening response waiting time.
It effectively shortens the waiting time from when a user speaks to when a response is executed, improving the real-time nature and efficiency of the response.
Smart Images

Figure CN122275787A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an automatic response device, an automatic response method, and a computer program product for automatically responding to statements made by occupants of a vehicle. Background Technology
[0002] A technique for controlling vehicle equipment via voice has been proposed (see Japanese Patent Application Publication No. 2006-137366). In this technique, the vehicle equipment control device converts input voice into text, investigates the correlation between the text and past operations by referring to operation history, estimates the device to be operated from multiple devices, and sets a priority for each device. Then, the vehicle equipment control device selects a device or a type of operation based on the priority and instructs the device to perform the operation corresponding to the text.
[0003] A large-scale language model (LLM) for automatically generating responses to questions was studied. However, due to the high computational cost of LLM, when using LLM to generate responses to statements from vehicle occupants, the waiting time from when the occupant speaks to when any response processing is executed can sometimes be too long, depending on the circumstances. Summary of the Invention
[0004] Therefore, the object of the present invention is to provide an automatic response device that can shorten the waiting time until a response is executed in accordance with a statement made by a passenger in a vehicle.
[0005] According to one embodiment, an automatic response device is provided. The automatic response device includes: a candidate determination unit that determines a plurality of candidate responses that are likely to be requested by a passenger based on at least one of past response history to speech made by a vehicle occupant, external information indicating the conditions surrounding the vehicle, and internal information indicating the conditions inside the vehicle; and for each of the plurality of candidate responses, determining a combination of hypothetical speech data representing the assumed speech content at the time of requesting the response and command data for executing the response; an identification unit that identifies the speech content of the passenger from voice signals collected by a microphone provided in the vehicle; a selection unit that selects a candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data most consistent with the identified speech content; and a response processing unit that executes a response following the command data regarding the selected candidate response.
[0006] In one implementation, the candidate determination unit determines multiple response candidates by inputting at least one of response history, external vehicle information, and internal vehicle information into a candidate determination model that has been pre-learned in order to determine the candidates for the response.
[0007] In one implementation, the selection unit calculates the degree of consistency between the assumed speech data of each of the multiple response candidates and the identified speech content. If the degree of consistency for any of the multiple response candidates is less than a predetermined selection threshold, none of the multiple response candidates are selected. If none of the multiple response candidates are selected, the response processing unit generates command data by inputting the identified speech content into a generative model that has been pre-learned to generate command data corresponding to the speech content, and executes a response that follows the generated command data.
[0008] According to another embodiment, an automatic response method is provided. The automatic response method includes: determining a plurality of candidate responses that are likely to be requested by the passenger, based on at least one of past response history to speech made by a occupant of a vehicle, external information indicating the conditions surrounding the vehicle, and internal information indicating the conditions inside the vehicle; generating, for each determined candidate response, a combination of hypothetical speech data representing assumed speech content at the time of requesting the response and command data for executing the response; identifying the passenger's speech content from voice signals collected via a microphone located in the vehicle; selecting a candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data most consistent with the identified speech content; and executing a response that follows the command data regarding the selected candidate response.
[0009] According to another embodiment, an automatic response computer program product is provided. The automatic response computer program product includes instructions for causing a computer to perform the following actions: determining a plurality of candidate responses that are likely to be requested by the passenger, based on at least one of past response history to speech made by a vehicle occupant, external information indicating the conditions surrounding the vehicle, and internal information indicating the conditions inside the vehicle; generating, for each determined candidate response, a combination of hypothetical speech data representing the assumed speech content at the time of requesting the response and command data for executing the response; identifying the passenger's speech content from voice signals collected via a microphone located in the vehicle; selecting a candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data most consistent with the identified speech content; and executing a response following the command data regarding the selected candidate response.
[0010] The automatic response device disclosed herein has the following effect: it can shorten the waiting time until a response is executed in accordance with the speech of the vehicle's occupants. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of a vehicle equipped with an automatic response system.
[0012] Figure 2 This is a hardware configuration diagram of an automatic response device.
[0013] Figure 3 This is a functional block diagram of the processor in an automatic response device.
[0014] Figure 4 This is a diagram illustrating an overview of the automatic response processing according to this embodiment.
[0015] Figure 5 This is a flowchart of the automatic response device's operation. Detailed Implementation
[0016] Hereinafter, with reference to the accompanying drawings, the automatic response device, the automatic response method executed by the automatic response device, and the computer program for automatic response will be described. The automatic response device predetermines a plurality of candidate responses that are likely to be requested by a vehicle occupant, and for each candidate response, determines a combination of hypothetical speech data representing the assumed speech content at the time of requesting the response and command data for executing the response. Furthermore, the automatic response device identifies the occupant's speech content from voice signals collected by a microphone located in the vehicle, and selects the candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data most consistent with the identified speech content. Then, the automatic response device shortens the waiting time from the occupant's speech to the execution of the response by executing a response that follows the command data regarding the selected candidate response.
[0017] Figure 1 This is a schematic diagram of a vehicle equipped with an automatic response device. In this embodiment, vehicle 1 includes at least one external sensor 2, at least one internal sensor 3, a microphone 4, a notification device 5, and an automatic response device 6. Each external sensor 2, each internal sensor 3, microphone 4, and notification device 5 is connected to the automatic response device 6 in a manner capable of communicating with each other. Furthermore, a wireless communication terminal (not shown) for wireless communication with other devices outside the vehicle 1 may also be installed in vehicle 1.
[0018] Each external sensor 2 is a sensor used to sense the conditions around the vehicle 1. Examples include external cameras that capture images of the area around the vehicle 1, or ranging sensors such as radar or LiDAR that measure the distance to objects around the vehicle 1. Each external sensor 2 may also include a thermometer that measures the temperature around the vehicle 1, a rain gauge that measures rainfall, or a positioning device that uses a satellite positioning system, such as a GPS (Global Positioning System) receiver, to determine the position of the vehicle 1. Each external sensor 2 generates external sensor signals representing the conditions around the vehicle 1 at predetermined intervals and outputs the generated external sensor signals to the automatic response device 6. It should be noted that the external sensor signals are just one example of external information representing the conditions around the vehicle 1.
[0019] Each in-vehicle sensor 3 is a sensor used to sense the conditions inside the passenger compartment of vehicle 1. Examples include an in-vehicle camera installed to capture images inside the passenger compartment of vehicle 1, or a thermometer to measure the temperature inside the vehicle. Each in-vehicle sensor 3 generates an in-vehicle sensor signal representing the conditions inside the passenger compartment of vehicle 1 at each predetermined cycle, and outputs the generated in-vehicle sensor signal to the automatic response device 6. It should be noted that the in-vehicle sensor signal is an example of in-vehicle information representing the conditions inside the passenger compartment of vehicle 1.
[0020] Microphone 4 is another example of an in-vehicle sensor, collecting the voice emitted by any passenger in vehicle 1 and outputting a voice signal representing that voice. Therefore, microphone 4 is installed inside the passenger compartment of vehicle 1. It should be noted that multiple microphones 4 can also be installed in vehicle 1. In this case, the multiple microphones 4 can be arranged in an array, or they can be installed around each seat in the passenger compartment of vehicle 1. Microphone 4 outputs the generated voice signal to the automatic response device 6. The voice signal generated by microphone 4 is another example of an in-vehicle sensor signal.
[0021] The notification device 5 is located inside the passenger compartment of vehicle 1 and notifies passengers of the response content represented by the response information generated by the automatic response device 6. Therefore, the notification device 5 includes at least one of a speaker or a display device. Then, when it receives a notification signal representing the response content to the passenger from the automatic response device 6, the notification device 5 notifies the driver of the response content via voice from the speaker or by displaying messages, images, or playing moving images on the display device.
[0022] The automatic response device 6 executes a response corresponding to the speech of any passenger from vehicle 1.
[0023] Figure 2 This is a hardware configuration diagram of the automatic response device 6. (For example...) Figure 2 As shown, the automatic response device 6 has a communication interface 21, a memory 22, and a processor 23. The communication interface 21, the memory 22, and the processor 23 can be configured as separate circuits, or they can be configured as a single integrated circuit.
[0024] The communication interface 21 has interface circuitry for connecting the automatic response device 6 to other devices within the vehicle. The communication interface 21 transmits external sensor signals received from the various external sensors 2 and internal sensor signals received from the various internal sensors 3 to the processor 23. Additionally, the communication interface 21 transmits voice signals received from the microphone 4 to the processor 23. Furthermore, the communication interface 21 transmits information received from external devices of the vehicle 1 via a wireless communication terminal (not shown) to the processor 23. Such information includes, for example, weather information or traffic information. Weather information and traffic information are another example of external information. Additionally, the communication interface 21 outputs notification signals received from the processor 23 to the notification device 5, or outputs control commands received from the processor 23 for any of the in-vehicle devices.
[0025] Memory 22 is an example of a storage unit, such as having volatile semiconductor memory and non-volatile semiconductor memory. Furthermore, memory 22 stores various data used in the automatic response processing executed by processor 23. Specifically, memory 22 stores response history representing responses executed in response to past statements made by occupants of vehicle 1, and various data for generating multiple response candidates. Regarding each response, the response history includes text data representing the content of the statement, the type of in-vehicle device associated with the executed response and the content of the response, and data representing in-vehicle and out-of-vehicle information at the time of execution of the response. Moreover, memory 22 may also temporarily store out-of-vehicle sensor signals received from each of the external sensors 2, in-vehicle sensor signals received from each of the internal sensors 3, and voice signals received from microphone 4. Additionally, memory 22 stores various data generated during the execution of the automatic response processing, such as command data and assumed statement data regarding each response candidate.
[0026] Processor 23 has one or more CPUs (Central Processing Units) and their peripheral circuitry. Processor 23 may also have other arithmetic circuitry such as logic units, numerical processing units, or graphics processing units. Furthermore, processor 23 performs automatic response processing.
[0027] Figure 3This is a functional block diagram of the processor 23 related to automatic response processing. The processor 23 has a candidate determination unit 31, an identification unit 32, a selection unit 33, and a response processing unit 34. These units of the processor 23 are, for example, functional modules implemented by a computer program that operates on the processor 23. Alternatively, these units of the processor 23 may also be dedicated arithmetic circuits provided on the processor 23.
[0028] The candidate determination unit 31 determines multiple candidate responses from the assumed response types that have the possibility of being requested by the occupants of vehicle 1.
[0029] It should be noted that the assumed response types include the operation of one or more in-vehicle devices, voice output or image display via notification device 5, information search via a wireless communication terminal network, or a combination thereof. Operations of in-vehicle devices may include, for example, changing the set temperature of the air conditioning unit, opening or closing any window, turning interior lights on or off, and playing or stopping video or audio content.
[0030] The candidate determination unit 31 determines candidate responses by referring to at least one of the response history, external vehicle information, and internal vehicle information. In this embodiment, the candidate determination unit 31 determines multiple candidate responses by inputting at least one of the response history, external vehicle information, and internal vehicle information into a candidate determination model that has been pre-learned based on a prescribed machine learning method to output multiple candidate responses. The candidate determination model may be, for example, a decision tree-based model, a deep neural network (DNN)-based model, or a boosting-based model. In the case where the candidate determination model is a DNN-based model, the candidate determination model is configured, for example, to have multiple fully coupled layers and an output layer sequentially from the input side. The output layer calculates a confidence value representing the type of each assumed response using a softmax operation. In this case, the candidate determination unit 31 converts the internal vehicle information, external vehicle information, and information from the response history input to the candidate determination model into vectors according to a prescribed transformation rule, and inputs these vectors into the candidate determination model. Furthermore, the candidate determination model can also be a model composed of multiple stacked blocks containing attention and feed-forward sublayers. In this case, a layer is also provided as the output layer to calculate the confidence value representing the type of each response using a softmax operation. In this case, the candidate determination unit 31 converts in-vehicle information, external vehicle information, and information input to the model from the response history into text data (e.g., the sensor names for in-vehicle information and the values of sensor signals used as in-vehicle information are presented as text data), and inputs this text data into the candidate determination model. The candidate determination unit 31 determines responses of types with a confidence value greater than or equal to a predetermined threshold, or responses of types with a predetermined number (wherein an integer greater than or equal to 2) of confidence values arranged from high to low, as candidate responses. By using such a candidate determination model, the candidate determination unit 31 can determine multiple candidate responses appropriate to the current state of the vehicle 1. Such candidate determination models follow a training learning method corresponding to the model (e.g., backpropagation of errors), and are pre-learned using training data that includes a combination of information input to the candidate determination model from a variety of external information, internal information, and response history, along with the content of the executed response.
[0031] Alternatively, for each type of response, a range of corresponding values for in-vehicle information, external information, and response history can be preset. Furthermore, the candidate determination unit 31 can also determine responses of types whose values for the latest in-vehicle information, external information, and response history within the most recent specified period are included within the preset range as response candidates.
[0032] When multiple response candidates are determined, for each of the multiple response candidates, the candidate determination unit 31 determines a combination of hypothetical speech data, which is text data representing the hypothetical speech content assumed when requesting the response, and command data for executing the response. Therefore, the candidate determination unit 31 determines the corresponding combination of hypothetical speech data and command data for each response candidate by referring to a table prepared in advance according to the type of each response, representing combinations of hypothetical speech data and command data. It should be noted that such a table is stored in advance in the memory 22. Furthermore, the command data represents the device that becomes the object of control through response processing and the control content in that response processing.
[0033] The candidate determination unit 31 performs the above-described processing at each predetermined cycle, when the vehicle 1 travels a predetermined distance each time, or when the value of the external sensor signal included in the external information or the value of the internal sensor signal included in the internal information changes by a predetermined update threshold each time, thereby appropriately updating the candidates for multiple responses. Alternatively, the candidate determination unit 31 may also update the candidates for multiple responses by performing the above-described processing each time a predetermined period has elapsed since the execution of the previous response processing.
[0034] The recognition unit 32 identifies the speech content of the occupants of vehicle 1 from the speech signal collected by microphone 4. Therefore, the recognition unit 32 inputs a speech signal generated by microphone 4 into a speech recognition model if the average volume of the speech signal over the most recent specified period exceeds a speech detection threshold, thereby recognizing the content of the speech signal and generating text data representing the speech content as actual speech data. Such a speech recognition model is configured, for example, as a DNN with an attention mechanism, or a recurrent neural network (RNN) or a long short-term memory network (LSTM) with a recursive structure. Alternatively, the speech recognition model can be configured as a GMM-HMM based on a mixture of normal distribution and hidden Markov model, or a DNN-HMM based on a DNN and a hidden Markov model. It should be noted that the recognition unit 32 can also segment the speech signal into frames of a specified time length, extract speech features for each frame, and input the features of each frame into the speech recognition model in chronological order, thereby recognizing the content of the speech signal. Furthermore, the feature quantities of each frame can be set as, for example, the specified elements of the cepstrum of that frame.
[0035] When the actual speech data representing the content of the speech is obtained, the recognition unit 32 outputs the actual speech data to the selection unit 33.
[0036] The selection unit 33 selects the candidate response from among multiple candidate responses that corresponds to the assumed speech data that is most consistent with the identified speech content. Therefore, the selection unit 33 calculates the degree of consistency between the actual speech data representing the speech content and the assumed speech data for each of the multiple candidate responses. Then, the selection unit 33 selects the candidate response corresponding to the assumed speech data with the highest degree of consistency as the data that is most consistent with the speech content.
[0037] For each of the multiple response candidates, the selection unit 33 calculates the Levenshtein distance between the assumed speech data and the actual speech data as the degree of consistency. In this case, the smaller the Levenshtein distance, the higher the degree of consistency. Alternatively, the selection unit 33 may also calculate other metrics representing the degree of consistency between two text data, such as edit distance, as the degree of consistency. Alternatively, the selection unit 33 may also calculate the degree of consistency between the actual speech data and each assumed speech data by inputting the actual speech data and each assumed speech data into a model that has been pre-learned in a manner that calculates the degree of consistency between the actual speech data and each assumed speech data. The model for calculating the degree of consistency is configured as an LLM consisting of multiple blocks stacked together, including attention sublayers and feedforward sublayers. In this case, a layer is provided as the output layer of the model to calculate a value representing the degree of consistency for each assumed speech data using a softmax operation or a sigmoid operation. It should be noted that the LLM for calculating the degree of consistency can employ a model with relatively less computational load compared to the LLM that generates command data from the actual speech data, as described later, due to the limitation of the types of output.
[0038] It should be noted that, for the assumed speech data with the highest degree of consistency, if this degree of consistency is also less than the predetermined selection threshold, the probability of finding a response whose content differs from all candidate responses is high. Therefore, in such cases, the selection unit 33 may not select any of the candidate responses. It should also be noted that, when calculating a value that decreases as consistency increases, such as the Levenstein distance, as an indicator of consistency, the selection unit 33 can compare the reciprocal of this indicator value or the reciprocal of the value obtained by adding a predetermined offset (e.g., 1) to the selection threshold.
[0039] The selection unit 33 notifies the response processing unit 34 of the selected response candidates and their corresponding command data. Furthermore, if no response candidate is selected, the selection unit 33 notifies the response processing unit 34 of the topic and actual speech data.
[0040] The response processing unit 34 executes a response to the command data corresponding to the selected response candidate, which is received from the selection unit 33. For example, if the command data indicates that the air conditioning unit is the controlled device and instructs that the set temperature be changed to a specified temperature as the response processing content, the response processing unit 34 outputs a control signal to the air conditioning unit via the communication interface 21 to change the set temperature to the specified temperature. Furthermore, if the command data indicates that the wireless communication terminal is the controlled device and instructs that the search for a specified item be performed as the response processing content, the response processing unit 34 outputs a signal requesting the search for the specified item to a search server (not shown) located outside the vehicle 1 via the communication interface 21 and the wireless communication terminal. Then, when the search result is received from the search server via the wireless communication terminal and the communication interface 21, the response processing unit 34 causes the notification device 5 to display the search result via the communication interface 21. Alternatively, if the command data indicates that the notification device 5 is the controlled device and instructs that the playback of video content be performed as the response processing content, the response processing unit 34 outputs a control signal to the notification device 5 to play the video content via the communication interface 21.
[0041] Furthermore, if the selection unit 33 is notified that none of the multiple response candidates have been selected, the response processing unit 34 generates command data by inputting the actual speech data into a generative model that has been pre-learned to generate command data corresponding to the speech content. Then, the response processing unit 34 performs response processing according to the generated command data. It should be noted that the generative model is configured as an LLM consisting of multiple stacked blocks containing attention sublayers and feedforward sublayers.
[0042] Figure 4 This diagram illustrates an overview of the response processing according to this embodiment. Figure 4 The system has three response candidates 401 to 403. Candidate 401 is a combination of command data indicating that the air conditioning unit's set temperature should be lowered by 2°C and the assumed statement "hot". Candidate 402 is a combination of command data indicating that a search for restaurants near the current location should be conducted via a wireless communication terminal and the assumed statement "I'm hungry". Candidate 403 is a combination of command data indicating that audio content should be played via a notification device and the assumed statement "Play some music". Then, when passenger 410 says "It's so hot", candidate 401, which corresponds to the assumed statement "hot" most closely to the passenger's statement, is selected, and the response process of lowering the air conditioning unit's set temperature by 2°C is executed.
[0043] When the response processing ends, the response processing unit 34 stores the response execution data, representing the content of the performed response processing, in the memory 22 as a response history. It should be noted that the response processing unit 34 may also include actual speech data associated with the response processing, as well as hypothetical speech data and command data regarding the selected response candidates, in the response execution data. Furthermore, the response processing unit 34 may also include the vehicle 1's location, date and time, in-vehicle information, and external information at the time the response processing was performed in the response execution data.
[0044] Figure 5 This is an action flowchart of the automatic response processing according to this embodiment. The processor 23 follows this action flowchart to perform the automatic response processing.
[0045] The candidate determination unit 31 determines multiple response candidates based on at least one of in-vehicle information, out-of-vehicle information, and response history, and determines a combination of command data and assumed speech data for each response candidate (step S101). The recognition unit 32 recognizes the speech content of any passenger based on the voice signal generated by the microphone 4 and generates actual speech data representing that speech content (step S102). Then, the selection unit 33 selects the candidate response from the multiple response candidates that corresponds to the assumed speech data that is most consistent with the passenger's speech content (step S103).
[0046] The response processing unit 34 determines whether the degree of consistency between the assumed speech data and the actual speech data of the selected response candidate is greater than or equal to a predetermined selection threshold (step S104). If the degree of consistency is greater than or equal to the selection threshold (step S104 - Yes), the response processing unit 34 performs response processing that follows the command data of the selected response candidate (step S105).
[0047] On the other hand, if the consistency level is less than the selection threshold (step S104 - No), the response processing unit 34 generates command data corresponding to the speech content by inputting the actual speech data into the generation model for command data generation (step S106). Then, the response processing unit 34 performs response processing that follows the generated command data (step S107).
[0048] As explained above, the automatic response device selects the candidate from a pool of pre-determined response candidates based on in-vehicle information, whose associated hypothetical speech data most closely matches the actual speech data representing the passenger's speech content. Then, the automatic response device executes response processing following the command data regarding the selected response candidate. Thus, the automatic response device can perform response processing corresponding to the content of the passenger's speech with relatively simple processing after the passenger speaks. As a result, the automatic response device can shorten the waiting time from the passenger's speech to the execution of the response.
[0049] Furthermore, given multiple candidate responses, assuming the consistency between the spoken data and the actual spoken data is less than a predetermined selection threshold, the automatic response device generates command data by inputting the actual spoken data into a generation model for command data generation. Therefore, even if the pre-prepared candidate responses do not match the passenger's spoken content, the automatic response device can still perform appropriate response processing corresponding to the spoken content.
[0050] According to a variation, the processor 23 of the automatic response device 6 may also include an update unit for updating the candidate determination model. The update unit uses a combination of command data stored as response history in memory 22 when none of the candidate responses have been selected, in-vehicle information, external information, and information input to the candidate determination model from the response history, as training data, and updates the candidate determination model following a training learning method corresponding to the candidate determination model. In this case, the update unit may also update a portion of the weight coefficients of the candidate determination model using the LoRA method. Alternatively, the update unit may replace the combination of command data associated with in-vehicle information and assumed speech data when none of the candidate responses have been selected with command data and actual speech data generated by the generative model. Thus, the algorithm for determining candidate responses using command data and actual speech data generated by the generative model when none of the pre-prepared candidate responses have been selected is updated, thereby including more suitable candidates among the candidate responses.
[0051] A computer program that implements automatic response processing according to the above-described embodiments or variations can be provided as a computer program product, for example, in the form of a computer-readable removable recording medium.
[0052] As described above, those skilled in the art can make various modifications within the scope of this invention and in accordance with the manner of implementation.
Claims
1. An automatic response device, comprising: The candidate determination unit determines a plurality of candidate responses that are likely to be requested by the passenger based on at least one of past response history to statements made by the occupants of the vehicle, external information indicating the conditions around the vehicle, and internal information indicating the conditions inside the vehicle. For each of the plurality of candidate responses, it determines a combination of hypothetical speech data representing the assumed speech content when requesting the response and command data for executing the response. The recognition unit identifies the content of the passenger's speech from the voice signals collected by the microphone installed in the vehicle; The selection unit selects the candidate response from the plurality of candidate responses that corresponds to the assumed speech data that is most consistent with the identified speech content; as well as The response processing unit executes a response that follows the command data regarding the selected candidate responses.
2. The automatic response device according to claim 1, wherein, The candidate determination unit determines the candidates for the plurality of responses by inputting at least one of the response history, the external information, and the internal information into a candidate determination model that has been pre-learned in a manner for determining the candidates for the response.
3. The automatic response device according to claim 1 or 2, wherein, The selection unit calculates the degree of consistency between the assumed speech data and the identified speech content for each of the plurality of response candidates. If the degree of consistency for any one of the plurality of response candidates is less than a predetermined selection threshold, then none of the plurality of response candidates is selected. If none of the multiple candidate responses are selected, the response processing unit generates the command data by inputting the identified speech content into a generative model that has been pre-learned to generate command data corresponding to the speech content, and executes a response that follows the generated command data.
4. An automatic response method, comprising: Based on at least one of past response history to statements made by the occupants of the vehicle, external information indicating the conditions around the vehicle, and internal information indicating the conditions inside the vehicle, a plurality of candidate responses that are likely to be requested by the occupants are determined. For each of the plurality of candidate responses, a combination of hypothetical speech data representing the assumed speech content at the time the response was requested and command data for executing the response is determined; Identify the occupant's speech from voice signals collected via microphones located in the vehicle; Select the candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data that is most consistent with the identified speech content; as well as Execute the response following the command data of the selected candidate responses.
5. A computer program product for automatic response, comprising instructions for causing a computer to perform the following actions: Based on at least one of past response history to statements made by the occupants of the vehicle, external information indicating the conditions around the vehicle, and internal information indicating the conditions inside the vehicle, a plurality of candidate responses that are likely to be requested by the occupants are determined. For each of the plurality of candidate responses, a combination of hypothetical speech data representing the assumed speech content at the time the response was requested and command data for executing the response is determined; Identify the occupant's speech from voice signals collected via microphones located in the vehicle; Select the candidate response from the plurality of candidate responses that corresponds to the hypothetical speech data that is most consistent with the identified speech content; as well as Execute the response following the command data of the selected candidate responses.
Citation Information
Patent Citations
Instrument control device for vehicle
JP2006137366A