Automatic response system, automatic response device, automatic response method, and computer program for automatic response
The automated response system addresses the challenge of providing appropriate vehicle responses by using combined vehicle and remote models to generate and deliver timely, detailed interactions with passengers.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2026-03-25
Smart Images

Figure 2026053118000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an automatic response system, an automatic response device, an automatic response method, and a computer program for automatic response that automatically respond to speech from vehicle occupants.
Background Art
[0002] An information providing method for providing information according to the preferences of vehicle passengers has been proposed (see Patent Document 1). This information providing method recognizes the behavior of the passengers based on the position information of the vehicle and / or the voice uttered by the passengers, and stores a behavior record that is a record related to the behavior of the passengers. Then, this information providing method estimates the preferences of the passengers regarding the information to be provided to the passengers based on the voice, and adjusts the content of the information to be provided to the passengers based on the estimated preferences when providing the information related to the behavior record to the passengers.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In order for a passenger to obtain a sense of satisfaction with the response to the speech from the vehicle occupant, it is required that the response includes appropriate content according to the speech content.
[0005] Therefore, an object of the present invention is to provide an automatic response system that can give an appropriate response to speech from a passenger in a vehicle.
Means for Solving the Problems
[0006] According to one embodiment, an automated response system is provided. This automated response system includes an automated response device mounted on a vehicle and a server located outside the vehicle. The automated response device includes a first response generation unit that generates first response information by inputting input information, including voice information representing the content of a utterance from a vehicle occupant, into a first generation model mounted on the vehicle and pre-trained to generate first response information for the content of the utterance, and also generates inquiry information representing the content of the utterance based on the input information; a transmission and reception processing unit that transmits the inquiry information to the server via a communication device mounted on the vehicle and receives second response information generated from the server based on the inquiry information; and a response processing unit that responds to the occupant based on at least one of the first response information and the second response information. The server also includes a second response generation unit that generates second response information by inputting inquiry information into a second generation model that is pre-trained to generate second response information and is larger than the first generation model.
[0007] In one embodiment, a first generation model is pre-trained to generate query information along with first response information in response to input information. The first response generation unit then generates query information by inputting the input information into the first generation model.
[0008] In one embodiment, the response processing unit responds to the crew based on the first response information generated, and then responds to the crew further based on the second response information received from the server after that response.
[0009] In one embodiment, the response processing unit generates a third response information by inputting the second response information into the first generation model, and responds to the occupant based on the generated third response information.
[0010] In one embodiment, the response processing unit notifies the occupant of a predetermined waiting response via a notification device installed in the vehicle during the waiting period from responding to the occupant based on the first response information until receiving the second response information.
[0011] According to another embodiment, an automated response method is provided. This automated response method includes: generating first response information by inputting input information, including voice information representing the content of a speech spoken by a vehicle occupant, into a first generative model installed in the vehicle and pre-trained to generate first response information for the content of the speech; generating query information representing the content of the speech based on the input information; generating second response information by inputting the query information into a second generative model, which is larger than the first generative model and installed outside the vehicle, and pre-trained to generate second response information for the query information; and responding to the occupant based on at least one of the first response information and the second response information.
[0012] In yet another embodiment, an automated response computer program is provided. This automated response computer program includes instructions for a computer to perform the following actions: input input information, including voice information representing the content of a utterance from a vehicle occupant, into a first generative model installed in the vehicle and pre-trained to generate first response information for the content of the utterance, thereby generating first response information; generate query information representing the content of the utterance based on the input information; input the query information into a second generative model, which is larger than the first generative model and installed outside the vehicle, and pre-trained to generate second response information for the query information, thereby generating second response information; and respond to the occupant based on at least one of the first response information and the second response information.
[0013] In yet another embodiment, an automatic response device is provided. This automatic response device includes: a first response generation unit that generates first response information by inputting input information, including voice information representing the content of a utterance from a vehicle occupant, into a first generation model mounted in the vehicle and pre-trained to generate first response information for the content of the utterance, and also generates inquiry information representing the content of the utterance based on the input information; a transmission and reception processing unit that transmits the inquiry information to a server located outside the vehicle via a communication device mounted in the vehicle and receives second response information generated based on the inquiry information from the server; and a response processing unit that responds to the occupant based on at least one of the first response information and the second response information. [Effects of the Invention]
[0014] The automated response system described herein has the effect of being able to provide appropriate responses to speech from occupants riding in a vehicle. [Brief explanation of the drawing]
[0015] [Figure 1] This is a schematic diagram of the automated response system. [Figure 2] This is a schematic diagram of a vehicle equipped with an automatic response system. [Figure 3] This is a hardware configuration diagram of an automated response system. [Figure 4] This is a functional block diagram of the processor in an automatic response system. [Figure 5] This is a hardware configuration diagram for the server. [Figure 6] This is a functional block diagram of the server's processor. [Figure 7] (a) and (b) are explanatory diagrams of the automated response process, respectively. [Figure 8] This is a flowchart illustrating the operation of the automated response process. [Modes for carrying out the invention]
[0016] Hereinafter, while referring to the drawings, an automatic response system, an automatic response method and an automatic response computer program executed by the automatic response system, and an automatic response device included in the automatic response system will be described. This automatic response system includes an automatic response device mounted on a vehicle and a server provided outside the vehicle. The automatic response device generates first response information using an in-vehicle first generation model for an utterance from any occupant of the vehicle, and transmits inquiry information corresponding to the content of the utterance to the server. Then, the server generates second response information for the received inquiry information by inputting it into a second generation model larger than the first generation model, and transmits the generated second response information to the vehicle. When the automatic response device receives the second response information from the server, it responds to the occupant based on at least one of the first response information and the second response information.
[0017] FIG. 1 is a schematic configuration diagram of the automatic response system. In the present embodiment, the automatic response system 1 includes an automatic response device 3 mounted on a vehicle 2 and a server 4. The automatic response device 3 is connected to the server 4 via a wireless base station 6 and a communication network 5 by accessing, for example, the wireless base station 6 connected to a communication network 5 to which the server 4 is connected and a gateway (not shown). In FIG. 1, only one vehicle 2 and one automatic response device 3 are shown, but the automatic response system 1 may include a plurality of vehicles 2 on which the automatic response device 3 is mounted. Similarly, a plurality of wireless base stations 6 may be connected to the communication network 5.
[0018] First, the vehicle 2 and the automatic response device 3 will be described. As described above, the automatic response system 1 may include a plurality of vehicles 2 on which the automatic response device 3 is mounted. However, since each vehicle 2 and each automatic response device 3 have the same configuration and execute the same process with respect to the automatic response process, hereinafter, one vehicle 2 and one automatic response device 3 will be described.
[0019] FIG. 2 is a schematic configuration diagram of a vehicle 2 equipped with an automatic response device 3. The vehicle 2 includes a camera 11, at least one microphone 12, a wireless communication terminal 13, a notification device 14, and the automatic response device 3. The camera 11, the microphone 12, the wireless communication terminal 13, the notification device 14, and the automatic response device 3 are communicably connected to each other.
[0020] The camera 11 is an example of an in-vehicle sensor, and is mounted inside the vehicle cabin facing the interior near the upper end of the windshield so that all passengers in the vehicle 2 are included in its imaging target area. Then, the camera 11 generates an image representing the interior of the vehicle 2 at each predetermined imaging cycle, and outputs the generated image to the automatic response device 3. Hereinafter, the image generated by the camera 11 is referred to as an in-vehicle image. Also, the in-vehicle image is an example of an in-vehicle sensor signal.
[0021] At least one microphone 12 is another example of an in-vehicle sensor, and picks up the voice emitted by any of the passengers in the vehicle 2 and outputs a voice signal representing the voice. For this purpose, each microphone 12 is mounted inside the vehicle cabin of the vehicle 2. A plurality of microphones 12 may be mounted in an array, or a microphone 12 may be mounted around each seat inside the vehicle cabin of the vehicle 2 for each individual seat. Each microphone 12 outputs the generated voice signal to the automatic response device 3. The voice signal generated by each microphone 12 is another example of an in-vehicle sensor signal.
[0022] The wireless communication terminal 13 is an example of a communication device, and is a device that performs wireless communication processing in accordance with a predetermined wireless communication standard. For example, by accessing the wireless base station 6, it connects to the server 4 via the wireless base station 6 and the communication network 5. That is, a communication line is established between the wireless communication terminal 13 and the server 4 via the wireless base station 6 and the communication network 5. The wireless communication terminal 13 then generates an uplink wireless signal containing the inquiry information received from the automatic response device 3 and transmits that wireless signal to the wireless base station 6. As a result, the inquiry information is transmitted to the server 4. The wireless communication terminal 13 also receives a downlink wireless signal containing second response information from the wireless base station 6 and outputs the second response information to the automatic response device 3. As a result, the second response information generated by the server 4 is transmitted to the automatic response device 3.
[0023] The notification device 14 is installed inside the vehicle 2 and notifies the occupant of the response content represented in the response information generated by the automatic response device 3 or the server 4. For this purpose, the notification device 14 has, for example, at least one of a speaker or a display device. When the notification device 14 receives a notification signal from the automatic response device 3 that represents the response content to the occupant, it notifies the driver of the response content by voice from the speaker or by displaying a message, image, or video on the display device. The display device or speaker of the notification device 14 may be mounted for each seat, facing the occupant seated in that seat.
[0024] The automated response device 3 generates first response information in response to a utterance from any of the occupants of the vehicle 2. Furthermore, the automated response device 3 generates inquiry information representing the content of the utterance and transmits the generated inquiry information to the server 4 via the wireless communication terminal 13. Furthermore, the automated response device 3 receives second response information in response to the inquiry information from the server 4 via the wireless communication terminal 13. The automated response device 3 then responds to the occupant based on at least one of the first response information and the second response information.
[0025] Figure 3 is a hardware configuration diagram of the automatic response device 3. As shown in Figure 3, the automatic response device 3 includes a communication interface 21, a memory 22, and a processor 23. The communication interface 21, the memory 22, and the processor 23 may each be configured as separate circuits, or they may be integrated as a single integrated circuit.
[0026] The communication interface 21 has an interface circuit for connecting the automatic response device 3 to other equipment inside the vehicle. The communication interface 21 passes the in-vehicle image received from the camera 11 and the audio signals received from each of the individual microphones 12 to the processor 23. The communication interface 21 also outputs the notification signal received from the processor 23 to the notification device 14, or outputs a control command received from the processor 23 for any of the in-vehicle equipment. Furthermore, the communication interface 21 outputs the inquiry information received from the processor 23 to the wireless communication terminal 13, and conversely outputs the second response information received from the wireless communication terminal 13 to the processor 23.
[0027] Memory 22 is an example of a storage unit and includes, for example, volatile semiconductor memory and non-volatile semiconductor memory. Memory 22 stores various data used in the automatic response processing performed by the processor 23. Specifically, memory 22 stores a parameter set that defines a first generation model for generating response information. Furthermore, memory 22 may temporarily store in-vehicle images received from camera 11 and audio signals received from individual microphones 12. Furthermore, memory 22 temporarily stores second response information. Moreover, memory 22 stores waiting response messages during the waiting period until the reception of the second response information is complete.
[0028] The processor 23 has one or more CPUs (Central Processing Units) and their peripheral circuits. The processor 23 may further have other arithmetic circuits such as a logic unit, a numerical unit, or a graphics processing unit. The processor 23 then performs automatic response processing.
[0029] Figure 4 is a functional block diagram of the processor 23 related to automatic response processing. The processor 23 has a first response generation unit 31, a transmission / reception processing unit 32, and a response processing unit 33. Each of these parts of the processor 23 is, for example, a functional module realized by a computer program running on the processor 23. Alternatively, each of these parts of the processor 23 may be a dedicated arithmetic circuit provided on the processor 23.
[0030] The first response generation unit 31 receives predetermined input information, including voice information representing the content of an utterance by any of the occupants of the vehicle 2, into the first generation model, thereby generating first response information for that utterance. Furthermore, the first response generation unit 31 generates inquiry information corresponding to the content of that utterance.
[0031] The first response generation unit 31 recognizes the content of the utterance expressed in the speech signal by inputting speech signals from the individual microphones 12 whose average volume over the most recent predetermined period exceeds a speech detection threshold, in order to generate speech information representing the content of the utterance, and generates a string representing the content of the utterance as speech information. Such a speech recognition model is configured as, for example, a deep neural network (DNN) with an attention mechanism, or a DNN with a recurrent structure such as an RNN or LSTM. Alternatively, the speech recognition model may be configured as a GMM-HMM based on a mixture of normal distributions and a hidden Markov model, or a DNN-HMM based on a DNN and a hidden Markov model. The first response generation unit 31 may also divide the speech signal into frames with a predetermined time length, extract speech features from each frame, and input the feature quantities for each frame into the speech recognition model in chronological order to recognize the content of the utterance expressed in the speech signal. The feature quantities for each frame can be, for example, predetermined elements of the cepstrum of that frame.
[0032] In this embodiment, the first generative model is configured as a Large Language Model (LLM). The LLM, which is the first generative model, is configured as a stack of multiple blocks, for example, each containing an attention layer and a feed-forward layer. The first response generation unit 31 inputs a string representing the content of the utterance to the LLM. As a result, the first generative model outputs text data representing the response content corresponding to the content of the utterance as the first response information.
[0033] Furthermore, the first response information is not limited to the response content notified to each occupant via the notification device 14, but may also include response content that controls the vehicle 2 itself or equipment installed in the vehicle 2.
[0034] The first generative model is a relatively small computational scale so that even when implemented in the processor 23 of the in-vehicle automatic response system 3, it can generate the first response information in a short time so that the occupant does not feel stressed by the occupant's speech. Therefore, the response content included in the first response information generated by the first generative model is relatively simpler than the response content included in the second response information generated by the second generative model.
[0035] In a modified version, the first response generation unit 31 may include, along with audio information, an in-vehicle image, or a sub-region on the in-vehicle image representing individual occupants, as input information to the first generation model. In this case, the first generation model is configured as a Vision Language Model (VLM). By receiving the in-vehicle image or sub-region representing occupants in this way, the first generation model can generate first response information by referring to the occupants' states.
[0036] When a subregion representing an occupant is input to the generative model, the first response generation unit 31 detects the subregion representing the occupant by inputting the in-vehicle image to a classifier that has been pre-trained to detect occupants. The classifier for occupant detection is configured as a convolutional neural network (CNN), such as a Single Shot MultiBox Detector or Faster R-CNN, or as a DNN with an attention mechanism, such as a Vision Transformer.
[0037] Furthermore, when a subregion representing an occupant is input to the first generation model, the first response generation unit 31 can either crop the subregion representing that occupant from the in-vehicle image for each occupant, or mask the areas outside the subregion by converting the values of individual pixels in the areas outside the subregion representing each occupant to predetermined pixel values.
[0038] The first response generation unit 31 may further include in its input information at least one of the following: the current position of the vehicle 2, the destination, sensor signals obtained from sensors installed in the vehicle 2 for detecting the behavior of the vehicle 2, the conditions inside the vehicle 2, or the conditions around the vehicle 2, the amount of operation of the vehicle 2 by the occupants, and signals representing the settings of the in-vehicle equipment. The sensors for detecting the behavior of the vehicle 2 are, for example, speed sensors or acceleration sensors. The sensors for detecting the conditions inside the vehicle 2 or the conditions around the vehicle 2 are, for example, thermometers, light meters, or rain sensors. The amount of operation of the vehicle 2 by the occupants are, for example, accelerator opening, brake amount, or steering angle. The settings of the in-vehicle equipment are, for example, air conditioning set temperature, set airflow, window open / closed status, or audio set volume. The current position of the vehicle 2 is determined by a positioning device (not shown) based on a satellite positioning system, such as a GPS receiver, installed in the vehicle 2. Furthermore, the destination of vehicle 2 can be obtained from a navigation device (not shown) installed in vehicle 2. Hereinafter, these signals and information will be referred to as vehicle status information. When vehicle status information is input to the first generation model, the first generation model can refer to the status of vehicle 2 or the in-vehicle equipment, and thus can generate more appropriate response information as the first response information.
[0039] The first response generation unit 31 can generate text data to input to the first generation model by converting the type and signal value of each sensor signal included in the vehicle state information into a string, and then combining the converted string with a string representing the audio information. Alternatively, the first generation model may have an input layer for inputting vehicle state information, separate from the block into which the audio information is input. In this case, only the string corresponding to the audio information is input to the input-side block of the stacked blocks, and the vehicle state information is input to the input layer for vehicle state information. The vehicle state information input to that input layer is then captured by a block in the stacked blocks of the generation model that has a cross-attention mechanism that calculates cross-attention between the vehicle state information and the output from the previous block. In this case, a tuning method such as LoRA may be applied to the learning of the first generation model regarding the acquisition of vehicle state information.
[0040] The first response generation unit 31 further generates query information. In this case, the first response generation unit 31 includes all the input information input to the first generative model in the query information. Alternatively, the first generative model may be configured to generate query information along with the first response information based on the input information. In this case, the first generative model is configured such that a stack of blocks branches off from the middle, and one or more blocks that generate the first response information and one or more blocks that generate query information are provided in parallel. Each block after branching is also configured to include an attention layer and a feed-forward layer. The first response information and query information are determined separately according to the output probability calculated by a softmax operation on the output from the corresponding final stage block. In this case as well, a tuning method such as LoRA may be applied to the learning of the first generative model regarding which of the first response information and query information is added to the base model. The first generative model is pre-trained to output text data in which the main points of the utterance expressed in the input speech information are clarified as query information. For example, if the occupant's utterance is "How much further until we reach XX?", the first generation model outputs text data as query information such as "What is the estimated time required to reach XX from △△?" (where △△ is the current location of vehicle 2). In this way, by the first generation model generating query information from the audio signal representing the occupant's utterance, the first response generation unit 31 can generate query information that more clearly captures the essence of the utterance.
[0041] The first response generation unit 31 outputs the first response information to the response processing unit 33 and also outputs the inquiry information to the transmission / reception processing unit 32.
[0042] When the transmission / reception processing unit 32 receives the query information, it includes the identification information of the vehicle 2 or the identification information of the wireless communication terminal 13 installed in the vehicle 2 in the query information. The transmission / reception processing unit 32 then transmits the query information, including the identification information of the vehicle 2 or the wireless communication terminal 13, to the server 4 via the wireless communication terminal 13.
[0043] Furthermore, when the transmission / reception processing unit 32 begins receiving the second response information from the server 4 via the wireless communication terminal 13, it sequentially outputs the received portion of the second response information to the response processing unit 33.
[0044] The response processing unit 33 performs a response process to respond to the occupants of the vehicle 2 based on at least one of the first response information and the second response information. In this embodiment, the response process includes not only making some kind of notification to the occupants via the notification device 14 according to the response information, but also controlling the vehicle 2 or various devices installed in the vehicle 2 according to the response information.
[0045] When the response processing unit 33 receives the first response information, it generates a notification signal representing the response content included in the first response information and outputs the generated notification signal to the notification device 14 via the communication interface 21. For example, the response processing unit 33 generates an audio signal representing the response content as a notification signal according to a predetermined speech synthesis method, based on the text data representing the response content included in the first response information. The response processing unit 33 then outputs the notification signal to the speaker of the notification device 14, causing the speaker to output audio representing the response content. Alternatively, the response processing unit 33 includes text data representing the response content in the notification signal. The response processing unit 33 then causes the display device of the notification device 14 to display the text data representing the response content.
[0046] Furthermore, the response processing unit 33 controls the device specified by the response content included in the first response information according to the response content. The response processing unit 33 determines the device to be controlled and the control command by referring to a control reference table that shows the correspondence between the text data representing the response content included in the first response information and the device to be controlled (including the vehicle 2 itself) and the control command for executing that control. Then, the response processing unit 33 outputs the specified control command to the electronic control unit (ECU) of the device to be controlled via the communication interface 21.
[0047] The controlled equipment may include, in addition to the air conditioning system, windows, door locks, interior lights, and seats. The response processing unit 33 controls the equipment according to the response content, such as opening and closing the windows, locking and unlocking the doors, turning the interior lights on or off, or adjusting the seat position of the seat where any occupant is seated.
[0048] Furthermore, if the text data representing the response does not contain any of the multiple words registered in the control reference table that identify the controlled device, the response information does not control the device. In this case, the response processing unit 33 does not output a control signal.
[0049] Once the response output based on the first response information is completed, the response processing unit 33 responds according to the second response information. In this case, the response processing unit 33 can perform the same processing for the second response information as it did for the response based on the first response information. If the response content included in the first response information matches the response content included in the second response information, the response processing unit 33 may omit the response according to the second response information.
[0050] Furthermore, the response processing unit 33 may generate third response information by inputting the second response information into the first generative model. In this case, the first generative model is not only pre-trained to generate first response information (and possibly query information in addition thereto) when the above input information is received, but is also pre-trained to generate a response output to the crew corresponding to the input when text data representing the response content, which is included in the second response information, is input. Therefore, the response processing unit 33 only needs to input the text data representing the response content, which is included in the second response information, into the first generative model. Then, the response processing unit 33 can perform response processing similar to the response processing based on the first response information, based on the third response information. By responding based on the third response information generated by reusing the first generative model used to generate the first response information, the response processing unit 33 can provide a more natural response to the crew. Furthermore, when generating the third response information, the response processing unit 33 may input text data obtained by combining the text data representing the response content included in the first response information with the text data representing the response content included in the second response information into the first generation model. This allows the response processing unit 33 to further improve the consistency between the response content included in the first response information and the response content included in the third response information.
[0051] In a modified example, the response processing unit 33 may not perform response processing based on the first response information until it has finished receiving the second response information, and once it has finished receiving the second response information, it may generate the third response information using the first generation model as described above. The response processing unit 33 may then perform response processing based on the third response information.
[0052] In another variation, the response processing unit 33 may not perform response processing based on the first response information, but may perform response processing based on the second response information once the reception of the second response information is complete.
[0053] In another variation, if the reception of the second response information has not been completed when the response processing based on the first response information is completed, the response processing unit 33 may notify the occupant of a predetermined waiting response message via the notification device 14 during the waiting period from the completion of the response processing based on the first response information until the completion of the reception of the second response information. The waiting response message can be, for example, a message informing the occupant that the response is not yet complete, such as "Please wait a moment" or "We are currently processing your inquiry." Notifying the occupant of such a waiting response message prevents a lack of response even if there is a certain time lag between the completion of the response processing based on the first response information and the receipt of the second response information. Therefore, the response processing unit 33 can provide a more natural response to the occupant.
[0054] Furthermore, if the response processing unit 33 is still receiving the second response information when the response processing based on the first response information is completed, it may generate a waiting response message by inputting a portion of the second response information received up to that point into the first generation model. If the portion of the second response information that has been received contains a flag indicating the end of the second response information, the response processing unit 33 will determine that it has finished receiving the second response information. If the flag is not present, the response processing unit 33 will determine that it is still continuing to receive the second response information. In this case, the response processing unit 33 may sequentially input the portion of the second response information it receives into the first generation model each time it receives a portion of the second response information. The text data output from the first generation model when the response processing based on the first response information is completed can then be used as the waiting response message. Alternatively, the response processing unit 33 may generate a waiting response message by re-inputting the text data output from the first generation model based on the portions of the second response information that have been sequentially input up to that point into the first generation model when the response processing based on the first response information is completed. Furthermore, if the length of the text data included in a portion of the second response information received at the time the response processing based on the first response information is completed is less than a predetermined lower threshold, the response processing unit 33 may send a standby response message, which is pre-stored in the memory 22, via the notification device 14. This prevents the notification of a standby response message that is meaningless because the portion of the second response information received at the time the response processing based on the first response information is completed is too small.
[0055] Next, let's describe Server 4. Server 4 has a second generative model implemented, and it uses this second generative model to generate a second response to the query information.
[0056] Figure 5 is a hardware configuration diagram of server 4. Server 4 includes a communication interface 41, a storage device 42, memory 43, and a processor 44. The communication interface 41, storage device 42, and memory 43 are connected to the processor 44 via signal lines.
[0057] The communication interface 41 is an example of a communication unit and has an interface circuit for connecting the server 4 to the communication network 5. The communication interface 41 is configured to communicate with the automatic response device 3 mounted on the vehicle 2 via the communication network 5, the wireless base station 6, and the wireless communication terminal 13 mounted on the vehicle 2. Specifically, the communication interface 41 passes query information received from the automatic response device 3 of the vehicle 2 via the wireless communication terminal 13, the wireless base station 6, and the communication network 5 to the processor 44. The communication interface 41 also transmits second response information received from the processor 44 to the automatic response device 3 of the vehicle 2 via the communication network 5, the wireless base station 6, and the wireless communication terminal 13 of the vehicle 2. Furthermore, the communication interface 41 passes various types of information received from other servers connected via the communication network 5 (for example, servers that distribute traffic information or weather information) to the processor 44.
[0058] The storage device 42 includes, for example, a solid-state drive, a hard disk drive, or an optical recording medium and its access device. The storage device 42 also stores a parameter set that defines a second generation model. The storage device 42 may further store identification information of the vehicle 2 or identification information of the wireless communication terminal 13 mounted on the vehicle 2. Furthermore, the storage device 42 may store a computer program that runs on the processor 44 for performing automatic response processing on the server 4 side. In addition, the storage device 42 may store various information received from other servers.
[0059] The memory 43 includes, for example, a non-volatile semiconductor memory and a volatile semiconductor memory. The memory 43 temporarily stores various data that are generated during the execution of the automatic response process or used in the automatic response process.
[0060] The processor 44 has one or more CPUs (Central Processing Units) and their peripheral circuits. The processor 44 may also have other arithmetic circuits such as logical operation units or numerical operation units. The processor 44 then performs automatic response processing on the server 4 side.
[0061] Figure 6 is a functional block diagram of the processor 44 related to automatic response processing on the server side. The processor 44 has a transmission / reception processing unit 51 and a second response generation unit 52. Each of these parts of the processor 44 is, for example, a functional module realized by a computer program running on the processor 44. Alternatively, each of these parts of the processor 44 may be a dedicated arithmetic circuit provided on the processor 44.
[0062] When the transmission / reception processing unit 51 receives inquiry information from the automatic response device 3 of the vehicle 2, it outputs the information included in the inquiry information (text data, in-vehicle images, vehicle status information, etc.) to be input to the second generation model to the second response generation unit 52. When the transmission / reception processing unit 51 receives the second response information from the second response generation unit 52, it identifies the vehicle 2 or the wireless communication terminal 13 installed in the vehicle 2 that sent the inquiry information by referring to the identification information of the vehicle 2 or the wireless communication terminal 13 included in the inquiry information. The transmission / reception processing unit 51 then transmits the second response information to the identified wireless communication terminal 13 of the vehicle 2 via the communication interface 41, the communication network 5, and the wireless base station 6. At this time, the transmission / reception processing unit 51 transmits the second response information after the generation of the second response information by the second generation model is complete. Alternatively, the transmission / reception processing unit 51 may transmit a part of the second response information (for example, text data with a predetermined number of characters or words) each time the second generation model outputs a part of the second response information.
[0063] The second response generation unit 52 generates second response information by inputting the text data included in the inquiry information into the second generation model. Furthermore, if the inquiry information includes an in-vehicle image or a portion of the in-vehicle image in which an occupant is represented, the second response generation unit 52 also inputs the in-vehicle image or the portion of the in-vehicle image into the second generation model. Similarly, if the inquiry information includes vehicle status information, the second response generation unit 52 also inputs the vehicle status information into the second generation model.
[0064] The second generative model is configured as an LLM (except for VLM if in-vehicle images or sub-regions are also input), similar to the first generative model. The second generative model is a larger-scale generative model compared to the first; for example, the number of blocks containing attention mechanisms and feed-forward layers in the second generative model is greater than the number of such blocks in the first generative model. Furthermore, the amount of data in the dataset used to train the second generative model is greater than the amount of data in the dataset used to train the first generative model. Therefore, although the computational load of the second generative model is greater than that of the first generative model, the second generative model can generate second response information that contains more detailed or accurate response content than the first response information generated by the first generative model. For example, if a passenger's utterance requests an explanation of a specific object or event, the second generative model can generate second response information that contains a more detailed or accurate explanation of that object or event than the explanation of that object or event contained in the first response information generated by the first generative model. For example, if a crew member says, "Tell me about the XX building," the first response information generated by the first generative model would include relatively simple information such as, "The XX building is located in △△." In contrast, the second response information generated by the second generative model would include more detailed information such as, "The XX building is located in △△ and is □□ meters tall. There is a famous restaurant called ×× there." Also, if a crew member says, "How much longer until we get to XX?", the first response information generated by the first generative model would query the navigation system installed in vehicle 2 for the estimated time required from vehicle 2's current location to XX, and the result of that query would include something like, "It will take △ minutes."In contrast, the second response information generated by the second generative model refers to the latest traffic information provided by the traffic information server and the route search results from the vehicle's current location to its destination, calculated by a route search algorithm implemented on the server 4. As a result, the second response information represents a more accurate estimated travel time to the destination.
[0065] When the second response information is generated by the second generation model, the second response generation unit 52 outputs the second response information to the transmission / reception processing unit 51. The second response generation unit 52 may also output a portion of the second response information to the transmission / reception processing unit 51 each time the second generation model outputs a portion of the second response information, for example, text data with a predetermined number of characters.
[0066] Figures 7(a) and 7(b) are explanatory diagrams of the automatic response process, respectively. In the example shown in Figure 7(a), the automatic response device 3 of the vehicle 2 receives input information 701, which includes voice information representing the content of the occupant's speech, into the first generation model 702, thereby generating inquiry information 703. The inquiry information 703 is then sent to the server 4, where it is input into the second generation model 704, generating second response information 705. When the automatic response device 3 receives the second response information 705 from the server 4, the automatic response device 3 performs response processing to the occupant based on the second response information 705. In this example, as described above, the automatic response device 3 may perform response processing according to the second response information 705 itself, or it may perform response processing according to the third response information 706 generated by the first generation model 702 after inputting the second response information 705 into the first generation model 702. In this example, the first response information generated when input information 701 is input to the first generation model 702 does not necessarily have to be used in response processing, or the first response information may be input to the first generation model 702 together with the second response information 705 to be used in generating the third response information 706. In this example, since response processing is performed according to the response information generated by the second generation model 704, which is larger than the first generation model 702, the automated response system 1 can provide a detailed or accurate response. In this example, it is more preferable that the first generation model 702 is used to generate the query information 703, or that response processing is performed according to the third response information 706 generated when the second response information 705 is input to the first generation model 702. As described above, by using the first generation model 702 to generate the inquiry information 703, the main points of the crew member's utterance become clearer in the inquiry information 703, and the response content of the second response information 705 generated based on the inquiry information 703 becomes more appropriate. Furthermore, by executing the response processing according to the third response information 706, the automated response system 1 can provide a more natural response as described above.
[0067] In the example shown in Figure 7(b), input information 711, which includes voice information representing the content of the occupant's speech, is input to the first generation model 712 in the vehicle 2's automatic response device 3, generating first response information 713 and inquiry information 714. The automatic response device 3 performs response processing based on the first response information 713 and sends inquiry information 714 to the server 4. The server 4 then inputs the inquiry information 714 to the second generation model 715, generating second response information 716. When the automatic response device 3 receives the second response information 716 from the server 4, it performs response processing based on the second response information 716, following the response processing based on the first response information 713. In this example as well, as described above, the automatic response device 3 may perform response processing according to the second response information 716 itself, or it may perform response processing according to the third response information 717 generated by inputting the second response information 716 into the first generation model 712. Also in this example, the first response information 713 may be used to generate the third response information 717 by being input into the first generation model 712 together with the second response information 716. In this example, since response processing is performed first according to the first response information 713 while the server 4 is generating the second response information 716, the waiting time from the crew's utterance to the first response is shortened, thereby reducing stress on the crew. Furthermore, since response processing according to the second response information 716 or the third response information 717 is performed after the response processing according to the first response information 713, the crew can receive a more detailed or accurate response.
[0068] Figure 8 is an operation flowchart of the automatic response process according to this embodiment. The processor 23 of the automatic response device 3 and the processor 44 of the server 4 execute the automatic response process according to this operation flowchart.
[0069] The first response generation unit 31 in the processor 23 of the automatic response device 3 generates first response information by inputting input information, including voice information representing the content of speech by any occupant of the vehicle 2, into a first generation model (step S101). Furthermore, the first response generation unit 31 generates inquiry information based on the input information (step S102). As described above, the input information may include an in-vehicle image, a partial region on the in-vehicle image representing an occupant, or vehicle status information.
[0070] The transmit / receive processing unit 32 in the processor 23 of the automatic response device 3 transmits inquiry information to the server 4 via the wireless communication terminal 13, the wireless base station 6, and the communication network 5 (step S103). The response processing unit 33 in the processor 23 of the automatic response device 3 starts response processing based on the first response information (step S104).
[0071] The second response generation unit 52 in the processor 44 of server 4 generates second response information by inputting query information into a second generation model (step S105). The transmission / reception processing unit 51 in the processor 44 of server 4 then transmits the second response information to the automatic response device 3 via the communication network 5, the wireless base station 6, and the wireless communication terminal 13 of the vehicle 2 (step S106). When the automatic response device 3 receives the second response information, the response processing unit 33 performs response processing based on the second response information, following the response processing based on the first response information (step S107). As described above, the response processing based on the first response information in step S104 may be omitted. Also, in step S107, the response processing unit 33 may perform response processing based on third response information generated by inputting the second response information into the first generation model.
[0072] As explained above, this automated response system can provide appropriate responses to occupant speech by utilizing a relatively large generative model located outside the vehicle. Furthermore, by using a relatively small generative model installed inside the vehicle in conjunction with the server-side generative model, this automated response system can maintain the appropriateness of the response while reducing the waiting time for a response to occupant speech.
[0073] The computer program that implements the automated response processing according to the above embodiment or modification may be provided as a computer program product, for example, in the form of being recorded on a computer-readable portable recording medium.
[0074] As described above, those skilled in the art can make various modifications within the scope of the present invention to suit the implemented form. [Explanation of Symbols]
[0075] 1 Automatic response system, 2 Vehicle, 3 Automatic response device, 4 Server, 5 Communication network, 6 Wireless base station, 11 Camera, 12 Microphone, 13 Wireless communication terminal, 14 Notification device, 21 Communication interface, 22 Memory, 23 Processor, 31 First response generation unit, 32 Transmit / receive processing unit, 33 Response processing unit, 41 Communication interface, 42 Storage device, 43 Memory, 44 Processor, 51 Transmit / receive processing unit, 52 Second response generation unit
Claims
1. An automated response system comprising an automated response device mounted on a vehicle and a server located outside the vehicle, The aforementioned automatic response device, A first response generation unit generates first response information by inputting input information, including voice information representing the content of a speech spoken by an occupant of the vehicle, into a first generation model installed in the vehicle and pre-trained to generate first response information for the content of the speech, and also generates inquiry information representing the content of the speech based on the input information. A transmission and reception processing unit that transmits the aforementioned inquiry information to the server via a communication device mounted on the vehicle and receives a second response information generated from the server based on the aforementioned inquiry information, A response processing unit that responds to the occupant based on at least one of the first response information and the second response information, It has, The aforementioned server, The system includes a second response generation unit that generates the second response information by inputting the query information into a second generation model that is pre-trained to generate the second response information and is larger than the first generation model. Automated response system.
2. The first generation model is pre-trained to generate the query information along with the first response information in response to the input information, The automated response system according to claim 1, wherein the first response generation unit generates the inquiry information by inputting the input information into the first generation model.
3. The automatic response system according to claim 1 or 2, wherein the response processing unit responds to the occupant based on the first response information generated, and further responds to the occupant based on the second response information received from the server after the response.
4. The automatic response system according to claim 1 or 2, wherein the response processing unit generates a third response information by inputting the second response information into the first generation model, and responds to the occupant based on the generated third response information.
5. The automatic response system according to claim 1 or 2, wherein the response processing unit notifies the occupant of a predetermined waiting response via a notification device installed in the vehicle during a waiting period from the time it responds to the occupant based on the first response information until it receives the second response information.
6. Input information including voice information representing the content of speech from the occupants of the vehicle is input to a first generative model installed in the vehicle and pre-trained to generate first response information for the content of the speech, thereby generating the first response information. Based on the input information, query information representing the content of the utterance is generated. The query information is input to a second generation model, which is pre-trained to generate a second response information to the query information, is larger in scale than the first generation model, and is located outside the vehicle, thereby generating the second response information. Responding to the occupant based on at least one of the first response information and the second response information, An automated response method that includes the following.
7. Input information including voice information representing the content of speech from the occupants of the vehicle is input to a first generative model installed in the vehicle and pre-trained to generate first response information for the content of the speech, thereby generating the first response information. Based on the input information, query information representing the content of the utterance is generated. The query information is input to a second generation model, which is pre-trained to generate a second response information to the query information, is larger in scale than the first generation model, and is located outside the vehicle, thereby generating the second response information. Responding to the occupant based on at least one of the first response information and the second response information, An automated computer program for instructing a computer to perform a specific action.
8. A first response generation unit generates the first response information by inputting input information, including voice information representing the content of a speech spoken by a vehicle occupant, into a first generation model installed in the vehicle and pre-trained to generate first response information for the content of the speech, and also generates inquiry information representing the content of the speech based on the input information. A transmission and reception processing unit transmits the aforementioned inquiry information to a server located outside the vehicle via a communication device mounted on the vehicle, and receives a second response information generated based on the inquiry information from the server. A response processing unit that responds to the occupant based on at least one of the first response information and the second response information, An automatic response device having the following features.
Citation Information
Patent Citations
Information provision method and information provision device
JP2023124286A